
AI Personas Beat Humans At Building Exploitable Trust: The Real Risk Is Scale
AI Personas Beat Humans At Building Exploitable Trust: The Real Risk Is Scale
A controlled study found an AI persona outperformed human scammers at the long, friendly prelude to fraud, hinting that social engineering will scale with cheap compute more than with flashy deepfakes.
AI has long promised to make work more efficient. One controlled study suggests it may also make the worst work more efficient, especially the months long trust building that precedes many online scams. The unsettling part is not a new cinematic deepfake. It is that amiable small talk can now be automated at volume.
What the study actually tested
Researchers from four universities pitted a generative AI chatbot against human romance scammers in a weeklong simulation focused on the relationship building phase of a fraud often called pig butchering. The experiment asked 22 unwitting test subjects to chat with both a human and an AI persona, then measured how willing they were to follow a simple request at the end of the week. According to the researchers, 46 percent of subjects downloaded an app suggested by the AI persona, while 18 percent downloaded a video game suggested by the human. The AI chatbot was also rated as more trustworthy by the subjects. The team used a Claude agent for the automated conversations and described the human counterpart as an expert scammer.
The study draws on interviews with 145 former scam workers, including human trafficking survivors from compounds in Cambodia, Myanmar, and Laos. Those interviews, along with transcripts and playbooks, informed a model the researchers call hook, line, and sinker. First comes an intriguing opener, then a long period of friendly or romantic rapport, and finally the pitch that steers a victim to a fake investment. The researchers hypothesized that the middle stage, mostly innocuous conversation, maps well to what a large language model already does.
“With relatively little effort, we are able to make an agent that can outperform a human at building this exploitable emotional trust.”
That quote belongs to Yisroel Mirsky of Ben Gurion University, who led the work. He argues that automated rapport at scale could hand off to a human scammer only at the end, which also sidesteps vendor safeguards that try to prevent models from directly facilitating fraud. In other words, AI handles the grind, a person handles the close.
Directional, not universal
A few caveats matter. The test ran for a week, not months. The final asks were proxies, not actual wire transfers. The requests were not symmetric, since the human asked for a game download while the AI asked to try an app the persona said it had coded. The sample was small. The researchers are explicit about these constraints. Still, the design isolates something familiar to anyone who has observed online fraud operations. The hardest part to staff is the long con, and those messages are mostly banal, friendly, and patient. That is exactly the kind of output that a chatbot can produce indefinitely.
There is also a human cost subtext. Scam compounds in Southeast Asia have relied on trafficked labor to fill the chat seats. If the front half of the hustle can be automated, that may change the labor economics of crime. It does not make the harm any less real. It just makes it cheaper to try more often.
Why scale beats spectacle
It is tempting to focus on deepfake voices and synthetic videos. The study suggests the bigger lever is repetition powered by low cost chat. A thousand plausible new friends can now talk for a week without sleep, in idiomatic text, while keeping a consistent persona. That changes the numerator for social engineering attempts. Even modest conversion rates can add up when the contact list is large and the patience is free.
Hook, line, and sinker becomes a workflow. The hook can be scripted outreach. The line is an LLM tuned to be attentive and warm, probing for interests and vulnerabilities. The sinker is a brief human intervention that steers the trusting target to a fake app or site. The researchers describe that handoff as a way to bypass model safety features, since the machine never issues the explicit fraud instruction. The vendor’s guardrails are obeyed, at least formally, while the scam proceeds.
Reading the risk for fintech
For financial platforms, the likely pressure point is volume. The study does not claim universality across user bases or products. It does, however, show that in a controlled setting an AI persona can elicit compliance more effectively than a human, and can do so reliably enough to matter. That is a directional signal. If the cost of patient persuasion drops, the attack surface is whatever depends on a user believing a friendly contact.
That can look mundane. More people nudged to install a fake app. More chats that slowly normalize off platform moves. More prompts to try a new investment site that seems to be endorsed by a helpful acquaintance. None of this requires high production value fakery. It just needs social stamina, which machines have in surplus.
This reframes the defense priority. It is not only about detecting a perfect clone of a public figure. It is also about spotting quiet, long running influence operations that prime a user to comply with a small request. The trigger event, whether an app download or a link click, happens at the end. The groundwork is where automation shines, and where current safeguards often do not look.
What to do before the sinker
Platform level friction can help, but it needs to meet the attack upstream. Safety features that only scan for explicit fraud requests at the point of transaction may miss the machine built relationship that made the request persuasive. Content and behavior signals across days matter. Repeated small talk from a fresh persona, unflagging responsiveness at odd hours, and scripted style consistency are all patterns you can score without reading the private content of messages.
User education also needs a refresh. Traditional advice focuses on the ask. In a world where the ask is delayed, the earlier red flags are about tempo and tone. If a new contact is endlessly available, unusually affirming, and gently steering the conversation toward finance adjacent topics, treat that as a prompt to slow down. The study’s hook, line, and sinker framing is memorable and practical.
Finally, vendors of AI models should expect more indirect use like the handoff described by the researchers. Safety work that only blocks direct scamming prompts is necessary but insufficient. Detection of patterns that indicate persona farming, along with rate limits and provenance in agent to user interactions, will matter as much as content filtering.
The headline finding is simple. In a controlled test, an AI persona built more exploitable trust than humans did, and did it at a scale that humans cannot match. The boring part of fraud looks like a bottleneck that automation can blow wide open. Plan for the line, not just the sinker.