BIP America News & Media Platform

collapse
Home / Daily News Analysis / Gut feeling does nothing against AI spear phishing texts

Gut feeling does nothing against AI spear phishing texts

Aug 08, 2026  Twila Rosenbaum  9 views
Gut feeling does nothing against AI spear phishing texts

Gut feeling is not a reliable defense against AI-powered spear phishing. In a pilot study at Brigham Young University, volunteers were shown a dozen personalized text messages and asked to sort them by how likely they were to click through. Half of the messages came from GPT-4; half were written by human students. The volunteers could not identify the AI-generated texts any better than chance, and their confidence in their own detection strategies did not help.

The study does not claim that AI has already surpassed expert social engineers. It pitted a single short prompt against students who had just completed coursework on phishing. The findings nevertheless offer an uncomfortable glimpse at what modern language models can do with even a small amount of personal information.

The setup: a pile of messages written from personal details

Twenty-five people volunteered for the experiment. Before the message-writing phase, each one filled out a survey that captured basic personal information: job title, employer, leisure activities, city, and something recently posted on social media. The researchers then turned those details into a prompt template and ran it once through GPT-4. The model generated six personalized messages for each participant. A group of undergraduates enrolled in a deception course received the same template and wrote the remaining messages by hand, working under a fifteen-minute time limit and producing up to four messages each. A review team that included two cybersecurity professors screened the students' output and discarded about a third of it because the messages were incomplete or unusable.

Once all the messages were ready, each participant sat down with twelve printed text messages, all written specifically for that person. They were asked to arrange them from most likely to elicit a click to least likely. They were also asked to draw a line in the pile: above the line, they would have clicked; below, they would not. The messages were printed on cards, so there were no sender numbers, no live links, and no other environmental clues. The measurement was self-reported intention, a standard proxy in phishing research but still a proxy.

AI messages did slightly better, but the margin was not conclusive

When the sorting results were tallied, GPT-4's messages landed above the click line 28% of the time. The student-written messages cleared the line 21.3% of the time. That looks like a 6.7 percentage point win for the AI, and it would be a tempting headline. The confidence interval tells a different story: it stretched from 2.9 points in favor of the human writers to 16.3 points in favor of GPT-4. In plain terms, the study could not determine which side had the edge.

The real concern is the asymmetry in effort. The AI side consisted of one short prompt filled in from a survey and run once per person. The human side involved students who had just taken phishing instruction, a review process, and the rejection of a third of their work. The AI matched that performance, and it can be run again and again for the price of an API call. At scale, that cost advantage changes the threat landscape even if the per-message click rate is only comparable.

Work-related messages were the most dangerous

The study gave every participant three flavors of message: one built around their job, one around a hobby, and one around a social media post. The differences were clear. Job-themed messages cleared the click line 38% of the time. Hobby-themed messages did so 19% of the time. Social-media-themed messages cleared it 17% of the time. This comparison survived statistical adjustment, making it the most reliable finding in the experiment.

Work messages worked because they hit a context where people expect urgent, impersonally personal communication. A banker paused on one message that resembled an internal fraud alert from their credit union, saying it 'literally looks like the alert we get [at work] when there's a fraud.' That kind of contextual resonance is difficult for a generic phishing message to achieve, but a model that has been given someone's job title and employer can produce it instantly.

Personalization also failed in instructive ways. One participant spotted a reference to a coworker who did not exist: 'there's no Mike at work.' Another knew the instructional design team at their university well enough to know there was no Sarah on it. A dance-related message went to the bottom of the pile for someone who had quit dancing. Wrong details are worse than no details because they give the target a reason to be suspicious.

People had theories about AI style, and all of them failed

After the sorting exercise, participants were told that at least one of the messages had been written by an AI and were asked to mark which ones they believed were artificial. Across 300 judgments, they were correct 52% of the time. Random guessing would produce 50%. The participants did not lack theories: some said AI writes too formally; others said it writes too generically. Perfect grammar was treated as a machine tell by some, while typos were read as proof of a human by others. One participant focused on punctuation, wondering why there were so many exclamation points. Twelve of the 25 admitted that they were mostly relying on instinct.

Emojis were a particularly revealing example. They were a real tell: 66% of GPT-4 messages contained at least one emoji, while only 2% of human-written messages did. Yet only five participants mentioned emojis at all, and they disagreed over what an emoji meant. Two saw emojis as evidence of AI; three saw them as evidence of a person, reasoning that a chatbot probably could not handle emojis. The human intuition was not merely weak; it was divided in the wrong direction.

Software found the pattern people missed

While human judgment hovered around chance, the researchers also tested a machine-learning approach. They converted all 300 messages into embeddings and trained a classifier to distinguish AI-written from human-written text. Under the strictest evaluation conditions, the classifier achieved 88.7% balanced accuracy. The strict conditions involved standardizing URLs, stripping emojis, flattening case, digits, and punctuation, and trimming each matched pair of messages to the length of the shorter one. The model was also tested only on people whose messages it had never encountered during training, so it was not memorizing individual targets.

The classifier's success is not the same as a ready-made detector. It was trained on one message set, from one model version, under one prompt design, and against one pool of student writers. There is no evidence yet that it would generalize to other phishing campaigns or other language models. Moreover, research cited in the paper shows that paraphrasing AI text while running a detector in the loop can defeat several of these classifiers. A determined attacker can adapt quickly.

What a small pilot can and cannot prove

This was a pilot study with 25 completed volunteers, and the authors are transparent about its limits. The messages were printed on cards, so the test could not account for sender IDs, phone notifications, link previews, or any of the other contextual signals that affect real-world phishing decisions. The comparison group consisted of novice students, not professional social engineers. To detect a 6.7 percentage point difference with confidence, the researchers calculate that roughly 100 completed targets would be needed rather than 25.

There is also a transparency gap in the paper. The exact GPT-4 snapshot and the API logs were never recorded, so while the messages themselves survive and the analysis can be reproduced, the exact generation run cannot be repeated. This is a common challenge in AI research, but it matters when the goal is to assess a moving target like a language model.

Practical defenses against AI-generated spear phishing

The study's practical advice does not depend on whether GPT-4 beat the human writers. The first line of defense is to stop trying to decide whether a message sounds like a robot. The study showed that people cannot do that reliably, and it may even create false confidence. Instead, check the sender, the channel, the link, and the request. Does the sender address match the person or institution it claims to be from? Did the message arrive through an expected channel? Does the link point to a familiar domain? Is the request something that person or organization would actually ask for?

Spear phishing succeeds when a message feels contextually plausible. AI makes it easier to generate that plausibility from public information and leaked personal data. The results here are a reminder that the weakness is not in the grammar or punctuation of the message; it is in the rush to respond. A deliberate pause, a separate verification step, and a healthy skepticism of any message that asks for credentials, payment, or urgent action are still the most practical tools available.


Source: Help Net Security News


Share:

Your experience on this site will be improved by allowing cookies Cookie Policy