Human Review Only Catches 18.5% of AI-Generated Images
In a study of 500 participants, TruthScan discovered that the average person correctly identifies only 22.9% of AI-generated images.
One image included the label “SIMULATED TRAINING IMAGE” and only 53.6% of participants detected it. Without it, the percentage of AI-generated images that the average participant correctly identifies falls to 18.5%.
Additionally, fewer than 1 in 5 people catch a synthetic image with no visible giveaway.
Never Worry About AI Fraud Again. TruthScan Can Help You:
- Detect AI generated images, text, voice, and video.
- Avoid major AI driven fraud.
- Protect your most sensitive enterprise assets.

In this report, we explore what these findings imply for businesses that rely on human review for user-submitted images, such as receipts, insurance claims, and identity documents.
What The Human Review Gap Costs Businesses
The cost of AI-enabled image fraud varies by industry, business size, and daily processing volume. Certain industries face higher rates of fraud than others, and the risk rises with the size of the workforce. The TruthScan cost calculator can provide a rough estimate of the daily, monthly, and yearly risk.
To calculate fraud risk, estimate the average cost per fraudulent image. Then multiply that value by your daily image processing volume, your industry’s fraud rate, and the risk multiplier associated with the size of your workforce.
Industry fraud rates, per the Merchant Risk Council 2024 Global Fraud Report
Risk Multiplier
Workforce risk multiplier applied to daily fraud risk.
Daily Fraud Risk = Cost Per Fraudulent Image x Daily Processing Volume x Industry Fraud Rate x Risk Multiplier
Illustrating the Cost
Fake Receipts
Today’s fraudsters use AI-generated receipts as a cheap and easy way to trick companies into paying fraudulent expense reimbursements. According to a study by AppZen, AI-generated receipts jumped from 0% fake receipts in 2025 to 70.8% in 2026.
The same study claims that the average cost of an AI-generated receipt is $101. A business with 150 employees that processes 100 reimbursements a day would then have a daily fraud risk worth $560.55, which amounts to a monthly risk of $16816.5.
Doctored Insurance Claims
Generative AI allows fraudsters to create realistic-looking accident scenes, exaggerate damages, and manipulate documents. When these doctored images bypass human review, they trick insurance companies into sending reimbursements.
In the UK, for example, 2024 data from ABI detected 98,400 fraudulent claims worth £1.4 billion, equivalent to approximately £11,600 per fraudulent claim. If we plugged that value into a small agency with 20 employees and a processing volume of 100 images per day, the daily fraud risk would amount to £92,800.
Return Fraud
Similarly, scammers use AI-generated images of damaged products to claim refunds and returns. This is especially prevalent for product categories that sellers don’t ask customers to return, such as groceries, beauty products, and fragile items. According to fraud detection company Forter, these cases rose by 15% from 2024 to 2025 alone.
Failing to detect these fraudulent images can cause significant losses, even if the returned products are relatively cheap. For example, a mid-sized retailer with 200 employees that processes 300 returns per day could face $833 in daily risk if the average return costs $50. This amounts to $24,975 per month, and $303,863 per year.
Is Human Review Reliable? A Study
What We Tested
TruthScan used Pollfish to survey a simple but urgent question: when people are shown AI-generated images and asked whether they are real, edited, fake, or AI-generated, how often do they actually recognize synthetic content?
Methodology
TruthScan’s study asked 500 United States-based participants to label 8 sets of prompts as real, edited, fake, or AI-generated. 6 prompts showed one image, while the remaining prompts showed image pairs. This design produced 4,000 question-level responses and 5,000 inferred image-level judgments across 10 tested visual stimuli.
All images shown to survey respondents were generated using ChatGPT. The images generated consisted of AI-generated scenes and ‘selfie’ photos, testing participants’ ability to identify AI signals. The report treats these images as synthetic survey stimuli and not as depictions of real people or real events.
All 10 images were AI-generated by ChatGPT. The study measures misses on synthetic imagery, but does not estimate false-positive rates on authentic images or full real-versus-AI classification accuracy.
Scoring Method
The survey used different answer choices across prompts. To streamline results, we placed answers under three classifications: strict detection, partial suspicion, and miss.
- Strict Detection: The respondent selected the option that directly identified the image or image pair as AI-generated, fake, or both fake.
- Partial suspicion: The respondent selected the option that identified the image or image pair as edited or manipulated.
- Miss: The respondent identified the image as real or authentic.
Findings
People are more likely to mislabel AI-generated images as real than explicitly identify them as synthetic.
All images in the study were synthetic. However, across all prompts, 56.7% of responses labeled the visual presented as authentic. In contrast, only 30.9% of responses correctly identified the visual as AI-generated. 9.8% of responses believed that the visual was edited but not explicitly AI-generated, while 2.6% of responses were unsure.
Verdict split across 4,000 question-level responses. A wrong answer lands overwhelmingly on “authentic.”
Very few people can detect AI-generated images reliably.
The typical participant struggled to identify AI-generated content consistently. More specifically, the average participant accurately detected only 1.83 (22.9%) of 8 prompts.
One image included the label “SIMULATED TRAINING IMAGE” and pulled 53.6% detection. Without it, the percentage of AI-generated images the average participant correctly identifies falls to 18.5%. Additionally, fewer than 1 in 5 people catch a synthetic image with no visible giveaway.
More than half of respondents, 261 people (52.2%), made zero or one strict detection. Nearly 3 quarters, 366 people (73.2%), made 2 or fewer. Only 6 respondents (1.2%) achieved a strict 8 of 8.
The most convincing image fooled 70.4% of respondents.
Q1 successfully fooled most respondents. 352 judgments (70.4%) labeled it real, 56 (11.2%) labeled it edited, while 92 (18.4%) correctly identified it as AI-generated.
Even clear warnings fail to produce universal recognition.
As mentioned, the only image to achieve a strict detection rate above 50% was Q7, which contained the explicit label “SIMULATED TRAINING IMAGE.” However, it is also noteworthy that even with this label, 13.8% of respondents incorrectly identified the image as real. 12.0% were uncertain, while 20.6% guessed that the image was edited.
All demographics struggle with AI detection
All participants struggled, regardless of age group, gender, educational attainment, or income. While some groups performed slightly better than others, these differences do not matter operationally. The best-performing subgroup, bachelor’s degree holders, still missed roughly 7 out of 10 AI-generated images.
Performance by Age
Performance by Gender
Performance by Educational Attainment
Performance by Income
What This Means For Businesses
Human detection does not work at review-queue speed
Manual verification is too slow and risky for daily processing needs. Most participants struggled to identify AI-generated images, even in the controlled environment of a survey. The task becomes even more difficult in the real world, where employees evaluate large volumes of images alongside captions, social cues, brand identities, and other contextual influences that shape their judgment.
Training and staff expansion will not close the gap
Training employees to spot AI-generated images is unlikely to solve the problem. As generative AI continues to improve, organizations would need to update employees’ skills continuously, making training both expensive and time consuming. At the same time, many businesses process more visual data than human reviewers can accurately evaluate.
While no published studies have quantified the specific cost of training employees to detect AI-generated content, organizations already spend an average of $1,254 per employee annually on workplace learning, according to the Association for Talent Development’s 2025 State of the Industry report. AI-detection training therefore represents an additional investment in employee time, instructor resources, and ongoing retraining as generative AI models evolve.
Additional detection safeguards are necessary
Businesses need more than human judgment to protect themselves from AI-enabled fraud. AI image detection tools can identify technical signals that people cannot see, including metadata, pixel-level artifacts, and frequency domain patterns. They can also analyze large volumes of images in seconds, giving organizations a scalable way to verify visual content before it reaches customers, employees, or decision makers.
What Automated Detection Solves
Automated AI image detection can flag content that human review fails to catch. With TruthScan, you can identify AI-generated, AI-manipulated, and digitally edited images quickly, accurately, and at scale.
What you get
What TruthScan detects
TruthScan’s image detection model scans content for AI signals that are normally invisible to the human eye. It evaluates multiple signals at once for increased credibility.
- Metadata detection: Examines embedded file information, such as the software used to create or edit the image, timestamps, and C2PA content credentials when available.
- Pixel-level analysis: Detects subtle pixel patterns and visual artifacts that AI models often introduce but people cannot reliably see.
- Frequency domain analysis: Analyzes hidden mathematical patterns in the image’s frequency spectrum to identify signatures associated with AI generation or digital editing.
How accurate is TruthScan?
TruthScan’s latest production model, tested on 250,000 pieces of content from open-source libraries, detected AI images with an average accuracy of 99.3% across 92 detectors, including Grok, Midjourney, and GPT-Image 1.5.
Accuracy differed across different content categories, with performance being strongest on chats, human images, and product photos.
Who needs AI image detection?
Get started with TruthScan
TruthScan’s AI image detector can reliably strengthen your image review process. With the free plan, you get enough credits for 25 image scans per month, enabling easy and commitment-free testing. To get started, you can create an account on our website, or contact X.
Appendix
Image-by-Image Findings
Every stimulus was AI-generated; the shading shows how many respondents accepted each as real.
The tendency to misclassify images reflected a consistent pattern across the survey rather than emerging from one weak prompt. 9 of the 10 inferred stimuli earned real or authentic labels from at least half of respondents.
Only Question 7 was classified as AI-generated by fewer than half of participants, but this image included text reading “SIMULATED TRAINING IMAGE.”
Limitations and Research Notes
This was a 500-person U.S. survey, which is large enough to reveal a clear pattern in the tested prompts but still represents a single sample and a single image set. These research notes should guide follow-up studies; they do not erase the central result that respondents repeatedly accepted AI-generated visuals as real.
- Sample scope: The survey included 500 U.S. respondents. Results should be read as a strong indicator of the tested population and prompts, not a universal measurement of every audience or country.
- Image set: The findings reflect the 10 visuals used in this study. Other image categories, quality levels, or subject matter could produce different rates.
- Prompt design: Answer choices varied by question, so this report uses a transparent scoring map and separates strict detections from partial suspicion.
- Pair prompts: Q3 and Q5 asked about two images at once. Prompt-level findings are direct; per-image rankings from those questions are inferred from the answer selected.
- Q7 cue: the Q7 image included a visible ‘SIMULATED TRAINING IMAGE’ label. Its 53.6% AI-generated response should be interpreted with that embedded cue in mind.
- Additional measures: The CSV does not include confidence scores, AI familiarity, media-literacy measures, or reasons for each answer. Future studies could add those fields.
- Weights: No survey weights were provided in the export, so calculations are unweighted.
- Ages: The oldest group is labeled 60-64 because the CSV age range ends at 64.
Exact Response Distributions
Q1: Is this image real, edited or AI-generated?
Synthetic survey stimulus. It does not depict a real person.
Q2: Is this image real, edited or AI-generated?
Synthetic survey stimulus. It does not depict a real person.
Q3: Which of these two images is real (if any)?
Synthetic pair shown to respondents. Both images are treated as AI-generated in this analysis.
Q4: How would you judge this image?
Synthetic survey stimulus. It does not depict a real person.
Q5: Which of these photos is real (if any)?
Synthetic pair shown to respondents. Both images are treated as AI-generated in this analysis.
Q6: Is this image clearly real, edited, or AI-generated?
Synthetic survey stimulus. It does not depict a real person.
Q7: Is this photo edited, real (unedited), or AI-generated?
Synthetic survey stimulus. It includes a visible simulated-training cue in the image.
Q8: Is this image real, edited, or fake?
Synthetic survey stimulus. It does not depict a real person.
Data Use and Media Disclosure
This data may be used in press articles, further reports, and for other informational pieces, provided this original report and its author (TruthScan) are credited and cited. For access to the full raw survey data and responses, email devan@truthscan.com.