User PDF

AI image detector: results, errors and limitations

The detectors flagged AI indicators in all 12 generated originals in this test. After JPEG compression, one of the 12 photos was incorrectly flagged. These results come from a small sample: they do not mean 100% accuracy or prove the origin of another image.

An analysis that also shows when it gets things wrong

  1. Review the counts by condition and the individual results to understand what was measured.
  2. When using the tool, prefer the original file before screenshots, cropping or compression.
  3. Read pixel signals, provenance credentials and limitations separately. An inconclusive result is not proof of a real photograph.

Sample and procedure

On September 17, 2026 we ran both local classifiers from engine 1.12 on 12 photos from scikit-image and the first 12 images in DiffusionDB part 000001, sorted by filename. Each original was evaluated as received, recompressed to JPEG quality 65 and reduced to at most 512 pixels per side. These are 24 originals and 72 correlated runs, not 72 independent images.

Models and input limits

We used AI-vs-Deepfake-vs-Real-ONNX and Community Forensics (commfor-model-384), with pinned engine weights. Counts reflect the conservative system verdict, not a raw model class: inputs with a shorter side below 384 pixels or an aspect ratio above 3:1 are restricted. This contributes to inconclusive results after resizing. Individual scores are not calibrated probabilities.

The observed error

The Hubble Deep Field photo recompressed as JPEG was flagged as possibly AI-generated. Among the resized generated images, only 7 of 12 received that flag. The other outputs, including inconclusive results, are in the spreadsheet. This error illustrates why a single score should not support accusations of fraud.

What this test does not validate

This convenience sample was not randomly selected to represent the internet. The generated images come from Stable Diffusion; we did not test ChatGPT generations or all recent models here. Training-data overlap was not checked. We did not measure probability calibration, AI editing or universal accuracy. Compression and resolution affect the result.

Pixels, C2PA and invisible watermarks are different evidence

This measurement isolates the pixel classifiers: the decision receives absent credentials and no watermark found. It does not validate C2PA verification or invisible watermark detection. In the full tool, these signals are presented separately. Missing credentials or watermarks do not prove human origin.

Results by condition

Each row uses the same 12 generated images and 12 photos. The inconclusive column combines both groups.
ConditionGenerated images flagged / 12Photos falsely flagged / 12Inconclusive / 24
Original 12 0 12
JPEG quality 65 11 1 12
Resized to at most 512 px 7 0 17

Files and measurements for this example

Sources and method