How it works
Methodology & measured accuracy
What GPTTrace checks, how it turns evidence into a probability, which models it runs on your device, and how well it actually performs on labelled data — including where it fails.
One model for evidence
Every GPTTrace detector produces a list of evidence items. Each item has a signed strength — positive leans toward AI, negative toward human or authentic — and a weight. The probability shown is a logistic combination: the weighted strengths are added to a baseline and passed through a sigmoid. The weights are fitted by L2-regularised logistic regression on labelled data, regularised toward hand-set starting values so that rare but meaningful signs (an AI disclaimer, a chat-interface citation token) are not erased just because the test set contains few of them. Each data source is weighted equally, so one large dataset can’t dominate.
After fitting, the baseline is shifted so that no more than 5% of human samples reach the “Likely AI” line at 70%. Provenance — a signed C2PA manifest recording AI generation, the IPTC trainedAlgorithmicMedia label, embedded diffusion settings or NovelAI’s alpha-channel record — bypasses the model and settles the verdict, because the file itself declares its origin.
Images
Provenance and metadata are read by GPTTrace’s own parsers: a C2PA reader that walks the JUMBF boxes in JPEG, PNG, WebP, HEIC, AVIF and MP4 files, decodes the CBOR claim and assertions and reads the signer’s X.509 certificate; EXIF, XMP and IPTC via the open-source exifr library; and PNG text chunks including compressed ones. The neural classifier is the Community Forensics ViT-Small model (Park & Owens, CVPR 2025), trained on 2.7 million images from 4,803 generators, run with ONNX Runtime WebAssembly in an int8 version of 22 MB. Images are preprocessed as the model specifies (shortest edge 440, centre crop 384); large images are also scored on four native-resolution crops. Forensic measurements — error-level analysis, the radial spectrum and a lattice-peak detector on a native-resolution crop — run in a Web Worker.
Models we tested and did not ship: a SigLIP-based “AI vs human” image classifier (Apache-2.0, 88 MB) scored an AUC of 0.64 on the same images and lowered the ensemble’s accuracy, so GPTTrace uses the Community Forensics model alone. Two other popular detectors were excluded because their licences forbid commercial use.
image check: AUC 0.938 (cross-validated)
679 labelled samples (399 AI, 280 human), run 2026-10-08. At the “Likely AI” line it caught 64% of AI samples and wrongly flagged 5% of human ones.
| Source | Truth | Samples | Result at “Likely AI” |
|---|---|---|---|
| gemini-nano-banana | AI | 40 | 40% caught |
| midjourney-v6 | AI | 40 | 68% caught |
| midjourney-v5 | AI | 40 | 73% caught |
| flux-dev | AI | 40 | 13% caught |
| flux-schnell | AI | 40 | 48% caught |
| sdxl | AI | 40 | 100% caught |
| gpt-image | AI | 40 | 30% caught |
| kling | AI | 39 | 97% caught |
| leonardo-stablecog | AI | 40 | 98% caught |
| bitmind-imagine-mix | AI | 40 | 80% caught |
| fullsize-photos | Human | 40 | 10% wrongly flagged |
| open-images-photos | Human | 40 | 5% wrongly flagged |
| lfw-faces | Human | 40 | 0% wrongly flagged |
| caltech-objects | Human | 40 | 3% wrongly flagged |
| coco-photos | Human | 40 | 0% wrongly flagged |
| ffhq-faces | Human | 40 | 0% wrongly flagged |
| celeba-faces | Human | 40 | 18% wrongly flagged |
Text
The standard check combines 22 signs derived from Wikipedia’s Signs of AI writing and an era-weighted vocabulary, sentence and paragraph statistics, and a compression-similarity measure using the browser’s deflate implementation. The optional deep scan adds the TMR RoBERTa detector, fine-tuned on the RAID benchmark, run locally after a 120 MB download. Evaluation uses the TextSight 2026 control set (pre-2022 Wikipedia, non-native-English academic writing, public-domain literature and genre-matched output from five 2026 models) and the HC3 corpus. Text numbers are five-fold cross-validated.
standard text check: AUC 0.795 (cross-validated)
2,467 labelled samples (1,258 AI, 1,209 human), run 2026-10-08. At the “Likely AI” line it caught 34% of AI samples and wrongly flagged 5% of human ones.
| Source | Truth | Samples | Result at “Likely AI” |
|---|---|---|---|
| claude-opus-5 | AI | 12 | 50% caught |
| gpt-4.1 | AI | 12 | 100% caught |
| gpt-oss-120b | AI | 12 | 100% caught |
| claude-haiku-4-5 | AI | 12 | 83% caught |
| qwen-3.8-27b | AI | 8 | 63% caught |
| pd-literature | Human | 306 | 1% wrongly flagged |
| human-pmc-esl | Human | 183 | 5% wrongly flagged |
| human-wikipedia-pre2022 | Human | 155 | 8% wrongly flagged |
| hc3-open_qa | Human | 7 | 14% wrongly flagged |
| chatgpt-3.5 | AI | 1202 | 32% caught |
| hc3-wiki_csai | Human | 558 | 6% wrongly flagged |
text check with deep scan: AUC 0.91 (cross-validated)
2,467 labelled samples (1,258 AI, 1,209 human), run 2026-10-08. At the “Likely AI” line it caught 56% of AI samples and wrongly flagged 5% of human ones.
| Source | Truth | Samples | Result at “Likely AI” |
|---|---|---|---|
| claude-opus-5 | AI | 12 | 42% caught |
| gpt-4.1 | AI | 12 | 100% caught |
| gpt-oss-120b | AI | 12 | 100% caught |
| claude-haiku-4-5 | AI | 12 | 100% caught |
| qwen-3.8-27b | AI | 8 | 75% caught |
| pd-literature | Human | 306 | 0% wrongly flagged |
| human-pmc-esl | Human | 183 | 1% wrongly flagged |
| human-wikipedia-pre2022 | Human | 155 | 1% wrongly flagged |
| hc3-open_qa | Human | 7 | 0% wrongly flagged |
| chatgpt-3.5 | AI | 1202 | 54% caught |
| hc3-wiki_csai | Human | 558 | 9% wrongly flagged |
Audio and video
Audio is decoded at its native sample rate after reading the container header, and measured over the first 60 seconds: spectral band limit relative to the codec’s expected low-pass, exact digital silence, noise floor, pitch micro-variation, loudness spread and steady upper-band tones. Speech-only measurements are skipped for music. Video is checked through its container metadata and Content Credentials, then eight frames are sampled and passed through the image classifier and spectral checks. Audio and video weights are currently hand-set; labelled evaluations for both are being built and their numbers will appear here when they are reliable enough to publish.