Voice & music

AI audio detector

Drop an audio file to check for synthetic speech or AI-generated music. GPTTrace reads its tags and Content Credentials, then measures the signal for traces that text-to-speech and music models leave behind.

Free · no sign-up · your file never leaves your device

Drop an audio file or click to choose

MP3 · WAV · M4A · FLAC · OGG · Opus — the first 60 seconds are analysed on your device

Why audio needed its own detector

Voice cloning went from a research demo to a phone scam in a few years. A few seconds of someone’s voice from a video or voicemail is now enough for consumer tools to produce convincing speech, and criminals use it for “family emergency” calls, fake executive instructions to finance teams, and bypassing voice authentication. Meanwhile music generators produce full songs with vocals, and streaming services report a rising share of uploads that are entirely AI-generated.

Most AI detectors ignore audio. GPTTrace analyses it with the same philosophy as its image checks: read what the file says about itself first, then measure the content, and show every reason.

What GPTTrace measures

Tags, headers and credentials

ID3 tags in MP3s, RIFF chunks in WAVs, MP4 atoms in M4As and Vorbis comments in FLAC and OGG files can name the software that produced them. GPTTrace reads them all, separates encoder fields from free-text titles, and checks C2PA manifests, which can be embedded in WAV, MP3 and MP4 audio.

Native sample rate and bandwidth

Many speech models generate audio at 16, 22.05 or 24 kHz. Delivered in a 44.1 or 48 kHz file, that leaves a hard ceiling in the spectrum at 8, 11 or 12 kHz, far below what the file’s encoder would keep. GPTTrace reads the native rate from the file header (the browser would otherwise hide it), compares the spectrum’s cut-off with what the codec and bitrate should allow, and names the model rate it matches.

Silence and noise floor

A microphone always records a room: hum, hiss, breath. Many TTS engines produce pauses of perfect digital zero. GPTTrace measures exact-zero gaps and the level of the quietest moments.

Pitch and loudness

Human voices waver in pitch from one moment to the next and vary in loudness. GPTTrace tracks the fundamental frequency through voiced speech and reports how smooth it is, along with how evenly loud the voice stays.

Spectral lines

Neural vocoders and upsamplers can leave narrow, constant tones in the upper band. The spectrogram view lets you see them alongside the rest of the signal.

Each measurement is weak alone; combined, they separate typical TTS output from typical recordings. The speech-specific checks are skipped for music.

Frequently asked questions

Can AI voices really be detected?
Sometimes with certainty, often only as a probability. Files from tools that label their output — in tags or Content Credentials — can be identified outright. Otherwise detection relies on signal traces such as band limits, unnaturally clean silence and smooth pitch, which the best voice clones increasingly avoid. Treat a “no AI signs” result with caution.
What audio formats are supported?
Anything your browser can decode: MP3, WAV, M4A/AAC, FLAC, OGG and Opus in most browsers. The first 60 seconds are analysed; tags and headers are read from the whole file.
Why does a phone call recording look suspicious?
Phone networks limit audio to a narrow band (often below 4 or 8 kHz), and noise suppression can create near-silent gaps. Both resemble traces of synthetic speech. GPTTrace recognises telephone bandwidth and weighs it lightly, but call recordings are inherently harder to judge.
Is my audio uploaded?
No. Decoding and analysis happen in your browser using the Web Audio API. Nothing is sent anywhere.
Does it detect AI music from Suno and Udio?
It reads any tags that name those services and measures the signal, including the band limits and upper-band artifacts common in generated tracks. Music is harder than speech for signal checks, so expect more inconclusive results; see our AI music detector page for details.