Voice & music
AI audio detector
Drop an audio file to check for synthetic speech or AI-generated music. GPTTrace reads its tags and Content Credentials, then measures the signal for traces that text-to-speech and music models leave behind.
Free · no sign-up · your file never leaves your device
Drop an audio file or click to choose
MP3 · WAV · M4A · FLAC · OGG · Opus — the first 60 seconds are analysed on your device
Why audio needed its own detector
Voice cloning went from a research demo to a phone scam in a few years. A few seconds of someone’s voice from a video or voicemail is now enough for consumer tools to produce convincing speech, and criminals use it for “family emergency” calls, fake executive instructions to finance teams, and bypassing voice authentication. Meanwhile music generators produce full songs with vocals, and streaming services report a rising share of uploads that are entirely AI-generated.
Most AI detectors ignore audio. GPTTrace analyses it with the same philosophy as its image checks: read what the file says about itself first, then measure the content, and show every reason.
What GPTTrace measures
Tags, headers and credentials
ID3 tags in MP3s, RIFF chunks in WAVs, MP4 atoms in M4As and Vorbis comments in FLAC and OGG files can name the software that produced them. GPTTrace reads them all, separates encoder fields from free-text titles, and checks C2PA manifests, which can be embedded in WAV, MP3 and MP4 audio.
Native sample rate and bandwidth
Many speech models generate audio at 16, 22.05 or 24 kHz. Delivered in a 44.1 or 48 kHz file, that leaves a hard ceiling in the spectrum at 8, 11 or 12 kHz, far below what the file’s encoder would keep. GPTTrace reads the native rate from the file header (the browser would otherwise hide it), compares the spectrum’s cut-off with what the codec and bitrate should allow, and names the model rate it matches.
Silence and noise floor
A microphone always records a room: hum, hiss, breath. Many TTS engines produce pauses of perfect digital zero. GPTTrace measures exact-zero gaps and the level of the quietest moments.
Pitch and loudness
Human voices waver in pitch from one moment to the next and vary in loudness. GPTTrace tracks the fundamental frequency through voiced speech and reports how smooth it is, along with how evenly loud the voice stays.
Spectral lines
Neural vocoders and upsamplers can leave narrow, constant tones in the upper band. The spectrogram view lets you see them alongside the rest of the signal.
Each measurement is weak alone; combined, they separate typical TTS output from typical recordings. The speech-specific checks are skipped for music.