Contents
Yes, and at rates meaningfully higher than the marketing suggests. Where an independent benchmark has checked a commercial AI-music detector’s false-positive claim, the measured rate on human-made music came out several times the advertised one. For a working musician, that is not an abstract statistic: a false positive is your own recording flagged as machine-made, with the takedown, demonetization, or reputational dispute that follows. A missed AI song is a detection failure; a false positive is a wrong accusation, and that is the error that lands on a real person. This article is about how often it happens and why the advertised rate is not the rate in the field.
What is a false positive?
A false positive is a real human recording flagged as AI-generated. In music that can happen for several reasons. The detector may read a production artifact that also appears in human mastering. It may be calibrated on a narrow benchmark. It may treat codec, resampling or platform processing as suspicious. Or it may be forced to answer AI or human when the honest answer is uncertain. That is different from the question in why did the music detector flag this?, which explains the possible causes behind one specific score. This article is about the rate: how often a detector wrongly flags real music at all.
Advertised versus measured
The clearest case on record involves IRCAM Amplify, whose AI-music detector is marketed with “99% accuracy” and “Less than 1% false positives.” When Cros Vila, Sturm, Casini and Dalmazzo benchmarked commercial detection for TISMIR in 2025, the measured false-positive rate on human music was 4.7 percent, several times the advertised bound. The same study found the field’s definition of “AI music” unsettled across tools, one detector missing 75.3 percent of Udio tracks in the other direction, and a commercial system whose verdict changed after a plain resample to 22.05 kHz. A tool that flips on a resample is not making a stable measurement of anything, and instability produces errors in both directions, including against real music.
What 4.7 percent means at platform scale
Rates that sound small stop being small at upload volume. Deezer’s own published figures describe over 13.4 million AI tracks detected and AI arriving at 44 percent of new uploads, which is a vendor’s operational claim rather than an audited number, but it fixes the scale: streaming platforms now screen enormous daily volumes. The arithmetic is unforgiving. Scan one million human uploads at a 1 percent false-positive rate and ten thousand real tracks are flagged; at 4.7 percent, that is forty-seven thousand. Those figures are a hypothetical illustration, not a measured platform total, but they show why a “small” rate is a queue of wrongly accused songs, concentrated on whoever’s production style happens to resemble what the detector learned to distrust. The musician on the receiving end does not experience a percentage; they experience a takedown notice with a confident score attached.
Why real music gets flagged
False positives are not random noise; they have causes. The largest is distribution shift: detectors calibrated on one benchmark misfire on audio from outside it. ArtifactNet, the strongest forensic detector published, holds a 0.09 percent false-positive rate on its home benchmark and 1.49 percent on an unseen one, a more than fifteenfold increase from changing nothing but the test set (Oh, 2026, arXiv:2604.16254). The same evaluation caught a rival system, CLAM, at a 69.26 percent false-positive rate cross-benchmark, flagging most real music it saw. Processing does the rest: platform codecs, resampling, and mastering move a real track’s signal toward what a fragile detector reads as suspicious, and broadcast conditions, where the best measured F1 falls to 61.1 percent (López-Ayala et al., 2026), degrade verdicts in both directions. Definition compounds it. A human performance with AI mastering, a real vocal over a generated backing, and a fully synthetic song are three different claims, and a tool that collapses them into one label can accuse a real musician for the wrong reason. A real song that has been compressed, remastered, or clipped short is a real song wearing exactly the distortions detectors mistake for evidence.
Why the advertised rate is not the field rate
A vendor’s false-positive figure is produced under conditions the vendor chose: their test set, their generator mix, their clean audio. Nothing in the commercial market currently forces those conditions to resemble deployment, and no major vendor publishes an independently audited error rate. The research community’s assessment of the commercial field is blunt: it is “closed-source and without any associated research publication” (Afchar, Meseguer-Brocal, Hennequin, 2025). Until a vendor’s number has been reproduced by someone who does not profit from it, the honest reading is that the advertised rate is a floor measured in favorable conditions, and the 4.7 percent measured by Cros Vila’s team is the only independent reference point the field has. A tool-by-tool audit of the named commercial detectors and what stands behind each claim is a separate article in preparation; this one owns the general finding and the measured rate.
If your own music was flagged
Treat the score as an accusation to be tested, not a verdict. First, check what the detector was actually reading: the why did the music detector flag this? walkthrough covers the common non-AI causes, and what does an AI music detector score mean? covers why a confidence number is not a probability that your song is fake. Second, re-run on your original export rather than the platform’s re-encoded copy, since compression itself shifts scores. Third, get a second opinion from a detector in a different family, because detectors disagree precisely on the marginal cases. Keep your project files, stems, and session history; they are provenance no detector can override. Separate “this file carries a trace this detector associates with AI” from “this artist faked the song”: the first may be a defensible technical statement, the second needs far more evidence. A single score from a tool whose error rate its own vendor will not publish independently is weak evidence against a documented human recording, and the broader case for skepticism across every medium is set out in are AI detectors accurate?.
Sources
- Cros Vila, Sturm, Casini, Dalmazzo (2025). The AI Music Arms Race: On the Detection of AI-Generated Music. TISMIR 8(1). DOI:10.5334/tismir.254.
- Oh (2026). ArtifactNet: Forensic Codec-Residual Detection of AI-Generated Music. arXiv:2604.16254.
- Afchar, Meseguer-Brocal, Hennequin (2025). AI-Generated Music Detection and its Challenges. arXiv:2501.10111.
- López-Ayala, Cabello, Zinemanas, Molina, Rocamora (2026). AI-Generated Music Detection for Broadcast Monitoring. ICASSP 2026. arXiv:2602.06823.