detectai.media

Can voice liveness detection be bypassed?

Research measures high bypass rates against verification and detection systems, but every measured bypass leaves its own physical evidence in the file.

By The detectai.media team
4 min read
Contents

Yes, at rates the research community has measured rather than guessed. A verification system with no liveness layer accepts most replays outright, replay laundering measurably degrades modern deepfake detectors, and simulated playback channels raise over-the-air attack success further still. The counterweight is that every measured bypass leaves its own physical evidence behind. This article reviews what the attack literature reports; it is a reliability review, not a procedure.

The baseline hole liveness was built to close

The oldest number is still the starkest. Testing far-field re-recordings against a speaker-verification system, Villalba and Lleida (BioID 2011) reported: “if we would choose the EER operating point as the decision threshold, we would accept 68% of the spoofing trials.” Without a liveness check, replay is not so much a bypass as a walk through an open door. The same experiments showed the channel cuts both ways: their GSM-channel replays were all rejected, because the mobile codec degraded the signal past the verifier’s own tolerance.

Replay also launders synthetic speech past AI detectors

The modern twist is that playback is itself an attack on deepfake detectors. Müller and colleagues at Fraunhofer AISEC (arXiv 2025) replayed synthetic speech through real loudspeakers and microphones and re-presented it to a state-of-the-art detector: the W2V2-AASIST system’s equal-error rate rose from 4.7% to 18.2%. The damage was asymmetric, which is the interesting part. Genuine speech barely moved, while per-generator accuracy collapsed:

GeneratorAccuracy before replayAfter replay
XTTS v2100%59.4%
Bark82.6%40.7%
VITS82.2%53.4%
Genuine speech98.3%97.7%

The same study found a correlation of 0.509 between detection accuracy and perceptual quality of the replay: the more aggressive the replay degradation, the worse the detector performed. The replay noise buries the synthesis artifacts the detector was reading.

Simulated channels make over-the-air attacks stronger

Attackers do not need real hardware chains to model one. Li, Wang, Xue, Wang and Wu (arXiv 2024) trained a neural replay simulator to make adversarial perturbations survive a real loudspeaker-to-microphone hop: over-the-air attack success against speaker-verification systems rose from a 53.3% baseline to 73.6% with the simulator in the loop. One detail points the other way: their raw-waveform target, RawNet, held attack success to 4.9%, suggesting feature choice changes how much a simulated channel can smuggle through.

Why a bypass is never free

The recurring finding across this literature is that laundering replaces evidence rather than erasing it. The replay that hides synthesis artifacts stamps loudspeaker distortion, low-frequency roll-off, and a second room onto the file, which is exactly the surface the four liveness method families read. Witkowski and colleagues (INTERSPEECH 2017) showed the extra digital-to-analog conversions concentrate a detectable anti-aliasing signature in the 6 to 8 kHz band. An attack that defeats the synthesis check hands evidence to the replay check, and vice versa; defeating both at once is the hard version of the problem.

Where the defense genuinely struggles

Two measured weaknesses temper that reassurance. Detectors generalize poorly across playback hardware: Ren, Fang and Liu (Multimedia Tools and Applications 2019) saw their true-positive rate fall from 97.78% to a floor near 89.7% when the loudspeaker changed, with the baseline method falling to 79.3%. And training on simulated replays transfers badly to real ones: Neri and Virtanen (arXiv 2025) report that replay detectors trained on synthetic room simulations reached equal-error rates above 50% in most environments when tested on real re-recorded audio, which is coin-flip territory. The bypass question therefore does not have a static answer. It is an arms race in which each side’s countermeasure feeds the other’s next move, and the version of the question that matters most in practice, whether a deepfake can ride a replay past both layers, is taken up in does voice liveness detection work against deepfake voice?.

Sources

  • Villalba, J. & Lleida, E. (2011). Detecting Replay Attacks From Far-Field Recordings on Speaker Verification Systems. BioID 2011.
  • Müller, N., Kawa, P., Choong, W.-H., Stan, A., Tirumala Bukkapatnam, A., Pizzi, K., Wagner, A. & Sperl, P. (2025). Replay Attacks Against Audio Deepfake Detection. arXiv:2505.14862.
  • Li, J., Wang, L., Xue, L., Wang, L. & Wu, Z. (2024). An Initial Investigation of Neural Replay Simulator for Over-the-Air Adversarial Perturbations to ASV. arXiv:2310.05354.
  • Witkowski, M., Kacprzak, S., Żelasko, P., Kowalczyk, K. & Gałka, J. (2017). Audio Replay Attack Detection Using High-Frequency Features. INTERSPEECH 2017. DOI: 10.21437/Interspeech.2017-776.
  • Ren, Fang & Liu (2019). Replay Attack Detection Based on Distortion by Loudspeaker for Voice Authentication. Multimedia Tools and Applications. DOI: 10.1007/s11042-018-6834-3.
  • Neri, M. & Virtanen, T. (2025). Acoustic Simulation Framework for Multi-Channel Replay Speech Detection. arXiv:2509.14789.
#audio#voice#liveness#replay#reliability
Last updated
22 July 2026
Category
Reliability