Preprint

Preprint Finds Single-Sentence Voice Fakes Hard to Spot

In a test of 82 IT professionals, selected ElevenLabs recordings were much harder to identify than recordings from two older speech-synthesis tools.

Listeners missed most one-sentence voice fakes in the study’s partial-spoof condition: they detected the manipulated sentence in only 22.65% of cases when it was embedded in otherwise bona fide, or authentic, speech. The synthetic sentence was classified as bona fide 77% of the time.

The finding comes from a questionnaire-based study of 82 IT professionals comparing selected speech-synthesis tools and automated detectors on the same recordings.

The gap was largest for the newer system

The researchers generated recordings with RTVC (2019), YourTTS (2022) and ElevenLabs (2024), using pretrained zero-shot models without fine-tuning them to the target speaker. Six pretrained detector systems were also tested on the same material.

Each speaker test set contained one bona fide recording, three full spoofs and three partial spoofs. Test-set and recording orders were randomized for each participant, and uncertain answers were excluded before scoring.

For full spoofs, the study’s F1 score—a combined measure of detection performance—was 48% for selected ElevenLabs recordings, compared with about 90% for older tools. At the recording level, 28% met the strict “All OK” measure and 33% met the “Over 50 Percent” measure. Mean per-sentence accuracy for ElevenLabs was 43%, with a reported 95% range of 37% to 50%.

Listeners also made false alarms: the false-alarm rate on bona fide recordings was 17.5%. The study’s signal-detection sensitivity score, d′, was 2.75 for RTVC and 2.88 for YourTTS, compared with 0.73 for ElevenLabs.

A single altered sentence was easier to miss

For the mixed ElevenLabs condition, in which a synthetic sentence was embedded in bona fide speech, mean F1 was 29%, versus almost 90% for older tools. Fewer than 10% of the mixed ElevenLabs recordings met the strict All OK measure.

The error was concentrated in the altered sentence. Participants detected manipulated sentences in 22.65% of cases while correctly classifying bona fide segments in 88.59%.

In a paired analysis of 73 participants, selected ElevenLabs detection was 43% for full spoofs versus 23% for partial spoofs. The full-to-partial decrease was also statistically significant for RTVC and YourTTS.

Automation did not settle the question

The automated comparison produced a mixed picture. On full spoofs, mean detector F1 was 96% for RTVC and 100% for YourTTS, but just 9% for ElevenLabs, compared with 48% for human listeners. On mixed ElevenLabs recordings, detectors scored 40% against human F1 of 29%, while detector All OK was 0%.

At the segment level, detectors identified manipulated speech in 51.9% to 77.8% of cases, while bona fide acceptance ranged from 63.0% to 100%. The benchmark used six pretrained systems and a fixed threshold of 0.5, so the paper presents these figures as indicative rather than calibrated performance.

What the preprint can—and cannot—show

The results describe this sample and stimulus set, not a general temporal trend in speech-synthesis deception. The comparison used three selected systems—RTVC, YourTTS and ElevenLabs—and the human sample consisted of IT professionals, so the study does not estimate how the general public would perform.

Because uncertain responses were excluded before scoring, the reported percentages are conditional on definite judgments. The paper is an arXiv version 1 preprint dated 20 August 2026, and its audio dataset is not publicly available because it contains deepfake speech from identifiable public figures, although it may be provided upon reasonable request.

Paper data and sources

Original title: Tracking the Trend in How Speech Synthesizers Deceive People
Authors: Milan Šalko, Anton Firc, Kamil Malinka et al.
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-20
DOI: Not available
Original paper · Full text

Versions and corrections

  1. Published automatically after legal-source, freshness, evidence, and independent-verification gates passed.