Turn-taking systems were broadly consistent when judging whether a speaker had reached the end of a turn, but their interruption alerts changed with the conversation style, a benchmark found. End-of-turn recall was largely stable across six interaction styles, while interruption false-positive rates were type-dependent. No evaluated system was simultaneously fast, selective and high-recall.
The benchmark asks whether conversation type can be a controllable variable associated with different human and system turn-taking behavior. It spans six interaction styles, and each conversation was triple-annotated.
A benchmark built around six ways of talking
The corpus contains approximately 30 hours of English conversations spread across 154 dialogues. They were recorded in a professional studio by 106 voice actors working in 53 pairs, with each dialogue assigned one of six styles: Casual, Task-Oriented, Instructional, Collaborative, Argumentative or Narrative. The average dialogue lasted approximately 11.7 minutes.
Researchers compared 14 turn-taking systems, including rule-based detectors, prompted full-duplex models, semantic and codec endpointers, voice-activity projection and a supervised predictor. The evaluation focused on two jobs: identifying the end of a turn and identifying an interruption. For each, the benchmark measured recall, false-positive rate and signed latency, meaning how early or late a system made its decision.
The release included an approximately 104-hour speaker-disjoint training set, meaning speakers in the training data were kept separate from those used for evaluation. The benchmark dialogues were divided into a speaker-disjoint, type-balanced development and test split: 38 labelled development dialogues and 116 audio-only test dialogues.
Human annotators marked the speech events used as the reference standard. Each segment received three annotations, and an event entered the majority-consensus gold set only when at least two annotators supplied the same canonical label and agreed on its endpoints within 200 milliseconds. The resulting gold set contains 8,197 end-of-turn anchors, 4,254 mid-turn pause negatives, 1,151 interruption anchors and 17,511 backchannel and non-content negatives.
The annotation process showed high agreement by the study's measures. Pairwise Cohen's kappa ranged from 0.77 to 0.80, Fleiss' kappa was 0.78, and event-onset boundary F1, a measure of how closely annotators placed the start of an event, ranged from 0.94 to 0.96 when matches within 200 milliseconds were accepted. Overall, 85.8% of annotator events survived the filtering used to create the consensus labels.
Humans leave clues that systems still miss
The dialogues differed in ways that could make turn-taking harder to model. Argumentative dialogues supplied the most consensus interruption events, with 377. Casual and Collaborative dialogues had the fastest exchanges and highest overlap. Instructional and Narrative dialogues had the longest turns and the fewest interruptions.
Human timing was not limited to waiting for one speaker to stop before another began. Across the corpus, turn transfers began a median 281 milliseconds before the current turn ended. When floor-taking interruptions were excluded, the median lead was 151 milliseconds. After an interruption began, the interrupted speaker continued for a median 1.48 seconds.
The study examines an association in a defined setting, not a causal claim about all conversations. Its corpus is English, studio-recorded and dyadic, and it was made with voice actors in 53 pairs, so the timing results describe that benchmark's setting.
The strongest system still faced a trade-off
Among the systems that stayed within the benchmark's false-positive-rate budget, voice-activity projection, or VAP, had the strongest reported operating point on both tasks. It reached an end-of-turn recall of 0.845 at a false-positive rate of 0.055, with a median latency of 368 milliseconds. On interruption detection, it reached recall of 0.945 at a false-positive rate of 0.107, with a median latency of 994 milliseconds.
The benchmark's leaderboard ranked test recall subject to a false-positive-rate ceiling of 0.15. Gold positives were matched in a window beginning 0.25 seconds before the annotated event and ending 3 seconds after it. Under those rules, no in-budget system approached the human reference timing of beginning a transfer a median 151 milliseconds before the turn ended.
The benchmark's central trade-off was between speed, selectivity and recall. No tested system achieved all three qualities at once, and Casual interruption false-positive rates exceeded Argumentative rates for every model.
A benchmark with clear boundaries
The document is an arXiv version 1 preprint dated 25 August 2026, and no journal venue was reported.
The work used ACCESS computing allocations CIS210014 and IRI120008P supported by five listed National Science Foundation grants.
Paper data and sources
Original title: TurnBench: A Multi-Domain Benchmark for Turn-Taking Dynamics in Spoken Dialogue
Authors: Freeman Jiang, Ramon Sanabria, Soham Deshmukh et al.
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-25
DOI: Not available
Original paper · Full text