Preprint

Two-Room Simulation Shows Howling Suppression and Speech Trade-Off

Preprint: A two-room, three-terminal verification simulation showed qualitative howling suppression, while over-muting fragmented desired speech and reduced intelligibility.

Controlled spectrograms in a two-room simulated conference showed qualitative howling suppression, but overlapping speech remained a weak point. During double-talk and triple-talk, similarity decisions became ambiguous, and over-muting fragmented desired speech and reduced intelligibility, or how easy the speech was to understand.

The work is a verification simulation designed to demonstrate feasibility and expose practical problems in sound identification and playback control. It tests whether sound-object identification can be used to address acoustic echo and howling in multi-terminal conferencing with complicated acoustic paths while limiting speech-quality degradation. It does not report human participants or a live deployment.

A gate that starts muted

The proposed policy keeps channels muted by default. It permits a signal to pass or play only when the current sound is judged different from recently observed sound objects. The processing has three parts: sound-object extraction and buffering, comparison with the buffered set, and playback control. When there are not enough suitable candidates, the safe-side decision is to mute.

For the verification, the author used an initial gate based on cosine similarity between magnitude spectra, which are patterns showing how signal energy is distributed across frequencies. This was not a final sound-object identifier, and its parameters were set empirically. On the receive side, pass gain was 1.0 only when similarity was 0.66 or lower; more similar signals were muted. On the transmit side, signals were muted at similarity of 0.68 or higher. Gain changes were smoothed with a factor of 0.88. A1 used receive and transmit muting, A2 used transmit muting only, and automatic muting was disabled for B1.

A narrow simulated test

The simulation placed two terminals, A1 and A2, in Room A and one terminal, B1, in Room B. A1 was hands-free and A2 was microphone-only, with each near a male talker; B1 was hands-free and near a female talker. The modeled call included same-room crosstalk and feedback from loudspeakers in one room to microphones in the other. Reverberation was set to 500 milliseconds, background noise was about -50 dBFS, and the inter-room delay was 200 milliseconds.

Howling appeared without the control

Without the proposed control, the simulation showed howling about 2 to 3 seconds after the call began. Controlled spectrograms showed howling suppression in both Rooms A and B. The supplied analysis reports no numerical suppression measure or uncertainty interval, making this a qualitative result for the modeled scenario.

After acoustic echo cancellation, or AEC, converged at about 13 seconds, single-talk identification errors were relatively few and suppression was stable. The analysis does not report an error rate or replicate variability for this observation.

Overlapping talk exposed the cost

The more difficult case was overlapping talk. Double-talk and triple-talk made similarity decisions ambiguous and control harder. Howling was still suppressed in the simulation, but over-muting fragmented desired speech and reduced intelligibility. No quantitative speech-quality or intelligibility measure was reported.

The evidence covers one two-room, three-terminal verification scenario. The gate was an initial cosine-similarity implementation with empirically chosen parameters, not a final sound-object identifier. The supplied analysis reports no replicate count, formal threshold-selection rationale, inferential statistical analysis, effect estimate, confidence interval, or quantitative speech-quality endpoint. It also reports no human subjective evaluation or real-world deployment.

What remains unresolved

The open questions are practical as well as technical. The analysis asks whether low-latency identity matching can remain accurate under noise, reverberation, codec distortion, clock mismatch, jitter, and nonlinear processing. It also leaves open how playback could reduce residual leakage without making over-muting audible when talkers overlap, and how thresholds, buffer lengths, and processing placement should be tuned across devices, rooms, codecs, and partial deployments. Partial deployment and loops involving terminals without the control are discussed as challenges, not tested conditions.

The manuscript is identified as an arXiv version 1 preprint and states that it was submitted to IEEE Signal Processing Letters. Its acknowledgment reports generative AI assistance with English writing and parts of simulation-software development, while stating that the author takes full responsibility for the manuscript and code. No funding source, data-availability statement, code-access statement, or supplementary-material statement is reported in the supplied text.

The authors interpret the simulation as showing howling suppression with a remaining speech-quality penalty from over-muting. They identify improved sound identification and playback control, especially with deep learning, as future work. For now, the work is best read as a proof of concept for the described simulated setup.

Paper data and sources

Original title: Acoustic Echo Control Based on Sound Object Identification for Suppressing Howling Caused by Complicated Acoustic Paths
Authors: Osamu Hoshuyama
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-26
DOI: Not available
Original paper · Full text

Versions and corrections

  1. Published automatically after legal-source, freshness, evidence, and independent-verification gates passed.