A smaller acoustic echo-control model had about the same overall performance as the medium CGGN16-M version while using 70% fewer parameters and 14% of its computational complexity, according to an arXiv preprint. The result came from a simulated benchmark, not a hardware deployment test, so it is a model-level comparison.
The study examines knowledge distillation, a training setup in which a larger model's output guides a smaller one. It asks whether that approach can mitigate performance loss after substantially downscaling the CGGN16 acoustic echo-control model and compares different strategies for balancing quality and computational complexity.
The comparison was built around CGGN16
CGGN16 is the baseline architecture, with its complexity adjusted through the feature-map setting F and the number of recurrent groups, g. The F = 64 version served as the teacher, and its enhanced output guided the smaller student. The design also included a two-step route in which the distilled student was fine-tuned with ground-truth labels.
The test mixtures covered three situations: double-talk, when near-end and far-end speech overlap; single-talk far-end speech; and single-talk near-end speech. The microphone signal contained near-end speech, background noise and echo. Signals were sampled at 16 kHz and processed in frames with a 128-sample shift.
Development data were close to, but disjoint from, the training data. The test set used different resources for a generalization check, drawing on TIMIT speakers, ETSI noise, an arctangent nonlinearity and Aachen room responses. Signal-to-echo and signal-to-noise ratios were varied across the mixtures.
The two-stage version ranked highest
In a development ablation, the small student was tested on Ddev double-talk with F = 8 base kernels and g = 2 GRU groups. Frequency-domain distillation alone ranked first three times among the single-step approaches. The version that added ground-truth fine-tuning recorded four overall first-place ranks, while feature matching did not improve on the no-distillation baseline.
On Dtest, the two-stage student reported a PESQ speech-quality score of 2.07, a PESQBB score of 3.56, a DT ERLEBB value of 11.19 dB, and DT O and DT E scores of 3.79 and 4.14. The analysis reported improvement over the CGGN16-S reference across near-end and echo measures, but it did not provide uncertainty estimates or formal statistical tests.
A benchmark result with a narrow reach
The subjective check used a crowd-sourced listening test on a noiseless version of Dtest, following P.808 and P.831 recommendations. Its scores favored CGGN16-S over the medium CGGN16-M, especially after the first distillation step. In comparisons with DLAC-Kalman, both distillation approaches used less than 4% of that model's computational requirements and scored better on DT O* and DT E*. Overall performance was reported as comparable with CRUSE-AEC.
The abstract reports less near-end speech distortion at 2% of the teacher's computational complexity and better overall performance than a ground-truth-trained model that was six times more complex. These comparisons remain tied to the paper's simulated training and test setup. They do not establish that the same results would hold in real rooms, on other devices or outside the reported test data.
The evidence also has several practical limits. Dataset sizes, speaker and mixture counts, and the number of listeners were not reported. The analysis gives no confidence intervals, standard errors, inferential tests or formal significance analyses. It also reports no hardware runtime, memory, energy-use or deployment measurements, and the subjective evaluation was limited to noiseless test material.
The authors state that the data are publicly available online, although TIMIT requires a small fee. Their software toolbox is described as including data-generation and evaluation scripts, along with FDKF and CGGN16 implementations.
The document is an arXiv version 1 preprint dated 26 August 2026. The supplied metadata identifies no journal or DOI publication.
Paper data and sources
Original title: Knowledge Distillation for Efficient Acoustic Echo Control
Authors: Ernst Seidel, Pejman Mowlaee, Tim Fingscheidt
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-26
DOI: Not available
Original paper · Full text