AI models recorded different emotion-recognition scores under matched views of the same simulated episode, according to the report. The main score was Macro-F1, a single measure of performance across the benchmark's five discrete emotion categories. Under the privileged reference condition, P-Ref, mean Macro-F1 was higher than under the stationary starting condition, P-Init: 15.56% versus 9.89% for open-source configurations and 30.78% versus 22.61% for closed-source models. An active condition, A-Obs, scored above P-Init in 21 of 24 configurations, but a substantial gap to P-Ref remained.
Same episode, different observation conditions
AffectSim separates emotion-expressive motion from observation conditions, allowing the same affective performance to be replayed under changes in distance, orientation, occlusion, scene geometry and agent initialization. In the matched comparisons, scene, motion, label, rendering, camera, clip duration and frame sampling stayed constant. P-Ref was a privileged reference trajectory, but the report does not treat it as an oracle or an upper bound.
The benchmark contains 27,647 episodes across 57 scenes: 26,207 single-person episodes from 1,259 motion assets and 1,440 dyadic episodes from 48 interpersonal motion assets. It covers five discrete emotion categories.
Source performances were kept within a single split to prevent leakage. The training, validation and test splits contained 19,662, 4,044 and 3,941 episodes, respectively. The evaluation used 19 open-source and five closed-source model configurations, and no recognizer was fine-tuned on AffectSim.
Reference and active conditions
Across the 19 open-source configurations, mean Macro-F1 was 9.89% under P-Init and 15.56% under P-Ref. Across the five closed-source models, the corresponding scores were 22.61% and 30.78%. P-Ref was the highest-scoring condition in 17 of 19 open-source configurations and in 22 of all 24 configurations.
The active baseline used two stages: a frozen ETPNav-based search with a handoff, followed by a lightweight controller that followed the performer using location and depth cues. In the A-Obs condition, mean Macro-F1 was 11.70% for open-source models and 24.26% for closed-source models, compared with 9.89% and 22.61% under P-Init. A-Obs outperformed P-Init in 21 of 24 configurations.
The active protocol acquired a new post-handoff observation in 1,881 episodes. In the remaining 2,060 episodes, evaluation fell back to the paired P-Init view. On the new-observation subset, Macro-F1 gains were 3.40 percentage points for open-source models and 2.91 points for closed-source models, compared with 1.81 and 1.64 points on the complete protocol.
Using Macro-F1, A-Obs recovered 32.0% of the mean P-Ref minus P-Init gap for open-source models and 20.1% for closed-source models. A substantial gap to P-Ref remained. No confidence intervals or other uncertainty estimates were reported for the main comparisons, so the averages show the size and direction of the reported differences without showing how precisely they were estimated.
The active run's diagnostics
The evaluation also used an episode-level Recovery Rate to record cases in which active observation corrected an initial error. Its mean was 3.10% for open-source models and 7.24% for closed-source models. Claude Opus 5 reached 11.27% in its group, while Emotion-LLaMA reached 7.88% in its group.
Operational diagnostics showed a stable handoff in 47.73% of test episodes and a mean path length of 23.23 metres per episode. The diagnostics also showed that 97.4% of locomotion occurred before handoff. The movement budget was exhausted in 52.27% of episodes, and 14.10% of attempted translations were clipped.
A separate path-aware efficiency measure, E-SPL, had a mean of 1.70% at a target distance of 1 metre, 0.95% at 3 metres and 0.78% at 5 metres across the 24 model configurations. The reported mean was lower at each longer distance.
What the benchmark leaves out
The benchmark relies primarily on acted expressions. Avatar rendering and motion retargeting introduce a simulation-to-reality gap, the label space contains only five discrete emotion categories, and simulated people follow predefined, nonreactive motions. The setup does not represent spontaneous, mixed, continuous or culturally dependent affective states, and it does not model fully closed-loop affective interaction.
The work is identified as a technical report dated 26 August 2026. Within the simulated benchmark, the comparison is condition-specific: P-Ref scored above P-Init, A-Obs scored above P-Init in most configurations, and A-Obs recovered only part of the P-Ref minus P-Init gap. The findings are confined to the simulated episodes, five-category label set and nonreactive motions described above.
Paper data and sources
Original title: AffectSim: A Controllable Interactive 3D Simulation Benchmark for Embodied Affective Perception
Authors: Ke Xing, Zhilong Wang, Zheng Lian et al.
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-26
DOI: Not available
Original paper · Full text