An arXiv preprint reports higher math and code benchmark averages for a reward-aligned trajectory filter, with total training time close to standard on-policy distillation.
A preprint presents AERA, a controller that estimates whether more reasoning may help and reports near-baseline GSM8K accuracy with sharply lower generation use.
A preprint reports that BERT had the strongest overall scores for assigning cyber threat intelligence events to sectors, with a caveat about unequal sector coverage.
CMGR outperformed three registration methods on right coronary artery motion in XCAT simulations and selected clinical pairs, while whole-volume results were less consistent.
An open-source harness for long-horizon coding agents scored above selected systems on two benchmarks, but the study did not isolate the harness from model, prompt, tool and implementation differences.
GAAT posted leading scores in several drone-imaging tests, including detection and segmentation, but competitors remained ahead on other reported measures.
A test-suite-free framework built around an AI model produced C candidates that often compiled and matched held-out input-output checks, but the benchmark also exposed false acceptances and weaker results on stripped binaries.
A method called DA3PO was reported to outperform GRPO, DAPO and GSPO on mathematical reasoning benchmarks, though the tests covered only two Qwen3 base models and did not report uncertainty estimates.
A preprint describes an AI system that retrieves evidence before choosing how its agents collaborate. It reports higher scores across seven reasoning benchmarks, but the evidence comes from in-silico evaluations rather than human users or real-world deployments.
A conceptual note argues that Monte Carlo Tree Search can be understood as every-visit Monte Carlo control when the comparison is limited to trajectory sampling and Monte Carlo action-value updating. The claim clarifies terminology and computational scope, but does not compare performance.
A theoretical preprint offers a direct Haar-based analysis of randomized quasi-Monte Carlo integration and derives an expected squared-error bound under specified function assumptions, but reports no experiments.
A secondary analysis of four-person conversations reported that combining speech intensity, gaze and perceived interpersonal closeness distinguished gaps from overlaps better than gaze alone, including in noisy conditions.
HubMixer sends recommendation features through compact learned hubs before mixing them and writing them back to individual tokens. The preprint reports the best offline AUC on four objectives, fewer parameters than RankMixer and TokenMixer, and a 5.48% online conversion improvement before full deployment in the tested business.
A preprint reports that GOD linked browser commands to execution records, replay checks and portable packs in 15 completed run slots. The evaluation focused on command, replay and artifact checks, not human or social validity.
A new arXiv preprint reports that a coding-inspired system kept average benchmark performance close to the uncompressed setup while using 25% of the visual-token budget on Qwen3-VL-8B.
A theoretical and synthetic analysis finds that conditional flow matching can replace endpoint likelihood calculations only when specific residual errors cancel. In on-policy training, the shortcut may still improve reward even when its likelihood ratios are far from exact.
An updated Magma benchmark separates reaching a bug, satisfying its trigger condition, and detecting the resulting fault. In reported 24-hour campaigns, 77 of 127 bugs were reached and 43 were triggered, while AFL++ recorded the highest observed trigger count.
A systematic review finds that most deep-learning brain MRI reconstruction studies did not compare image-fidelity scores with radiologist assessments on the same data, leaving key safety questions unanswered.
A preprint introduces egRUE, a method that combines uncertainty scoring with feature-level explanations; its reported gains came from benchmark tests and a small BloodMNIST expert study.
A new benchmark finds that language models often handle the meaning of Chinese internet neologisms better than the sounds, characters and source forms used to create them.
A new preprint reports a method that builds lower and upper disturbance-reserve bounds, reaches exactness after finite expansions and delivers benchmark speedups over warm-started full solves.
An arXiv preprint presents a unified Rényi framework for composite binary testing, combining finite-sample bounds with asymptotic results that separate two Type II error regimes.
CF-YOLO, a YOLOv11-based detector built to refine context and fine detail, scored higher than YOLOv11n on CTDD but delivered mixed results on external NEU-DET data.
Benchmark tests on three RVV 1.0 RISC-V CPU configurations found that juFFTe's gains varied by hardware, with a reported threefold average speedup on the SG2044 multi-core comparison, while AMD Zen 5 led across the full range.
A theoretical preprint reports that fixed-order Khatri-Rao sketches can achieve a near-linear subspace-embedding dimension in the subspace size, improving the stated dependence on that size over earlier bounds. The result is a proof under sub-Gaussian assumptions, not an experimental performance claim.
A tiny two-recording test found that a parameter-free color extrapolator beat copying the last frame and every trained comparator on held-out copper in both directions. The air-to-chamber advantage was individually separated from zero, while the chamber-to-air interval included zero.
An evaluation in simulation and on a physical robot found that one shared interface could support three instruction modes, with gesture-based modes often performing better when wording, surfaces or objects changed.
H-Scale uses calibration activations to choose hardware-valid NVFP4 scales, with higher average scores reported across Qwen3 and LLaMA tests but a narrow evaluation scope.
Gibbs-family methods led the aggregate scores in Traffic Hourly, Electricity Hourly and Solar Weekly, but classical baselines remained on top in some M4 frequency and disagreement groups.
VICT uses a task's terminal verifier to trace credit through long action sequences instead of using it only as a final pass-or-fail signal. In ALFWorld and WebShop tests, it beat GRPO and achieved higher validation AUC over the same 300 updates, while the authors caution that verifier-defined links do not establish causal necessity.
An arXiv preprint proposes a 3D-grounded test for robot video models. Cosmos led composite scores, but detailed checks found weaker object localization and trajectory accuracy, underscoring the gap between plausible footage and executable behavior.
A theoretical result forces the number of cells in a bent partition to be a prime power and narrows its exponent using the dimension of the underlying finite space.
A video AI method called Token-Budget Distillation retained strong benchmark scores after aggressive visual-token compression, although its training still depended on the uncompressed teacher model.
An anatomy-aware system called CheXtriev reported stronger case-retrieval scores than global and local comparison methods on selected chest radiographs. The gains were especially notable for several lower-prevalence findings.
An arXiv preprint tests a multi-critic training method for robots that push and transport objects and open a dishwasher through contact. It reports 94.1% simulation success and 69.0% success in 58 trials on four unseen objects, while the dishwasher test is qualitative.
A preprint examining French and Egyptian Arabic movie dialogue finds that six AI systems align more closely with humans on visible social cues than on subtle power relationships, while multimodal results are limited by incomplete coverage.
Across 120 slots per condition, first post-edit re-verification appeared in 78.3% of cadence-guided slots and 26.7% of cadence-omitted slots. The same descriptive comparison showed fewer cadence violations and more bounded final successes with the guidance.