A computational evaluation reports that HarnessLens improved average held-out performance by 7.6% to 13.6% across three agent harnesses and four benchmarks, while using a smaller configured interaction budget than the baselines.
A preprint reports that KLOD can update targeted facts inside language models while keeping edit success near 100 percent in benchmark tests, although broader generalization and preservation of unrelated predictions remain in tension.
A training method for autoregressive spiking language models improved reported benchmark averages over matched knowledge-distillation checkpoints and stayed non-collapsed in a 10-seed stress test. The work remains an early computational evaluation, with no hardware measurements or evidence about human language performance.
A preprint describes an offline model-based reinforcement-learning framework for cost-sensitive incentive allocation. It reports higher net profit in offline and online comparisons, while diagnostics show weaker evaluation under policy drift.
A score learned from ICU trajectories distinguished survivors from non-survivors across four baseline-severity groups, yet its patient-level behavior varied between two hospital systems.
A preprint reports that MAIL, an automated literature-based system for generating chemistry hypotheses, posted the strongest listed overlap scores on the 51-paper TOMATO-Chem benchmark. The result shows reference-idea recovery, not demonstrated chemical discovery.
A live edge testbed evaluation recorded zero wrongful actuation and complete benign success for the full Edge Skillguard policy, with the outcome measured at the software handoff to device adapters.
A preprint reports higher MO-IKE results than comparison methods in several frozen-model benchmarks. Results vary by benchmark, with retention remaining weaker than other measures in key tests.
A computational preprint tested one internal direction across five language models, reporting a wide range of tool-call rates, transfer to held-out tools and a live-search cost-accuracy trade-off.
A seven-design evaluation found that DEVICES was more accurate than two one-shot prompting baselines and used an 8.6-fold smaller input context, but it tested only documentation-level compatibility.
A retrospective registry evaluation found that a hierarchical model combining lobe-aware CT analysis with selected EHR groups reached a mean AUC of 0.8750, slightly above imaging-only REN, but the difference was not statistically significant.
MICCAI Machine Learning in Medical Imaging (MLMI 2026)5 min read
A modeling study of five signalized intersections found that pooled training produced lower 10-second trajectory errors than site-specific models at every site. Performance was less consistent at shorter horizons, and fine-tuning helped some held-out sites but not all.
A benchmark of 13 frontier AI models found partial success on prescribed computational-biology workflows, with performance lower on the deepest analyses and largest raw-data tasks.
A preprint introduces a benchmark that tests financial language models on distinct review operations and on decisions about when to request more evidence.
An arXiv preprint reports a gap between retaining required values and placing them correctly in complex JSON and tables, then tests a reward aimed at structural placement.
An AI coding benchmark linked stripped-down task specifications and thinking effort to different costs, while a low-cost probe was linked to better estimates for a held-out task.
A controlled 2D benchmark found that multimodal agents made most reconstruction gains early, often lost accuracy later, and remained well below a data-policy reference.
A preprint describes an open simulator for beyond-visual-range air combat that supports mixed aircraft, high-throughput testing and standard multi-agent interfaces. Its reported evidence comes from simulated benchmarks, aggregate policy-transfer results and a limited module-level missile case, not operational testing.
A computational preprint compared reinforcement-learning designs for language-model auditors. One calibrated pairwise checkpoint scored above an untrained baseline on the reported benchmark measures, but results varied by reward strategy and a concerningness-focused run paired a higher production-discovery score with sharply lower false-positive calibration.
PonsRAG links character and plot evidence through a separate bridge layer and led the study's reported single-step and multi-step multiple-choice comparisons.
An arXiv preprint reports that CaSKG recorded higher benchmark scores and fewer mean environment steps than Graph-of-Skills across six language-model backbones.
An arXiv preprint tests RLHEV on Unity asset classification and generation, then examines transfer, development traces and embodied diagnostics. Full RLHEV posted the highest reported Unity classification score and higher generation quality than an engine-only comparison, while related experiments produced positive transfer and ranking signals.
A preprint compares six AI systems with 18 workplace activities through shared cognitive-capability profiles, ranking Gemini 3.1 Pro highest under neutral settings while warning that the scores are comparative, not task-success probabilities.
The tested routing cases favored a monitor that fused localization and task-execution evidence, while planning horizon and cooperative fleet capacity shaped recovery.
A preprint reports lower annotated control-failure incidence than a four-baseline average, alongside benchmark success rates for a GUI-agent architecture.
The study combines spoken exchanges, surveillance and onboard data into timed rules, reporting exact synthetic verdicts and early detections in two reconstructed accidents.
A trajectory-based analysis finds that vehicle behavior often changes either nearly together or in sequence, pointing to a mix of interaction models rather than one universal pattern.
A preprint reports that SymTrace Replay reproduced recorded failures more often than full reruns, while Suspicious-Node Intervention achieved a higher task-repair rate. The comparison used benchmark failures from three multi-agent systems and different attempt budgets.
An AI system adjusted selective laser sintering settings across three polymer materials on one Inova Mk1 platform. Later batches recorded stronger results, especially for a PA12 blend, while print-bed variation and an uneven manual comparison limit what the findings can show.
A four-benchmark preprint found that the rankings of tested automated fact-checking systems varied by dataset, while tested language models scored higher with gold-annotated than TF-IDF-retrieved evidence.
A preprint reports Praxist results across machine-learning, rocket, trading, SLAM, fusion and scheduling tests, with conclusions limited to the stated evaluation setups.
A proposed exact method for robust Markov decision processes matched a baseline's values across the tested benchmarks and scaled better on Garnet and Inventory Management, while the baseline led on Frozen Lake.
A 2,527-sample benchmark reports that models often separate scientific correctness from instruction following, with the largest weaknesses in chemistry, fine-grained rules and some multimodal comparisons.
A computational evaluation found that sentence-level context paired with document-wide context was associated with the strongest tested scores for a multimodal document-questioning system.
An AI preprint reports near-full-context benchmark accuracy at lower compute, with performance varying by compression, model pairing and implementation details.
A preprint tests whether proof trajectories can teach a theorem prover to choose valid search steps across unseen problems, finding a trade-off between coverage, proof retention and search length.
A proof-of-concept study argues that LLM data agents should be judged not only by whether their answers are correct, but also by whether their underlying computations can be inspected and validated.