A new arXiv preprint describes EnCore, a method that feeds solver-generated early solutions into a predictor of which integer assignments will persist in a full-budget solution. Across four benchmark families, the method reported lower gaps in 11 of 12 Gurobi settings and transferred to SCIP and an eleven-instance MIPLIB IIS test.
A synthetic evaluation of financial-compliance agents found that trader models submitted actions rejected by the execution layer, while language-model monitors scored below rule-based and logistic baselines. The study also found that monitoring performance changed sharply with the evidence available to the model.
An arXiv preprint reports that Best Prefix Selection reached 0.73 measured task success in its benchmark, compared with 0.20–0.52 for the systems tested, while using 28% fewer tokens than the strongest released router.
A new arXiv preprint examines whether direct preference optimization can make continuous-time flow models score better on preferences while drifting away from their pretrained data manifold. Its proposed winner-anchored objective performed strongly in toy and image-generation tests, but broader transfer remains untested.
A conceptual preprint proposes a framework for classifying possible AI agency, separating legal from moral questions and keeping human responsibility in view.
An arXiv preprint reports a large held-out-composition gap between four-cell societies whose cells saw only assigned evidence and matched societies that saw all evidence. One globally visible model also performed well, while the study’s complete preregistered battery formally fell just short of its threshold.
A computational preprint describes an AI method that designs compact communication graphs for language-model agents. In tests across six benchmarks, it reported similar accuracy to ARG-Designer with lower average token use.
An arXiv preprint reports that DECOWAM improved future-video and whole-body action prediction in robot tests. It also used a much smaller Stage-2 parameter update, although the evidence came from a fixed replay slice and one physical platform.
An arXiv preprint reports that its V3 reward variant had the strongest combined lane-tracking and curve-progress metrics in CARLA simulation. The comparison is descriptive and does not establish real-world safety or causal superiority.
A methods preprint proposes software built around storage, a large model and an agent. In a production-scheduling prototype, direct model proposals violated constraints, while storage rejected them and solver delegation produced correct reported results in limited trials.
A benchmark found that two AI models scored lower under every evaluated memory strategy than without retrieved memories. A prompt-based safeguard improved reported scores in selected comparisons, but the study does not establish how the result translates to real-world use.
A new preprint benchmarked 10 language models on 44 English-language contracts and 3,014 lawyer-annotated tasks. Performance was uneven, suggesting a cautious expert-assisted role rather than automatic replacement of legal reviewers.
A preprint reports that tuned machine-learning models classified critical versus non-critical electronic navigational chart changes with 90% and 94% accuracy across two operational datasets. The benchmark does not establish a real-world safety or workload effect.
A new arXiv preprint reports that ten language models struggled to identify legally important omissions in incomplete legal queries. The study also found a calibration problem: models tended either to flag too much or to answer incomplete questions without acknowledging missing information.
A preprint benchmark study reports higher rule-following scores for Disentangled Multimodal Planning than for direct supervised fine-tuning in two synthetic maze environments. The results are limited to controlled computational tests, not noisy or continuous real-world settings.
A new arXiv preprint reports strong results from a hybrid neural network designed to authenticate X-band synthetic-aperture-radar signals from ICEYE satellites. The system reached 96.9% accuracy on held-out data using 10% of the collected corpus, but the quantum circuit was simulated classically and the study covered one constellation.
An arXiv preprint reports that a 1.5-billion-parameter model trained to choose NoThink, Short or Long reasoning used 41% fewer response tokens than its base model on MATH-500, with slightly lower accuracy. The routing pattern tracked problem difficulty, but the evidence covers one model, one mathematical training distribution and three seeds.
A preprint audit of language-model self-training found that measurement choices could create apparent gains and losses. In a 48-run comparison, external distillation reached more low-base problems than the tested self-training arms.
An arXiv preprint reports higher tool-use scores for Qwen3-4B and Qwen3-8B configurations that included dedicated mid-training before later training stages. The gains appeared across three benchmarks but did not extend to the MCP-Universe web-search subset, and no confidence intervals or run-to-run variability estimates were reported.
AI4AI-Bench tested six AI systems across 10 research repositories and found that agents more often changed run-level features than the procedures governing how models learn. Learning-side edits were associated with higher scores, while the benchmark did not establish recursive self-improvement.
An arXiv preprint tests a multi-agent workflow that combines chatbot survey collection with conventional, machine-learning and large-language-model prediction. Weather images produced the highest five-choice accuracy, but the study is a small, within-sample methods case study rather than evidence about population-wide travel behavior.
A new arXiv preprint reports that CAMA, a method for resolving conflicting memories in long-term multi-agent systems, posted the best displayed task and correlation-aware results among the tested methods. The evaluation used benchmark data and controlled variants, so the findings do not establish performance in naturally occurring memory logs.
A new arXiv preprint describes EXIMO, a three-stage method that uses a vision-language model to guide exploration before fine-tuning and reinforcement learning. In the reported benchmark, the approach improved success and data efficiency, while a supplementary analysis found that one distilled residual policy learned more slowly once online training began.
A preprint reports higher benchmark success when an external verifier tracks policy workflows for customer-service AI agents. It also reports lower attack success in a selected airline test and stronger procedural traces, while stressing that the evidence comes from simulated English tasks.
A preprint tests a benchmark that measures both object trajectories and recovery of mass, friction and restitution. PhyODE reported its clearest advantage on long-horizon forecasting, but trajectory accuracy and property recovery did not always align.
A computational preprint reports substantial storage savings for proof-of-concept HCAS and VCAS collision-avoidance implementations. The combined systems matched lookup-table ground truth across their complete discretized input spaces, with the guarantee stopping at the grid.
35th Congress of the International Councilof the Aeronautical Sciences (ICAS) 20263 min read
A methods preprint outlines a framework for small scientific communities to record AI-agent work, trace evidence and make trust judgments for particular purposes. The paper describes an implementation and synthetic biology examples, not an empirical evaluation.
An arXiv preprint tests whether early liquidity and trading patterns can flag memecoins that meet the study’s one-hour rug-pull rule. The models show a signal within platforms, but a sharp cross-platform drop limits what the results say about a detector used in practice.
A proposed two-tier system starts with a cheap estimate and pays for a more accurate one only when the expected value clears a cost threshold. In computational tests, the router nearly matched always-expensive routing at a low query cost and matched cheap-only routing without costly queries at a high one; the bidding version could improve a specialist’s own outcome while worsening efficiency.
An arXiv preprint evaluates Brain Researcher, a researcher-governed AI harness for neuroimaging analysis. In a paired tool-calling benchmark, the platform condition selected the correct first route or tool far more often, but verified grounding reached only 22.0% and most evidence rows still failed.
An arXiv preprint describes DARS, a reinforcement-learning approach that gives an image-editing planner structured instructions about what to change and preserve. Its authors report the best score in five benchmark regimes.