A preprint describes a sarcasm detector that adapts how it combines visual and textual clues. It reports the best F1 among listed methods on MMSD and similar gains on MMSD2.0, while its component analysis contains an internal reporting mismatch.
A new arXiv preprint describes EnCore, a method that feeds solver-generated early solutions into a predictor of which integer assignments will persist in a full-budget solution. Across four benchmark families, the method reported lower gaps in 11 of 12 Gurobi settings and transferred to SCIP and an eleven-instance MIPLIB IIS test.
A preprint evaluates G-MARK, a provenance-aware knowledge graph for cooperative-driving reasoning. Reported gains were strongest in occlusion, hidden-object and motion tasks, while communication use in the future-trajectory comparison was 0.0159 MB per sample.
An arXiv preprint reports that PruhaNLP/USER2-1C-code led an offline test of natural-language retrieval for Russian 1C/BSL code. Its ranking remained similar after an exact-match and 13-gram overlap audit, though the study did not measure human search success or production performance.
An arXiv preprint reports that 3D-CurvSegFlow led the paper’s main segmentation measures across its tested datasets. The result suggests cross-dataset promise, but does not establish clinical benefit or performance beyond the evaluated imaging settings.
A synthetic evaluation of financial-compliance agents found that trader models submitted actions rejected by the execution layer, while language-model monitors scored below rule-based and logistic baselines. The study also found that monitoring performance changed sharply with the evidence available to the model.
A new preprint benchmark shows that language resource level and the quality of translated test material can materially change how medical language models perform.
An arXiv preprint reports that Best Prefix Selection reached 0.73 measured task success in its benchmark, compared with 0.20–0.52 for the systems tested, while using 28% fewer tokens than the strongest released router.
An arXiv preprint reviews selected neural-computation literature and identifies a forward–backward disconnect: the audited configurations span several kinds of forward dynamics, but scalable learning evidence clusters around global or gradient-derived error propagation. The review warns that its counts are descriptive, not estimates of field-wide prevalence.
An arXiv preprint reports that SATS led several benchmark comparisons, including tests on datasets held out from pretraining, while using fewer parameters than key baselines.
A new arXiv preprint examines whether direct preference optimization can make continuous-time flow models score better on preferences while drifting away from their pretrained data manifold. Its proposed winner-anchored objective performed strongly in toy and image-generation tests, but broader transfer remains untested.
A retrospective comparison found different strengths between two forecasting models: Chronos-2 had the lowest aggregate error, while TabPFN-TS was better calibrated and remained among the leading models in a second network.
An AI workflow applied to Google Street View imagery in the north-eastern periphery of Nice found that 10% of the mapped network met adopted thresholds for three streetscape measures. The map showed sharp local contrasts, but coverage was incomplete and each image was scored once per task.
A conceptual preprint proposes a framework for classifying possible AI agency, separating legal from moral questions and keeping human responsibility in view.
A preprint testing FedCurv-DR, a method for federated continual learning, found higher final accuracy and less-negative forgetting scores than FedAvg in a simulated WikiArt benchmark. FedCurv-based methods also showed lower client-level disparity, while FedAvg used the least measured energy.
An arXiv preprint describes a dual-stream forecasting model that reports accuracy and efficiency gains on benchmark data, while leaving broader questions about generalisation, calibration and interpretability unanswered.
A modeling study reports that smaller proxy runs can inform learning-rate choices for much larger mixture-of-experts models. Retrospective checks were close, but the predicted setting for a 10-trillion-token run was not tested in a full-scale sweep.
An arXiv preprint reports a large held-out-composition gap between four-cell societies whose cells saw only assigned evidence and matched societies that saw all evidence. One globally visible model also performed well, while the study’s complete preregistered battery formally fell just short of its threshold.
A new arXiv preprint describes two gravity-aware geometric solvers for estimating camera pose and focal length. The methods were faster than selected comparison solvers and performed favorably in synthetic tests and Cambridge Landmarks and Aachen Day-Night evaluations, although the study does not establish energy savings or broad real-world superiority.
A methods preprint compares a JEPA with separate prediction branches against standard and other benchmark baselines across five systems, with results favoring the factorized design within the tested settings.
A preprint reports that V-REX, a veterinary-radiology model trained from scratch, matched or exceeded a listed larger fine-tuned model on report-generation scores in some comparisons. The evidence comes from offline experiments on proprietary veterinary X-ray and report data.
An arXiv preprint reports that SABET-QA, a model for answering questions involving facts and dates, scored above comparison systems on several temporal question-answering benchmarks. The paper also reports lower scores for versions missing key components.
A computational preprint describes an AI method that designs compact communication graphs for language-model agents. In tests across six benchmarks, it reported similar accuracy to ARG-Designer with lower average token use.
A modeling preprint tests a replay-free Deep artificial immune network on four grayscale class-incremental image streams. In a sklearn-digits trajectory, initial-class retention was 0.978 at the final step, while scores varied with the dataset and external readout.
A preprint describes HandMvNet, a multi-view system that reconstructs 3D hand joints and mesh vertices. The authors report lower relative errors than comparison methods across the evaluated public datasets and the highest frame rate in a figure-based comparison, while performance was weaker on the smaller HO3D-MV benchmark.
Proceedings of the 20th International Joint Conference on Computer Vision, Imaging and Computer Graphics Theory and Applications - Volume 2: VISAPP (2025), pp. 555-5623 min read
A 3,266-question wine benchmark found a sharp drop from entry-level to expert items. It also reported an uneven comparison between reasoning and standard configurations, question-family self-preference patterns and a wide cost-accuracy spread.
A flow-matching reconstruction method produced lower reported noise than MLEM, MAPEM and DDS in simulated low-count brain PET data, while showing greater visual contrast than PET-FlowDPS. The evaluation also found a favorable balance between uptake accuracy and spatial variability, but it was based on a small computational test rather than clinical outcomes.
Researchers introduce BeyondMasks, a benchmark pairing object-present videos with clean background references, and CORE, a scoring system that separately measures object disappearance and the removal of object-induced after-effects. Qualitative examples showed that methods often removed visible objects while leaving shadows, reflections, illumination changes or other traces.
An arXiv preprint reports that open-weight instruction-tuned language models use different strategies when textual, numerical and external-tool evidence conflict. Some conditions produced near-zero or below-chance accuracy.
An arXiv preprint reports that DECOWAM improved future-video and whole-body action prediction in robot tests. It also used a much smaller Stage-2 parameter update, although the evidence came from a fixed replay slice and one physical platform.
A new machine-learning method adds geographic information to sparse autoencoders and turns their internal features into rule-based explanations. The preprint reports close computational agreement between those rules and the underlying features, while leaving expert validation and operational use untested.
A computational preprint describes inference-time methods that steer a discrete diffusion language model toward sequence-level rewards without retraining. Its main comparison favored the nested methods on the paper’s automated measures, within a narrow evaluation.
An arXiv preprint reports that its V3 reward variant had the strongest combined lane-tracking and curve-progress metrics in CARLA simulation. The comparison is descriptive and does not establish real-world safety or causal superiority.
Candidate features tracked through a Vision Transformer shifted mainly toward earlier layers during training, while deeper migration peaked at 4%. Deeper layers stabilized earlier and more strongly. The result is a descriptive map of one model’s training trajectory, with technical limits on feature matching and generalization.
An arXiv preprint presents DPC-Net, a single image-restoration system tested on denoising, deraining, dehazing, deblurring and low-light enhancement. Its tables report average PSNR/SSIM pairs of 33.01/0.922 in one benchmark and 31.00/0.923 in the expanded five-degradation comparison.
A preprint reports that PelviNeXt achieved 92.00% accuracy on a deduplicated pelvic-ultrasound benchmark after the PCOSGen image pool was reduced from 4,668 images to 225. The model also reached 87.33% accuracy on a 150-image pelvic X-ray dataset.
A preprint benchmark found a wide gap between language models’ ability to describe proof strategies and their ability to translate theorems into machine-checkable form. In a separate experiment, six of 64 generated claims survived expert review and proof verification.
An arXiv preprint reports that G3Ego achieved the highest compared Macro-F1 on MECCANO action recognition and the highest average mean accuracy across EGTEA Gaze+ splits. Its graphs were much smaller than full graphs, although the pipeline still relied on a computationally heavy vision-language component.