A preprint describes a sarcasm detector that adapts how it combines visual and textual clues. It reports the best F1 among listed methods on MMSD and similar gains on MMSD2.0, while its component analysis contains an internal reporting mismatch.
An arXiv preprint reports that PruhaNLP/USER2-1C-code led an offline test of natural-language retrieval for Russian 1C/BSL code. Its ranking remained similar after an exact-match and 13-gram overlap audit, though the study did not measure human search success or production performance.
A new preprint benchmark shows that language resource level and the quality of translated test material can materially change how medical language models perform.
An arXiv preprint reports that SABET-QA, a model for answering questions involving facts and dates, scored above comparison systems on several temporal question-answering benchmarks. The paper also reports lower scores for versions missing key components.
A 3,266-question wine benchmark found a sharp drop from entry-level to expert items. It also reported an uneven comparison between reasoning and standard configurations, question-family self-preference patterns and a wide cost-accuracy spread.
An arXiv preprint reports that open-weight instruction-tuned language models use different strategies when textual, numerical and external-tool evidence conflict. Some conditions produced near-zero or below-chance accuracy.
A preprint benchmark found a wide gap between language models’ ability to describe proof strategies and their ability to translate theorems into machine-checkable form. In a separate experiment, six of 64 generated claims survived expert review and proof verification.
An arXiv preprint reports Task-CoEvolve, a method for comparing AI-agent harnesses with an adaptive subset of validation tasks. It reported stronger online text-classification results than a fixed-subset approach and 67–80% lower Terminal-Bench search costs, while averaging 51.7% versus 52.8% for full search.
A computational arXiv preprint reports that a three-stage training recipe separating document injection, answer-only QA alignment and post-hoc model merging was associated with higher scores than direct fine-tuning in its main comparisons. The authors describe IAR as a setting-dependent operating point, not a universal recipe.
Task Model Induction turns visual and input-event traces into task models that describe both what computer work is trying to achieve and how it is carried out.
A benchmark study found that current AI unlearning methods can reduce harmful memorization but do not reliably preserve useful responses involving the same concepts.
An audit of synthetic text across 11 languages found that watermark detection and quality results changed sharply with the model regime, threshold and metric. In the instruction-tuned panel, mean detection was below 0.7 in 16 of 18 cells, although AUC remained above chance in 197 of 198 entries.
An arXiv preprint describes UniLang, a language model extended with machine-native tokens. It led the reported MovieLens-20M and LePaRD benchmark comparisons, while leaving natural-language quality unexamined.
An arXiv preprint reports that LLM agents supplied with longitudinal life-event memories generally matched several patterns in human survey data more closely than static-profile and other baseline systems. The comparison covered two longitudinal surveys, three model backbones and sampled respondents, with important limits on generalization.
An arXiv preprint reports large measured speedups for FlashPrefill V2, a block-sparse prefill method, while preserving benchmark scores near full attention. The evaluation covered three model configurations and synthetic serving tests, with no confidence intervals reported.