Preprint finds language-model watermarks can look strong in one test and weak in another
An audit across 11 languages found detection fell sharply in an instruction-tuned panel, even when detector signal often remained above chance.
441–460
An audit across 11 languages found detection fell sharply in an instruction-tuned panel, even when detector signal often remained above chance.
PEtab SciML is designed to describe models that combine mechanistic equations with machine-learning components and estimate their parameters from time-series data.
A mathematical and simulation study predicts that changing reproduction and survival selection strengths can bend fitness landscapes, while a plant-data example remains only a proof of concept.
A model closely matched simulated mean rates, but it did not predict when any individual flip would occur and was less reliable near the low-energy edge.
A symmetry-based model predicts that emission becomes more cavity-like as excitation rises, but the result has not been experimentally validated.
A water-device test reported an approximately fivefold rise over its pre-switch baseline; larger Galinstan and stacked results were numerical projections.
Double-J stents were linked to fewer urinary infections and shorter hospital stays, but more ureteric stenosis; five of the six studies were retrospective.
CVSD-Reg was evaluated on KITTI, nuScenes and held-out HeLiPR scan pairs, with its strongest reported result on cross-sensor tests.
The formal construction works with balanced 2-term L∞-algebras, but remains local and has not been tested with data or computation.
UniLang posted the highest displayed scores on the reported MovieLens-20M recommendation metrics and all LePaRD legal-precedent metrics, but its effect on ordinary language output was not tested.
Two voluntary surveys of active Fledge.Love users measured stated attitudes, not real-world agent use or dating outcomes.
The single-lepton analysis finds a 6.1-standard-deviation excess, while emphasizing that its toponium interpretation depends on approximate models.
An arXiv study reports expected accuracy patterns and very small conservation errors across four numerical tests of a 1D-2V hybrid model.
In 16 mini-split runs, cooling-capacity error averaged 7.37%; a separate predictive-control comparison lasted 90 seconds in simulation.
A 41-choice comparison reports different benchmark scores and a higher real-world success rate for one latent-action-tuned model, but its nonrandomized tests do not establish cause.
An adaptive time-window approach matched full-time adjoint calculations while reporting lower runtime and storage in computational tests.
A formal analysis gives the randomized rule a tight 2k worst-case cost bound, while an optimized two-facility mixture reaches a reported ratio of about 3.519.
Five automated approaches traded places across error measures on a small network, while all were judged insufficient at reproducing stop-and-go waves.
A dynamic 4D reward was linked to better reconstruction and preference scores across three streaming models, but the evidence rests largely on learned judges and a small human test.
The laboratory system followed wave-speed changes in a rat artery and produced three-dimensional scans of a rabbit eye and one human finger, while the evidence remains limited to technical demonstrations.