An arXiv preprint reports that 3D-CurvSegFlow led the paper’s main segmentation measures across its tested datasets. The result suggests cross-dataset promise, but does not establish clinical benefit or performance beyond the evaluated imaging settings.
An AI workflow applied to Google Street View imagery in the north-eastern periphery of Nice found that 10% of the mapped network met adopted thresholds for three streetscape measures. The map showed sharp local contrasts, but coverage was incomplete and each image was scored once per task.
A new arXiv preprint describes two gravity-aware geometric solvers for estimating camera pose and focal length. The methods were faster than selected comparison solvers and performed favorably in synthetic tests and Cambridge Landmarks and Aachen Day-Night evaluations, although the study does not establish energy savings or broad real-world superiority.
A preprint reports that V-REX, a veterinary-radiology model trained from scratch, matched or exceeded a listed larger fine-tuned model on report-generation scores in some comparisons. The evidence comes from offline experiments on proprietary veterinary X-ray and report data.
A modeling preprint tests a replay-free Deep artificial immune network on four grayscale class-incremental image streams. In a sklearn-digits trajectory, initial-class retention was 0.978 at the final step, while scores varied with the dataset and external readout.
A preprint describes HandMvNet, a multi-view system that reconstructs 3D hand joints and mesh vertices. The authors report lower relative errors than comparison methods across the evaluated public datasets and the highest frame rate in a figure-based comparison, while performance was weaker on the smaller HO3D-MV benchmark.
Proceedings of the 20th International Joint Conference on Computer Vision, Imaging and Computer Graphics Theory and Applications - Volume 2: VISAPP (2025), pp. 555-5623 min read
A flow-matching reconstruction method produced lower reported noise than MLEM, MAPEM and DDS in simulated low-count brain PET data, while showing greater visual contrast than PET-FlowDPS. The evaluation also found a favorable balance between uptake accuracy and spatial variability, but it was based on a small computational test rather than clinical outcomes.
Researchers introduce BeyondMasks, a benchmark pairing object-present videos with clean background references, and CORE, a scoring system that separately measures object disappearance and the removal of object-induced after-effects. Qualitative examples showed that methods often removed visible objects while leaving shadows, reflections, illumination changes or other traces.
Candidate features tracked through a Vision Transformer shifted mainly toward earlier layers during training, while deeper migration peaked at 4%. Deeper layers stabilized earlier and more strongly. The result is a descriptive map of one model’s training trajectory, with technical limits on feature matching and generalization.
An arXiv preprint presents DPC-Net, a single image-restoration system tested on denoising, deraining, dehazing, deblurring and low-light enhancement. Its tables report average PSNR/SSIM pairs of 33.01/0.922 in one benchmark and 31.00/0.923 in the expanded five-degradation comparison.
A preprint reports that PelviNeXt achieved 92.00% accuracy on a deduplicated pelvic-ultrasound benchmark after the PCOSGen image pool was reduced from 4,668 images to 225. The model also reached 87.33% accuracy on a 150-image pelvic X-ray dataset.
An arXiv preprint reports that G3Ego achieved the highest compared Macro-F1 on MECCANO action recognition and the highest average mean accuracy across EGTEA Gaze+ splits. Its graphs were much smaller than full graphs, although the pipeline still relied on a computationally heavy vision-language component.
A preprint reports that RoMAN-Flow’s one-step policy reduced action-generation latency on LIBERO-Long from about 697 milliseconds to 81.5 milliseconds while maintaining similar success. Broader tests showed stronger results for NF-IQL than imitation learning, but the evidence was limited to the listed tasks.
An arXiv preprint describes a glasses-removal system that combines synthetic training data, simulated lens effects and a temporal-stability objective. The authors report benchmark scores and participant preferences, while acknowledging that the results do not establish broad real-world performance.
A preprint reports that two prompt-guided AI models generally outperformed comparison systems across four 2D medical-image benchmarks. PROMISE-Txformer also showed closer agreement with reference cardiac measurements in one analysis, but the findings do not establish clinical benefit or deployment readiness.
A preliminary arXiv study found that forecasting three future surgical steps at a time outperformed one-shot prediction on visual and instrument-trajectory measures. Both approaches deteriorated over longer horizons.
An arXiv preprint describes CalcSeg, a model for segmenting myocardial scar in LGE-CMR images. It reports benchmark scores on multi-center data and falling estimated uncertainty across stages for six clinician-flagged challenging cases.
An arXiv preprint describes an offline system that reconstructs both hands in first-person video and reports lower error and higher measured throughput than listed baselines.
A new preprint reports competitive benchmark scores for 3B and 6B versions of the Swift-Image image-generation and editing system. The findings come from model-output comparisons, not human testing or real-world deployment.
A new camera-free LiDAR registration method reported high success across several benchmark protocols, including 99.3% overall strict success on held-out HeLiPR cross-sensor pairs. The arXiv preprint also reports large differences in an ablation comparing the full two-stage system with a Stage-2-only control, but the comparisons are protocol-specific and include no uncertainty estimates.
A new preprint reports that Stream4D produced more consistent, motion-preserving video outputs than several comparison methods, while noting that the tests do not establish physically accurate geometry or broad real-world performance.
A methods preprint describes S²GS, a sparse framework for reconstructing video from changing viewpoints. Its reported efficiency gains extended from an RTX 4090 test to a Jetson AGX Orin evaluation, but the study also reports limits when later frames need new spatial support.
A computational preprint describes a method that turns a single end-diastolic biventricular mesh into a full-cycle motion sequence. Across 666 subjects, it reported improved benchmark geometry and agreement with reference left- and right-ventricular ejection measures, while relying on fitted reference meshes.
An arXiv preprint describes a two-stage 3D model that compares brain scans with mirrored versions to segment ischemic-stroke lesions across several imaging datasets. It reported its clearest advantage on AISD NCCT scans, while results on MRI and perfusion data were competitive and uncertainty estimates were not clinically validated.
A two-phase federated method using image-level labels reported higher segmentation scores than several weakly supervised comparisons on three public brain-tumor MRI datasets. The computational results do not establish a clinical benefit.
An arXiv preprint describes Core-KAN, a vision operator that separates geometric scale control from content-dependent mixing and reports higher scores across three computer-vision benchmarks.
An arXiv preprint reports higher benchmark scores for Co-3DGT, a system designed to discover and detect 3D objects from novel categories. The reported gains appeared on two datasets and in a separate annotation-free evaluation, but the evidence remains limited to benchmark comparisons.
A preprint reports higher seen–unseen balance scores for a CLIP-based method that uses diffusion features during training and removes that branch at inference.
An arXiv preprint describes STEP, a video anomaly detector that uses whitened PCA pose representations, confidence weighting and noise-scale conditioning. It reports 90.1% AUROC on UBnormal’s full test set and 90.9% on its Human-Related split across 20 training runs. Results were lower on the broader MSAD benchmark, highlighting the limits of relying on tracked human posture alone.
An arXiv preprint describes a differentiable renderer that reconstructs object geometry from posed images when direct lighting is known. Its strongest results came in a controlled synthetic test, and the study does not establish how the method will perform on real captures.
An arXiv preprint describes UPAL, a single-network computer-vision system that jointly extracts points, line segments and feature descriptors. Across image benchmarks, the authors report strong results alongside low parameter count and latency. The experiments are descriptive and do not report confidence intervals or run-to-run variability.
A preprint describes a video-search system that combines an image of a specific instance with a text description to locate the matching action. Its reported scores were higher than a named baseline on the proposed benchmarks, including a Web test, but the study remains limited to benchmark comparisons.
A computer-vision preprint reports that 4DAnyone produced stronger results than listed comparison methods for multiview video consistency and downstream 4D Gaussian Splatting reconstruction. The evidence comes from computational benchmarks, ablation tests and qualitative examples rather than human-participant research.
A version 1 arXiv preprint describes WithEveryone, a unified model for generating group images from multiple reference identities. Its reported results were stronger than GPT-Image 2 on one automated benchmark, though the evaluation was limited to a single dataset.