Pith. sign in

cs.CV

Computer Vision and Pattern Recognition

Covers image processing, computer vision, pattern recognition, and scene understanding. Roughly includes material in ACM Subject Classes I.2.10, I.4, and I.5.

Papers reviewed in the last 7 days lead, then the papers readers actually read. Ranking is not a quality score.

sort pith recommended most recent

New language converts single images into simulation-ready sewing patterns

PatternGSL encodes panel boundaries, seams and stitch topology without templates; a VLM predicts it and deterministic rules decode to valid

· “PatternGSL: A Structured Specification Language for Template-Free and Simulation-Ready 3D Garments”

open re-runnable review →
Figure from the paper

Iterative diffusion produces identifiable suspect faces from multi-modal inputs

Multi-modal controls and step-by-step refinement raise identity retrieval rates over standard text-to-image diffusion in crime scenarios.

· “IdentiFace: Multi-Modal Iterative Diffusion Framework for Identifiable Suspect Face Generation in Crime Investigations”

open re-runnable review →
Figure from the paper

Hybrid field fuses rigid and deformable fields for cervical CT-MRI

MSR aligns each vertebra rigidly then applies gated Mamba-Swin deformable correction before fusing into one hybrid field and releases an new

· “MSR:Hybrid Field Modeling for CT-MRI Rigid-Deformable Registration of the Cervical Spine with an Annotated Dataset”

open re-runnable review →
Figure from the paper

Fusing token and verbal confidence improves MLLM calibration

Monotone merge of dual signals plus mean alignment yields more reliable reliability estimates while preserving selective prediction trade-o

· “Instinct vs. Reflection: Unifying Token and Verbalized Confidence in Multimodal Large Models”

open re-runnable review →
Figure from the paper

New benchmark tests segmentation on 661 transmission corridor scenes

TowerDataset supplies 22 fine classes and a fusion framework to handle long corridors and rare components in point cloud inspection.

· “TowerDataset: A Heterogeneous Benchmark for Transmission Corridor Segmentation with a Global-Local Fusion Framework”

open re-runnable review →
Figure from the paper

PBE-UNet beats prior methods on ultrasound lesion segmentation

Scale-aware aggregation and progressive widening of boundary attention maps improve accuracy on four benchmark datasets while keeping the U-

· “PBE-UNet: A light weight Progressive Boundary-Enhanced U-Net with Scale-Aware Aggregation for Ultrasound Image Segmentation”

open re-runnable review →
Figure from the paper

Wavelet regularization stabilizes sparse-view 3D Gaussian Splatting

Local and global frequency penalties yield sharper, more consistent reconstructions from limited input images

· “LGDWT-GS: Local and Global Discrete Wavelet-Regularized 3D Gaussian Splatting for Sparse-View Scene Reconstruction”

open re-runnable review →
Figure from the paper

Expert-log regularization cuts collisions in offline driving RL

Using pseudo ground-truth trajectories stabilizes training on fixed simulator data and raises route completion over imitation baselines.

· “Pseudo-Expert Regularized Offline RL for End-to-End Autonomous Driving in Photorealistic Closed-Loop Environments”

open re-runnable review →
Figure from the paper

Four lost Iraqi mounds found by AI on 1960s spy-satellite imagery

A retrained segmentation model reads pre-destruction CORONA photos, confirming sites invisible on modern maps.

· “AI-ming backwards: Vanishing archaeological landscapes in Mesopotamia and automatic detection of sites on CORONA imagery”

open re-runnable review →
Figure from the paper

Self-supervised video model plans robot actions zero-shot on new arms

Pretrained on over a million hours of web video plus under 62 hours of unlabeled robot data, it enables pick-and-place planning without any

· “V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning”

open re-runnable review →

Training-free memory bank beats fine-tuning TTA for medical segmentation

Curating reliable image-text predictions and matching prototypes yields up to +12.2% DSC over fine-tuning on optic disc and lung benchmarks.

· “Memory-Supported Synergistic Adaptation for Training-Free Test-Time Medical Image Segmentation”

open re-runnable review →
Figure from the paper

Real video clips fix long-rollout drift in streaming video AI

Swapping real footage into a teacher's memory cache during training gives streaming video models dense corrective targets with zero added推理

· “OPSD-V: On-Policy Self-Distillation for Post-Training Few-Step Autoregressive Video Generators”

open re-runnable review →
Figure from the paper

Automated matching creates region labels for better person search

ROGLE mines pseudo region-sentence pairs from existing data to add local alignment, raising accuracy on detailed natural-language queries.

· “ROGLE: Robust Global-Local Alignment with Automated Region Supervision for Text-Based Person Search”

open re-runnable review →
Figure from the paper

Generic metrics fail to predict success in grounded presentation tasks

UniPPTBench experiments show performance varies sharply by input setting and that visual appeal scores do not ensure content grounding or跨源s

· “UniPPTBench: A Unified Benchmark for Presentation Generation Across Diverse Input Settings”

open re-runnable review →
Figure from the paper

Training-free axis detects CT artifacts and grades severity

Radiology text prompts plus spectral cues enable joint type and severity prediction on single and mixed degradations, beating baselines.

· “CT-DegradBench: A Physics-Informed Benchmark for CT Degradation Detection and Severity Estimation”

open re-runnable review →
Figure from the paper

Handover during denoising fixes virtual try-on structure-texture trade-off

A single switch from structure-biased to texture-biased control inside one diffusion process improves perceptual quality while keeping align

· “LPH-VTON: Resolving the Structure-Texture Dilemma of Virtual Try-On via Latent Process Handover”

open re-runnable review →
Figure from the paper

Mixture of inverse models turns robot video predictions into actions

By extracting latent actions from semantic, depth, and flow transitions, MoLA improves success and consistency over direct use of generated帧

· “From Imagined Futures to Executable Actions: Mixture of Latent Actions for Robot Manipulation”

open re-runnable review →
Figure from the paper

Semantic codebook creates style-matched co-speech gestures

By organizing motion codes according to gesture semantics and applying reference prompts, the system produces motions faithful to both word,

· “PersonaGest: Personalized Co-Speech Gesture Generation with Semantic-Guided Hierarchical Motion Representation”

open re-runnable review →
Figure from the paper

Masks supply precise text to boost remote sensing change detection

Transcribing where-what-how-how-many details from ground-truth labels yields better accuracy than large language model methods on a new Gaza

· “Masks Can Talk: Extracting Structured Text Information from Single-Modal Images for Remote Sensing Change Detection”

open re-runnable review →
Figure from the paper

Few-shot CNN pipeline classifies monkeypox lesions from limited data

Frozen backbones plus nearest-centroid classification deliver high accuracy and stable binary transfer across three public datasets.

· “Few-Shot Learning Pipeline for Monkeypox Skin Disease Classification Using CNN Feature Extractors”

open re-runnable review →
Figure from the paper

Network selects key bands to fuse multi-source remote sensing data

RSCNet uses cross-source guidance for spectral selection and adaptive fusion to outperform prior methods with lower complexity on benchmark

· “Representative Spectral Correlation Network for Multi-source Remote Sensing Image Classification”

open re-runnable review →
Figure from the paper

Probabilistic model resists label noise for fruit maturity

Treating ripening as a continuous distribution rather than discrete classes improves reliability when annotator labels contain boundary fuzz

· “FruitProM-V2: Robust Probabilistic Maturity Estimation and Detection of Fruits and Vegetables”

open re-runnable review →
Figure from the paper

Shot boundary detection recast as relational query prediction

OmniShotCut uses a Transformer with shot queries to jointly model ranges and relations, trained on fully synthetic transitions for exact, sc

· “OmniShotCut: Holistic Relational Shot Boundary Detection with Shot-Query Transformer”

open re-runnable review →
Figure from the paper

Mask-free local edits now work in frozen diffusion transformers

Lightweight adapters factorize instructions from spatial regions and let a learned predictor supply the mask, beating both mask-free and or-

· “Edit Where You Mean: Region-Aware Adapter Injection for Mask-Free Local Image Editing”

open re-runnable review →
Figure from the paper

Hierarchical consistency fixes label errors in open-vocabulary detection

A new calibration technique plus an objectness token in CLIP deliver reliable pseudo labels and set new performance records on COCO and LVIS

· “Exploring Hierarchical Consistency and Unbiased Objectness for Open-Vocabulary Object Detection”

open re-runnable review →
Figure from the paper

browse all of cs.CV → full archive · search · sub-categories