EPO is a trackless, edge-map-alignment framework that refines pose estimates from 3D foundation models and matches or exceeds bundle-adjustment performance with substantially lower runtime and memory use.
hub Mixed citations
In: Proceedings of the IEEE conference on computer vision and pattern recognition
Mixed citation behavior. Most common role is method (56%).
hub tools
citation-role summary
citation-polarity summary
years
2026 38representative citing papers
Introduces the ASTAD task and training-free ASTModel framework for semantically consistent asymmetric style transfer using labeled synthetic content and unlabeled real references.
DiSI disentangles stochastic interpolants into separate generation and regression paths, allowing controllable transitions between regression and generative image restoration with a unified few-step sampler.
Next-acceleration-scale autoregressive prediction in discrete latent space with on-policy privileged information distillation yields improved MRI reconstructions from sparse measurements on the fastMRI benchmark.
OP4KSR enables efficient one-step 4K super-resolution without patches by adapting Flux with RoPE rescaling and periodicity loss to suppress artifacts.
Flow of Truth is the first proactive temporal forensics framework for image-to-video generation that uses a learnable forensic template following pixel motion and a template-guided flow module to decouple motion from content.
Video diffusion models can be adapted into permutation-invariant generators for sparse novel view synthesis by treating the problem as video completion and removing temporal order cues.
Flow Divergence Sampler refines flow matching by computing velocity field divergence to correct ambiguous intermediate states during inference, improving fidelity in text-to-image and inverse problem tasks.
HairOrbit leverages video generation priors and a neural orientation extractor to achieve state-of-the-art strand-level 3D hair reconstruction from single-view portraits in visible and invisible regions.
Rule-VLN injects 177 regulatory signs into Touchdown-scale urban graphs; SNRM’s VLM perception plus mental-map detours cuts constraint violations ~19% and raises task completion ~6% zero-shot.
A real multi-city, multi-kilometer surround-view driving dataset plus an urban-tailored 3DGS baseline shows that city-scale reconstruction still degrades with scale, off-trajectory views, and real-world noise.
WING generates synthetic CT from MRI/CBCT by decomposing the target into lung, soft-tissue, and bone windows, fusing them via a differentiable soft-fusion operator, and refining with a Transformer, achieving state-of-the-art results on SynthRAD2025.
RTE-FM-Dehazer trains a flow-matching model with an RTE-derived diffusion-absorption regularizer on a new 50k real-haze dataset and reports leading results on five real-world dehazing benchmarks.
Adapting diffusion models causes hidden damage to unrelated concepts detectable via sparse autoencoders and zero-shot classification, and DriftScope provides a prompt-level token-drift diagnostic.
RosettaSim adapts frozen LLMs via structured autoregressive modeling of scene topology and agent states to reach SOTA short- and long-term traffic simulation on WOSAC, paired with RTE evaluation that correlates better with human-like fidelity.
SyncCache accelerates DiT-based audio-driven portrait animation up to 4.12x via spatially-asymmetric probing and modality-decoupled caching while preserving near-lossless quality and audio sync.
SEAR introduces a dual-process agentic framework for image restoration that combines pruning-aware MCTS planning with self-evolving episodic memory to address greedy search and episodic amnesia limitations.
StreamEdit enables high-quality training-free video editing by adapting streaming video generation models with dual-branch fast sampling, self-attention bridge, cross-attention grounding, source-oriented guidance, and visual prompting, outperforming prior methods in few-step regimes.
DS-DiT decouples LR and Ref conditions in a Siamese diffusion transformer, adds patch-level weighting, and uses autoguidance to improve reference-based super-resolution for remote sensing images.
UniFixer is a universal reference-guided framework that fixes spatial, temporal, and backbone-related degradations in diffusion-based view synthesis via coarse-to-fine modules and achieves zero-shot SOTA results on novel view synthesis and stereo conversion.
MULTI uses two-stage textual inversion to disentangle camera lens, sensor, view, and domain factors for novel image generation, supporting dataset extension and ControlNet modifications on the new DF-RICO benchmark.
SVG synthesis is cast as step-wise generation conditioned on intermediate rendered canvases, trained with Visual Self-Feedback and filtered by Render-and-Verify, claiming gains on MMSVGBench.
Empty-prompt (unconditional) inversion into a native text-to-3D model avoids "sink traps" and reconstructs and edits out-of-distribution 3D shapes more faithfully than text-guided inversion.
A single freehand sketch can generate a full orbit of photorealistic views in one pass, trained on a 9k synthetic sketch-to-multiview dataset with camera-aware adapters and SfM-supervised correspondences.
citing papers explorer
-
EPO: Boosting 3D Foundation Models with Edge-based Pose Optimization
EPO is a trackless, edge-map-alignment framework that refines pose estimates from 3D foundation models and matches or exceeds bundle-adjustment performance with substantially lower runtime and memory use.
-
ASTAD: Asymmetric Style Transfer for Synthetic-to-Real Adaptation in Autonomous Driving
Introduces the ASTAD task and training-free ASTModel framework for semantically consistent asymmetric style transfer using labeled synthetic content and unlabeled real references.
-
Disentangling Generation and Regression in Stochastic Interpolants for Controllable Image Restoration
DiSI disentangles stochastic interpolants into separate generation and regression paths, allowing controllable transitions between regression and generative image restoration with a unified few-step sampler.
-
Next-Acceleration-Scale Prediction for Autoregressive MRI Reconstruction
Next-acceleration-scale autoregressive prediction in discrete latent space with on-policy privileged information distillation yields improved MRI reconstructions from sparse measurements on the fastMRI benchmark.
-
OP4KSR: One-Step Patch-Free 4K Super-Resolution with Periodic Artifact Suppression
OP4KSR enables efficient one-step 4K super-resolution without patches by adapting Flux with RoPE rescaling and periodicity loss to suppress artifacts.
-
Flow of Truth: Proactive Temporal Forensics for Image-to-Video Generation
Flow of Truth is the first proactive temporal forensics framework for image-to-video generation that uses a learnable forensic template following pixel motion and a template-guided flow module to decouple motion from content.
-
Novel View Synthesis as Video Completion
Video diffusion models can be adapted into permutation-invariant generators for sparse novel view synthesis by treating the problem as video completion and removing temporal order cues.
-
Training-Free Refinement of Flow Matching with Divergence-based Sampling
Flow Divergence Sampler refines flow matching by computing velocity field divergence to correct ambiguous intermediate states during inference, improving fidelity in text-to-image and inverse problem tasks.
-
HairOrbit: Multi-view Aware 3D Hair Modeling from Single Portraits
HairOrbit leverages video generation priors and a neural orientation extractor to achieve state-of-the-art strand-level 3D hair reconstruction from single-view portraits in visible and invisible regions.
-
Rule-VLN: Bridging Perception and Compliance via Semantic Reasoning and Geometric Rectification
Rule-VLN injects 177 regulatory signs into Touchdown-scale urban graphs; SNRM’s VLM perception plus mental-map detours cuts constraint violations ~19% and raises task completion ~6% zero-shot.
-
WildCity: A Real-World City-Scale Testbed for Rendering, Simulation, and Spatial Intelligence
A real multi-city, multi-kilometer surround-view driving dataset plus an urban-tailored 3DGS baseline shows that city-scale reconstruction still degrades with scale, off-trajectory views, and real-world noise.
-
WING: A Window-Prior-Based Generative Network with Gated Inception for Cross-Modality CT Synthesis
WING generates synthetic CT from MRI/CBCT by decomposing the target into lung, soft-tissue, and bone windows, fusing them via a differentiable soft-fusion operator, and refining with a Transformer, achieving state-of-the-art results on SynthRAD2025.
-
RTE-FM-Dehazer: Radiative Transfer Equation Inspired Flow Matching for Real-World Image Dehazing
RTE-FM-Dehazer trains a flow-matching model with an RTE-derived diffusion-absorption regularizer on a new 50k real-haze dataset and reports leading results on five real-world dehazing benchmarks.
-
DriftScope: Measuring The Hidden Effects of Diffusion Model Adaptation
Adapting diffusion models causes hidden damage to unrelated concepts detectable via sparse autoencoders and zero-shot classification, and DriftScope provides a prompt-level token-drift diagnostic.
-
Long-term Traffic Simulation via Structured Autoregressive Modeling
RosettaSim adapts frozen LLMs via structured autoregressive modeling of scene topology and agent states to reach SOTA short- and long-term traffic simulation on WOSAC, paired with RTE evaluation that correlates better with human-like fidelity.
-
SyncCache: Exploiting Asymmetric Dynamics for Fast Audio-Driven Portrait Animation
SyncCache accelerates DiT-based audio-driven portrait animation up to 4.12x via spatially-asymmetric probing and modality-decoupled caching while preserving near-lossless quality and audio sync.
-
Self-Evolving Agentic Image Restoration via Deliberate Planning and Intuitive Execution
SEAR introduces a dual-process agentic framework for image restoration that combines pruning-aware MCTS planning with self-evolving episodic memory to address greedy search and episodic amnesia limitations.
-
StreamEdit: Training-Free Video Editing via Few-Step Streaming Video Generation
StreamEdit enables high-quality training-free video editing by adapting streaming video generation models with dual-branch fast sampling, self-attention bridge, cross-attention grounding, source-oriented guidance, and visual prompting, outperforming prior methods in few-step regimes.
-
Learning to Balance: Decoupled Siamese Diffusion Transformer for Reference-Based Remote Sensing Image Super-Resolution
DS-DiT decouples LR and Ref conditions in a Siamese diffusion transformer, adds patch-level weighting, and uses autoguidance to improve reference-based super-resolution for remote sensing images.
-
UniFixer: A Universal Reference-Guided Fixer for Diffusion-Based View Synthesis
UniFixer is a universal reference-guided framework that fixes spatial, temporal, and backbone-related degradations in diffusion-based view synthesis via coarse-to-fine modules and achieves zero-shot SOTA results on novel view synthesis and stereo conversion.
-
MULTI: Disentangling Camera Lens, Sensor, View, and Domain for Novel Image Generation
MULTI uses two-stage textual inversion to disentangle camera lens, sensor, view, and domain factors for novel image generation, supporting dataset extension and ControlNet modifications on the new DF-RICO benchmark.
-
Render-in-the-Loop: Vector Graphics Generation via Visual Self-Feedback
SVG synthesis is cast as step-wise generation conditioned on intermediate rendered canvases, trained with Visual Self-Feedback and filtered by Render-and-Verify, claiming gains on MMSVGBench.
-
Beyond Prompts: Unconditional 3D Inversion for Out-of-Distribution Shapes
Empty-prompt (unconditional) inversion into a native text-to-3D model avoids "sink traps" and reconstructs and edits out-of-distribution 3D shapes more faithfully than text-guided inversion.
-
Geometrically Consistent Multi-View Scene Generation from Freehand Sketches
A single freehand sketch can generate a full orbit of photorealistic views in one pass, trained on a 9k synthetic sketch-to-multiview dataset with camera-aware adapters and SfM-supervised correspondences.
-
HOIGS: Human-Object Interaction Gaussian Splatting
HOIGS adds a cross-attention HOI module to Gaussian Splatting that combines HexPlane human features with Cubic Hermite Spline object features to model interaction-induced deformations.
-
Improving Sparse-View 3DGS Generalization via Flat Minima Optimization
Adapts flat minima optimization to 3DGS via anisotropy-aware perturbations and periodic reinitialization to improve generalization under sparse-view supervision.
-
DVG-WM: Disentangled Video Generation Enables Efficient Embodied World Model for Robotic Manipulation
DVG-WM disentangles dynamics learning from visual synthesis via flow matching and latent degradation to deliver faster, higher-quality video predictions for robotic manipulation.
-
OTCache: Optimal Transport for Geometry-Aware Caching in Diffusion Models
OTCache uses optimal transport to interpolate caching schedules between a graph-based reference and an Optuna-optimized anchor, delivering 3.66x-4.7x speedups on FLUX.1, Qwen-Image and HunyuanVideo with improved fidelity.
-
Stochastic Optimal Control Sampling for Diffusion Inverse Problems
SOCS derives per-step closed-form control signals from stochastic optimal control to steer diffusion sampling trajectories toward measurements while preserving the generative prior.
-
PSCT-Net: Geometry-Aware Pediatric Skull CT Reconstruction via Differentiable Back-Projection and Attention-Guided Refinement
PSCT-Net introduces a geometry-aware neural framework that uses differentiable back-projection and attention-guided 3D refinement to reconstruct pediatric skull CT from bi-planar X-rays.
-
SmartPhotoCrafter: Unified Reasoning, Generation and Optimization for Automatic Photographic Image Editing
SmartPhotoCrafter performs automatic photographic image editing by coupling an Image Critic module that identifies deficiencies with a Photographic Artist module that generates edits, trained via multi-stage pretraining, reasoning supervision, and reinforcement learning.
-
Allo{SR}$^2$: Rectifying One-Step Super-Resolution to Stay Real via Allomorphic Generative Flows
AlloSR² claims state-of-the-art one-step real-world super-resolution by SNR-guided trajectory init, velocity regularization (FATC), and allomorphic self-adversarial distillation that preserves flow-matching generative priors.
-
UniGeo: Unifying Geometric Guidance for Camera-Controllable Image Editing via Video Models
UniGeo improves camera-controllable image editing by injecting point cloud geometry into a video diffusion model at the representation, architecture, and loss levels, achieving state-of-the-art geometric consistency on RE10K, DL3DV, and Tanks benchmarks.
-
ComSim: Building Scalable Real-World Robot Data Generation via Compositional Simulation
Compositional Simulation generates scalable real-world robot training data by combining classical simulation with neural simulation in a closed-loop real-sim-real augmentation pipeline.
-
EEG2Vision: A Multimodal EEG-Based Framework for 2D Visual Reconstruction in Cognitive Neuroscience
EEG2Vision reconstructs images from EEG using diffusion models plus LLM-guided boosting, with reconstruction quality holding up reasonably as electrode count drops from 128 to 24 channels.
-
CAMEO: A Conditional and Quality-Aware Multi-Agent Image Editing Orchestrator
A closed-loop multi-agent image editor (CAMEO) reports ~20% higher average win rates than strong one-shot editors on anomaly insertion and pose switching.
-
PRISM: Rethinking Atmospheric Scattering Reconstruction as a Unified Understanding and Restoration Model for Real-world Dehazing
PRISM couples proximal atmospheric-scattering reconstruction with selective self-distillation and online non-uniform haze synthesis to dehaze real images without paired clean targets.
- GeoRect4D: Geometry-Compatible Generative Rectification for Dynamic Sparse-View 3D Reconstruction