ReActor jointly optimizes motion retargeting and RL policy training with an approximate gradient to generate physically consistent robot motions from human references using only sparse body correspondences.
hub Canonical reference
In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp
Canonical reference. 100% of citing Pith papers cite this work as background.
hub tools
citation-role summary
citation-polarity summary
years
2026 12roles
background 4polarities
background 4representative citing papers
BERAG applies Bayesian ensemble weighting of individual documents via token-by-token posterior updates in retrieval-augmented generation, yielding gains on knowledge-based visual QA tasks.
Point&Grasp probabilistically integrates pointing and grasp gestures for out-of-reach object selection in MR, trained on a new ORG dataset, and outperforms single-cue baselines in user studies.
DINORANKCLIP outperforms CLIP and RANKCLIP on fine-grained and out-of-distribution tasks by injecting DINOv3 local structure and using third-order ranking consistency trained on Conceptual Captions 3M.
MooD introduces continuous valence-arousal modeling with VA-aware retrieval and perception-enhanced guidance for efficient, controllable affective image editing, plus a new AffectSet dataset.
S²VAE replaces Gaussian bottlenecks with hyperspherical Power Spherical latents in a VAE on VGGT features, yielding better results on depth estimation, camera pose recovery, and point cloud reconstruction especially at high compression.
The submitted paper simultaneously claims that oracle bounding-box cropping degrades medical VQA and that it improves it, with the abstract evaluating a different model set than the body.
ZID-Net decouples diffusion-based priors into a training-only head to create an efficient feed-forward network for single-image dehazing, reporting 40.75 dB PSNR on RESIDE and 19 ms inference.
Late fusion of asynchronous vehicle predictions improves trajectory success rate (TSR_0.5) by 1.22-1.69% on real-world V2V4Real data compared to single-vehicle forecasting.
FGINet uses a band-masked frequency encoder and layer-wise gated injection to fuse frequency artifacts with vision foundation model semantics, plus hyperspherical compactness learning, to achieve better generalization in AI-generated image detection.
ACPO uses anchor-based regularization with NR-IQA guidance to enable stable perceptual quality improvements in diffusion model fine-tuning.
Watermark removal leaves detectable forensic artifacts, so no current method balances attack success, perceptual quality, and undetectability.
citing papers explorer
-
ReActor: Reinforcement Learning for Physics-Aware Motion Retargeting
ReActor jointly optimizes motion retargeting and RL policy training with an approximate gradient to generate physically consistent robot motions from human references using only sparse body correspondences.
-
BERAG: Bayesian Ensemble Retrieval-Augmented Generation for Knowledge-based Visual Question Answering
BERAG applies Bayesian ensemble weighting of individual documents via token-by-token posterior updates in retrieval-augmented generation, yielding gains on knowledge-based visual QA tasks.
-
Point & Grasp: Flexible Selection of Out-of-Reach Objects Through Probabilistic Cue Integration
Point&Grasp probabilistically integrates pointing and grasp gestures for out-of-reach object selection in MR, trained on a new ORG dataset, and outperforms single-cue baselines in user studies.
-
DINORANKCLIP: DINOv3 Distillation and Injection for Vision-Language Pretraining with High-Order Ranking Consistency
DINORANKCLIP outperforms CLIP and RANKCLIP on fine-grained and out-of-distribution tasks by injecting DINOv3 local structure and using third-order ranking consistency trained on Conceptual Captions 3M.
-
MooD: Perception-Enhanced Efficient Affective Image Editing via Continuous Valence-Arousal Modeling
MooD introduces continuous valence-arousal modeling with VA-aware retrieval and perception-enhanced guidance for efficient, controllable affective image editing, plus a new AffectSet dataset.
-
Beyond Gaussian Bottlenecks: Topologically Aligned Encoding of Vision-Transformer Feature Spaces
S²VAE replaces Gaussian bottlenecks with hyperspherical Power Spherical latents in a VAE on VGGT features, yielding better results on depth estimation, camera pose recovery, and point cloud reconstruction especially at high compression.
-
Why Does Grounding Hurt Medical VQA? Benchmarking, Diagnosis, and Fine-Tuning of Vision-Language Models
The submitted paper simultaneously claims that oracle bounding-box cropping degrades medical VQA and that it improves it, with the abstract evaluating a different model set than the body.
-
ZID-Net: Zero-Inference Diffusion Prior Decoupling Network for Single Image Dehazing
ZID-Net decouples diffusion-based priors into a training-only head to create an efficient feed-forward network for single-image dehazing, reporting 40.75 dB PSNR on RESIDE and 19 ms inference.
-
Collaborative Trajectory Prediction via Late Fusion
Late fusion of asynchronous vehicle predictions improves trajectory success rate (TSR_0.5) by 1.22-1.69% on real-world V2V4Real data compared to single-vehicle forecasting.
-
Frequency-Aware Semantic Fusion with Gated Injection for AI-generated Image Detection
FGINet uses a band-masked frequency encoder and layer-wise gated injection to fuse frequency artifacts with vision foundation model semantics, plus hyperspherical compactness learning, to achieve better generalization in AI-generated image detection.
-
ACPO: Anchor-Constrained Perceptual Optimization for Diffusion Models with No-Reference Quality Guidance
ACPO uses anchor-based regularization with NR-IQA guidance to enable stable perceptual quality improvements in diffusion model fine-tuning.
-
The Forensic Cost of Watermark Removal: From Dedicated Attacks to Image Editing
Watermark removal leaves detectable forensic artifacts, so no current method balances attack success, perceptual quality, and undetectability.