CineMatte uses a cross-attention design on a Siamese DINOv3 ViT plus a pretrained upsampler to produce robust mattes for virtual production, backed by a new non-synthetic 4K VP dataset that supports camera motion.
Featup: A model- agnostic framework for features at any resolution
9 Pith papers cite this work. Polarity classification is still indexing.
years
2026 9verdicts
UNVERDICTED 9representative citing papers
LLMind uses bio-inspired non-uniform sampling via a Mobius module and closed-loop semantic feedback to retain 82-97% of full-resolution VLM performance with only 1-5% of pixels on VQA benchmarks.
CoIn introduces a multi-stage pipeline using diffusion models for initial 2D inpainting, Reference Adaptive GS for 3D reconstruction, and GS-based warping plus a discriminator for multi-view consistent 3D scene inpainting supporting removal and insertion.
ART-VS achieves 95.4% convergence under perturbation in ViT-based visual servoing by adaptively using coarse then tiled high-resolution features, reducing error 53% versus standard ViT while running faster than full-resolution.
Proposes GPS representation for articulated parts, uses VR to annotate 41K frames across 234 objects, trains an RGB-D model, and achieves 73% success in heuristic manipulation policies on 9 objects.
ViCrop-Det uses spatial attention entropy from the decoder to dynamically crop and refine small-object regions in transformer detectors during inference.
PEPA is a post-encoder adapter combining target-conditioned snake upsampling and adaptive differentiable thresholding that improves topological metrics over region overlap when added to frozen-encoder curvilinear segmentation baselines.
GlowGS improves 3D Gaussian Splatting in nighttime glow scenes via semantic feature generation from diffusion models and novel-view semantic learning with vision foundation models.
PANC augments Normalized Cut with anchor-augmented token graphs using priors to steer spectral partitions, yielding mIoU gains of 2.3-8.7% over baselines on DUTS-TE, DUT-OMRON, and CrackForest.
citing papers explorer
-
CineMatte: Background Matting for Virtual Production and Beyond
CineMatte uses a cross-attention design on a Siamese DINOv3 ViT plus a pretrained upsampler to produce robust mattes for virtual production, backed by a new non-synthetic 4K VP dataset that supports camera motion.
-
LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models
LLMind uses bio-inspired non-uniform sampling via a Mobius module and closed-loop semantic feedback to retain 82-97% of full-resolution VLM performance with only 1-5% of pixels on VQA benchmarks.
-
CoIn: Comprehensive 2D-3D Inpainting with Gaussian Splatting Guidance
CoIn introduces a multi-stage pipeline using diffusion models for initial 2D inpainting, Reference Adaptive GS for 3D reconstruction, and GS-based warping plus a discriminator for multi-view consistent 3D scene inpainting supporting removal and insertion.
-
ART-VS: Adaptive Resolution Tiling for Vision Transformer Visual Servoing
ART-VS achieves 95.4% convergence under perturbation in ViT-based visual servoing by adaptively using coarse then tiled high-resolution features, reducing error 53% versus standard ViT while running faster than full-resolution.
-
Revisiting Articulated Parts Perception in Robot Manipulation
Proposes GPS representation for articulated parts, uses VR to annotate 41K frames across 234 objects, trains an RGB-D model, and achieves 73% success in heuristic manipulation policies on 9 objects.
-
ViCrop-Det: Spatial Attention Entropy Guided Cropping for Training-Free Small-Object Detection
ViCrop-Det uses spatial attention entropy from the decoder to dynamically crop and refine small-object regions in transformer detectors during inference.
-
From Reconstruction to Decision: A Post-Encoder Plug-in Adapter for Curvilinear Segmentation
PEPA is a post-encoder adapter combining target-conditioned snake upsampling and adaptive differentiable thresholding that improves topological metrics over region overlap when added to frozen-encoder curvilinear segmentation baselines.
-
GlowGS: Generative Semantic Feature Learning for 3D Gaussian Splatting in Nighttime Glow Scenes
GlowGS improves 3D Gaussian Splatting in nighttime glow scenes via semantic feature generation from diffusion models and novel-view semantic learning with vision foundation models.
-
PANC: Prior-Aware Normalized Cut via Anchor-Augmented Token Graphs
PANC augments Normalized Cut with anchor-augmented token graphs using priors to steer spectral partitions, yielding mIoU gains of 2.3-8.7% over baselines on DUTS-TE, DUT-OMRON, and CrackForest.