Pith. sign in

REVIEW 10 cited by

FeatUp: A Model-Agnostic Framework for Features at Any Resolution

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.10516 v2 pith:KKWLZGBL submitted 2024-03-15 cs.CV cs.AIcs.IRcs.LG

FeatUp: A Model-Agnostic Framework for Features at Any Resolution

classification cs.CV cs.AIcs.IRcs.LG
keywords featuresfeatupresolutiondeepimagepredictionsegmentationapproaches
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Deep features are a cornerstone of computer vision research, capturing image semantics and enabling the community to solve downstream tasks even in the zero- or few-shot regime. However, these features often lack the spatial resolution to directly perform dense prediction tasks like segmentation and depth prediction because models aggressively pool information over large areas. In this work, we introduce FeatUp, a task- and model-agnostic framework to restore lost spatial information in deep features. We introduce two variants of FeatUp: one that guides features with high-resolution signal in a single forward pass, and one that fits an implicit model to a single image to reconstruct features at any resolution. Both approaches use a multi-view consistency loss with deep analogies to NeRFs. Our features retain their original semantics and can be swapped into existing applications to yield resolution and performance gains even without re-training. We show that FeatUp significantly outperforms other feature upsampling and image super-resolution approaches in class activation map generation, transfer learning for segmentation and depth prediction, and end-to-end training for semantic segmentation.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. CineMatte: Background Matting for Virtual Production and Beyond

    cs.CV 2026-05 unverdicted novelty 7.0

    CineMatte uses a cross-attention design on a Siamese DINOv3 ViT plus a pretrained upsampler to produce robust mattes for virtual production, backed by a new non-synthetic 4K VP dataset that supports camera motion.

  2. LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models

    cs.CV 2026-03 unverdicted novelty 7.0

    LLMind uses bio-inspired non-uniform sampling via a Mobius module and closed-loop semantic feedback to retain 82-97% of full-resolution VLM performance with only 1-5% of pixels on VQA benchmarks.

  3. VLRC: Vision-Language Reprojection Consistency as a scalable signal for better feed-forward 3D pretraining

    cs.CV 2026-07 conditional novelty 6.5

    Enforcing multi-view consistency of dense frozen VLM features via predicted 3D reprojection improves feed-forward depth, camera estimates and zero-shot open-vocabulary 3D segmentation.

  4. CoIn: Comprehensive 2D-3D Inpainting with Gaussian Splatting Guidance

    cs.CV 2026-06 unverdicted novelty 6.0

    CoIn introduces a multi-stage pipeline using diffusion models for initial 2D inpainting, Reference Adaptive GS for 3D reconstruction, and GS-based warping plus a discriminator for multi-view consistent 3D scene inpain...

  5. ART-VS: Adaptive Resolution Tiling for Vision Transformer Visual Servoing

    cs.RO 2026-06 unverdicted novelty 6.0

    ART-VS achieves 95.4% convergence under perturbation in ViT-based visual servoing by adaptively using coarse then tiled high-resolution features, reducing error 53% versus standard ViT while running faster than full-r...

  6. Revisiting Articulated Parts Perception in Robot Manipulation

    cs.RO 2026-06 unverdicted novelty 6.0

    Proposes GPS representation for articulated parts, uses VR to annotate 41K frames across 234 objects, trains an RGB-D model, and achieves 73% success in heuristic manipulation policies on 9 objects.

  7. ViCrop-Det: Spatial Attention Entropy Guided Cropping for Training-Free Small-Object Detection

    cs.CV 2026-04 unverdicted novelty 6.0

    ViCrop-Det uses spatial attention entropy from the decoder to dynamically crop and refine small-object regions in transformer detectors during inference.

  8. From Reconstruction to Decision: A Post-Encoder Plug-in Adapter for Curvilinear Segmentation

    cs.CV 2026-06 unverdicted novelty 5.0

    PEPA is a post-encoder adapter combining target-conditioned snake upsampling and adaptive differentiable thresholding that improves topological metrics over region overlap when added to frozen-encoder curvilinear segm...

  9. GlowGS: Generative Semantic Feature Learning for 3D Gaussian Splatting in Nighttime Glow Scenes

    cs.CV 2026-05 unverdicted novelty 5.0

    GlowGS improves 3D Gaussian Splatting in nighttime glow scenes via semantic feature generation from diffusion models and novel-view semantic learning with vision foundation models.

  10. PANC: Prior-Aware Normalized Cut via Anchor-Augmented Token Graphs

    cs.CV 2026-02 unverdicted novelty 5.0

    PANC augments Normalized Cut with anchor-augmented token graphs using priors to steer spectral partitions, yielding mIoU gains of 2.3-8.7% over baselines on DUTS-TE, DUT-OMRON, and CrackForest.