Pith. sign in

REVIEW 2 cited by

EVP: Enhanced Visual Perception using Inverse Multi-Attentive Feature Refinement and Regularized Image-Text Alignment

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.08548 v1 pith:KA4RD346 submitted 2023-12-13 cs.CV

classification cs.CV
keywords featuredepthalignmentarchitecturedecoderdiffusionenhancedestimation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This work presents the network architecture EVP (Enhanced Visual Perception). EVP builds on the previous work VPD which paved the way to use the Stable Diffusion network for computer vision tasks. We propose two major enhancements. First, we develop the Inverse Multi-Attentive Feature Refinement (IMAFR) module which enhances feature learning capabilities by aggregating spatial information from higher pyramid levels. Second, we propose a novel image-text alignment module for improved feature extraction of the Stable Diffusion backbone. The resulting architecture is suitable for a wide variety of tasks and we demonstrate its performance in the context of single-image depth estimation with a specialized decoder using classification-based bins and referring segmentation with an off-the-shelf decoder. Comprehensive experiments conducted on established datasets show that EVP achieves state-of-the-art results in single-image depth estimation for indoor (NYU Depth v2, 11.8% RMSE improvement over VPD) and outdoor (KITTI) environments, as well as referring segmentation (RefCOCO, 2.53 IoU improvement over ReLA). The code and pre-trained models are publicly available at https://github.com/Lavreniuk/EVP.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Accuracy Does Not Guarantee Human-Likeness: Cross-Domain Human-Centered Benchmark in Monocular Depth Estimation

    cs.CV 2025-12 conditional novelty 6.0 of 10

    Across 69 monocular depth estimators, human-likeness of error patterns peaks near human-level accuracy and declines for the most accurate models: accuracy does not guarantee human-like depth perception.

  2. DiffRIS: Enhancing Referring Remote Sensing Image Segmentation with Pre-trained Text-to-Image Diffusion Models

    cs.CV 2025-06 conditional novelty 6.0 of 10

    DiffRIS combines frozen Stable Diffusion and CLIP encoders with a new text adapter and decoder to set a new state-of-the-art mean IoU on three referring remote sensing image segmentation benchmarks.

Pith tools