Pith. sign in

REVIEW 2 cited by

Discrete Latent Perspective Learning for Segmentation and Detection

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.10475 v1 pith:5DALZMGJ submitted 2024-06-15 cs.CV

classification cs.CV
keywords learningperspectivedlplimagesdiscretelatentdetectionframework
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this paper, we address the challenge of Perspective-Invariant Learning in machine learning and computer vision, which involves enabling a network to understand images from varying perspectives to achieve consistent semantic interpretation. While standard approaches rely on the labor-intensive collection of multi-view images or limited data augmentation techniques, we propose a novel framework, Discrete Latent Perspective Learning (DLPL), for latent multi-perspective fusion learning using conventional single-view images. DLPL comprises three main modules: Perspective Discrete Decomposition (PDD), Perspective Homography Transformation (PHT), and Perspective Invariant Attention (PIA), which work together to discretize visual features, transform perspectives, and fuse multi-perspective semantic information, respectively. DLPL is a universal perspective learning framework applicable to a variety of scenarios and vision tasks. Extensive experiments demonstrate that DLPL significantly enhances the network's capacity to depict images across diverse scenarios (daily photos, UAV, auto-driving) and tasks (detection, segmentation).

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Not Every Patch is Needed: Towards a More Efficient and Effective Backbone for Video-based Person Re-identification

    cs.CV 2025-01 conditional novelty 6.0 of 10

    A patch-selection backbone prunes redundant video patches using GOP motion and residual cues, then adds pseudo global context, matching ViT-B accuracy at roughly 26% of its FLOPs.

  2. Generating Negative Samples for Multi-Modal Recommendation

    cs.IR 2025-01 conditional novelty 6.0 of 10

    NegGen masks and replaces key attributes of items via a multi-modal LLM, generating hard negative descriptions, and uses a contrastive causal module to improve multi-modal recommendation across four Amazon datasets.

Pith tools