Pith. sign in

REVIEW 2 cited by

Which2comm: An Efficient Collaborative Perception Framework for 3D Object Detection

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.17175 v2 pith:52VCKLWZ submitted 2025-03-21 cs.CV

classification cs.CV
keywords detectionperceptionobjectperformancecollaborativecommunicationfeaturessparse
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Collaborative perception allows real-time inter-agent information exchange and thus offers invaluable opportunities to enhance the perception capabilities of individual agents. However, limited communication bandwidth in practical scenarios restricts the inter-agent data transmission volume, consequently resulting in performance declines in collaborative perception systems. This implies a trade-off between perception performance and communication cost. To address this issue, we propose Which2comm, a novel multi-agent 3D object detection framework leveraging object-level sparse features. By integrating semantic information of objects into 3D object detection boxes, we introduce semantic detection boxes (SemDBs). Innovatively transmitting these information-rich object-level sparse features among agents not only significantly reduces the demanding communication volume, but also improves 3D object detection performance. Specifically, a fully sparse network is constructed to extract SemDBs from individual agents; a temporal fusion approach with a relative temporal encoding mechanism is utilized to obtain the comprehensive spatiotemporal features. Extensive experiments on the V2XSet and OPV2V datasets demonstrate that Which2comm consistently outperforms other state-of-the-art methods on both perception performance and communication cost, exhibiting better robustness to real-world latency. These results present that for multi-agent collaborative 3D object detection, transmitting only object-level sparse features is sufficient to achieve high-precision and robust performance.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Resilience of Task-Oriented V2X Networks to Incomplete Information Sharing

    cs.NI 2026-02 conditional novelty 5.0 of 10

    Task-oriented V2X networks tolerate content-selection mistakes because multiple vehicles independently observe and broadcast the same relevant objects, so omitted information is usually retransmitted by neighbors.

  2. Is Intermediate Fusion All You Need for UAV-based Collaborative Perception?

    cs.CV 2025-04 conditional novelty 5.0 of 10

    A late-intermediate fusion method that transmits only 2D and 3D detection boxes and confidence scores among UAVs, then injects them into the receiver's BEV features, achieves 72.1% mAP on UAV3D with minimal bandwidth.

Pith tools