Pith. sign in

Paper Citation Record · LEDGER

EVT: Efficient View Transformation for Multi-Modal 3D Object Detection

As of 15 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 0 inbound Pith citation observations for arXiv:2411.10715.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.10715 v4

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T19:28:32.866993Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

60 of 60 outbound references displayed

  • verified exact0
  • verified fuzzy33
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cf34b8cf-97eb-463f-9cd9-2f91a96750f6 · outbound

This paper cites Transfusion: Robust lidar-camera fusion for 3d object detection with transform- ers.

EVT: Efficient View Transformation for Multi-Modal 3D Object Detection Transfusion: Robust lidar-camera fusion for 3d object detection with transform- ers

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:28:34.356826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:28:31.869339Z digest=sha256:3d238f492bcced2dae0e79495b4d79c6cbc540804627ce5c3bd0d0a3f7231c51

Observation 831a497f-5d72-44a3-8a01-048fa561c219 · outbound

This paper cites nuscenes: A multi- modal dataset for autonomous driving.

EVT: Efficient View Transformation for Multi-Modal 3D Object Detection nuscenes: A multi- modal dataset for autonomous driving

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:31.902701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:28:31.902701Z digest=sha256:cca67a4eaed67868cd80665aa3b0918c1cd3f99f2786ebb22708759c6bdb949b

Observation 9cdc39d3-77ac-4cd9-837f-fe535a365661 · outbound

This paper cites BEVFusion4D: Learning LiDAR-Camera Fusion Under Bird's-Eye-View via Cross-Modality Guidance and Temporal Aggregation.

EVT: Efficient View Transformation for Multi-Modal 3D Object Detection BEVFusion4D: Learning LiDAR-Camera Fusion Under Bird's-Eye-View via Cross-Modality Guidance and Temporal Aggregation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:31.945210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:28:31.945210Z digest=sha256:ce28da0c620c0ecbc2cd686e62c269ebf514f8f5134ceefaec6aa64ae3622d37

Observation ba345c00-23e6-4c50-84dc-0c33205e9177 · outbound

This paper cites Objectfusion: Multi-modal 3d object detection with object-centric fusion.

EVT: Efficient View Transformation for Multi-Modal 3D Object Detection Objectfusion: Multi-modal 3d object detection with object-centric fusion

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:28:34.324462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:28:31.952456Z digest=sha256:720a607db8dda524de83a1d525c24def14e7866dbeb66e3aa345a915e44691c0

Observation df852ae9-8fd4-47a0-8677-3426972f01fd · outbound

This paper cites Futr3d: A unified sensor fusion framework for 3d detection.

EVT: Efficient View Transformation for Multi-Modal 3D Object Detection Futr3d: A unified sensor fusion framework for 3d detection

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:31.960051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:28:31.960051Z digest=sha256:54373553828f4839bd13978e2a68f003db3cd85933a8cacf141a7cff42c5a153

Observation d326804a-495d-447c-91f9-b19b9c4e3102 · outbound

This paper cites Focal- former3d: focusing on hard instance for 3d object detection.

EVT: Efficient View Transformation for Multi-Modal 3D Object Detection Focal- former3d: focusing on hard instance for 3d object detection

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:28:34.281120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:28:31.966404Z digest=sha256:34401177e24c803d54f06fd243fd60456b373516b67de16e2bfb8a123495bb4c

Observation 1cc38717-a2b4-40fb-91d2-b16553222e95 · outbound

This paper cites Deformable feature aggregation for dynamic multi-modal 3d object detection.

EVT: Efficient View Transformation for Multi-Modal 3D Object Detection Deformable feature aggregation for dynamic multi-modal 3d object detection

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:31.973386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:28:31.973386Z digest=sha256:f030f871645131858ec872f72ebf0e71067125e4d61baa7412e498357ba53dc4

Observation 000040bf-6d9c-4652-a0e4-886a086bbae2 · outbound

This paper cites AutoAlign: Pixel-Instance Feature Aggregation for Multi-Modal 3D Object Detection.

EVT: Efficient View Transformation for Multi-Modal 3D Object Detection AutoAlign: Pixel-Instance Feature Aggregation for Multi-Modal 3D Object Detection

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:31.979698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:28:31.979698Z digest=sha256:81c77bb88b526ab08386577882582619c2674a5de907cfd081a329aaf1352d31

Observation 80d41623-02eb-4ff9-ab5e-091d93f53cba · outbound

This paper cites Li3detr: A li- dar based 3d detection transformer.

EVT: Efficient View Transformation for Multi-Modal 3D Object Detection Li3detr: A li- dar based 3d detection transformer

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:28:34.247997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:28:31.986071Z digest=sha256:409fe7799d1c14c02671572b217c0174d94781b3e6a72f95a6c93cee7d4de6e3

Observation 1b2aa19d-a697-4fc9-bee0-4efde8581b5a · outbound

This paper cites Adamixer: A fast-converging query-based object detector.

EVT: Efficient View Transformation for Multi-Modal 3D Object Detection Adamixer: A fast-converging query-based object detector

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:28:34.221067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:28:31.992416Z digest=sha256:b405842c8795df34a4bb74e701f8d5ade82368c20d235b2ee441aa2eb8e6965e

Observation cc1897e8-9e8f-4d78-b1fc-36d58d4269e4 · outbound

This paper cites Deep residual learning for image recognition.

EVT: Efficient View Transformation for Multi-Modal 3D Object Detection Deep residual learning for image recognition

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:28:34.196184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:28:31.997847Z digest=sha256:cf5dca781dfab607d62c722a7d3a03514ade4da2c4049bda0eae5b23874c3f5c

Observation b5f09690-b9c6-4449-851b-f0a9da5c29b9 · outbound

This paper cites FusionFormer: A Multi-sensory Fusion in Bird's-Eye-View and Temporal Consistent Transformer for 3D Object Detection.

EVT: Efficient View Transformation for Multi-Modal 3D Object Detection FusionFormer: A Multi-sensory Fusion in Bird's-Eye-View and Temporal Consistent Transformer for 3D Object Detection

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:32.026565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:28:32.026565Z digest=sha256:46f0ebfb5b7a3975807b8911ca9fa24b6d60456d70cbe25bdaf9f1a3df105743

Observation c57c0f88-cab8-4ab1-bac7-078e928635af · outbound

This paper cites EA-LSS: Edge-aware Lift-splat-shot Framework for 3D BEV Object Detection.

EVT: Efficient View Transformation for Multi-Modal 3D Object Detection EA-LSS: Edge-aware Lift-splat-shot Framework for 3D BEV Object Detection

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:32.073565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:28:32.073565Z digest=sha256:dae144f2c5c7a0c167ceafd58ee659e68a97989530599569049ebd76fc48eb75

Observation 8624b5a5-52ff-43fc-b83e-7b8d73ff62ab · outbound

This paper cites BEVDet4D: Exploit Temporal Cues in Multi-camera 3D Object Detection.

EVT: Efficient View Transformation for Multi-Modal 3D Object Detection BEVDet4D: Exploit Temporal Cues in Multi-camera 3D Object Detection

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:32.100561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:28:32.100561Z digest=sha256:e86423b1bbee3e72962d0082d1ec31ce7e6fdae7cc68eb2d8ee84419be0fa7ae

Observation 529056fc-d4a2-4640-9363-78d345cb6783 · outbound

This paper cites BEVDet: High-performance Multi-camera 3D Object Detection in Bird-Eye-View.

EVT: Efficient View Transformation for Multi-Modal 3D Object Detection BEVDet: High-performance Multi-camera 3D Object Detection in Bird-Eye-View

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:32.110177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:28:32.110177Z digest=sha256:b79ea9d0bf3d4a1cfa8b27e70c962e1a6ca058a1307a255cf6e5bbdaa07eaeca

Observation 90c89b00-7150-4bd5-9216-75d36552350c · outbound

This paper cites Far3d: Expanding the horizon for surround-view 3d object detec- tion.

EVT: Efficient View Transformation for Multi-Modal 3D Object Detection Far3d: Expanding the horizon for surround-view 3d object detec- tion

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:28:34.177090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:28:32.117315Z digest=sha256:32d6b2be19ec3d6e47cc85dd913823f3635aa5c1b95135e77e0c01734d3ab212

Observation 41f62e9f-b0e6-4a1a-8721-3ba7d64d5c6b · outbound

This paper cites Msmdfusion: Fusing lidar and camera at multiple scales with multi-depth seeds for 3d ob- ject detection.

EVT: Efficient View Transformation for Multi-Modal 3D Object Detection Msmdfusion: Fusing lidar and camera at multiple scales with multi-depth seeds for 3d ob- ject detection

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:28:34.152651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:28:32.123896Z digest=sha256:41a6275029da76dea3a58648ee2dfd6f708a7a703534b52580f64f602b2ac4c5

Observation dd3d40dd-fbd1-4db6-bd2a-a50017551b09 · outbound

This paper cites Centermask: Real-time anchor-free instance segmentation.

EVT: Efficient View Transformation for Multi-Modal 3D Object Detection Centermask: Real-time anchor-free instance segmentation

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:28:34.129462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:28:32.132234Z digest=sha256:40d07d7e1b254274f6c070a0dfe0b4a89a37be48ef37e890a4f66f4a7c63568f

Observation d540e775-3042-4f19-8dd5-9d4307478b73 · outbound

This paper cites Dn-detr: Accelerate detr training by intro- ducing query denoising.

EVT: Efficient View Transformation for Multi-Modal 3D Object Detection Dn-detr: Accelerate detr training by intro- ducing query denoising

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:28:34.107666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:28:32.139659Z digest=sha256:c162099a372fd4af1a8996d7cb9402ccce398a85c2d45cb6706ed6e6d96b834e

Observation eda1fe5b-6c09-458c-8853-54738a770d88 · outbound

This paper cites Gafusion: Adaptive fusing lidar and camera with multi- ple guidance for 3d object detection.

EVT: Efficient View Transformation for Multi-Modal 3D Object Detection Gafusion: Adaptive fusing lidar and camera with multi- ple guidance for 3d object detection

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:28:34.087414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:28:32.147120Z digest=sha256:3d0a073e6245f9779410fdf6c0202c1430f2ac8b85f4e74cca919ef1b95cfbcb

Observation 633861bd-206c-4d43-9e82-cab98f7642cc · outbound

This paper cites Unifying voxel-based representation with transformer for 3d object detection.

EVT: Efficient View Transformation for Multi-Modal 3D Object Detection Unifying voxel-based representation with transformer for 3d object detection

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:28:34.062138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:28:32.187755Z digest=sha256:1b8a220237aef4781bc51b74a17bd129e9324e990308f143fa960b137b2dfffd

Observation d26534f8-9184-474a-a38b-09bb6f801027 · outbound

This paper cites Bevdepth: Acquisition of reliable depth for multi-view 3d object detec- tion.

EVT: Efficient View Transformation for Multi-Modal 3D Object Detection Bevdepth: Acquisition of reliable depth for multi-view 3d object detec- tion

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:28:34.040921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:28:32.252018Z digest=sha256:c9c0b4726539d318f463196091dbc96030a4f43f8f4c4f5dc8af5dfcf880027c

Observation 195c2a61-2a56-487a-a9d3-efa5a5848897 · outbound

This paper cites Fast-bev: A fast and strong bird’s- eye view perception baseline.

EVT: Efficient View Transformation for Multi-Modal 3D Object Detection Fast-bev: A fast and strong bird’s- eye view perception baseline

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:28:34.020290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:28:32.258150Z digest=sha256:9223f52adcfa08d33032d5213c983319ad626a9ebd6b41a848207790783e7e0b

Observation 9ef83d42-a8b3-4692-ac9a-4dd59905c3be · outbound

This paper cites Bevformer: Learning bird’s-eye-view representation from multi-camera images via spatiotemporal transformers.

EVT: Efficient View Transformation for Multi-Modal 3D Object Detection Bevformer: Learning bird’s-eye-view representation from multi-camera images via spatiotemporal transformers

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:32.264382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:28:32.264382Z digest=sha256:6949058d88b17e4893e43ef5c5a783e5c1e0c5b9905f2e684852d980b61fddcf

Observation e6f6b9fe-c088-4b14-879b-a74451b72c3d · outbound

This paper cites Bevfusion: A simple and robust lidar-camera fusion framework.

EVT: Efficient View Transformation for Multi-Modal 3D Object Detection Bevfusion: A simple and robust lidar-camera fusion framework

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:28:33.972462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:28:32.270371Z digest=sha256:f5b72ea2a3d803315d2a7c3644c175d25f802ee76f1b54877c05dde3621e5690

Observation 923fb104-c073-41ce-a7c0-94a155be2e4b · outbound

This paper cites Feature pyra- mid networks for object detection.

EVT: Efficient View Transformation for Multi-Modal 3D Object Detection Feature pyra- mid networks for object detection

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:32.276545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:28:32.276545Z digest=sha256:e2d1125fbee703bd43fb53f5388c30902425495f2241225f575c05d460bcb6db

Observation 6331c775-a5b6-4b73-a6d9-ff63cd19df36 · outbound

This paper cites Sparsebev: High-performance sparse 3d object de- tection from multi-camera videos.

EVT: Efficient View Transformation for Multi-Modal 3D Object Detection Sparsebev: High-performance sparse 3d object de- tection from multi-camera videos

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:28:33.933417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:28:32.288586Z digest=sha256:eb1bccba6359fce86897ca5405ecadd4f64a047cefcac390e305cd7b864cce36

Observation d4560ec2-925c-499d-9383-79c3165e1b88 · outbound

This paper cites DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETR.

EVT: Efficient View Transformation for Multi-Modal 3D Object Detection DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETR

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:32.349200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:28:32.349200Z digest=sha256:f16e050f374655985de80519a28681ce8843cb94e7557be271d083cd656377fd

Observation e4e42334-4b28-4ebf-a877-f52f50e8dead · outbound

This paper cites Petr: Position embedding transformation for multi-view 3d object detection.

EVT: Efficient View Transformation for Multi-Modal 3D Object Detection Petr: Position embedding transformation for multi-view 3d object detection

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:28:33.916200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:28:32.389505Z digest=sha256:fa4a0f37ed6486a182ff07a21eaadc9278cf5c03729d42b3d54fa3e33ba40266

Observation d9c64888-8943-46c3-b6b3-26b7948b1dd2 · outbound

This paper cites Petrv2: A unified framework for 3d perception from multi-camera images.

EVT: Efficient View Transformation for Multi-Modal 3D Object Detection Petrv2: A unified framework for 3d perception from multi-camera images

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:28:33.898541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:28:32.396866Z digest=sha256:1c351a8d155f369916e0bebe7f33f0bfe012b94d34b58a25ce6f9076c60dd416

Observation 9a4c979a-4e49-4e10-82c4-c4b8b8e122b3 · outbound

This paper cites Bevfusion: Multi- task multi-sensor fusion with unified bird’s-eye view repre- sentation.

EVT: Efficient View Transformation for Multi-Modal 3D Object Detection Bevfusion: Multi- task multi-sensor fusion with unified bird’s-eye view repre- sentation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:28:33.880251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:28:32.403024Z digest=sha256:8dabca199ebe81aa8c330bb418bcdc72b094cfaf171400e7f22635b4697ce5d0

Observation 3df95c01-9555-4966-a908-6f2004176533 · outbound

This paper cites Decoupled weight decay regularization.

EVT: Efficient View Transformation for Multi-Modal 3D Object Detection Decoupled weight decay regularization

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:32.409589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:28:32.409589Z digest=sha256:f90eae87d7eeb7e6b164ff8302eed90065575ce12a0d2b2458b813463e8ad8b1

Observation 6e2743bd-466a-4abd-ade4-bc393fa33d7b · outbound

This paper cites DETR4D: Direct Multi-View 3D Object Detection with Sparse Attention.

EVT: Efficient View Transformation for Multi-Modal 3D Object Detection DETR4D: Direct Multi-View 3D Object Detection with Sparse Attention

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:32.417165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:28:32.417165Z digest=sha256:0d5d1499e03b05ac450a789d662c269ede47e1c9b3d22686b6b4400418bca988

Observation a6aa273d-edaa-4cac-a3e3-dc0251ef65e8 · outbound

This paper cites Lift, splat, shoot: Encod- ing images from arbitrary camera rigs by implicitly unpro- jecting to 3d.

EVT: Efficient View Transformation for Multi-Modal 3D Object Detection Lift, splat, shoot: Encod- ing images from arbitrary camera rigs by implicitly unpro- jecting to 3d

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:32.473306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:28:32.473306Z digest=sha256:4d1c2d08c8139f0a29f9683c2ccbd14fd6a07f5ed064a31378c06245d3a9fac5

Observation 713536b3-5403-487e-b5ab-765cdf112c45 · outbound

This paper cites Categorical depth distribution network for monocular 3d object detection.

EVT: Efficient View Transformation for Multi-Modal 3D Object Detection Categorical depth distribution network for monocular 3d object detection

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:32.526518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:28:32.526518Z digest=sha256:d9ab1b622889dc057952f80ee05b803fd761d7750c1235f693b4eb2e6b7df3f1

Observation 39402c44-1e5e-4d69-aa7a-e5655a14d85f · outbound

This paper cites Focal loss for dense ob- ject detection.

EVT: Efficient View Transformation for Multi-Modal 3D Object Detection Focal loss for dense ob- ject detection

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:32.533375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:28:32.533375Z digest=sha256:431b34a170d968e7fb55531c4b25999a5c9aca2705c8603e9f1c2566e5710b1b

Observation 3c7976a6-4263-4681-9bca-286c31ecf2d5 · outbound

This paper cites Cyclical learning rates for training neural networks.

EVT: Efficient View Transformation for Multi-Modal 3D Object Detection Cyclical learning rates for training neural networks

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:32.540886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:28:32.540886Z digest=sha256:c5af62a714ab81699d7e177f51a1177070a5217107761a590e54b46b5ac8ed4e

Observation 01a03ecf-58fd-43a4-91ce-6c44abcfb3fb · outbound

This paper cites Attention is all you need.

EVT: Efficient View Transformation for Multi-Modal 3D Object Detection Attention is all you need

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:28:33.782414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:28:32.547697Z digest=sha256:429d5d17e11382d4ea0e44cd57db3e460b483df0d3dea458118535518a7c35e4

Observation ca45c5d4-86e2-4fc5-ba7a-b8b2a9ef9295 · outbound

This paper cites Pointpainting: Sequential fusion for 3d object de- tection.

EVT: Efficient View Transformation for Multi-Modal 3D Object Detection Pointpainting: Sequential fusion for 3d object de- tection

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:32.553384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:28:32.553384Z digest=sha256:acc0a0113b3f7fcae7fcf769676bf0148aadd7584aa1fc7deeaf370dc4f86db0

Observation 38f4380e-b09d-4b7c-949b-cd9e077f00db · outbound

This paper cites Unitr: A unified and efficient multi-modal transformer for bird’s-eye-view repre- sentation.

EVT: Efficient View Transformation for Multi-Modal 3D Object Detection Unitr: A unified and efficient multi-modal transformer for bird’s-eye-view repre- sentation

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:28:33.738673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:28:32.559169Z digest=sha256:b85062ea2bfc08e38899a22de1123f20fa196c5a03265aa9c5bb7b973b8099d0

Observation 9433cc5e-d441-4f00-a1bc-487f80013518 · outbound

This paper cites an unresolved cited work.

EVT: Efficient View Transformation for Multi-Modal 3D Object Detection Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-12T19:28:33.716692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:28:32.564250Z digest=sha256:c7b82f161be32508516ef016873d8117e84509138698182fd629ca1667bc6d0a

Observation 31193498-e793-4c76-b2cc-cdb91510b183 · outbound

This paper cites Exploring object-centric temporal modeling for efficient multi-view 3d object detection.

EVT: Efficient View Transformation for Multi-Modal 3D Object Detection Exploring object-centric temporal modeling for efficient multi-view 3d object detection

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:28:33.694585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:28:32.570535Z digest=sha256:f7d41c71b875a602b08766d92299f7833f567b4a9db9290699b6975302544f1b

Observation 5da9b462-c8e6-47cd-a085-3d26e0abc278 · outbound

This paper cites Detr3d: 3d object detection from multi-view images via 3d-to-2d queries.

EVT: Efficient View Transformation for Multi-Modal 3D Object Detection Detr3d: 3d object detection from multi-view images via 3d-to-2d queries

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:28:33.675157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:28:32.576703Z digest=sha256:e7b3acf2681c0a6153f7186970d9a67a4838c3ce0b42ba8b2ee285299a2f744c

Observation dbd26f5a-ede3-45d9-9a24-5a8c516ac9d6 · outbound

This paper cites MV2DFusion: Leveraging Modality-Specific Object Semantics for Multi-Modal 3D Detection.

EVT: Efficient View Transformation for Multi-Modal 3D Object Detection MV2DFusion: Leveraging Modality-Specific Object Semantics for Multi-Modal 3D Detection

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:32.581814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:28:32.581814Z digest=sha256:8b1f0f9302849f527c5b31f774971f1fc629169669a7c439cbbcc85e3a9c40df

Observation 34b09d61-d230-4752-9ab5-acd1aceee89c · outbound

This paper cites Sparsefusion: Fusing multi-modal sparse rep- resentations for multi-sensor 3d object detection.

EVT: Efficient View Transformation for Multi-Modal 3D Object Detection Sparsefusion: Fusing multi-modal sparse rep- resentations for multi-sensor 3d object detection

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:28:33.651370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:28:32.636500Z digest=sha256:eef102621ed95dbdd1afcf83a0fddfdff351857c2d259437adbe61c37756040e

Observation 545362d3-1372-4ecd-94ec-8cea2e99e7ef · outbound

This paper cites Cross modal trans- former: Towards fast and robust 3d object detection.

EVT: Efficient View Transformation for Multi-Modal 3D Object Detection Cross modal trans- former: Towards fast and robust 3d object detection

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:28:33.624715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:28:32.681023Z digest=sha256:58a6c48bc466283ad5b5ee36a916496972422004f55ec4cf7229d61e537744ae

Observation e0ed78f5-633e-4b80-9302-1e313eb20f62 · outbound

This paper cites Second: Sparsely embed- ded convolutional detection.

EVT: Efficient View Transformation for Multi-Modal 3D Object Detection Second: Sparsely embed- ded convolutional detection

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:28:33.601168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:28:32.686676Z digest=sha256:aadbaf9797d4464de2f191b0774c4cbf193b247fff9a1f1d41cac16ff4505217

Observation acfb5e71-4fb1-4d32-800e-5829fb21f471 · outbound

This paper cites Bevformer v2: Adapting modern image backbones to bird’s-eye-view recognition via perspective su- pervision.

EVT: Efficient View Transformation for Multi-Modal 3D Object Detection Bevformer v2: Adapting modern image backbones to bird’s-eye-view recognition via perspective su- pervision

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:32.694155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:28:32.694155Z digest=sha256:c8c1d311471cc94716d5b1d0d7e41de13d8c3ef5aac32906c248e687106d718a

Observation dfef9795-ebcb-4d7d-8300-72ca824feeb4 · outbound

This paper cites Deepinteraction: 3d object detection via modality interaction.

EVT: Efficient View Transformation for Multi-Modal 3D Object Detection Deepinteraction: 3d object detection via modality interaction

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:28:33.555150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:28:32.700956Z digest=sha256:ab3d068ce491c47e1f9a84a9ea4df9395756f1217dbe7f5ea5a5eff8bcb73002

Observation 88b54127-fd66-471d-9912-1d16bbf095de · outbound

This paper cites DeepInteraction++: Multi-Modality Interaction for Autonomous Driving.

EVT: Efficient View Transformation for Multi-Modal 3D Object Detection DeepInteraction++: Multi-Modality Interaction for Autonomous Driving

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:32.705803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:28:32.705803Z digest=sha256:7a949855cef7e62887a7ad95e0f0152508957ff86d35c8db747b14b486e168c5

Observation 2b76b5ee-3b94-4e1e-a582-7b4dd311ea82 · outbound

This paper cites Is-fusion: Instance-scene collaborative fusion for multimodal 3d ob- ject detection.

EVT: Efficient View Transformation for Multi-Modal 3D Object Detection Is-fusion: Instance-scene collaborative fusion for multimodal 3d ob- ject detection

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:28:33.530527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:28:32.714103Z digest=sha256:b3762c57fe4fbe21003c20e8e88c147f3761c987615865e0de75de27ac36c044

Observation 6781b3c6-4158-41da-a7cd-142cc06a06fb · outbound

This paper cites Center- based 3d object detection and tracking.

EVT: Efficient View Transformation for Multi-Modal 3D Object Detection Center- based 3d object detection and tracking

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:28:33.509180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:28:32.721909Z digest=sha256:9c37f77f966b34e35b560e9b6085d3321ea196118dddb21634670590d31b01c7

Observation 58607b24-2237-4a45-a873-7aacf3030b23 · outbound

This paper cites Multi- modal virtual point 3d detection.

EVT: Efficient View Transformation for Multi-Modal 3D Object Detection Multi- modal virtual point 3d detection

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:28:33.482646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:28:32.783175Z digest=sha256:79bdb5a283c4c231a23a62af6153de259639ae7911f614d227fad88119215127

Observation 431aa5a5-6523-46d0-869d-d4cf80606bb9 · outbound

This paper cites DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection.

EVT: Efficient View Transformation for Multi-Modal 3D Object Detection DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:32.828663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:28:32.828663Z digest=sha256:3c876ad25c00d5a844f7eda602de0bc36bfe4175503c5de456f9ccf5db363e04

Observation 1b0c9fab-649f-4a49-96ac-7b9ebb21924b · outbound

This paper cites Sparselif: High-performance sparse lidar- camera fusion for 3d object detection.

EVT: Efficient View Transformation for Multi-Modal 3D Object Detection Sparselif: High-performance sparse lidar- camera fusion for 3d object detection

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:28:33.459457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:28:32.833803Z digest=sha256:97df827ccba102dd5b77b8a051299faedbc249ac12c7552600c59eec6fd5a26b

Observation d4b743e8-a98a-4ffe-90d4-43c4a94fd101 · outbound

This paper cites SimpleBEV: Improved LiDAR-Camera Fusion Architecture for 3D Object Detection.

EVT: Efficient View Transformation for Multi-Modal 3D Object Detection SimpleBEV: Improved LiDAR-Camera Fusion Architecture for 3D Object Detection

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:32.841383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:28:32.841383Z digest=sha256:15b36686410788ff8644b9e23ac28d975e12c560643149d0a6c63fffb34703c8

Observation 6e9daa43-c357-4cbc-a104-ca6d2336b3f7 · outbound

This paper cites V oxelnet: End-to-end learning for point cloud based 3d object detection.

EVT: Efficient View Transformation for Multi-Modal 3D Object Detection V oxelnet: End-to-end learning for point cloud based 3d object detection

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:28:33.333466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:28:32.847930Z digest=sha256:a5b70edadbcb304e302d4a1578ab786450ec703c469a8e5d77053fd8b84b3dda

Observation f77c4e70-c88f-41a6-8587-eb0658698e35 · outbound

This paper cites Centerformer: Center-based transformer for 3d object detection.

EVT: Efficient View Transformation for Multi-Modal 3D Object Detection Centerformer: Center-based transformer for 3d object detection

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:28:33.275324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:28:32.854555Z digest=sha256:2231c17c6e3164cdab4d98fe6085f4df7154baa32bddf40eb002c4a13c5e5e59

Observation 44122285-e031-4b7f-9445-9c6d9fc1437d · outbound

This paper cites Class-balanced Grouping and Sampling for Point Cloud 3D Object Detection.

EVT: Efficient View Transformation for Multi-Modal 3D Object Detection Class-balanced Grouping and Sampling for Point Cloud 3D Object Detection

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:32.860314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:28:32.860314Z digest=sha256:42b59b7763ab909197d7fccb0e51c2bc77835b04eaeb6fbfd96b81d5ffdd6df1

Observation c2c59f5a-8caa-48e3-8951-fb90bccff40d · outbound

This paper cites Deformable DETR: Deformable Transformers for End-to-End Object Detection.

EVT: Efficient View Transformation for Multi-Modal 3D Object Detection Deformable DETR: Deformable Transformers for End-to-End Object Detection

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:32.866993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:28:32.866993Z digest=sha256:9917e5752635afb8f7552bfecc9e217fe8191929f45949afa58b45f9d2ec9260

Pith citing papers

No inbound Pith citation observations are available.