Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T12:33:45.234194Z
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 66 of 66 outbound references and 0 inbound Pith citation observations for arXiv:2504.12576.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T12:33:45.234194Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
66 of 66 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 8ba16635-3d8c-4301-b6a1-6f41a79519c6 · outbound
CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Multimae: Multi-modal multi-task masked autoen- coders
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5703e613-8c86-4795-86d5-1de470fba762 · outbound
CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Beit: Bert pre-training of image transformers
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation df0011cb-5772-4953-90df-ea39ecaf7b04 · outbound
CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Is Space-Time Attention All You Need for Video Understanding?
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a40b3fa2-2778-4886-8096-0e6779e508ee · outbound
CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework DeepSeek LLM: Scaling Open-Source Language Models with Longtermism
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9860ca14-a597-4263-b046-560ef9db0abc · outbound
CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Lan- guage models are few-shot learners
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa93e432-08d9-44cd-b267-138a72124aeb · outbound
CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework End-to- end object detection with transformers
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c08d217c-1bb1-4b4a-8664-9067cb276728 · outbound
CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Weakly misalignment-free adaptive feature alignment for uavs- based multimodal object detection
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation dfc6f7f1-7290-4ce9-ac27-f489d5702f41 · outbound
CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework An empirical study of training self-supervised vision transformers
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 2c7e24f8-b6e3-4255-b13a-092e665e96e1 · outbound
CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Segment any event streams via 11 weighted adaptation of pivotal tokens
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 3439abeb-f364-49b7-b0f9-4d799dd7798b · outbound
CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Unihcp: A unified model for human-centric perceptions
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 38d2fb5d-d24e-432b-b290-d20ae97bc3c3 · outbound
CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Satmae: Pre-training transformers for tem- poral and multi-spectral satellite imagery
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 98ac94c9-fdf6-4ef3-8dc5-d253015a5fa4 · outbound
CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Deepseekmoe: Towards ultimate expert spe- cialization in mixture-of-experts language models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f59cbf1d-53eb-4070-8f65-f78cd35eb765 · outbound
CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Sfod: Spiking fusion object detector
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9980f297-1682-430c-8ef0-1241ade1e83f · outbound
CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Hypergraph-based multi-view action recognition using event cameras
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation de460341-6127-4e57-8104-119f7d72db4d · outbound
CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Multimodal Masked Autoencoders Learn Transferable Representations
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe194490-520d-45cd-81c3-7fae7ea3d348 · outbound
CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework The Llama 3 Herd of Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60b8eedb-36d3-441c-97e2-7c3a88c27f93 · outbound
CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e0369dd-4a64-4df6-b92c-11afd59e42b5 · outbound
CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Momentum contrast for unsupervised visual rep- resentation learning
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8bd98e9c-45d6-4423-a31f-055b5ce28c91 · outbound
CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Masked autoencoders are scalable vision learners
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 034413c5-16c7-4253-b51d-d1856d61227a · outbound
CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Data-efficient Event Camera Pre-training via Disentangled Masked Modeling
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed5fa96b-ccab-45d1-8890-cdbb3cf69b19 · outbound
CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework N-imagenet: Towards robust, fine-grained object recognition with event cameras
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation fefa5436-bb6e-4ac9-a0fb-110eb4d3a2f5 · outbound
CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Spiking-yolo: spiking neural network for energy- efficient object detection
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 800ecb3d-ae9e-4fe9-a87e-f5854ef98554 · outbound
CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Segment any- thing
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2ff80af-8f69-4d7e-bc4b-cbc8b470212a · outbound
CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Masked event modeling: Self-supervised pretraining for event cameras
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de3313fb-2b40-487a-a9d7-94727cda921c · outbound
CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Openess: Event-based semantic scene understanding with open vocabularies
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation bd9ab563-4775-4dfa-ba6b-41e4b9ad8aab · outbound
CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Mulfs-cap: Multimodal fusion- supervised cross-modality alignment perception for unreg- istered infrared-visible image fusion
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3869773-b7c9-4d8a-987f-c2c7e5b13175 · outbound
CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Coupled mamba: Enhanced multimodal fusion with coupled state space model
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 97e04996-459f-4629-879e-6143eabe9f48 · outbound
CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Exploring plain vision transformer backbones for object de- tection
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 55f2a1b4-1763-4618-aab0-8d7753b2d2f9 · outbound
CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0309dcdf-261d-4b8a-b9b7-223f75749363 · outbound
CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Visual instruction tuning
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f68ca61a-21d8-465b-9239-dd19f5b72050 · outbound
CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Pixmim: Rethinking pixel reconstruction in masked image modeling
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 02ce43ba-edb4-4727-be67-8050ee285cda · outbound
CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Decoupled weight de- cay regularization
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation cfa01b67-b41b-4160-90f4-b4e40bbcf0a2 · outbound
CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Integer-valued training and spike-driven inference spiking neural network for high-performance and energy-efficient object detection
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c3370488-1ea8-492f-9a40-ae424cf4d65b · outbound
CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Event-based moving object 12 detection and tracking
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e3dc42bd-b708-467e-86eb-359e73d63f8a · outbound
CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Rethinking transformers pre-training for multi- spectral satellite imagery
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5b8fe7b3-96cb-40c8-9a09-78a03e83ce5e · outbound
CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Dinov2: Learning robust visual features without supervision
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20971ba1-d35b-46ab-99ed-729cfdd53ed2 · outbound
CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Pytorch: An im- perative style, high-performance deep learning library
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f8113807-ab7e-474d-9622-52d29421a468 · outbound
CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework BEiT v2: Masked Image Modeling with Vector-Quantized Visual Tokenizers
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 222c146c-c03c-48cb-a699-0c35812d253f · outbound
CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Detectors: Detecting objects with recursive feature pyramid and switch- able atrous convolution
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0105fd7e-2943-4883-9f02-d8cbedb24a00 · outbound
CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Improving language understanding by gen- erative pre-training
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65c96e97-f43b-4d3d-bd3c-f7c288d743cb · outbound
CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Language models are unsu- pervised multitask learners
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ed4d3bd-0aa9-4d44-b1f5-6f872cb9e01f · outbound
CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Learning transferable visual models from natural language supervi- sion
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3674552d-2517-4313-9bdd-5f19b8022cee · outbound
CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Scale-mae: A scale-aware masked autoencoder for multiscale geospatial representation learning
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation af078688-5acc-4713-a44c-21535e3343a3 · outbound
CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Revisiting Color-Event based Tracking: A Unified Network, Dataset, and Metric
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0f732fa-39b3-469c-96f1-09333daf40be · outbound
CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Humanbench: Towards general human- centric perception with projector assisted pretraining
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 4ce8beff-01ba-4ca3-abd0-03ee8f555aa6 · outbound
CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a94de740-76e3-47cc-bd6c-d0cdec7f82f6 · outbound
CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework LLaMA: Open and Efficient Foundation Language Models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b3114b7-928e-4e90-af68-6b680b606807 · outbound
CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 114a39bf-a593-480c-8fb7-a6c1635fbd4d · outbound
CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Videomae v2: Scaling video masked autoencoders with dual masking
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 2d48d814-d7c5-4c57-ba50-76c3c3e1a92b · outbound
CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Image as a foreign language: Beit pretraining for vision and vision- language tasks
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 659f7318-309d-43d2-bc62-70506a10c14d · outbound
CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Vi- sevent: Reliable object tracking via collaboration of frame and event flows
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 6cabab66-b653-495d-b316-1007dbcc9457 · outbound
CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Object Detection using Event Camera: A MoE Heat Conduction based Detector and A New Benchmark Dataset
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d60a012c-cff1-4059-a3cb-6635e5bc8211 · outbound
CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Pre-training on High Definition X-ray Images: An Experimental Study
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c82d5b01-70e9-4563-b000-c083b4671f72 · outbound
CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Event stream-based visual object tracking: A high-resolution benchmark dataset and a novel baseline
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5d19998c-fecb-40ec-a7f8-bb8dfc73f81b · outbound
CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Structural information guided multimodal pre-training for vehicle-centric percep- tion
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 519c5d77-222a-47be-a7be-b6cb6c73b803 · outbound
CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Hardvs: Re- visiting human activity recognition with dynamic vision sen- sors
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9f5af015-e868-42e2-9939-4bd60da1cd75 · outbound
CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Multipath event-based network for low-power human action recognition
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 6619c4a3-f1c4-4597-b17f-458372397f34 · outbound
CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Event camera data pre-training
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 015ff7b0-d559-4ef7-998a-84cc6be01d65 · outbound
CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Florence: A New Foundation Model for Computer Vision
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f939e99e-2774-471e-8bab-d66d1273ef23 · outbound
CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Dino: Detr with improved denoising anchor boxes for end-to-end object de- tection
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 33f32570-e839-459e-96b9-e2096d899de8 · outbound
CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Odtrack: Online dense temporal token learning for visual tracking
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1c57396-d39c-4ebe-93e8-8329ad2606f8 · outbound
CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Image bert pre-training with online tokenizer
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation da8c1892-5d29-40d3-9215-cc601b1c1f3b · outbound
CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Ex- act: Language-guided conceptual reasoning and uncertainty estimation for event-based action recognition and more
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b4b144e9-967a-4cd5-9f00-5ef8ae27502a · outbound
CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Event-free moving object segmentation from moving ego vehicle
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f30c6a45-9e2b-49b5-b6c3-1ba9152998fd · outbound
CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Segment everything everywhere all at once
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4967cac0-fc44-4460-be2f-93c159ec9ba6 · outbound
CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework PLIP: Language-Image Pre-training for Person Representation Learning
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.