Pith. sign in

Paper Citation Record · LEDGER

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework

As of 19 August 2026, this Paper Citation Record lists 66 of 66 outbound references and 0 inbound Pith citation observations for arXiv:2504.12576.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.12576 v1

Coverage vector

measured 66 of 66 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:33:45.234194Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

66 of 66 outbound references displayed

  • verified exact0
  • verified fuzzy34
  • unresolved32
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8ba16635-3d8c-4301-b6a1-6f41a79519c6 · outbound

This paper cites Multimae: Multi-modal multi-task masked autoen- coders.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Multimae: Multi-modal multi-task masked autoen- coders

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:33:46.159804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:33:44.920575Z digest=sha256:16dc40c6cd512dbbe1f75ed55f48350373af0b39d7158361f2e723e373efbf23

Observation 5703e613-8c86-4795-86d5-1de470fba762 · outbound

This paper cites Beit: Bert pre-training of image transformers.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Beit: Bert pre-training of image transformers

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:33:46.144879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:33:44.925773Z digest=sha256:574929ebd1a8cad06bb70f3659bc5f61840bf99022f84b3d0175eb74bdf624e3

Observation df0011cb-5772-4953-90df-ea39ecaf7b04 · outbound

This paper cites Is Space-Time Attention All You Need for Video Understanding?.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Is Space-Time Attention All You Need for Video Understanding?

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T12:33:44.930487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:33:44.930487Z digest=sha256:f4d9846418123bfcf9348f5769b56a7249f9f4c94081fd578e7b5c062526cddd

Observation a40b3fa2-2778-4886-8096-0e6779e508ee · outbound

This paper cites DeepSeek LLM: Scaling Open-Source Language Models with Longtermism.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework DeepSeek LLM: Scaling Open-Source Language Models with Longtermism

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T12:33:44.935504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:33:44.935504Z digest=sha256:84f9d249706c029674fdfe4682c05f2e14fc2cd85d6b93d8e7fa2c8812cea074

Observation 9860ca14-a597-4263-b046-560ef9db0abc · outbound

This paper cites Lan- guage models are few-shot learners.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Lan- guage models are few-shot learners

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T12:33:44.940844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:33:44.940844Z digest=sha256:5c1a1d685a3ebdaae39aff60e8ece9b8c47067744802c02ae8d78db213b754c3

Observation aa93e432-08d9-44cd-b267-138a72124aeb · outbound

This paper cites End-to- end object detection with transformers.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework End-to- end object detection with transformers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T12:33:44.946536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:33:44.946536Z digest=sha256:5e9a878aecdf2eb8684c5fe574c3dc99bd8b3f46ad5ac1e81e2690d5311be76e

Observation c08d217c-1bb1-4b4a-8664-9067cb276728 · outbound

This paper cites Weakly misalignment-free adaptive feature alignment for uavs- based multimodal object detection.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Weakly misalignment-free adaptive feature alignment for uavs- based multimodal object detection

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:33:46.110267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:33:44.951657Z digest=sha256:0342fa7f61ab5582fecd1fe9ba07ccd16fa5d0512b1ff071e56b189c9d84489f

Observation dfc6f7f1-7290-4ce9-ac27-f489d5702f41 · outbound

This paper cites An empirical study of training self-supervised vision transformers.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework An empirical study of training self-supervised vision transformers

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:33:46.093955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:33:44.956835Z digest=sha256:d35ccb0e2324caca22e3ee8fb1ea2bb436612ecc06607dace832cd3752cb171e

Observation 2c7e24f8-b6e3-4255-b13a-092e665e96e1 · outbound

This paper cites Segment any event streams via 11 weighted adaptation of pivotal tokens.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Segment any event streams via 11 weighted adaptation of pivotal tokens

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:33:46.077155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:33:44.961205Z digest=sha256:c9c024b0d31e072e8e3af7e1a6fb216752ecac9c503cc9946deb5f49fc33213a

Observation 3439abeb-f364-49b7-b0f9-4d799dd7798b · outbound

This paper cites Unihcp: A unified model for human-centric perceptions.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Unihcp: A unified model for human-centric perceptions

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:33:46.061113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:33:44.965721Z digest=sha256:c62151ae45dcb8a0f40f51ae3510076c095fe49e63cae8d45d6b82c2f8b22d61

Observation 38d2fb5d-d24e-432b-b290-d20ae97bc3c3 · outbound

This paper cites Satmae: Pre-training transformers for tem- poral and multi-spectral satellite imagery.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Satmae: Pre-training transformers for tem- poral and multi-spectral satellite imagery

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:33:46.046607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:33:44.970391Z digest=sha256:642ece0b3320d28c630692f2ef32f0dc822191a9b5b3f3cf8ddd69f3595bd257

Observation 98ac94c9-fdf6-4ef3-8dc5-d253015a5fa4 · outbound

This paper cites Deepseekmoe: Towards ultimate expert spe- cialization in mixture-of-experts language models.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Deepseekmoe: Towards ultimate expert spe- cialization in mixture-of-experts language models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:33:46.030964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:33:44.974774Z digest=sha256:4e11151a86bd2626cf31b4a3869cbcc14442349e973b9a739e103ac4ae1c9f32

Observation f59cbf1d-53eb-4070-8f65-f78cd35eb765 · outbound

This paper cites Sfod: Spiking fusion object detector.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Sfod: Spiking fusion object detector

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:33:46.015958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:33:44.980632Z digest=sha256:2d63fef17b6573ca79e04df5722d45cad503248273b83a3a2e3d447a5541bc45

Observation 9980f297-1682-430c-8ef0-1241ade1e83f · outbound

This paper cites Hypergraph-based multi-view action recognition using event cameras.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Hypergraph-based multi-view action recognition using event cameras

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:33:46.000459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:33:44.985177Z digest=sha256:c68f94dcbb1fb9bccd6d5d5cfc3f39a7a312b39596e8c3c0e98abec0eecb5af7

Observation de460341-6127-4e57-8104-119f7d72db4d · outbound

This paper cites Multimodal Masked Autoencoders Learn Transferable Representations.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Multimodal Masked Autoencoders Learn Transferable Representations

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T12:33:44.989651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:33:44.989651Z digest=sha256:20cc9aee83eeade75cb77e94a6fdf4291efeff525310ddb027c0f2098ecd65a9

Observation fe194490-520d-45cd-81c3-7fae7ea3d348 · outbound

This paper cites The Llama 3 Herd of Models.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework The Llama 3 Herd of Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T12:33:44.995396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:33:44.995396Z digest=sha256:9e26f63f5cd4fe659483171947b400a7a5b95135d62487421fd487c618251e33

Observation 60b8eedb-36d3-441c-97e2-7c3a88c27f93 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T12:33:45.000275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:33:45.000275Z digest=sha256:60b5d1b2a1a98bbcbf613704ca1bfdb965e88f8607f9057e5a665ebdfb1451d6

Observation 4e0369dd-4a64-4df6-b92c-11afd59e42b5 · outbound

This paper cites Momentum contrast for unsupervised visual rep- resentation learning.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Momentum contrast for unsupervised visual rep- resentation learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T12:33:45.006252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:33:45.006252Z digest=sha256:7eb5aec2b6f41527a3fd934a4ef3838e937adca1d761368e88594b7f125cbb2e

Observation 8bd98e9c-45d6-4423-a31f-055b5ce28c91 · outbound

This paper cites Masked autoencoders are scalable vision learners.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Masked autoencoders are scalable vision learners

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:33:45.976142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:33:45.011021Z digest=sha256:550aa76aec7cb64bc75f813235d8bc0be7fd4d1c02e6173accae4789a9645a57

Observation 034413c5-16c7-4253-b51d-d1856d61227a · outbound

This paper cites Data-efficient Event Camera Pre-training via Disentangled Masked Modeling.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Data-efficient Event Camera Pre-training via Disentangled Masked Modeling

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T12:33:45.016104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:33:45.016104Z digest=sha256:f03cd388840ca23e30585f891fc33897a0001fc4fe2d18f3587d95a989f03af6

Observation ed5fa96b-ccab-45d1-8890-cdbb3cf69b19 · outbound

This paper cites N-imagenet: Towards robust, fine-grained object recognition with event cameras.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework N-imagenet: Towards robust, fine-grained object recognition with event cameras

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:33:45.961258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:33:45.021864Z digest=sha256:e14ef3c26b6d3bdcf634fd1045ac065dff29d4ca823427a4b13690316d747c24

Observation fefa5436-bb6e-4ac9-a0fb-110eb4d3a2f5 · outbound

This paper cites Spiking-yolo: spiking neural network for energy- efficient object detection.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Spiking-yolo: spiking neural network for energy- efficient object detection

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T12:33:45.026355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:33:45.026355Z digest=sha256:770cbb15b2cb082923d886ff05baa88edf6883baee568c9506c14cc908db8ba6

Observation 800ecb3d-ae9e-4fe9-a87e-f5854ef98554 · outbound

This paper cites Segment any- thing.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Segment any- thing

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T12:33:45.030964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:33:45.030964Z digest=sha256:b87d27832acca772819c77facfa5f49ea3a8f5b2c43950cd7c325d47743f2f45

Observation c2ff80af-8f69-4d7e-bc4b-cbc8b470212a · outbound

This paper cites Masked event modeling: Self-supervised pretraining for event cameras.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Masked event modeling: Self-supervised pretraining for event cameras

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T12:33:45.035526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:33:45.035526Z digest=sha256:f8685d1f36b5d80dcb8834436b140ba5fc55ba88a79a78ba82e9184fb3bab08a

Observation de3313fb-2b40-487a-a9d7-94727cda921c · outbound

This paper cites Openess: Event-based semantic scene understanding with open vocabularies.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Openess: Event-based semantic scene understanding with open vocabularies

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:33:45.918387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:33:45.040195Z digest=sha256:7865d6fe9ad1e13cd06b9860bbc8b1cefe159a3dfe0fa5fb674250e4f93dadd6

Observation bd9ab563-4775-4dfa-ba6b-41e4b9ad8aab · outbound

This paper cites Mulfs-cap: Multimodal fusion- supervised cross-modality alignment perception for unreg- istered infrared-visible image fusion.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Mulfs-cap: Multimodal fusion- supervised cross-modality alignment perception for unreg- istered infrared-visible image fusion

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T12:33:45.044826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:33:45.044826Z digest=sha256:6353d3b000b76e2df0f6c7f6c189ed85b0b71ac09ce5fd86f3c5335617ae2078

Observation e3869773-b7c9-4d8a-987f-c2c7e5b13175 · outbound

This paper cites Coupled mamba: Enhanced multimodal fusion with coupled state space model.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Coupled mamba: Enhanced multimodal fusion with coupled state space model

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:33:45.893764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:33:45.049239Z digest=sha256:e98df650eccc9629b58851c8d7bdc32eea252f05075be9f0816b78b9476abbc3

Observation 97e04996-459f-4629-879e-6143eabe9f48 · outbound

This paper cites Exploring plain vision transformer backbones for object de- tection.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Exploring plain vision transformer backbones for object de- tection

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:33:45.878991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:33:45.054058Z digest=sha256:04e19469c48fc55929daeb25a2a0bd1bfc42de4cf00dc18c8b3c45436378ee86

Observation 55f2a1b4-1763-4618-aab0-8d7753b2d2f9 · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T12:33:45.058883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:33:45.058883Z digest=sha256:373042b35c83c78be7c0a383c013732acce658dd2d5fbf6b4fd6241de8dad593

Observation 0309dcdf-261d-4b8a-b9b7-223f75749363 · outbound

This paper cites Visual instruction tuning.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Visual instruction tuning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T12:33:45.063794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:33:45.063794Z digest=sha256:aa2f150f3b073faf80194fe33a2d76893f18908dad1cc428caf6590a8e98a1b0

Observation f68ca61a-21d8-465b-9239-dd19f5b72050 · outbound

This paper cites Pixmim: Rethinking pixel reconstruction in masked image modeling.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Pixmim: Rethinking pixel reconstruction in masked image modeling

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:33:45.854175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:33:45.068364Z digest=sha256:53889ab4e78628cfddc20dc4ee04c97e2ed016828b8bfcd2b3c54721f4c7b8a7

Observation 02ce43ba-edb4-4727-be67-8050ee285cda · outbound

This paper cites Decoupled weight de- cay regularization.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Decoupled weight de- cay regularization

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:33:45.840121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:33:45.073164Z digest=sha256:d78559a330b09c81b33652335e350de0e63a61c37b651e2b6133dba38334daeb

Observation cfa01b67-b41b-4160-90f4-b4e40bbcf0a2 · outbound

This paper cites Integer-valued training and spike-driven inference spiking neural network for high-performance and energy-efficient object detection.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Integer-valued training and spike-driven inference spiking neural network for high-performance and energy-efficient object detection

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:33:45.824766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:33:45.077761Z digest=sha256:17ef03b4b174347d68f510feee1a79049e0dc8e1ff36a0a66d82de9a2b67fce9

Observation c3370488-1ea8-492f-9a40-ae424cf4d65b · outbound

This paper cites Event-based moving object 12 detection and tracking.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Event-based moving object 12 detection and tracking

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:33:45.809465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:33:45.082367Z digest=sha256:40115c3d83838aca213d6a97f0e3a86ca0ad1e958fb4e0745c97ab2d7f566bac

Observation e3dc42bd-b708-467e-86eb-359e73d63f8a · outbound

This paper cites Rethinking transformers pre-training for multi- spectral satellite imagery.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Rethinking transformers pre-training for multi- spectral satellite imagery

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:33:45.794835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:33:45.087166Z digest=sha256:8eff4a1a1a1be4c0e8a2fa4818d94bd946dd4ed2cca9749bfa9f120426a3ee8d

Observation 5b8fe7b3-96cb-40c8-9a09-78a03e83ce5e · outbound

This paper cites Dinov2: Learning robust visual features without supervision.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Dinov2: Learning robust visual features without supervision

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T12:33:45.092015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:33:45.092015Z digest=sha256:c96eaf7a010a589694eacb92545a2b320678cafd4390c1c2f0f1a85d809479ae

Observation 20971ba1-d35b-46ab-99ed-729cfdd53ed2 · outbound

This paper cites Pytorch: An im- perative style, high-performance deep learning library.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Pytorch: An im- perative style, high-performance deep learning library

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:33:45.770086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:33:45.097788Z digest=sha256:d829c24da9ae8aa7f05cef36b00224cb29da2836791e08ef6054568a9f34e304

Observation f8113807-ab7e-474d-9622-52d29421a468 · outbound

This paper cites BEiT v2: Masked Image Modeling with Vector-Quantized Visual Tokenizers.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework BEiT v2: Masked Image Modeling with Vector-Quantized Visual Tokenizers

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T12:33:45.102760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:33:45.102760Z digest=sha256:67dd547be855490040aac5389a4f48fd9155d5f21b1094d7bd834f789f4f1678

Observation 222c146c-c03c-48cb-a699-0c35812d253f · outbound

This paper cites Detectors: Detecting objects with recursive feature pyramid and switch- able atrous convolution.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Detectors: Detecting objects with recursive feature pyramid and switch- able atrous convolution

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T12:33:45.107543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:33:45.107543Z digest=sha256:71a2c2488426fa0de52be0bc3401e0c93187600fec2b0d0d610d7838546c877e

Observation 0105fd7e-2943-4883-9f02-d8cbedb24a00 · outbound

This paper cites Improving language understanding by gen- erative pre-training.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Improving language understanding by gen- erative pre-training

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T12:33:45.112156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:33:45.112156Z digest=sha256:e5bb30d2a441b5341b77b8d3934fa8ca35ef1f8b7f721a5d732cfbc10df9cf58

Observation 65c96e97-f43b-4d3d-bd3c-f7c288d743cb · outbound

This paper cites Language models are unsu- pervised multitask learners.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Language models are unsu- pervised multitask learners

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T12:33:45.116617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:33:45.116617Z digest=sha256:2fb3e8b0019700b5a85d4c308547f08023f59a2241e096001d6e037c16da6af4

Observation 1ed4d3bd-0aa9-4d44-b1f5-6f872cb9e01f · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Learning transferable visual models from natural language supervi- sion

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T12:33:45.121084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:33:45.121084Z digest=sha256:277fe82712747508999cc7ee489f049fdf551c4a4419f9c2dde04f756395d79c

Observation 3674552d-2517-4313-9bdd-5f19b8022cee · outbound

This paper cites Scale-mae: A scale-aware masked autoencoder for multiscale geospatial representation learning.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Scale-mae: A scale-aware masked autoencoder for multiscale geospatial representation learning

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:33:45.717949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:33:45.125427Z digest=sha256:83d5edeefbef5d244a33078a168a11fa8156003c500ab018ab1db0c06b8e2696

Observation af078688-5acc-4713-a44c-21535e3343a3 · outbound

This paper cites Revisiting Color-Event based Tracking: A Unified Network, Dataset, and Metric.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Revisiting Color-Event based Tracking: A Unified Network, Dataset, and Metric

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T12:33:45.130092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:33:45.130092Z digest=sha256:cfb4a75c64b3184e8e98cc5f6a08c818c888bc3c5fff33380b8b1c44eb62c928

Observation e0f732fa-39b3-469c-96f1-09333daf40be · outbound

This paper cites Humanbench: Towards general human- centric perception with projector assisted pretraining.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Humanbench: Towards general human- centric perception with projector assisted pretraining

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:33:45.703106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:33:45.135573Z digest=sha256:acc10f5a4304823d097f5fe976177021a92ba91ce836c6597abeaf6bdc7ebedb

Observation 4ce8beff-01ba-4ca3-abd0-03ee8f555aa6 · outbound

This paper cites VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-16T12:33:45.140311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:33:45.140311Z digest=sha256:a86c635faf42ca029bdc60d8ef341cc7f2f34216c23fe2d1b30db5457e6a0c90

Observation a94de740-76e3-47cc-bd6c-d0cdec7f82f6 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework LLaMA: Open and Efficient Foundation Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-16T12:33:45.145257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:33:45.145257Z digest=sha256:d18257bd259633e75f0661fdd4b1147e928c097d1ad459f967e36f31c8a19856

Observation 0b3114b7-928e-4e90-af68-6b680b606807 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-16T12:33:45.149913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:33:45.149913Z digest=sha256:e65f417aa71d96df786ebea81403c3e369cbaa93c283c4c3719836531eb7c994

Observation 114a39bf-a593-480c-8fb7-a6c1635fbd4d · outbound

This paper cites Videomae v2: Scaling video masked autoencoders with dual masking.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Videomae v2: Scaling video masked autoencoders with dual masking

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:33:45.688661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:33:45.154973Z digest=sha256:5d5cbccab9c127c1000381741f9093b6ce0599790fbabb74c239f514d4b00505

Observation 2d48d814-d7c5-4c57-ba50-76c3c3e1a92b · outbound

This paper cites Image as a foreign language: Beit pretraining for vision and vision- language tasks.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Image as a foreign language: Beit pretraining for vision and vision- language tasks

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:33:45.673472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:33:45.159467Z digest=sha256:9ca9a94fce0f1666aba739e2a4f0ee5608adc2e0040769f2861ab5d99fed95c1

Observation 659f7318-309d-43d2-bc62-70506a10c14d · outbound

This paper cites Vi- sevent: Reliable object tracking via collaboration of frame and event flows.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Vi- sevent: Reliable object tracking via collaboration of frame and event flows

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:33:45.656566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:33:45.164092Z digest=sha256:8aaa9f264c6134e0470935b28aca708f85444f6bfec64294e88d9d56755a996a

Observation 6cabab66-b653-495d-b316-1007dbcc9457 · outbound

This paper cites Object Detection using Event Camera: A MoE Heat Conduction based Detector and A New Benchmark Dataset.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Object Detection using Event Camera: A MoE Heat Conduction based Detector and A New Benchmark Dataset

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-16T12:33:45.168675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:33:45.168675Z digest=sha256:17dac7cf520c7e1401b35986433d946cbc1714b3d9c84eebab094951c24fd295

Observation d60a012c-cff1-4059-a3cb-6635e5bc8211 · outbound

This paper cites Pre-training on High Definition X-ray Images: An Experimental Study.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Pre-training on High Definition X-ray Images: An Experimental Study

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-16T12:33:45.173911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:33:45.173911Z digest=sha256:223672bb7944eebc5c8966ef44ec6ce45a302ae2fdc3606778249d8ce709e76e

Observation c82d5b01-70e9-4563-b000-c083b4671f72 · outbound

This paper cites Event stream-based visual object tracking: A high-resolution benchmark dataset and a novel baseline.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Event stream-based visual object tracking: A high-resolution benchmark dataset and a novel baseline

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:33:45.641198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:33:45.178551Z digest=sha256:a4dad2256a56deaba1df3f36f617bc0710b21069f33290ee55cab74f34c89c0e

Observation 5d19998c-fecb-40ec-a7f8-bb8dfc73f81b · outbound

This paper cites Structural information guided multimodal pre-training for vehicle-centric percep- tion.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Structural information guided multimodal pre-training for vehicle-centric percep- tion

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:33:45.626716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:33:45.183525Z digest=sha256:2d11732c7a7b8036a25256ed77e14e3b2d6006badc8a64f0d78e05c9b1105de2

Observation 519c5d77-222a-47be-a7be-b6cb6c73b803 · outbound

This paper cites Hardvs: Re- visiting human activity recognition with dynamic vision sen- sors.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Hardvs: Re- visiting human activity recognition with dynamic vision sen- sors

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:33:45.611488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:33:45.188277Z digest=sha256:9030bfff120df095330ad174a84e7d1080367d8bf76c84fc063007ae75f8d79f

Observation 9f5af015-e868-42e2-9939-4bd60da1cd75 · outbound

This paper cites Multipath event-based network for low-power human action recognition.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Multipath event-based network for low-power human action recognition

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:33:45.596237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:33:45.192890Z digest=sha256:8e938964dd9fd65680d74ea98f801c413d30d2d732d085c616d450a89300104b

Observation 6619c4a3-f1c4-4597-b17f-458372397f34 · outbound

This paper cites Event camera data pre-training.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Event camera data pre-training

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-16T12:33:45.197522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:33:45.197522Z digest=sha256:dbf2151ea6324c4c1d81826e40d0ce4351cea3d5d2c0b0c2e7036763f30d752c

Observation 015ff7b0-d559-4ef7-998a-84cc6be01d65 · outbound

This paper cites Florence: A New Foundation Model for Computer Vision.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Florence: A New Foundation Model for Computer Vision

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-16T12:33:45.202194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:33:45.202194Z digest=sha256:df35bb6f2e3d7d2a8d24138ec6ca9cdcecba313463054ff1de2daadfd5bfbc7e

Observation f939e99e-2774-471e-8bab-d66d1273ef23 · outbound

This paper cites Dino: Detr with improved denoising anchor boxes for end-to-end object de- tection.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Dino: Detr with improved denoising anchor boxes for end-to-end object de- tection

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:33:45.571713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:33:45.207077Z digest=sha256:017aeb577000fad7edd76a94c42aed8f77422de0b76560a6dae6e9536fe2d087

Observation 33f32570-e839-459e-96b9-e2096d899de8 · outbound

This paper cites Odtrack: Online dense temporal token learning for visual tracking.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Odtrack: Online dense temporal token learning for visual tracking

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-16T12:33:45.211899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:33:45.211899Z digest=sha256:2e699ccf06f1481d76f67bdd0dd4de5416e403b3431d1cfa61a8cf9f6a668559

Observation a1c57396-d39c-4ebe-93e8-8329ad2606f8 · outbound

This paper cites Image bert pre-training with online tokenizer.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Image bert pre-training with online tokenizer

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:33:45.547965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:33:45.216184Z digest=sha256:04cea1287b18b6482807458ea2a4b04f8ad05d6a80473326879c80e4b0f116ad

Observation da8c1892-5d29-40d3-9215-cc601b1c1f3b · outbound

This paper cites Ex- act: Language-guided conceptual reasoning and uncertainty estimation for event-based action recognition and more.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Ex- act: Language-guided conceptual reasoning and uncertainty estimation for event-based action recognition and more

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:33:45.533474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:33:45.220773Z digest=sha256:067d45f4f1f71d8ca812bc6b2c6e0c095a1dfe6dafc22b9388574d2feb2f9bee

Observation b4b144e9-967a-4cd5-9f00-5ef8ae27502a · outbound

This paper cites Event-free moving object segmentation from moving ego vehicle.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Event-free moving object segmentation from moving ego vehicle

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:33:45.518562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:33:45.225157Z digest=sha256:2f0f2c6ae69c844015c29a5848d3dfaf9cb28d1e098bf8b5923444dd26d5bf14

Observation f30c6a45-9e2b-49b5-b6c3-1ba9152998fd · outbound

This paper cites Segment everything everywhere all at once.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Segment everything everywhere all at once

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-16T12:33:45.229598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:33:45.229598Z digest=sha256:8d4a0af9830f2b8c40d714055f90017e38938a5236f83a134364e0220d624781

Observation 4967cac0-fc44-4460-be2f-93c159ec9ba6 · outbound

This paper cites PLIP: Language-Image Pre-training for Person Representation Learning.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework PLIP: Language-Image Pre-training for Person Representation Learning

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-16T12:33:45.234194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:33:45.234194Z digest=sha256:5f668241d695882dae6b0bd276ff5bcaab405dac30fd1ef234c67612f808c4b3

Pith citing papers

No inbound Pith citation observations are available.