Pith. sign in

Paper Citation Record · LEDGER

EventGPT: Event Stream Understanding with Multimodal Large Language Models

As of 12 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 3 inbound Pith citation observations for arXiv:2412.00832.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.00832 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T05:01:12.829064Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:35:01.366536Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T20:18:57.044501Z

Reference resolution

51 of 51 outbound references displayed

  • verified exact2
  • verified fuzzy26
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f27a5cf4-8386-4842-b18b-319ef0b3c0ea · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:01:15.376332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.146077Z digest=sha256:4f131dfaef921abec252cfabc35eefe12e8f50d3af8ca7d30c5be88a6f4d3aaa

Observation cda93b88-d0ba-4f75-8fc4-b88416fabf8d · outbound

This paper cites The (r) evolution of multi- modal large language models: A survey.

EventGPT: Event Stream Understanding with Multimodal Large Language Models The (r) evolution of multi- modal large language models: A survey

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T05:01:12.191004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:01:12.191004Z digest=sha256:9fbd54829a234143e44c2cc3db5b861de2193e55a2561fc867b5988e64057cd5

Observation 877ca8af-dd48-45b6-a512-3ebb13ac7acc · outbound

This paper cites First Place Solution to the CVPR'2023 AQTC Challenge: A Function-Interaction Centric Approach with Spatiotemporal Visual-Language Alignment.

EventGPT: Event Stream Understanding with Multimodal Large Language Models First Place Solution to the CVPR'2023 AQTC Challenge: A Function-Interaction Centric Approach with Spatiotemporal Visual-Language Alignment

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-08-12T05:01:14.030175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.213248Z digest=sha256:1cc448be53ca52b7c28f729b8b3121aa7bf1ae6ccbb67a2ad90aeea080e4ae1a

Observation d4b315fa-d1df-4ebc-a4fb-60c6b9f3e872 · outbound

This paper cites Dress: Instructing large vision-language models to align and interact with humans via natural lan- guage feedback.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Dress: Instructing large vision-language models to align and interact with humans via natural lan- guage feedback

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:01:15.345875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.221664Z digest=sha256:19e70758a66a9dd6e09f214ec03af3afecc02f4537205a3ad1793e33c7d1e471

Observation 3f724c32-951e-477f-9df4-0600bbbe027b · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:01:15.311023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.233969Z digest=sha256:6f83ae3dce9a471abed14971bee7cb95de9273be8c508000b7a729203e9becab

Observation 2bc6fd32-242c-49f5-a715-bb0e4893d8a9 · outbound

This paper cites Reproducible scal- ing laws for contrastive language-image learning.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Reproducible scal- ing laws for contrastive language-image learning

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:01:15.275105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.246718Z digest=sha256:f1dcd8e552252fd4e54c1255dc7fbf89d7bf871341f95324c35615dd7bf3ed41

Observation 05333c3c-b55a-4dcb-91f5-368cdc6e3118 · outbound

This paper cites Event-based vision: A survey.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Event-based vision: A survey

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:01:15.224190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.264858Z digest=sha256:4b6c2c8024f148b0997527b2adc3f4cb038597df9b6a56cd892e8976710f59fe

Observation 3a8da49b-83ad-49ee-bb36-20e4c81e0fea · outbound

This paper cites Low-latency auto- motive vision with event cameras.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Low-latency auto- motive vision with event cameras

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T05:01:12.272493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:01:12.272493Z digest=sha256:0e62e5373638404b4b70442bd35b444cd82db6ec93b3769872b713ba2a9934da

Observation dddee9e6-87e8-4149-bb7e-7306ef5af691 · outbound

This paper cites Eklt: Asynchronous photometric feature tracking using events and frames.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Eklt: Asynchronous photometric feature tracking using events and frames

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:01:15.147226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.299020Z digest=sha256:199d65e326f2831e6def0fbefbbb0f37d4674b41841dd56e1622ed474eea9ad3

Observation 6c17c68a-1f27-4629-90f8-718b53456217 · outbound

This paper cites Recurrent vision transformers for object detection with event cameras.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Recurrent vision transformers for object detection with event cameras

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:01:15.094373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.307967Z digest=sha256:af278386455c4120511e248461bb9ddc5ada78761baf6714974d2e107cee1457

Observation fe9d2bdf-22a2-4961-a51e-a9d697ac4c09 · outbound

This paper cites Dsec: A stereo event camera dataset for driving scenarios.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Dsec: A stereo event camera dataset for driving scenarios

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:01:15.058056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.319807Z digest=sha256:1ca14840189de13c95a6c6154c7fdb69039b605f34ce683a91fcaabab384bf3c

Observation d9220487-9589-44c5-9e3a-af4cfab38a60 · outbound

This paper cites Event-based Simultaneous Localization and Mapping: A Comprehensive Survey.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Event-based Simultaneous Localization and Mapping: A Comprehensive Survey

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T05:01:12.325692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:01:12.325692Z digest=sha256:233a8eddc5d4389db391c86505d1e8a5e3ef23d395a5a6bf9bf765b34301adda

Observation d87a4c97-c318-4cce-b567-56c9cff7cff2 · outbound

This paper cites Img-Diff: Contrastive Data Synthesis for Multimodal Large Language Models.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Img-Diff: Contrastive Data Synthesis for Multimodal Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T05:01:12.370094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:01:12.370094Z digest=sha256:42e978db1689f84656e85b443d8c2144bd7419a2aba10e6df38d822d23170ea7

Observation 8696cb31-8981-4de7-8ad8-0e9c6db56ac8 · outbound

This paper cites Real-time 3d reconstruction and 6-dof tracking with an event camera.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Real-time 3d reconstruction and 6-dof tracking with an event camera

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:01:15.013695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.380301Z digest=sha256:535efb62c60e58a591a1a10d3c294dc39a4e1129e9c80c457a9182e5fdfa9952

Observation 9a76f3d8-4de0-4523-a974-9f71ad70a25b · outbound

This paper cites N-imagenet: Towards robust, fine-grained object recognition with event cameras.

EventGPT: Event Stream Understanding with Multimodal Large Language Models N-imagenet: Towards robust, fine-grained object recognition with event cameras

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:01:14.969940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.389515Z digest=sha256:ac9aa9db3b8639127244395570796eefe07a079b69e009720ead075c9d0920fc

Observation 164a4524-4963-4a2c-ac7e-28b526a83323 · outbound

This paper cites Sodformer: Streaming object detection with transformer using events and frames.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Sodformer: Streaming object detection with transformer using events and frames

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:01:14.921551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.407322Z digest=sha256:ff6847c9b25f4eb3d5cb29d1cedd8017986b14625ff49d6816bff5a6552a940c

Observation 2d517bce-5fa1-43e7-9785-691d9a70c3d4 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T05:01:12.419221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:01:12.419221Z digest=sha256:4c2a3c8ed8954db158a25e3270f82dc495f6372341ddee388389d148d7fef199

Observation 33a74984-41e2-4f56-931e-d1f364193f75 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:01:14.861678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.432463Z digest=sha256:8135d3ad73dddf543745a6db04621108ea280eb52ef8b317539089bfb1c3c4cc

Observation 8f8b143d-538b-4088-bcd4-f14b415b7773 · outbound

This paper cites Vila: On pre-training for visual language models.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Vila: On pre-training for visual language models

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:01:14.815717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.451299Z digest=sha256:92b4be568400100f495ac51da30fb304f96640c2124972c6d3a91fdaa72e2ac6

Observation f159e167-0cad-49fc-9fc7-b59820cb659b · outbound

This paper cites Improved baselines with visual instruction tuning.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Improved baselines with visual instruction tuning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T05:01:12.463039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:01:12.463039Z digest=sha256:669e2cc6036c2fbdd949c53521f0cd82f4cde2010eced588fff979d662213229

Observation de61ff03-66d6-4813-9971-792a65dd3047 · outbound

This paper cites Visual instruction tuning.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Visual instruction tuning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:01:14.748627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.471738Z digest=sha256:73b8fe9ee81778e568613f7746e93419f8a38fc7e8477c7ec7786dd78c94c412

Observation ae5221da-37bc-420b-875f-228503fba171 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T05:01:12.482716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:01:12.482716Z digest=sha256:1a93d9d3ce75eba2b08aa91d8e3b4aae74bd8c7916cfc2c64bb53e9cf36f431d

Observation 79b9cb82-abac-45af-85a3-abd0e06f9dd0 · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

EventGPT: Event Stream Understanding with Multimodal Large Language Models DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T05:01:12.494304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:01:12.494304Z digest=sha256:48daa8fbc7e6abcd598a6d7f59edd689a93c5318952037c0f2362e8f501d08dd

Observation 33e8d8d2-9aad-4e1b-8346-d4062f9c4626 · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T05:01:12.504000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:01:12.504000Z digest=sha256:0a7a49e160eb1a62d77ff6895f756d353ad6024f9dd6adf8ad7db2bae217fa18

Observation 1e1285f9-2f14-4f3a-9e42-ac08ff1bbb71 · outbound

This paper cites Data-driven feature tracking for event cameras.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Data-driven feature tracking for event cameras

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:01:14.712020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.516781Z digest=sha256:157e0d678615f29eb3296069c5d91e79481f7a65a7e762f3613d0fa9bfd7ce95

Observation 36c26336-66e5-4419-a6ec-152428026790 · outbound

This paper cites Esl: Event-based structured light.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Esl: Event-based structured light

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:01:14.660582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.526710Z digest=sha256:1c0c8799b9fa8f13a0209496eccad72093eada0d91143f782c1fb1affbe3a934

Observation 159208c6-faaf-4660-80bb-fd0348074fdb · outbound

This paper cites Cross-modal Attention Congruence Regularization for Vision-Language Relation Alignment.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Cross-modal Attention Congruence Regularization for Vision-Language Relation Alignment

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T05:01:12.536142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:01:12.536142Z digest=sha256:2ef259978b1733b8b754a22328c8b1ecd6b28a6d8da82a75968d45712a591294

Observation f13b0614-776c-4379-9b42-f765ae47f85d · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Learn- ing transferable visual models from natural language super- vision

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T05:01:12.546821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:01:12.546821Z digest=sha256:a5e53c586496ddbd5675e6e3293770c47e82080743068d44a8afda21c4f6554d

Observation 968532f8-6901-4f09-a2a3-5980694820d2 · outbound

This paper cites Emvs: Event-based multi-view stereo—3d 9 reconstruction with an event camera in real-time.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Emvs: Event-based multi-view stereo—3d 9 reconstruction with an event camera in real-time

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:01:14.580424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.559931Z digest=sha256:5a8b97b4818a3987631c706848415324c0fab6accad76c523ab8fbb7219b441f

Observation 53c8eda1-4a95-47ef-957c-7d272039650d · outbound

This paper cites Events-to-video: Bringing modern computer vision to event cameras.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Events-to-video: Bringing modern computer vision to event cameras

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:01:14.535641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.576078Z digest=sha256:158b3dbc5d674001c252f10d5f0e83e8250ed0b3fa055bc05d696709829c2b7d

Observation 6528a202-e7dc-44f5-9e4f-6df6a237dd1f · outbound

This paper cites Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T05:01:12.585257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:01:12.585257Z digest=sha256:1f4ab7bafd7af4f9835658353333cf3fe8b859d8de32bc58b9a075a2a68f986f

Observation bf38fdca-0f41-4b83-94e2-aa4adf5b7bc6 · outbound

This paper cites Aligning and prompting everything all at once for univer- sal visual perception.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Aligning and prompting everything all at once for univer- sal visual perception

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:01:14.504155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.592175Z digest=sha256:dd6d2958670b4be803f0a71349a1c8c498f43b249caad239daefbbe55591ab6f

Observation c169dbb5-d2d0-4f4b-b4bc-83397472b7aa · outbound

This paper cites BlinkTrack: Feature Tracking over 80 FPS via Events and Images.

EventGPT: Event Stream Understanding with Multimodal Large Language Models BlinkTrack: Feature Tracking over 80 FPS via Events and Images

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-08-12T05:01:13.614829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.601049Z digest=sha256:4d0420dfbb578999e46002ad23c8e8f1f1b0fd4f68668ae2adcfac262434cd7b

Observation 4ef5e9f7-7c74-4902-9b55-dca0f2c8a767 · outbound

This paper cites Flava: A foundational language and vision alignment model.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Flava: A foundational language and vision alignment model

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:01:14.470618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.615369Z digest=sha256:5c510278d5706e0f2ebb15ecef5aa1d8aedf30988313316032a98e63745b9fb4

Observation 7c054af0-cdea-4ef8-ad61-d2753fc46196 · outbound

This paper cites Cloud-device collaborative learning for multimodal large language models.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Cloud-device collaborative learning for multimodal large language models

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:01:14.410556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.629470Z digest=sha256:64ca5166c26e9daba40a15a167556da2b3f83700b280110c8b4824515cb06082

Observation 53e4539d-0540-4ab5-8523-50f11d46a391 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T05:01:12.644835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:01:12.644835Z digest=sha256:374601e923a8b63e6750f75f9415c37cf4f4678369cf568b4bff8c5ecfecf2cc

Observation a051b0d7-b2bb-4400-8e40-870824ed9ee4 · outbound

This paper cites NExT-GPT: Any-to-Any Multimodal LLM.

EventGPT: Event Stream Understanding with Multimodal Large Language Models NExT-GPT: Any-to-Any Multimodal LLM

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T05:01:12.652501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:01:12.652501Z digest=sha256:b5a890f3f660cb5fe0eac76f87cecbe0f0db075d5ec38e556d8fa814f1fe5b2b

Observation f09726e2-4787-4789-964c-78d9265993ee · outbound

This paper cites EventCLIP: Adapting CLIP for Event-based Object Recognition.

EventGPT: Event Stream Understanding with Multimodal Large Language Models EventCLIP: Adapting CLIP for Event-based Object Recognition

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T05:01:12.664201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:01:12.664201Z digest=sha256:b55885433ef3d009354148c2bd39b995695bf3bf408c1f716a5d573e7d1179d8

Observation 177c1df0-cd5d-4dfd-ac9e-bfeacf6a6b63 · outbound

This paper cites Leod: Label-efficient object detection for event cameras.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Leod: Label-efficient object detection for event cameras

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:01:14.377574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.675218Z digest=sha256:c24857627a09db68b261ce652ec350267185ce44508df7a67e789463fa1f86b4

Observation e795f0ce-43e0-4479-91b5-ba390758a38f · outbound

This paper cites xgen-mm (blip-3): A family of open large multimodal models.

EventGPT: Event Stream Understanding with Multimodal Large Language Models xgen-mm (blip-3): A family of open large multimodal models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T05:01:12.682307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:01:12.682307Z digest=sha256:fec6ec41f35f459066d456c4ab08378fc67b006e40a6ec05604d00ba50cac4ff

Observation ce77d6d9-3793-4b57-a355-011819caa0ce · outbound

This paper cites A Survey on Multimodal Large Language Models.

EventGPT: Event Stream Understanding with Multimodal Large Language Models A Survey on Multimodal Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T05:01:12.701293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:01:12.701293Z digest=sha256:ddc199332d56a79ed9220e1200fb906419b0ba89b7ccc994fea1439e6371b8e7

Observation 1f35d5f7-59f1-43ca-8419-a13b448ebc7d · outbound

This paper cites Eventps: Real-time photometric stereo using an event camera.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Eventps: Real-time photometric stereo using an event camera

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:01:14.326105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.711889Z digest=sha256:04501031061dc8f71e0641adac095ad5d2490cfb719a5638e528ab9ee3e1d010

Observation e8e3c9be-af2f-4cef-8bed-155f1dfaf6bb · outbound

This paper cites AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling.

EventGPT: Event Stream Understanding with Multimodal Large Language Models AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T05:01:12.725243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:01:12.725243Z digest=sha256:91f48bf41ba95680309419c3f79e17d48489b2d200758b291ab466b472e9abc0

Observation 173d49c1-b66d-4209-a7f9-6e1e7aa1f175 · outbound

This paper cites MM-LLMs: Recent Advances in MultiModal Large Language Models.

EventGPT: Event Stream Understanding with Multimodal Large Language Models MM-LLMs: Recent Advances in MultiModal Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T05:01:12.746090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:01:12.746090Z digest=sha256:7a153a1ce8f1011e94c091075a128f2b12e8bbb2303cf5bca954c7fa401e601a

Observation b418c2d4-6d9e-4efc-829c-9033f9a779fa · outbound

This paper cites Spiking transform- ers for event-based single object tracking.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Spiking transform- ers for event-based single object tracking

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:01:14.289514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.767968Z digest=sha256:c2b21b96f7485ec7f14d22ff009da09fba6865cf6590d87a748613551c9c093e

Observation 93adb63c-5752-49c7-90b4-9f28c8c7bfb3 · outbound

This paper cites OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding.

EventGPT: Event Stream Understanding with Multimodal Large Language Models OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T05:01:12.778547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:01:12.778547Z digest=sha256:5ecde7761d122e05d5f55656a3d8a214f6a08e48db61d8b72b70ef5ae922095d

Observation e8fd68e8-126b-481c-884d-b0955e5130d9 · outbound

This paper cites Deep Learning for Event-based Vision: A Comprehensive Survey and Benchmarks.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Deep Learning for Event-based Vision: A Comprehensive Survey and Benchmarks

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T05:01:12.790227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:01:12.790227Z digest=sha256:b32419e3fe770120e500f9afc54d8e992a99283364e9563d6bdb236f14d0f1a6

Observation b86f87bf-1754-44f4-86c0-e7b6f28dd19f · outbound

This paper cites E- clip: Towards label-efficient event-based open-world under- standing by clip.

EventGPT: Event Stream Understanding with Multimodal Large Language Models E- clip: Towards label-efficient event-based open-world under- standing by clip

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:01:14.250690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.799667Z digest=sha256:e23b77a82014af062748c1f281e802f706298cc360926d04bdc378961d30dfb6

Observation 74cf5c2a-b2db-4702-b0a2-d23a41979662 · outbound

This paper cites Ex- act: Language-guided conceptual reasoning and uncertainty estimation for event-based action recognition and more.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Ex- act: Language-guided conceptual reasoning and uncertainty estimation for event-based action recognition and more

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:01:14.224985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.811621Z digest=sha256:6e6cf921e15706fd7925c2adb402d25a9932e407fe5798979ad86713e37c0687

Observation 1e43ee78-5e59-493c-99cb-f23f3c94085b · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

EventGPT: Event Stream Understanding with Multimodal Large Language Models MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T05:01:12.822819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:01:12.822819Z digest=sha256:445c1861e63be887fd0f0c18aed9c9f0ce476a670f907ec19e3e6682fb25a59e

Observation 517844bd-825e-4b7e-bd36-135d6f994670 · outbound

This paper cites VL-GPT: A Generative Pre-trained Transformer for Vision and Language Understanding and Generation.

EventGPT: Event Stream Understanding with Multimodal Large Language Models VL-GPT: A Generative Pre-trained Transformer for Vision and Language Understanding and Generation

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T05:01:12.829064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:01:12.829064Z digest=sha256:ac938ce5e0782d7268a268af4b8fd1b2076f6f4154e712249328bff02a9d29ee

Pith citing papers

Observation e341bfd4-a294-4028-a135-048efe39d61c · inbound

Event-Priori-Based Vision-Language Model for Efficient Visual Understanding cites this paper.

Event-Priori-Based Vision-Language Model for Efficient Visual Understanding EventGPT: Event Stream Understanding with Multimodal Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:01.366536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:01.366536Z digest=sha256:5da7bb1fa618754ae84e1ca91ac6ea1d7830012e4af0af848a03d3f324eb4d41

Observation cee59036-540e-4c27-8a3f-e3af7a893453 · inbound

EventDrive: Event Cameras for Vision-Language Driving Intelligence cites this paper.

EventDrive: Event Cameras for Vision-Language Driving Intelligence EventGPT: Event Stream Understanding with Multimodal Large Language Models

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:18:57.048104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-27T01:26:43.752335Z digest=sha256:c5ef6e8b892fb0bcc1b3d92ff961544068e0916956b4f85855901885f09de41f

Observation 7ec6b07c-ae39-4d48-b2f3-a86ee3c2bfd6 · inbound

DynFly: Dynamic-Aware Continuous Trajectory Generation for UAV Vision-Language Navigation in Urban Environments cites this paper.

DynFly: Dynamic-Aware Continuous Trajectory Generation for UAV Vision-Language Navigation in Urban Environments EventGPT: Event Stream Understanding with Multimodal Large Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-01T10:25:41.246070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-01T05:35:32.213470Z digest=sha256:3f69439529f249b3d270b223ff5ac7730e0c9d7d4487fc979a7c3cfac347bbf7