Pith. sign in

Paper Citation Record · LEDGER

EventGPT: Event Stream Understanding with Multimodal Large Language Models

As of 12 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 3 inbound Pith citation observations for arXiv:2412.00832.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.00832 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T05:01:12.829064Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:35:01.366536Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T20:18:57.044501Z

Reference resolution

51 of 51 outbound references displayed

  • verified exact2
  • verified fuzzy26
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f27a5cf4-8386-4842-b18b-319ef0b3c0ea · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:01:15.376332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.146077Z digest=sha256:fb54ddcc158d8308f8a7f9d5f294b7fd35e6df4c857e68b60f229850ea2b862d

Observation cda93b88-d0ba-4f75-8fc4-b88416fabf8d · outbound

This paper cites The (r) evolution of multi- modal large language models: A survey.

EventGPT: Event Stream Understanding with Multimodal Large Language Models The (r) evolution of multi- modal large language models: A survey

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T05:01:12.191004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:01:12.191004Z digest=sha256:23658a6664f3719e900ff3d7d67412b2e37601125967843969279afcf1a46fb3

Observation 877ca8af-dd48-45b6-a512-3ebb13ac7acc · outbound

This paper cites First Place Solution to the CVPR'2023 AQTC Challenge: A Function-Interaction Centric Approach with Spatiotemporal Visual-Language Alignment.

EventGPT: Event Stream Understanding with Multimodal Large Language Models First Place Solution to the CVPR'2023 AQTC Challenge: A Function-Interaction Centric Approach with Spatiotemporal Visual-Language Alignment

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-08-12T05:01:14.030175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.213248Z digest=sha256:fffc34033a74e4e4082d8e9c556c99094341d99cb88e444b49099278542531cb

Observation d4b315fa-d1df-4ebc-a4fb-60c6b9f3e872 · outbound

This paper cites Dress: Instructing large vision-language models to align and interact with humans via natural lan- guage feedback.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Dress: Instructing large vision-language models to align and interact with humans via natural lan- guage feedback

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:01:15.345875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.221664Z digest=sha256:9a0f2c34e9dcc27f530c2726a4aaae16a61b4df6083b09a4bab28fa44d5fc88a

Observation 3f724c32-951e-477f-9df4-0600bbbe027b · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:01:15.311023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.233969Z digest=sha256:fe23e194d1cd232c323f619d77b212fa697d46ea73ebdc95038a165163ed83f1

Observation 2bc6fd32-242c-49f5-a715-bb0e4893d8a9 · outbound

This paper cites Reproducible scal- ing laws for contrastive language-image learning.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Reproducible scal- ing laws for contrastive language-image learning

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:01:15.275105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.246718Z digest=sha256:f04a8aa0c0d6c02b6305f1c4df0b57ee9a78f2d445dbf27c63be04dd5e02e1b0

Observation 05333c3c-b55a-4dcb-91f5-368cdc6e3118 · outbound

This paper cites Event-based vision: A survey.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Event-based vision: A survey

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:01:15.224190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.264858Z digest=sha256:9eef235320ba6bd9cc9757bd9bcb20444926a1b75928fdcfc742ba45332a7aef

Observation 3a8da49b-83ad-49ee-bb36-20e4c81e0fea · outbound

This paper cites Low-latency auto- motive vision with event cameras.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Low-latency auto- motive vision with event cameras

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T05:01:12.272493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:01:12.272493Z digest=sha256:d8dd482e8f3662cf0a5bbb575b38b793bd5d0a1dc3bbb389fd2d1a92f74debde

Observation dddee9e6-87e8-4149-bb7e-7306ef5af691 · outbound

This paper cites Eklt: Asynchronous photometric feature tracking using events and frames.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Eklt: Asynchronous photometric feature tracking using events and frames

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:01:15.147226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.299020Z digest=sha256:2fa569a0a5e98d56a9982848ab3859f3854e53b05c0fd3605d4fc99a63540569

Observation 6c17c68a-1f27-4629-90f8-718b53456217 · outbound

This paper cites Recurrent vision transformers for object detection with event cameras.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Recurrent vision transformers for object detection with event cameras

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:01:15.094373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.307967Z digest=sha256:8cd898d31398812dfb7b615ba0a44290f4e6b6dfe38e07c73acbce386ef4a88e

Observation fe9d2bdf-22a2-4961-a51e-a9d697ac4c09 · outbound

This paper cites Dsec: A stereo event camera dataset for driving scenarios.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Dsec: A stereo event camera dataset for driving scenarios

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:01:15.058056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.319807Z digest=sha256:9f2b77626bb24fa6dd0b9b5b57d940a1fc44092eb54889af505d34cbf74a2c77

Observation d9220487-9589-44c5-9e3a-af4cfab38a60 · outbound

This paper cites Event-based Simultaneous Localization and Mapping: A Comprehensive Survey.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Event-based Simultaneous Localization and Mapping: A Comprehensive Survey

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T05:01:12.325692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:01:12.325692Z digest=sha256:a554a7d4022ca292884cf83231ceddce4250bea517022859b93d44537c6ecc2b

Observation d87a4c97-c318-4cce-b567-56c9cff7cff2 · outbound

This paper cites Img-Diff: Contrastive Data Synthesis for Multimodal Large Language Models.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Img-Diff: Contrastive Data Synthesis for Multimodal Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T05:01:12.370094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:01:12.370094Z digest=sha256:c3245c9905a16867c2868f5d8ec3f9b37583ee0dae401b2113cdc65385f48c2c

Observation 8696cb31-8981-4de7-8ad8-0e9c6db56ac8 · outbound

This paper cites Real-time 3d reconstruction and 6-dof tracking with an event camera.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Real-time 3d reconstruction and 6-dof tracking with an event camera

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:01:15.013695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.380301Z digest=sha256:21d0055846bbd5a37d0e753a402e926bc30ead60a8ad7337bb538fce73395b8d

Observation 9a76f3d8-4de0-4523-a974-9f71ad70a25b · outbound

This paper cites N-imagenet: Towards robust, fine-grained object recognition with event cameras.

EventGPT: Event Stream Understanding with Multimodal Large Language Models N-imagenet: Towards robust, fine-grained object recognition with event cameras

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:01:14.969940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.389515Z digest=sha256:fb15495607733f8f9389b0875fc352beea1dd30c7806de57dfee7c3f1e63128b

Observation 164a4524-4963-4a2c-ac7e-28b526a83323 · outbound

This paper cites Sodformer: Streaming object detection with transformer using events and frames.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Sodformer: Streaming object detection with transformer using events and frames

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:01:14.921551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.407322Z digest=sha256:d73dd3c07a5795bce2d8b4451836af97b87c72d0d744c71261575a0c2399c05e

Observation 2d517bce-5fa1-43e7-9785-691d9a70c3d4 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T05:01:12.419221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:01:12.419221Z digest=sha256:f71232b432d76766f31eb78948396ca8bdb6bedc2cdaafe62eeb301a2569fe79

Observation 33a74984-41e2-4f56-931e-d1f364193f75 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:01:14.861678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.432463Z digest=sha256:cbde0848c400074d8489801b7af618715c6741425f8bbb49ac1a7f2452cde06b

Observation 8f8b143d-538b-4088-bcd4-f14b415b7773 · outbound

This paper cites Vila: On pre-training for visual language models.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Vila: On pre-training for visual language models

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:01:14.815717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.451299Z digest=sha256:5399b18f6a90ba1852d8709554f8c35fd5050beaf1bc4c8c649e88b846454d11

Observation f159e167-0cad-49fc-9fc7-b59820cb659b · outbound

This paper cites Improved baselines with visual instruction tuning.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Improved baselines with visual instruction tuning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T05:01:12.463039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:01:12.463039Z digest=sha256:9d092b8e5a636a80357ae1dd7aeeadcc51efd8d500ae531d964cb256ff7cbf13

Observation de61ff03-66d6-4813-9971-792a65dd3047 · outbound

This paper cites Visual instruction tuning.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Visual instruction tuning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:01:14.748627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.471738Z digest=sha256:df9dafc6716e87492847b0c9b79fc76be06a81a39ab65c030139e93d31271467

Observation ae5221da-37bc-420b-875f-228503fba171 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T05:01:12.482716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:01:12.482716Z digest=sha256:fdd596002d15a30901e10e64206e463553872b5d9994c62fd96388fc791b9812

Observation 79b9cb82-abac-45af-85a3-abd0e06f9dd0 · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

EventGPT: Event Stream Understanding with Multimodal Large Language Models DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T05:01:12.494304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:01:12.494304Z digest=sha256:800a3b173145ad43ad54e32e8f25a45d61332613ed3d9465002878aaa3a1de9a

Observation 33e8d8d2-9aad-4e1b-8346-d4062f9c4626 · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T05:01:12.504000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:01:12.504000Z digest=sha256:3b0c2974c21fb02f9a44016cf30b4ed66a3acb1d70d8f45cb8f222b24534a739

Observation 1e1285f9-2f14-4f3a-9e42-ac08ff1bbb71 · outbound

This paper cites Data-driven feature tracking for event cameras.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Data-driven feature tracking for event cameras

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:01:14.712020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.516781Z digest=sha256:df54b7d03a546377d2524aa9041b9d68d21aabf05b48ba1f78b78cc0c790e030

Observation 36c26336-66e5-4419-a6ec-152428026790 · outbound

This paper cites Esl: Event-based structured light.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Esl: Event-based structured light

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:01:14.660582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.526710Z digest=sha256:66e8173a4473002b79c632a8953c059d06824e86314666796ba353e8e356d830

Observation 159208c6-faaf-4660-80bb-fd0348074fdb · outbound

This paper cites Cross-modal Attention Congruence Regularization for Vision-Language Relation Alignment.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Cross-modal Attention Congruence Regularization for Vision-Language Relation Alignment

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T05:01:12.536142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:01:12.536142Z digest=sha256:5e7edb74a4cdb42891fe8defe5eb01f4ec3a5c44a37db9c00aa6a96320ea10b7

Observation f13b0614-776c-4379-9b42-f765ae47f85d · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Learn- ing transferable visual models from natural language super- vision

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T05:01:12.546821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:01:12.546821Z digest=sha256:6f3b2a21a3fd5580e3b6a717880d7a8bb0c110a3ba38f02c4f529b6b7b5eb1f7

Observation 968532f8-6901-4f09-a2a3-5980694820d2 · outbound

This paper cites Emvs: Event-based multi-view stereo—3d 9 reconstruction with an event camera in real-time.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Emvs: Event-based multi-view stereo—3d 9 reconstruction with an event camera in real-time

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:01:14.580424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.559931Z digest=sha256:3e7cfd131e3d88a776ab1ad7cf5a31c25afb4f02223933a72c3adfa29b5abfb3

Observation 53c8eda1-4a95-47ef-957c-7d272039650d · outbound

This paper cites Events-to-video: Bringing modern computer vision to event cameras.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Events-to-video: Bringing modern computer vision to event cameras

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:01:14.535641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.576078Z digest=sha256:4ea9bb331218a23970b94fcb02781a497382aa6866ee21d0b69bb2e57b869d05

Observation 6528a202-e7dc-44f5-9e4f-6df6a237dd1f · outbound

This paper cites Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T05:01:12.585257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:01:12.585257Z digest=sha256:d95e4a90e913f10a1e577b541a420965a53c99ae9c902020b9529df77ce5e603

Observation bf38fdca-0f41-4b83-94e2-aa4adf5b7bc6 · outbound

This paper cites Aligning and prompting everything all at once for univer- sal visual perception.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Aligning and prompting everything all at once for univer- sal visual perception

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:01:14.504155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.592175Z digest=sha256:3b4da5c534036427afcf22b151d369fdace3f9ef5ff6e3472bbf152b9a2b3edf

Observation c169dbb5-d2d0-4f4b-b4bc-83397472b7aa · outbound

This paper cites BlinkTrack: Feature Tracking over 80 FPS via Events and Images.

EventGPT: Event Stream Understanding with Multimodal Large Language Models BlinkTrack: Feature Tracking over 80 FPS via Events and Images

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-08-12T05:01:13.614829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.601049Z digest=sha256:3956b502b8c2ed2321644c22d545fdad9180d5585f3b40a8a152603c675bee02

Observation 4ef5e9f7-7c74-4902-9b55-dca0f2c8a767 · outbound

This paper cites Flava: A foundational language and vision alignment model.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Flava: A foundational language and vision alignment model

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:01:14.470618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.615369Z digest=sha256:9c65a92b1142f53621464c3f39a983f43cddf9ba765910d487ff26ef0544b960

Observation 7c054af0-cdea-4ef8-ad61-d2753fc46196 · outbound

This paper cites Cloud-device collaborative learning for multimodal large language models.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Cloud-device collaborative learning for multimodal large language models

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:01:14.410556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.629470Z digest=sha256:05824297c21a49dd44e512e8f246e26feae00cfd6c2c4195097aca24cbd8bdf0

Observation 53e4539d-0540-4ab5-8523-50f11d46a391 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T05:01:12.644835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:01:12.644835Z digest=sha256:9ff88317f0091510cf86c8cc7e18573bff3c857778546d55d27f3067c39b1a2e

Observation a051b0d7-b2bb-4400-8e40-870824ed9ee4 · outbound

This paper cites NExT-GPT: Any-to-Any Multimodal LLM.

EventGPT: Event Stream Understanding with Multimodal Large Language Models NExT-GPT: Any-to-Any Multimodal LLM

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T05:01:12.652501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:01:12.652501Z digest=sha256:c25ffe84e67348b138bc5d65e4a8d713850e23a7d6590cb194fce722b7d28453

Observation f09726e2-4787-4789-964c-78d9265993ee · outbound

This paper cites EventCLIP: Adapting CLIP for Event-based Object Recognition.

EventGPT: Event Stream Understanding with Multimodal Large Language Models EventCLIP: Adapting CLIP for Event-based Object Recognition

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T05:01:12.664201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:01:12.664201Z digest=sha256:265578b2592989a61759d877089533919b9cdd264e22f2c4597ba38e84400257

Observation 177c1df0-cd5d-4dfd-ac9e-bfeacf6a6b63 · outbound

This paper cites Leod: Label-efficient object detection for event cameras.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Leod: Label-efficient object detection for event cameras

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:01:14.377574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.675218Z digest=sha256:29d5d839ad4abe86b5223094ae7908d752617e287118bb51f60670047f421d20

Observation e795f0ce-43e0-4479-91b5-ba390758a38f · outbound

This paper cites xgen-mm (blip-3): A family of open large multimodal models.

EventGPT: Event Stream Understanding with Multimodal Large Language Models xgen-mm (blip-3): A family of open large multimodal models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T05:01:12.682307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:01:12.682307Z digest=sha256:2d5e2c6b66e117a04536e387c5cce036494ab2b8b5ca6df9ea1e619fd35831db

Observation ce77d6d9-3793-4b57-a355-011819caa0ce · outbound

This paper cites A Survey on Multimodal Large Language Models.

EventGPT: Event Stream Understanding with Multimodal Large Language Models A Survey on Multimodal Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T05:01:12.701293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:01:12.701293Z digest=sha256:f4b7b3bdca1b3c69d2edf0beeb4cb144e120ebc2867021e760bcd6bfa5db012f

Observation 1f35d5f7-59f1-43ca-8419-a13b448ebc7d · outbound

This paper cites Eventps: Real-time photometric stereo using an event camera.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Eventps: Real-time photometric stereo using an event camera

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:01:14.326105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.711889Z digest=sha256:f4436b68c2b4d5483876233a51226ce315cbc89e04770951d2df6f65d275ba25

Observation e8e3c9be-af2f-4cef-8bed-155f1dfaf6bb · outbound

This paper cites AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling.

EventGPT: Event Stream Understanding with Multimodal Large Language Models AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T05:01:12.725243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:01:12.725243Z digest=sha256:e7dd7feb1d6916cbff7230e704e031c58a8bb3a5d80ea0032a3124a4778369f2

Observation 173d49c1-b66d-4209-a7f9-6e1e7aa1f175 · outbound

This paper cites MM-LLMs: Recent Advances in MultiModal Large Language Models.

EventGPT: Event Stream Understanding with Multimodal Large Language Models MM-LLMs: Recent Advances in MultiModal Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T05:01:12.746090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:01:12.746090Z digest=sha256:6977b639ba5cedc33a8072f74672d6e62583f637bd0e7a21da347292bed14166

Observation b418c2d4-6d9e-4efc-829c-9033f9a779fa · outbound

This paper cites Spiking transform- ers for event-based single object tracking.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Spiking transform- ers for event-based single object tracking

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:01:14.289514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.767968Z digest=sha256:22a28af51b80ab559ea83d784b7a0240bda2018123b1fdb9e311189ea6d5ae31

Observation 93adb63c-5752-49c7-90b4-9f28c8c7bfb3 · outbound

This paper cites OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding.

EventGPT: Event Stream Understanding with Multimodal Large Language Models OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T05:01:12.778547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:01:12.778547Z digest=sha256:3919cbb905a3f9aedf509cbf1ec20c52456346e427b4748787a8cbec0ccf35d9

Observation e8fd68e8-126b-481c-884d-b0955e5130d9 · outbound

This paper cites Deep Learning for Event-based Vision: A Comprehensive Survey and Benchmarks.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Deep Learning for Event-based Vision: A Comprehensive Survey and Benchmarks

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T05:01:12.790227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:01:12.790227Z digest=sha256:37b77a41977007ac1d9898f5b885222ebfb36e4f929ced4b2574fd8ee998f4e3

Observation b86f87bf-1754-44f4-86c0-e7b6f28dd19f · outbound

This paper cites E- clip: Towards label-efficient event-based open-world under- standing by clip.

EventGPT: Event Stream Understanding with Multimodal Large Language Models E- clip: Towards label-efficient event-based open-world under- standing by clip

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:01:14.250690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.799667Z digest=sha256:fa5a7288b5280947aec6bccc3c384ebd01378fdd1f4ca3fe03d8cba8d26b8d6a

Observation 74cf5c2a-b2db-4702-b0a2-d23a41979662 · outbound

This paper cites Ex- act: Language-guided conceptual reasoning and uncertainty estimation for event-based action recognition and more.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Ex- act: Language-guided conceptual reasoning and uncertainty estimation for event-based action recognition and more

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:01:14.224985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.811621Z digest=sha256:28e4340f56c8d705af0516fb8dc3bfc0ced8b1e63502a85976a8d5d7fc0a39f3

Observation 1e43ee78-5e59-493c-99cb-f23f3c94085b · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

EventGPT: Event Stream Understanding with Multimodal Large Language Models MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T05:01:12.822819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:01:12.822819Z digest=sha256:7d861a44a8226e2ede856de0024d38ed7ce5f10dc3d63d1687a6dac27d19d5b5

Observation 517844bd-825e-4b7e-bd36-135d6f994670 · outbound

This paper cites VL-GPT: A Generative Pre-trained Transformer for Vision and Language Understanding and Generation.

EventGPT: Event Stream Understanding with Multimodal Large Language Models VL-GPT: A Generative Pre-trained Transformer for Vision and Language Understanding and Generation

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T05:01:12.829064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:01:12.829064Z digest=sha256:1ace9fbe8503ff11e63dd60bb2a60dfd09c3ff9855001ff9e618d09612321400

Pith citing papers

Observation e341bfd4-a294-4028-a135-048efe39d61c · inbound

Event-Priori-Based Vision-Language Model for Efficient Visual Understanding cites this paper.

Event-Priori-Based Vision-Language Model for Efficient Visual Understanding EventGPT: Event Stream Understanding with Multimodal Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:01.366536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:01.366536Z digest=sha256:db3d220cfe3db3f13bc8d5c55ddc7fb6e4a794ee0f1d196a03ba49951a60200a

Observation cee59036-540e-4c27-8a3f-e3af7a893453 · inbound

EventDrive: Event Cameras for Vision-Language Driving Intelligence cites this paper.

EventDrive: Event Cameras for Vision-Language Driving Intelligence EventGPT: Event Stream Understanding with Multimodal Large Language Models

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:18:57.048104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-27T01:26:43.752335Z digest=sha256:fb513053bf6d107dbdb537cc7f318580f6e742479e1bb50d8025de1c26585d6f

Observation 7ec6b07c-ae39-4d48-b2f3-a86ee3c2bfd6 · inbound

DynFly: Dynamic-Aware Continuous Trajectory Generation for UAV Vision-Language Navigation in Urban Environments cites this paper.

DynFly: Dynamic-Aware Continuous Trajectory Generation for UAV Vision-Language Navigation in Urban Environments EventGPT: Event Stream Understanding with Multimodal Large Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-01T10:25:41.246070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-01T05:35:32.213470Z digest=sha256:d5259038ff00f193dee1ffbfac8f8db715a8123241d0ed396f7ea455575d736c