Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T19:15:55.834763Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 0 inbound Pith citation observations for arXiv:2507.06072.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T19:15:55.834763Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
49 of 49 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 89028ee0-0b70-4494-ba94-28430c64dfda · outbound
MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Spatio-temporal dynamics and se- mantic attribute enriched visual encoding for video caption- ing
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fa785253-d847-4c1f-8860-af2cfc55b373 · outbound
MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Spice: Semantic propositional image cap- tion evaluation
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 51361623-4519-4913-b172-1cb203345ccb · outbound
MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Covla: Comprehensive vision-language-action dataset for autonomous driving
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33c76992-50e8-44ba-9246-b3c7cc996789 · outbound
MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f37597d-ed3d-43c9-bc72-34c6bcaf9ed3 · outbound
MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Meteor: An automatic metric for mt evaluation with improved correlation with hu- man judgments
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b4c892c-c7b2-431e-9b3c-45e983b49fee · outbound
MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Egocentric vehicle dense video captioning
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 92989898-f2b5-49af-b013-f11327f9883a · outbound
MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Tem- adapter: Adapting image-text pretraining for video question answer
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4d5ebcdb-3054-438d-b7e9-c1dca0430c33 · outbound
MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Llcp: Learning latent causal processes for reasoning-based video question answer
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b8382273-60d4-45b3-9dad-33c9b164ae7b · outbound
MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding H-MBA: Hierarchical MamBa Adaptation for Multi-Modal Video Understanding in Autonomous Driving
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9fcc3e3e-ffe9-4c38-b258-121821e08c59 · outbound
MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Spatial-temporal trans- former for dynamic scene graph generation
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2ca91166-6193-4900-88b1-f2b7fb1b3c92 · outbound
MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Trafficvlm: A controllable visual lan- guage model for traffic video captioning
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b0e785c1-0a5e-446b-9148-c0991612fcf1 · outbound
MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Reversed in time: A novel temporal-emphasized benchmark for cross-modal video-text retrieval
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 75cccf7c-c386-4424-b487-f15c47ba9a51 · outbound
MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6bfba898-4c05-4dd7-b574-29ff23bc40e8 · outbound
MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Hierarchical representation net- work with auxiliary tasks for video captioning and video question answering
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 907367e7-4c4c-46a9-97d5-b78322f75558 · outbound
MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Text with knowledge graph aug- mented transformer for video captioning
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation adec4662-b589-4bc5-bbd6-503d136867c5 · outbound
MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Video re- cap: Recursive captioning of hour-long videos
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1d7ba409-7d9f-4509-a8d6-dd0c11050aa3 · outbound
MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Adapt: Action-aware driving caption transformer
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a2f44626-fccb-441c-934d-61c85406d0c8 · outbound
MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Cladder: Assessing causal reasoning in language models
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 66869762-d884-4ced-a7c7-cd5668e87800 · outbound
MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding ROAD-Waymo: A Large-Scale Action Awareness Dataset for Autonomous Driving
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 312eff15-ef43-4eab-8270-9bbf460de6ca · outbound
MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Textual explanations for self-driving ve- hicles
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 036c0c0a-005f-49a5-a080-cdd862b711ae · outbound
MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Learning hierarchical modular networks for video captioning
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ef5a778b-e72d-4a5f-9f8d-252a2bcc39e7 · outbound
MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Automatic evaluation of machine translation quality using longest common sub- sequence and skip-bigram statistics
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1b6d1d58-7bd1-42ce-9ef8-4a50ab2f3e20 · outbound
MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Swin- bert: End-to-end transformers with sparse attention for video 9 captioning
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation da708e17-d4ef-481b-b5c5-ccfa7c2724e2 · outbound
MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Cross-modal causal relational reasoning for event-level visual question answer- ing
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8b5cf959-bf80-4a92-a2e5-e81aae9ab809 · outbound
MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Spatio-temporal pixel- level contrastive learning-based source-free domain adapta- tion for video semantic segmentation
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8df80fa6-841c-4ac0-9d8d-eed51ab677ad · outbound
MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Llavilo: Boosting video moment retrieval via adapter-based multimodal modeling
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fa48fc39-b120-4253-8ffe-447cd0d85699 · outbound
MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Icsvr: Investigating compositional and syntactic understanding in video retrieval models
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 75acad3b-7d49-4586-be85-06eccf90de14 · outbound
MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Drama: Joint risk localization and captioning in driving
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af2edada-4990-401f-8ce5-0bc77a322194 · outbound
MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Lingoqa: Visual question answering for autonomous driv- ing
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8eaed610-8a7f-4637-a4ab-dfc195e1a606 · outbound
MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Query-dependent video representa- tion for moment retrieval and highlight detection
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0de4e240-8c09-44ba-a2a6-58d77b00490a · outbound
MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Bleu: a method for automatic evaluation of machine translation
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b958ed7-0d40-4415-9aea-ac001b0267ce · outbound
MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Causality
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0fbc6361-e44d-4f1e-8c4d-95bc4571c14d · outbound
MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Modular Learning of Deep Causal Generative Models for High-dimensional Causal Inference
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6dc8d77b-3255-4585-92c4-b2bff229ab42 · outbound
MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Clip4caption: Clip for video caption
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation feb098ee-39da-4e81-bdcf-6c8f634b6d7d · outbound
MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Cider: Consensus-based image description evalua- tion
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e73d360-1c84-4fc5-ac8d-65a3e63e56be · outbound
MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Sequence to sequence-video to text
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2e0f0f3b-4662-4617-a249-6721fd31ee13 · outbound
MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Deconfounding causal inference for zero-shot action recognition
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 30b8abb7-4683-404d-9561-ee063b4d9ca9 · outbound
MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Weakly- supervised video object grounding via causal intervention
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 82606f18-851c-4af0-8918-e1f7b9e75b31 · outbound
MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Rac3: Retrieval-augmented corner case comprehen- sion for autonomous driving with vision-language models
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9affc6dc-4edd-442c-bce0-b7523d5c5af6 · outbound
MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Visual causal scene refinement for video question an- swering
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 95e67e85-3683-40c8-9eb2-4c7234110ca1 · outbound
MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Bridging the gap: A unified video comprehension framework for mo- ment retrieval and highlight detection
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bfd72214-3e30-428f-b3f6-de1783938c66 · outbound
MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Retrieval-augmented egocentric video captioning
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 462a7aec-368e-4131-b654-957751a05e3d · outbound
MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Drivegpt4: Interpretable end-to-end autonomous driving via large language model
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a8756662-d0f3-4695-a37d-f6a3fb4bab33 · outbound
MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Prompt learns prompt: Exploring knowledge-aware generative prompt collaboration for video captioning
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 61f4aff8-1ddd-4279-a54d-022e0a998088 · outbound
MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Causal attention for vision-language tasks
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f5aee295-bc9f-4513-a243-d486342fcb1c · outbound
MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Rag-driver: Gen- eralisable driving explanations with retrieval-augmented in- context learning in multi-modal large language model
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73eec33c-2fd0-4186-9717-925de6e64b7d · outbound
MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Vision-language models for vision tasks: A survey
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e312ab5e-73c3-4ef7-8f16-19a13624324a · outbound
MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Causal inference with latent variables: 10 Figure 7
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ea6f08e4-87bf-4d4c-b4ba-01dc38dd7cad · outbound
MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Unresolved cited work
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
No inbound Pith citation observations are available.