Pith. sign in

Paper Citation Record · LEDGER

MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding

As of 19 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 0 inbound Pith citation observations for arXiv:2507.06072.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.06072 v1

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:15:55.834763Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

49 of 49 outbound references displayed

  • verified exact3
  • verified fuzzy36
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 89028ee0-0b70-4494-ba94-28430c64dfda · outbound

This paper cites Spatio-temporal dynamics and se- mantic attribute enriched visual encoding for video caption- ing.

MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Spatio-temporal dynamics and se- mantic attribute enriched visual encoding for video caption- ing

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:16:02.118591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T19:15:51.693713Z digest=sha256:13aa113d064292492aadb8ec786804e7bd46c32ab154b70bb03d724ee9be429d

Observation fa785253-d847-4c1f-8860-af2cfc55b373 · outbound

This paper cites Spice: Semantic propositional image cap- tion evaluation.

MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Spice: Semantic propositional image cap- tion evaluation

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:16:01.978664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T19:15:51.767952Z digest=sha256:4f0aa50be1a6117c8f18887ae2429d07e43c099051f759ad9cce1d9a8dc0044a

Observation 51361623-4519-4913-b172-1cb203345ccb · outbound

This paper cites Covla: Comprehensive vision-language-action dataset for autonomous driving.

MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Covla: Comprehensive vision-language-action dataset for autonomous driving

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:51.856249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:51.856249Z digest=sha256:bb7a24300f109f603652374049a015a68923ea43f9ac1ee3c6754c421cedc523

Observation 33c76992-50e8-44ba-9246-b3c7cc996789 · outbound

This paper cites MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens.

MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:51.936727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:51.936727Z digest=sha256:dff72e470caded636f1143ae87a733102267308b5ad0d5c507f5312c27eb5a89

Observation 1f37597d-ed3d-43c9-bc72-34c6bcaf9ed3 · outbound

This paper cites Meteor: An automatic metric for mt evaluation with improved correlation with hu- man judgments.

MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Meteor: An automatic metric for mt evaluation with improved correlation with hu- man judgments

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:52.022860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:52.022860Z digest=sha256:28fdccf86a7b523df08ff03e200069e461df449c1c667e9efb4c7ca3b1bd110c

Observation 9b4c892c-c7b2-431e-9b3c-45e983b49fee · outbound

This paper cites Egocentric vehicle dense video captioning.

MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Egocentric vehicle dense video captioning

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:16:01.823687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T19:15:52.082107Z digest=sha256:eb4c4bd1b8c5b1d58a7ad06f62734d3e656050641ec4a1634769b9929fd53994

Observation 92989898-f2b5-49af-b013-f11327f9883a · outbound

This paper cites Tem- adapter: Adapting image-text pretraining for video question answer.

MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Tem- adapter: Adapting image-text pretraining for video question answer

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:16:01.665888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T19:15:52.215115Z digest=sha256:4a0ea559895a7a6494fbcc094cb89e3cb2c4a48fe9d274f1c5b8eace4059e1ea

Observation 4d5ebcdb-3054-438d-b7e9-c1dca0430c33 · outbound

This paper cites Llcp: Learning latent causal processes for reasoning-based video question answer.

MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Llcp: Learning latent causal processes for reasoning-based video question answer

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:16:01.502522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T19:15:52.328341Z digest=sha256:80e87f7442e85d4b02119d998a7ff334995e0823015735a0b9bba1ca97c38fa0

Observation b8382273-60d4-45b3-9dad-33c9b164ae7b · outbound

This paper cites H-MBA: Hierarchical MamBa Adaptation for Multi-Modal Video Understanding in Autonomous Driving.

MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding H-MBA: Hierarchical MamBa Adaptation for Multi-Modal Video Understanding in Autonomous Driving

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:52.438880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:52.438880Z digest=sha256:59d0a6fab99e888ce5912016edde1c073044006634a251bbe1a04e52f1c92ecd

Observation 9fcc3e3e-ffe9-4c38-b258-121821e08c59 · outbound

This paper cites Spatial-temporal trans- former for dynamic scene graph generation.

MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Spatial-temporal trans- former for dynamic scene graph generation

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:16:01.400190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T19:15:52.522416Z digest=sha256:fd7b7449fc6a60b7d170a56a3e1d3518ec9e3fef2d05f342eb72d82834993b4a

Observation 2ca91166-6193-4900-88b1-f2b7fb1b3c92 · outbound

This paper cites Trafficvlm: A controllable visual lan- guage model for traffic video captioning.

MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Trafficvlm: A controllable visual lan- guage model for traffic video captioning

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:16:01.286301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T19:15:52.611203Z digest=sha256:85f6c0b61113c1d316a63ed811b79b26a0c30347232891efcc4c8af3e632b23e

Observation b0e785c1-0a5e-446b-9148-c0991612fcf1 · outbound

This paper cites Reversed in time: A novel temporal-emphasized benchmark for cross-modal video-text retrieval.

MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Reversed in time: A novel temporal-emphasized benchmark for cross-modal video-text retrieval

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:16:01.220112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T19:15:52.730614Z digest=sha256:4d4a3ee88effaec1dc28b28f15da657512823dfa520f57706ee58aa82b69303d

Observation 75cccf7c-c386-4424-b487-f15c47ba9a51 · outbound

This paper cites MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios.

MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:52.844188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:52.844188Z digest=sha256:064e7ce6b91f8f54aceb03ce630192657222d29dd1729a5a9cd73c645cd3b09a

Observation 6bfba898-4c05-4dd7-b574-29ff23bc40e8 · outbound

This paper cites Hierarchical representation net- work with auxiliary tasks for video captioning and video question answering.

MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Hierarchical representation net- work with auxiliary tasks for video captioning and video question answering

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:16:01.137806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T19:15:52.979990Z digest=sha256:b1bdde6dc76867750d25ef8f4c1ca559f1e2624af0a5afa6c8b013a722de9e34

Observation 907367e7-4c4c-46a9-97d5-b78322f75558 · outbound

This paper cites Text with knowledge graph aug- mented transformer for video captioning.

MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Text with knowledge graph aug- mented transformer for video captioning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:16:00.984175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T19:15:53.060296Z digest=sha256:fe1dfe429d92950024041b37f74cef11b3858d9c70635fc077fa0ac3dae8b075

Observation adec4662-b589-4bc5-bbd6-503d136867c5 · outbound

This paper cites Video re- cap: Recursive captioning of hour-long videos.

MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Video re- cap: Recursive captioning of hour-long videos

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:16:00.830894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T19:15:53.164839Z digest=sha256:04e7d3313579acc6f877ced2357ff1b75c57808e6ca1fdfe0934f86d65d7469a

Observation 1d7ba409-7d9f-4509-a8d6-dd0c11050aa3 · outbound

This paper cites Adapt: Action-aware driving caption transformer.

MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Adapt: Action-aware driving caption transformer

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:16:00.681506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T19:15:53.291205Z digest=sha256:8cdd23f92d929bc645be76ec94fecd4438168d5ff99a0aead0c916a7965100a6

Observation a2f44626-fccb-441c-934d-61c85406d0c8 · outbound

This paper cites Cladder: Assessing causal reasoning in language models.

MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Cladder: Assessing causal reasoning in language models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:16:00.532441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T19:15:53.415582Z digest=sha256:6934431d4686303049d38fd171d7bc88b41ce0328ce3dda468b218dec6a2918f

Observation 66869762-d884-4ced-a7c7-cd5668e87800 · outbound

This paper cites ROAD-Waymo: A Large-Scale Action Awareness Dataset for Autonomous Driving.

MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding ROAD-Waymo: A Large-Scale Action Awareness Dataset for Autonomous Driving

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:15:56.454572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T19:15:53.530074Z digest=sha256:f6c5eebf61048d543e5dd90efb0294c796176e93c819872960708eff15afba2e

Observation 312eff15-ef43-4eab-8270-9bbf460de6ca · outbound

This paper cites Textual explanations for self-driving ve- hicles.

MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Textual explanations for self-driving ve- hicles

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:16:00.407176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T19:15:53.649717Z digest=sha256:b75efbfa0bc5bf88ed4b24b3a4379154f24af2201e26d179836c41f5c8aeaa8a

Observation 036c0c0a-005f-49a5-a080-cdd862b711ae · outbound

This paper cites Learning hierarchical modular networks for video captioning.

MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Learning hierarchical modular networks for video captioning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:16:00.252180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T19:15:53.699498Z digest=sha256:796aa04984dd6d56df77e59fae0861fa27c72bf32f4abef0253257d7e064eb9a

Observation ef5a778b-e72d-4a5f-9f8d-252a2bcc39e7 · outbound

This paper cites Automatic evaluation of machine translation quality using longest common sub- sequence and skip-bigram statistics.

MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Automatic evaluation of machine translation quality using longest common sub- sequence and skip-bigram statistics

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:16:00.099477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T19:15:53.772761Z digest=sha256:2d11c6fd1d7e55c4b16581e2e380ce4ccd4e190df1bf3454f2b79e0a961b5b06

Observation 1b6d1d58-7bd1-42ce-9ef8-4a50ab2f3e20 · outbound

This paper cites Swin- bert: End-to-end transformers with sparse attention for video 9 captioning.

MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Swin- bert: End-to-end transformers with sparse attention for video 9 captioning

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:15:59.925750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T19:15:53.820496Z digest=sha256:0aec51850849a6c788d4bee47997fee4e304e5c9ccb6ff050858853bb61e10a0

Observation da708e17-d4ef-481b-b5c5-ccfa7c2724e2 · outbound

This paper cites Cross-modal causal relational reasoning for event-level visual question answer- ing.

MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Cross-modal causal relational reasoning for event-level visual question answer- ing

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:15:59.744926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T19:15:53.892898Z digest=sha256:6895da0176df0e016e28df9134d987d086ad41db721c357e8eb1c62baa4e784f

Observation 8b5cf959-bf80-4a92-a2e5-e81aae9ab809 · outbound

This paper cites Spatio-temporal pixel- level contrastive learning-based source-free domain adapta- tion for video semantic segmentation.

MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Spatio-temporal pixel- level contrastive learning-based source-free domain adapta- tion for video semantic segmentation

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:15:59.592783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T19:15:53.957464Z digest=sha256:3f47e944c8f4655b271f660eb949efcd4666905183f27f045a2ad0e42d382c61

Observation 8df80fa6-841c-4ac0-9d8d-eed51ab677ad · outbound

This paper cites Llavilo: Boosting video moment retrieval via adapter-based multimodal modeling.

MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Llavilo: Boosting video moment retrieval via adapter-based multimodal modeling

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:15:59.434867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T19:15:54.018598Z digest=sha256:a6e48df65abcfef77b8bed70626cd74629179c029eae7c14879b85194386c872

Observation fa48fc39-b120-4253-8ffe-447cd0d85699 · outbound

This paper cites Icsvr: Investigating compositional and syntactic understanding in video retrieval models.

MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Icsvr: Investigating compositional and syntactic understanding in video retrieval models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:15:59.280031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T19:15:54.081336Z digest=sha256:36630ff54a9bba71a71e4ad544b895a16ba962e4d354e2ff4326af987d67713d

Observation 75acad3b-7d49-4586-be85-06eccf90de14 · outbound

This paper cites Drama: Joint risk localization and captioning in driving.

MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Drama: Joint risk localization and captioning in driving

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:54.142836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:54.142836Z digest=sha256:2e080a2ab7ca94bfd0ec2c1b09b80acc918ce693d9d372689d4d9f12e90833bc

Observation af2edada-4990-401f-8ce5-0bc77a322194 · outbound

This paper cites Lingoqa: Visual question answering for autonomous driv- ing.

MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Lingoqa: Visual question answering for autonomous driv- ing

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:15:59.146754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T19:15:54.205076Z digest=sha256:b7d3d14f5dd19da1810bf897c5faa7dd48459997330f8d164da701e2c0cead7c

Observation 8eaed610-8a7f-4637-a4ab-dfc195e1a606 · outbound

This paper cites Query-dependent video representa- tion for moment retrieval and highlight detection.

MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Query-dependent video representa- tion for moment retrieval and highlight detection

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:15:58.999449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T19:15:54.268006Z digest=sha256:6acbe6320765438894ad0ae0e374bda34f0f7289451ecb0d37f25db2a13daace

Observation 0de4e240-8c09-44ba-a2a6-58d77b00490a · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation.

MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Bleu: a method for automatic evaluation of machine translation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:54.333943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:54.333943Z digest=sha256:fa0446b2bcf4744485e9d6a87aff2727af4130eee7098a48541f4b7c468bdbfe

Observation 6b958ed7-0d40-4415-9aea-ac001b0267ce · outbound

This paper cites Causality.

MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Causality

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:15:58.815478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T19:15:54.392034Z digest=sha256:0f1586bc424eb03b55a1eb341f22b4fbf6d368f4a3af2f71ff8d933d8eecf2e2

Observation 0fbc6361-e44d-4f1e-8c4d-95bc4571c14d · outbound

This paper cites Modular Learning of Deep Causal Generative Models for High-dimensional Causal Inference.

MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Modular Learning of Deep Causal Generative Models for High-dimensional Causal Inference

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:15:56.313290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T19:15:54.452774Z digest=sha256:485ec684c465bafc429f169b1ce507fb925c8368b230657ac409e2d04ac23105

Observation 6dc8d77b-3255-4585-92c4-b2bff229ab42 · outbound

This paper cites Clip4caption: Clip for video caption.

MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Clip4caption: Clip for video caption

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:15:58.668315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T19:15:54.513741Z digest=sha256:6d83460b7ea93fed4713bf647e98ae9b164bb1efb5179a994b53f4d2ae317a1a

Observation feb098ee-39da-4e81-bdcf-6c8f634b6d7d · outbound

This paper cites Cider: Consensus-based image description evalua- tion.

MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Cider: Consensus-based image description evalua- tion

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:54.574799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:54.574799Z digest=sha256:f772c7b0714ac44a94dec80f5e19538576bb6e305927645b050099621a9f5ef3

Observation 8e73d360-1c84-4fc5-ac8d-65a3e63e56be · outbound

This paper cites Sequence to sequence-video to text.

MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Sequence to sequence-video to text

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:15:58.512611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T19:15:54.639691Z digest=sha256:234c2125c627cd9b431591d4516130eb05334bd1814ec2267c77a08910384089

Observation 2e0f0f3b-4662-4617-a249-6721fd31ee13 · outbound

This paper cites Deconfounding causal inference for zero-shot action recognition.

MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Deconfounding causal inference for zero-shot action recognition

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:15:58.367986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T19:15:54.699040Z digest=sha256:b95a3e4c3db29b4779d72c234bc710b53ab0b34c095015645c48eb326f2fe866

Observation 30b8abb7-4683-404d-9561-ee063b4d9ca9 · outbound

This paper cites Weakly- supervised video object grounding via causal intervention.

MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Weakly- supervised video object grounding via causal intervention

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:15:58.232681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T19:15:54.765633Z digest=sha256:5c17f218843f8e6b16c8e6afeaa58b1380fd85f4c2ee36ea6e71ce812f979e79

Observation 82606f18-851c-4af0-8918-e1f7b9e75b31 · outbound

This paper cites Rac3: Retrieval-augmented corner case comprehen- sion for autonomous driving with vision-language models.

MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Rac3: Retrieval-augmented corner case comprehen- sion for autonomous driving with vision-language models

Reference 39

Resolution
verified exact
raw_fallback, observed 2026-08-06T19:15:56.152954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T19:15:54.840851Z digest=sha256:f25326f2eac26c0bdba5912a4e3798fc1d33ae2bafca65c2e1659fdea56cb900

Observation 9affc6dc-4edd-442c-bce0-b7523d5c5af6 · outbound

This paper cites Visual causal scene refinement for video question an- swering.

MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Visual causal scene refinement for video question an- swering

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:15:58.107868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T19:15:54.905112Z digest=sha256:6bcdf7da2263b84b9f5158ecc59ea5f5099cd37b429548a1c080fbfa6c30cd87

Observation 95e67e85-3683-40c8-9eb2-4c7234110ca1 · outbound

This paper cites Bridging the gap: A unified video comprehension framework for mo- ment retrieval and highlight detection.

MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Bridging the gap: A unified video comprehension framework for mo- ment retrieval and highlight detection

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:15:57.954692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T19:15:54.986106Z digest=sha256:c3f12815bcde79c3a9dcd0221a9931a98a4fb65429ed9db817dbe3ce5a6537a4

Observation bfd72214-3e30-428f-b3f6-de1783938c66 · outbound

This paper cites Retrieval-augmented egocentric video captioning.

MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Retrieval-augmented egocentric video captioning

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:15:57.808237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T19:15:55.043092Z digest=sha256:4245cc5f3bd2ff485baf0d6730a1748e8c0df13b6e841212d7cbb1b77fc35e11

Observation 462a7aec-368e-4131-b654-957751a05e3d · outbound

This paper cites Drivegpt4: Interpretable end-to-end autonomous driving via large language model.

MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Drivegpt4: Interpretable end-to-end autonomous driving via large language model

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:15:57.672526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T19:15:55.194858Z digest=sha256:85c58c5c1aca0bf5b5e79d246e29a0d30ee7d6c6a476404787af8bb2f58eff3e

Observation a8756662-d0f3-4695-a37d-f6a3fb4bab33 · outbound

This paper cites Prompt learns prompt: Exploring knowledge-aware generative prompt collaboration for video captioning.

MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Prompt learns prompt: Exploring knowledge-aware generative prompt collaboration for video captioning

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:15:57.555070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T19:15:55.301118Z digest=sha256:21c3b9ded1dc577e8dd62ab022ff9cee166ca07991e13c8756ce5dffdc382c1b

Observation 61f4aff8-1ddd-4279-a54d-022e0a998088 · outbound

This paper cites Causal attention for vision-language tasks.

MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Causal attention for vision-language tasks

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:15:57.412920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T19:15:55.419088Z digest=sha256:af544eb08b51cf487a5e2bd4780d20f3dae123995e97e4d40c7491d3f15589e2

Observation f5aee295-bc9f-4513-a243-d486342fcb1c · outbound

This paper cites Rag-driver: Gen- eralisable driving explanations with retrieval-augmented in- context learning in multi-modal large language model.

MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Rag-driver: Gen- eralisable driving explanations with retrieval-augmented in- context learning in multi-modal large language model

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:55.528946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:55.528946Z digest=sha256:005708f36406279f37f139d73ecab50ec0e3e88a66aea2c1b8737ecd0f54e83a

Observation 73eec33c-2fd0-4186-9717-925de6e64b7d · outbound

This paper cites Vision-language models for vision tasks: A survey.

MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Vision-language models for vision tasks: A survey

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:15:57.203793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T19:15:55.655076Z digest=sha256:c530e05e3a7dc49649e7a434754b950bdcd7a3d1adf48db92f08a0f2f0a115e2

Observation e312ab5e-73c3-4ef7-8f16-19a13624324a · outbound

This paper cites Causal inference with latent variables: 10 Figure 7.

MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Causal inference with latent variables: 10 Figure 7

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:15:56.951227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T19:15:55.734302Z digest=sha256:42d03051a4c2caeb938970649524e250a0c8561b9b981a49bb49d39899dba466

Observation ea6f08e4-87bf-4d4c-b4ba-01dc38dd7cad · outbound

This paper cites an unresolved cited work.

MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:15:56.683905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T19:15:55.834763Z digest=sha256:9288f60bef04db312481fa97d92a6765c61ad5262289392a2a2ecef80ed9d273

Pith citing papers

No inbound Pith citation observations are available.