Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-29T22:27:37.631870Z
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 1 inbound Pith citation observation for arXiv:2605.26104.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-29T22:27:37.631870Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-31T04:55:06.976425Z
A source-named dated measurement, never combined with another source.
Source: cited_works
59 of 59 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 59683863-dfe6-4790-90cf-304ea96d036e · outbound
EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Slot-guided adaptation of pre-trained diffusion models for object-centric learning and compositional generation
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee1662a8-c807-4b1d-9574-9151eccb8ddc · outbound
EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding DEVIAS: Learning disentangled video representations of action and scene
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42fb294e-730e-4ae0-bbd1-03ff9f58cb1b · outbound
EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Qwen2.5-VL Technical Report
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ca6b291a-1416-42c1-92f3-b15c87336c8c · outbound
EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Learning sample importance for cross-scenario video temporal grounding
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59ff1312-5ae6-4bc5-9370-b7f69fd956df · outbound
EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding End-to-end object detection with transformers
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbceec8d-df5a-4211-b19e-211b24a6ccb7 · outbound
EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Towards a complete benchmark on video moment localization
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf929a28-c1f6-430a-ab45-9b10808689cd · outbound
EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Slot-MLLM: Object-Centric Visual Tokenization for Multimodal LLM
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 47b534db-961d-412f-97cc-59f28e61395b · outbound
EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Learning phrase representations using RNN encoder–decoder for statistical machine translation
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8223d8e-54e0-4287-af27-dfc9a7d6fb93 · outbound
EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Why can’t i dance in the mall? learning to mitigate scene bias in action recognition
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fbd0d8e-10ef-4f2e-9984-be8c0e61d8e7 · outbound
EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Tall: Temporal activity localization via language query
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c773432e-e090-455a-a8f9-3f7d01ba95db · outbound
EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding TRACE: Temporal Grounding Video LLM via Causal Event Modeling
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 32172b78-bf7a-4188-a47f-1577be444464 · outbound
EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Can shuffling video benefit temporal bias problem: A novel training framework for temporal grounding
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e01abf9-6027-409d-be35-a7e42d180116 · outbound
EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Localizing moments in video with natural language
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc866930-3d4d-4710-bbb8-627b48dc0657 · outbound
EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Lora: Low-rank adaptation of large language models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 628cb1d9-ecfd-4888-841d-53da34c13bf6 · outbound
EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Vtimellm: Empower llm to grasp video moments
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8934bb89-a70e-4a81-abf0-8b0d91106f7a · outbound
EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Knowing where to focus: Event-aware transformer for video grounding
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 723fb3c1-3723-4537-9fdc-98c5f7f3ef81 · outbound
EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Transferable video moment localization by moment-guided query prompting
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f37e56a-b779-4a55-9f1d-30b5eff385c8 · outbound
EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Map the flow: Revealing hidden pathways of information in videollms
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f65293df-07ed-496b-af53-97c37d82d410 · outbound
EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Conditional object-centric learning from video
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90325556-c64b-4282-aec6-8767828413cc · outbound
EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding The hungarian method for the assignment problem.Naval research logistics quarterly, 2 (1-2):83–97, 1955
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9563e9a-d80b-4fe2-838e-1ebef6ced8e5 · outbound
EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Curriculum multi-negative augmentation for debiased video grounding
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e7df4e1-3df4-4e26-b6fc-7ebe2ee5a836 · outbound
EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Detecting moments and highlights in videos via natural language queries
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a692466-f2df-4439-ad39-88ed50ffa580 · outbound
EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Revealing single frame bias for video-and-language learning
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a930184-552d-4f14-af62-27d77950007d · outbound
EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding CORE: Compact object-centric representations as a new paradigm for token merging in lvlms
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69139d4f-7d97-4955-b23b-2f525af80212 · outbound
EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Compositional temporal grounding with structured variational cross-graph correspondence learning
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 375b91c7-9253-4990-8f70-83a7d18aff17 · outbound
EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4fde34e-033f-4457-b21e-3363373f1183 · outbound
EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Resound: Towards action recognition without representation bias
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a58bed83-be5d-4fbb-b9da-adf6d596d1ad · outbound
EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Universal video temporal grounding with generative multi-modal large language models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3ae0a3c-449e-4365-b218-ec2a3f80d7f8 · outbound
EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Univtg: Towards unified video-language temporal grounding
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf46a071-579e-480b-8709-c60a25f820e6 · outbound
EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding VideoMind: A chain-of-lora agent for temporal-grounded video reasoning
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc227c4b-e38b-488f-a021-43c351df4d89 · outbound
EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Object-centric learning with slot attention
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf459cf9-1452-4311-bebd-9f50e9d0e620 · outbound
EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Decoupled weight decay regularization
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 379328b2-8db1-419d-8dd3-8b967bdd6c36 · outbound
EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Chrono: A simple blueprint for representing time in mllms
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b763c0e-541e-4203-ac15-37ad7f6b8df7 · outbound
EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Correlation-Guided Query-Dependency Calibration for Video Temporal Grounding
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation b9b0e86b-e7df-4266-9ba5-d48756fc7982 · outbound
EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Query-dependent video representation for moment retrieval and highlight detection
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 903db503-eafe-4923-ac4d-c891d729331b · outbound
EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Interventional video grounding with dual contrastive learning
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7dc02467-a97f-4efd-90a7-b3d1fd6c2fc8 · outbound
EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding DINOv2: Learning Robust Visual Features without Supervision
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation b51c7ff8-b196-4780-89cf-a98431081e1b · outbound
EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Uncovering Hidden Challenges in Query-Based Video Moment Retrieval
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation cf729709-b1e9-4dee-9b22-c2b0d07d5c1d · outbound
EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Bias-conflict sample synthesis and adversarial removal debias strategy for temporal sentence grounding in video
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9f3256e-b22b-4bc8-ba66-51bf5bc7c03f · outbound
EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Timechat: A time-sensitive multimodal large language model for long video understanding
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8cef6afb-f290-4e59-85f3-8bf5e7a37f8e · outbound
EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Bridging the gap to real-world object-centric learning
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation daea9e2a-da62-4bde-ab80-9c74af55d450 · outbound
EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Tr-detr: Task-reciprocal transformer for joint moment retrieval and highlight detection
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5c6d972-d062-4d48-b9e8-16059f2fc499 · outbound
EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Time-R1: Post-training large vision language model for temporal video grounding
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25b9edde-1c18-4302-88f3-8f11480aea48 · outbound
EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding HawkEye: Training Video-Text LLMs for Grounding Text in Videos
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 17ffea0d-fa8c-4efd-988c-c7971b4bb70e · outbound
EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Slotformer: Unsupervised visual dynamics simulation with object-centric models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b828490-9ea3-452f-9aad-61a9a16dc715 · outbound
EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Slot-VLM: Object-event slots for video-language modeling
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02957ce7-fc70-4ba5-9778-d81679611481 · outbound
EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Qwen3 Technical Report
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation a5deb9da-0e8e-4827-a1cd-22a3a734134b · outbound
EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding AIM: Adapting image models for efficient video understanding
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f070caf-f161-4b0c-bdc8-da23cb92c6ad · outbound
EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Deconfounded video moment retrieval with causal intervention
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8257b31-6366-4abc-b33d-c584e10d86e2 · outbound
EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding A closer look at temporal sentence grounding in videos: Dataset and metric
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4f3e902-df82-4209-b22c-5d6b2bd34537 · outbound
EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding TimeSuite: Improving MLLMs for long video understanding via grounded tuning
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f3acf0e-1cc0-4948-88b0-115147b5747c · outbound
EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Timelens: Rethinking video temporal grounding with multimodal llms
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7b498b9-2a23-447d-9ab2-c9eb44e80cb0 · outbound
EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation fe8f4150-5d4c-4048-80d7-47c811c70704 · outbound
EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Unresolved cited work
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 396009c7-bbe6-4f50-90ec-c873073f8782 · outbound
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 553b4387-f131-4fa3-a626-0405bd14c439 · outbound
EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Unresolved cited work
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a3d6e56-0b37-4848-88e0-2b8ae83a6e55 · outbound
EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Unresolved cited work
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5343485c-65d7-43cf-8324-fe9160be530d · outbound
EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Unresolved cited work
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19d6cd0d-e4df-4919-8ba1-f864f3c51868 · outbound
EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding a person opens the refrigerator in the kitchen
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93dea148-fa54-486e-8d3b-e4901bff6b7c · inbound
Beyond Frame Selection: Generative Latent Evidence Aggregation for Long-Video Understanding EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.