Pith. sign in

Paper Citation Record · LEDGER

EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding

As of 19 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 1 inbound Pith citation observation for arXiv:2605.26104.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.26104 v1

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-29T22:27:37.631870Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-31T04:55:06.976425Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

59 of 59 outbound references displayed

  • verified exact8
  • verified fuzzy0
  • unresolved50
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 59683863-dfe6-4790-90cf-304ea96d036e · outbound

This paper cites Slot-guided adaptation of pre-trained diffusion models for object-centric learning and compositional generation.

EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Slot-guided adaptation of pre-trained diffusion models for object-centric learning and compositional generation

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-29T22:27:37.631870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:27:37.631870Z digest=sha256:c1ef8d28d3e92d7565606457f7e041a3bb28ca9a91262259317275829061e0b3

Observation ee1662a8-c807-4b1d-9574-9151eccb8ddc · outbound

This paper cites DEVIAS: Learning disentangled video representations of action and scene.

EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding DEVIAS: Learning disentangled video representations of action and scene

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-29T22:27:37.631870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:27:37.631870Z digest=sha256:1b5c54acfeb1bd80cb6930b802b1513b682314996f7fe93a46cd77fee2ef9697

Observation 42fb294e-730e-4ae0-bbd1-03ff9f58cb1b · outbound

This paper cites Qwen2.5-VL Technical Report.

EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Qwen2.5-VL Technical Report

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:34:01.977971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-29T22:27:37.631870Z digest=sha256:ff354d1043277f51d80ad421221fa5f50ead4257c669434d35e5db45bad87220

Observation ca6b291a-1416-42c1-92f3-b15c87336c8c · outbound

This paper cites Learning sample importance for cross-scenario video temporal grounding.

EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Learning sample importance for cross-scenario video temporal grounding

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-29T22:27:37.631870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:27:37.631870Z digest=sha256:5164d5ad2ea9e735004a327c03b7fac6e7b353da591b97296affd698bb24b00f

Observation 59ff1312-5ae6-4bc5-9370-b7f69fd956df · outbound

This paper cites End-to-end object detection with transformers.

EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding End-to-end object detection with transformers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-29T22:27:37.631870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:27:37.631870Z digest=sha256:4c851e694aa3be2d7058f4dd4cfba7eb299181516c7ad530f7a5d8ad99d76f9c

Observation fbceec8d-df5a-4211-b19e-211b24a6ccb7 · outbound

This paper cites Towards a complete benchmark on video moment localization.

EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Towards a complete benchmark on video moment localization

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-29T22:27:37.631870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:27:37.631870Z digest=sha256:ab326dec6070ee9489be0463594c7e6972b575c8f00c5e9b66fe4f11cf66e120

Observation bf929a28-c1f6-430a-ab45-9b10808689cd · outbound

This paper cites Slot-MLLM: Object-Centric Visual Tokenization for Multimodal LLM.

EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Slot-MLLM: Object-Centric Visual Tokenization for Multimodal LLM

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:34:01.955333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-29T22:27:37.631870Z digest=sha256:96b84529ae3d468f7a063def4656996a7531ee7fa0c5e1172d814ff2ea0f3f43

Observation 47b534db-961d-412f-97cc-59f28e61395b · outbound

This paper cites Learning phrase representations using RNN encoder–decoder for statistical machine translation.

EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Learning phrase representations using RNN encoder–decoder for statistical machine translation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-29T22:27:37.631870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:27:37.631870Z digest=sha256:834fef19eb3e68bcde9a7112c14f41f30d6de994fada37afcd99e21fdb6545da

Observation d8223d8e-54e0-4287-af27-dfc9a7d6fb93 · outbound

This paper cites Why can’t i dance in the mall? learning to mitigate scene bias in action recognition.

EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Why can’t i dance in the mall? learning to mitigate scene bias in action recognition

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-29T22:27:37.631870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:27:37.631870Z digest=sha256:673d87749d6fbf778613b000d079b02dbc388c855dafc08804824b4a010b0895

Observation 2fbd0d8e-10ef-4f2e-9984-be8c0e61d8e7 · outbound

This paper cites Tall: Temporal activity localization via language query.

EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Tall: Temporal activity localization via language query

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-29T22:27:37.631870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:27:37.631870Z digest=sha256:9ec27bf0888880bd68e40b2540be4dab730e3f400bce597c64fe52d7378e0ba2

Observation c773432e-e090-455a-a8f9-3f7d01ba95db · outbound

This paper cites TRACE: Temporal Grounding Video LLM via Causal Event Modeling.

EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding TRACE: Temporal Grounding Video LLM via Causal Event Modeling

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:34:01.960099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-29T22:27:37.631870Z digest=sha256:04b55e26ab6c7a09fcf0dd4d0b368321aac720b41a0b70595ec43234f1c24cee

Observation 32172b78-bf7a-4188-a47f-1577be444464 · outbound

This paper cites Can shuffling video benefit temporal bias problem: A novel training framework for temporal grounding.

EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Can shuffling video benefit temporal bias problem: A novel training framework for temporal grounding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-29T22:27:37.631870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:27:37.631870Z digest=sha256:9c4898f3b10c7a2bcd8f9fbaec70d4dd2ce8a350e56a95af224b7a351282b31a

Observation 6e01abf9-6027-409d-be35-a7e42d180116 · outbound

This paper cites Localizing moments in video with natural language.

EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Localizing moments in video with natural language

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-29T22:27:37.631870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:27:37.631870Z digest=sha256:d65115fd42a6129ea053528adf15a11a8d56ff2d34aa43d0234c208865d0c6bf

Observation fc866930-3d4d-4710-bbb8-627b48dc0657 · outbound

This paper cites Lora: Low-rank adaptation of large language models.

EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Lora: Low-rank adaptation of large language models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-29T22:27:37.631870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:27:37.631870Z digest=sha256:95f0a067c276fc0dec8850af3247f3a57d851a924008fc6d8b10cec27e64f49e

Observation 628cb1d9-ecfd-4888-841d-53da34c13bf6 · outbound

This paper cites Vtimellm: Empower llm to grasp video moments.

EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Vtimellm: Empower llm to grasp video moments

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-29T22:27:37.631870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:27:37.631870Z digest=sha256:8658dc86d7167354499215c014d5c631e435bc5c51534e344c525ba463488b89

Observation 8934bb89-a70e-4a81-abf0-8b0d91106f7a · outbound

This paper cites Knowing where to focus: Event-aware transformer for video grounding.

EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Knowing where to focus: Event-aware transformer for video grounding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-29T22:27:37.631870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:27:37.631870Z digest=sha256:610ed65ca7a8d049de63f8abaa0a97ecce1577c5ed5e42009de284991d21cb03

Observation 723fb3c1-3723-4537-9fdc-98c5f7f3ef81 · outbound

This paper cites Transferable video moment localization by moment-guided query prompting.

EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Transferable video moment localization by moment-guided query prompting

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-29T22:27:37.631870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:27:37.631870Z digest=sha256:cbeb9e9297ad4e38098379cb3391c8883973bb32d014a95ecdfc621b848cee32

Observation 7f37e56a-b779-4a55-9f1d-30b5eff385c8 · outbound

This paper cites Map the flow: Revealing hidden pathways of information in videollms.

EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Map the flow: Revealing hidden pathways of information in videollms

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-29T22:27:37.631870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:27:37.631870Z digest=sha256:23fe9e7c0e9066c142892f6dc05f8a72ad42a419674c150d9c915892a89a06fb

Observation f65293df-07ed-496b-af53-97c37d82d410 · outbound

This paper cites Conditional object-centric learning from video.

EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Conditional object-centric learning from video

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-29T22:27:37.631870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:27:37.631870Z digest=sha256:65feeeef331a7692ae58e6387aa15a1339ed1e12f15971a02c14e2823828ef94

Observation 90325556-c64b-4282-aec6-8767828413cc · outbound

This paper cites The hungarian method for the assignment problem.Naval research logistics quarterly, 2 (1-2):83–97, 1955.

EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding The hungarian method for the assignment problem.Naval research logistics quarterly, 2 (1-2):83–97, 1955

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-29T22:27:37.631870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:27:37.631870Z digest=sha256:01b8c8b8a74ba16b1b2357afa60d3deb902e7f57508de2990571edd883ad94b1

Observation a9563e9a-d80b-4fe2-838e-1ebef6ced8e5 · outbound

This paper cites Curriculum multi-negative augmentation for debiased video grounding.

EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Curriculum multi-negative augmentation for debiased video grounding

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-29T22:27:37.631870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:27:37.631870Z digest=sha256:39be088842cfe46110d041cf711fb0285a24a4e24d8e86622fe74e988ecbe4c4

Observation 8e7df4e1-3df4-4e26-b6fc-7ebe2ee5a836 · outbound

This paper cites Detecting moments and highlights in videos via natural language queries.

EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Detecting moments and highlights in videos via natural language queries

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-29T22:27:37.631870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:27:37.631870Z digest=sha256:e6065da4887f9c245e8c33b84da46fa5de3090135e8f9208e7a713bdc93271a1

Observation 7a692466-f2df-4439-ad39-88ed50ffa580 · outbound

This paper cites Revealing single frame bias for video-and-language learning.

EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Revealing single frame bias for video-and-language learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-29T22:27:37.631870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:27:37.631870Z digest=sha256:7c3a9773cf2bb67c2950abe5b9be57cceabb279d062859fbedec2ea517234e7c

Observation 6a930184-552d-4f14-af62-27d77950007d · outbound

This paper cites CORE: Compact object-centric representations as a new paradigm for token merging in lvlms.

EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding CORE: Compact object-centric representations as a new paradigm for token merging in lvlms

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-29T22:27:37.631870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:27:37.631870Z digest=sha256:1c7cf3b88f4ab3d9ed626875e97547d75188b50905de5c924b2e9de01a411db2

Observation 69139d4f-7d97-4955-b23b-2f525af80212 · outbound

This paper cites Compositional temporal grounding with structured variational cross-graph correspondence learning.

EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Compositional temporal grounding with structured variational cross-graph correspondence learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-29T22:27:37.631870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:27:37.631870Z digest=sha256:bf0f411c436d9e2725c1d4072b87ac7effb619b9408d8e0cc716420a9f18960e

Observation 375b91c7-9253-4990-8f70-83a7d18aff17 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-06-29T22:27:37.631870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:27:37.631870Z digest=sha256:6300f2ef56a92ef89441c017dfa0f8803d809431b93cbf53bef361d2a16f3b2f

Observation b4fde34e-033f-4457-b21e-3363373f1183 · outbound

This paper cites Resound: Towards action recognition without representation bias.

EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Resound: Towards action recognition without representation bias

Reference 27

Resolution
unresolved
no resolver link, observed 2026-06-29T22:27:37.631870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:27:37.631870Z digest=sha256:a808f5bf5fab2b2111ded0b045d5f64af0dc1c42b259152da7964fb71d43d353

Observation a58bed83-be5d-4fbb-b9da-adf6d596d1ad · outbound

This paper cites Universal video temporal grounding with generative multi-modal large language models.

EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Universal video temporal grounding with generative multi-modal large language models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-29T22:27:37.631870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:27:37.631870Z digest=sha256:6857d8dd91d06fdb3172e030275d2428f0f8a43d485501fd9733db45b5b89748

Observation d3ae0a3c-449e-4365-b218-ec2a3f80d7f8 · outbound

This paper cites Univtg: Towards unified video-language temporal grounding.

EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Univtg: Towards unified video-language temporal grounding

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-29T22:27:37.631870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:27:37.631870Z digest=sha256:0bb434a6eac22205b9b7e99191d6569ad79b4a67b393f6b0a3979c3b86164553

Observation cf46a071-579e-480b-8709-c60a25f820e6 · outbound

This paper cites VideoMind: A chain-of-lora agent for temporal-grounded video reasoning.

EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding VideoMind: A chain-of-lora agent for temporal-grounded video reasoning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-06-29T22:27:37.631870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:27:37.631870Z digest=sha256:1ccb507c3617792c01b0fa56b17c75d5dcf76a145cf30d513e8cc8cb091bfa20

Observation dc227c4b-e38b-488f-a021-43c351df4d89 · outbound

This paper cites Object-centric learning with slot attention.

EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Object-centric learning with slot attention

Reference 31

Resolution
unresolved
no resolver link, observed 2026-06-29T22:27:37.631870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:27:37.631870Z digest=sha256:2fc2e359747eff6ca6d26ab0246af09a8bf23f0651a8005e7603cf067cb24cf2

Observation cf459cf9-1452-4311-bebd-9f50e9d0e620 · outbound

This paper cites Decoupled weight decay regularization.

EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Decoupled weight decay regularization

Reference 32

Resolution
unresolved
no resolver link, observed 2026-06-29T22:27:37.631870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:27:37.631870Z digest=sha256:bea23e9751f66b0ebf9098402c2833c1137f7bd100d2a97e534aeb847e0be9c9

Observation 379328b2-8db1-419d-8dd3-8b967bdd6c36 · outbound

This paper cites Chrono: A simple blueprint for representing time in mllms.

EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Chrono: A simple blueprint for representing time in mllms

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-29T22:27:37.631870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:27:37.631870Z digest=sha256:6fa70b063d9ad09450f7073bb92f443ea56bb8593c4bb0ed6bb95dcd0f03333d

Observation 2b763c0e-541e-4203-ac15-37ad7f6b8df7 · outbound

This paper cites Correlation-Guided Query-Dependency Calibration for Video Temporal Grounding.

EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Correlation-Guided Query-Dependency Calibration for Video Temporal Grounding

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:34:01.973699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-29T22:27:37.631870Z digest=sha256:6a9d73d5c3d513aedcb5d9cf926150860be0c6350d15fa4489cbe1b86bde7057

Observation b9b0e86b-e7df-4266-9ba5-d48756fc7982 · outbound

This paper cites Query-dependent video representation for moment retrieval and highlight detection.

EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Query-dependent video representation for moment retrieval and highlight detection

Reference 35

Resolution
unresolved
no resolver link, observed 2026-06-29T22:27:37.631870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:27:37.631870Z digest=sha256:edc8ae5abc1c7a5a2c3f3046f04eca7eb1d4adbc771dbfe18347f3244f6a0018

Observation 903db503-eafe-4923-ac4d-c891d729331b · outbound

This paper cites Interventional video grounding with dual contrastive learning.

EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Interventional video grounding with dual contrastive learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-06-29T22:27:37.631870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:27:37.631870Z digest=sha256:7b9766bb3a0f68a67331d10eab13b130d07752341ee3bb66d130389c7596cdda

Observation 7dc02467-a97f-4efd-90a7-b3d1fd6c2fc8 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding DINOv2: Learning Robust Visual Features without Supervision

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:34:01.950495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-29T22:27:37.631870Z digest=sha256:b96c506e3f99dcc3fbe68094c144fafd787fa953ca0a710813394e2f63a784da

Observation b51c7ff8-b196-4780-89cf-a98431081e1b · outbound

This paper cites Uncovering Hidden Challenges in Query-Based Video Moment Retrieval.

EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Uncovering Hidden Challenges in Query-Based Video Moment Retrieval

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:34:01.976448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-29T22:27:37.631870Z digest=sha256:969deee82532243ae4d4015da7826f041ace519bec0dc9bd03414807834aad60

Observation cf729709-b1e9-4dee-9b22-c2b0d07d5c1d · outbound

This paper cites Bias-conflict sample synthesis and adversarial removal debias strategy for temporal sentence grounding in video.

EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Bias-conflict sample synthesis and adversarial removal debias strategy for temporal sentence grounding in video

Reference 39

Resolution
unresolved
no resolver link, observed 2026-06-29T22:27:37.631870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:27:37.631870Z digest=sha256:fd53d676e2209591a54796d55cf56c5268437ff20560e5d509ff9d497674a722

Observation e9f3256e-b22b-4bc8-ba66-51bf5bc7c03f · outbound

This paper cites Timechat: A time-sensitive multimodal large language model for long video understanding.

EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Timechat: A time-sensitive multimodal large language model for long video understanding

Reference 40

Resolution
unresolved
no resolver link, observed 2026-06-29T22:27:37.631870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:27:37.631870Z digest=sha256:83f9c086a7367be8356d2902eca2d71d57de8e90983d3223eca319afb32304b4

Observation 8cef6afb-f290-4e59-85f3-8bf5e7a37f8e · outbound

This paper cites Bridging the gap to real-world object-centric learning.

EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Bridging the gap to real-world object-centric learning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-06-29T22:27:37.631870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:27:37.631870Z digest=sha256:fc2125adbdcc186fd8f03c4035f37863a3a2f27f791a0e30d4a1cbc60ff9af05

Observation daea9e2a-da62-4bde-ab80-9c74af55d450 · outbound

This paper cites Tr-detr: Task-reciprocal transformer for joint moment retrieval and highlight detection.

EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Tr-detr: Task-reciprocal transformer for joint moment retrieval and highlight detection

Reference 42

Resolution
unresolved
no resolver link, observed 2026-06-29T22:27:37.631870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:27:37.631870Z digest=sha256:eb26c3a53ad042ed3075e695298f0838ac72b4e7c6229a2849992bb1e8057203

Observation e5c6d972-d062-4d48-b9e8-16059f2fc499 · outbound

This paper cites Time-R1: Post-training large vision language model for temporal video grounding.

EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Time-R1: Post-training large vision language model for temporal video grounding

Reference 43

Resolution
unresolved
no resolver link, observed 2026-06-29T22:27:37.631870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:27:37.631870Z digest=sha256:4d5b6632dd1bac885f22ec751ec017afc02b644023fed09070203fb42f8faa23

Observation 25b9edde-1c18-4302-88f3-8f11480aea48 · outbound

This paper cites HawkEye: Training Video-Text LLMs for Grounding Text in Videos.

EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:34:01.972853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-29T22:27:37.631870Z digest=sha256:e85723535b8d11b81def9af3cad35e57a17264d858f3cb1cbd8e17f7a6acd8b3

Observation 17ffea0d-fa8c-4efd-988c-c7971b4bb70e · outbound

This paper cites Slotformer: Unsupervised visual dynamics simulation with object-centric models.

EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Slotformer: Unsupervised visual dynamics simulation with object-centric models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-06-29T22:27:37.631870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:27:37.631870Z digest=sha256:0c87ca187a61eb5e112d75f81ba9329a436ebf52bffd8595a3068871184c3bec

Observation 1b828490-9ea3-452f-9aad-61a9a16dc715 · outbound

This paper cites Slot-VLM: Object-event slots for video-language modeling.

EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Slot-VLM: Object-event slots for video-language modeling

Reference 46

Resolution
unresolved
no resolver link, observed 2026-06-29T22:27:37.631870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:27:37.631870Z digest=sha256:fec9280dd19351c312f651878cd761b4902f9b37008fd90390bb11ff2206e0df

Observation 02957ce7-fc70-4ba5-9778-d81679611481 · outbound

This paper cites Qwen3 Technical Report.

EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Qwen3 Technical Report

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:34:01.968738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-29T22:27:37.631870Z digest=sha256:302909db547c18ce9efa96c4a0fa09506f96c6a55c3249cd78cc8e7f61d16647

Observation a5deb9da-0e8e-4827-a1cd-22a3a734134b · outbound

This paper cites AIM: Adapting image models for efficient video understanding.

EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding AIM: Adapting image models for efficient video understanding

Reference 48

Resolution
unresolved
no resolver link, observed 2026-06-29T22:27:37.631870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:27:37.631870Z digest=sha256:7cb8960eb02bbafed9c56363eec75302aa0ff601d24a16591d1d885d471537cf

Observation 1f070caf-f161-4b0c-bdc8-da23cb92c6ad · outbound

This paper cites Deconfounded video moment retrieval with causal intervention.

EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Deconfounded video moment retrieval with causal intervention

Reference 49

Resolution
unresolved
no resolver link, observed 2026-06-29T22:27:37.631870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:27:37.631870Z digest=sha256:eb16a408806f40348ae299d69ade0bdcdf7d671eb0bc9fb9985c91b686e9e15c

Observation d8257b31-6366-4abc-b33d-c584e10d86e2 · outbound

This paper cites A closer look at temporal sentence grounding in videos: Dataset and metric.

EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding A closer look at temporal sentence grounding in videos: Dataset and metric

Reference 50

Resolution
unresolved
no resolver link, observed 2026-06-29T22:27:37.631870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:27:37.631870Z digest=sha256:052fad14e73edc756af5513c9ca06dbe96b989bda3c740dc00770d7812271bb5

Observation f4f3e902-df82-4209-b22c-5d6b2bd34537 · outbound

This paper cites TimeSuite: Improving MLLMs for long video understanding via grounded tuning.

EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding TimeSuite: Improving MLLMs for long video understanding via grounded tuning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-06-29T22:27:37.631870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:27:37.631870Z digest=sha256:2760a96b2d44c94d17a1dbaa1b1f4268d17bd42f1be9c1116a84de3c37627cbb

Observation 3f3acf0e-1cc0-4948-88b0-115147b5747c · outbound

This paper cites Timelens: Rethinking video temporal grounding with multimodal llms.

EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Timelens: Rethinking video temporal grounding with multimodal llms

Reference 52

Resolution
unresolved
no resolver link, observed 2026-06-29T22:27:37.631870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:27:37.631870Z digest=sha256:48bfcde47c298643b695ad5579796e5d802edc842907d21377cf008365a14bdc

Observation c7b498b9-2a23-447d-9ab2-c9eb44e80cb0 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 53

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T22:34:01.967411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-29T22:27:37.631870Z digest=sha256:92d00b6ddc3d230c0ac2cc3d38a40758c49f984b863d4893aa6b6207833c12d7

Observation fe8f4150-5d4c-4048-80d7-47c811c70704 · outbound

This paper cites an unresolved cited work.

EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Unresolved cited work

Reference 54

Resolution
unresolved
no resolver link, observed 2026-06-29T22:27:37.631870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:27:37.631870Z digest=sha256:6e59d6cd24ba00aecb9f8195f8386415354cca967a7b5a8677ab930844c8b0f5

Observation 396009c7-bbe6-4f50-90ec-c873073f8782 · outbound

This paper cites Subject:.

EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Subject:

Reference 55

Resolution
unresolved
no resolver link, observed 2026-06-29T22:27:37.631870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:27:37.631870Z digest=sha256:7b38db0ea087737bf1642593f63d6b1fb660791e9ec7def075b4090574b729fd

Observation 553b4387-f131-4fa3-a626-0405bd14c439 · outbound

This paper cites an unresolved cited work.

EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Unresolved cited work

Reference 56

Resolution
unresolved
no resolver link, observed 2026-06-29T22:27:37.631870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:27:37.631870Z digest=sha256:ad46e44356fcf925711acd67e2c5917079da331f0ec50875187b15324f1423c8

Observation 0a3d6e56-0b37-4848-88e0-2b8ae83a6e55 · outbound

This paper cites an unresolved cited work.

EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Unresolved cited work

Reference 57

Resolution
unresolved
no resolver link, observed 2026-06-29T22:27:37.631870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:27:37.631870Z digest=sha256:87dd41e056933f1a12a1bd7a4bf5dc50f39c675ba6aa1ca1964a9469a03f95be

Observation 5343485c-65d7-43cf-8324-fe9160be530d · outbound

This paper cites an unresolved cited work.

EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding Unresolved cited work

Reference 58

Resolution
unresolved
no resolver link, observed 2026-06-29T22:27:37.631870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:27:37.631870Z digest=sha256:0eaf7ac46eb060bcbf60236b3851af6508746489ab97c27eba7b592de70529c8

Observation 19d6cd0d-e4df-4919-8ba1-f864f3c51868 · outbound

This paper cites a person opens the refrigerator in the kitchen.

EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding a person opens the refrigerator in the kitchen

Reference 59

Resolution
unresolved
no resolver link, observed 2026-06-29T22:27:37.631870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:27:37.631870Z digest=sha256:7261b2b648141ea7a8f9664aae796c88ccbdf441871deba1b44fb49cd786519d

Pith citing papers

Observation 93dea148-fa54-486e-8d3b-e4901bff6b7c · inbound

Beyond Frame Selection: Generative Latent Evidence Aggregation for Long-Video Understanding cites this paper.

Beyond Frame Selection: Generative Latent Evidence Aggregation for Long-Video Understanding EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-31T04:55:06.976425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T04:55:06.976425Z digest=sha256:36f94bb8f7eb648bc28fe99077155c41736f1ef58eceaafd3033860b2571e3ae