Pith. sign in

Paper Citation Record · LEDGER

A Survey on Video Temporal Grounding with Multimodal Large Language Model

As of 20 August 2026, this Paper Citation Record lists 100 of 164 outbound references and 3 inbound Pith citation observations for arXiv:2508.10922.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.10922 v1

Coverage vector

measured 100 of 164 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T23:32:17.911414Z

measured 103 of 103 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T13:27:58.123306Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T03:29:31.264199Z

Reference resolution

100 of 164 outbound references displayed

  • verified exact6
  • verified fuzzy0
  • unresolved94
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d47efee2-bed4-4c74-8608-84941efa5a9a · outbound

This paper cites an unresolved cited work.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.435157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.435157Z digest=sha256:33a5cac67020746da5b42a57c4e31da9dd1a0747dbcd3f629b16ec57a08d3e6a

Observation b6867581-5b02-4758-9364-ed0c7a825638 · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

A Survey on Video Temporal Grounding with Multimodal Large Language Model InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.441242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.441242Z digest=sha256:9f65b43247d9f9dedad570e70c313531ea4e9f9e26946b0a674d77c535fac3db

Observation f1c7ca98-9471-4fc4-aba7-1f20aa7d5cbd · outbound

This paper cites an unresolved cited work.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.446542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.446542Z digest=sha256:1f2b90f7b99fa8c3570ed435be22617e35b7452cc4a8e428caea46b5f7567ae5

Observation 710efd66-7ff8-480c-b180-e74d17545b68 · outbound

This paper cites an unresolved cited work.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.451571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.451571Z digest=sha256:73eccde0bfed9086674f226a2a1caa53a16a023413d5e61604a593f7eb45daa9

Observation d3dd2b83-3a52-41f0-b953-5241e969b2e0 · outbound

This paper cites an unresolved cited work.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.458333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.458333Z digest=sha256:28a81fa5af4a4ec762aecb93b70c7d3fb022feda4c9ec2b743cec48068cb8a69

Observation 95f37fb3-323d-4daf-9765-a459f1fadd24 · outbound

This paper cites Regneri, M.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Regneri, M

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.464199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.464199Z digest=sha256:c4a2c98818d5dc4b8b83395bae669b71da729d450a0f3a3661599b62caa8dcdd

Observation b3dd524e-b921-401e-99ae-cc67967b3084 · outbound

This paper cites Anne Hendricks, O.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Anne Hendricks, O

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.469635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.469635Z digest=sha256:2efcae295ae4f8d922f49cd4585e3ce196f21c55672fc1cddd62bb57acf7872a

Observation 161fa926-d0a6-4d94-81f0-88a26f8fd460 · outbound

This paper cites Krishna, K.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Krishna, K

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.475076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.475076Z digest=sha256:0cb54c08b594c4c53740725b343958b102e601365e0a35a49fb330e31d36e314

Observation f0bdd02f-48d0-4c82-86ba-4d65edb917a7 · outbound

This paper cites an unresolved cited work.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.479960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.479960Z digest=sha256:9ce4c805174074ce7c66e06f14bfa1497f865d7601b69f61099598b724ca2228

Observation 49472454-c85a-4843-8d08-80e58e3c1e08 · outbound

This paper cites an unresolved cited work.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.484812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.484812Z digest=sha256:7e63d01176c59bd88dc8d1a76e48101e0d49688ab6f282ee86e583878a026195

Observation b64cca6f-fa04-4422-bbd6-4e97073754eb · outbound

This paper cites Xiong, Y.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Xiong, Y

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.489327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.489327Z digest=sha256:9a2e97c28f03649121081e083e1894c51e797dad21966b2574c3d4424d492078

Observation 07245b3a-70d3-41fa-9714-4e67115b95c7 · outbound

This paper cites an unresolved cited work.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.493949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.493949Z digest=sha256:907f892266370ca4423fe82529b6dcc2bba17a9db1847d42cf0c1d0c1018a72f

Observation 066bf953-70a3-4a21-b9b4-ef8ba83ced6e · outbound

This paper cites Chen, Y.-C.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Chen, Y.-C

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.500829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.500829Z digest=sha256:38f1ca617524f58161161d575e189cb3a7e87e2dc447c4f3629e7b2dcf4ef0df

Observation 3e5bc3c0-2e9a-4d2e-b5e2-a760029593be · outbound

This paper cites Zhang, X.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Zhang, X

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.505641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.505641Z digest=sha256:b6ecddebcbf5e9cfbac7c9c198e0ca95a630d958c96ca68eca213deafc515789

Observation 0d5c2cb1-6af6-4817-8dd1-d749ecdf7b34 · outbound

This paper cites an unresolved cited work.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.510006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.510006Z digest=sha256:beccbfb4273180161ac971c7d7b74375ac38d49946436ba8c1483a99c634a5db

Observation d7347c27-3b3c-47fa-8d06-1442d2c1b9b8 · outbound

This paper cites an unresolved cited work.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.514435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.514435Z digest=sha256:2cc12bc2c081724f0278b7c148a42a0e35add465eb00976067a296892a7822f7

Observation 1964705f-3d3e-41aa-b950-c773e4027ee4 · outbound

This paper cites Zhang, Y.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Zhang, Y

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.519062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.519062Z digest=sha256:0d53b344a9ac8b00663782a16a01567c97bc96372109f9938d9d112b1ae51ba8

Observation 358ddc0a-c1fb-498f-8ba1-e81a049dc7bd · outbound

This paper cites an unresolved cited work.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.523349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.523349Z digest=sha256:a610f0f55ba028c1bf4e966564dc72d2ae9c3228824a22461aa8c22b25db6892

Observation 1cfede5a-2869-407c-a0a5-f8716337bc1f · outbound

This paper cites an unresolved cited work.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.527721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.527721Z digest=sha256:e0f139d5499d59c91d548e1104b1c91188a24ba1b1ef35d7a76c5fe911789a0c

Observation 65bd9630-41a3-46b9-916e-bed277a27ef1 · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

A Survey on Video Temporal Grounding with Multimodal Large Language Model OPT: Open Pre-trained Transformer Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.532947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.532947Z digest=sha256:608c3953509e5282c6cb9a3f9d951113b44a8677775b77ec59bedcb89c74fd21

Observation 2bcbd8b2-c144-4457-a91c-3df27b2dc296 · outbound

This paper cites an unresolved cited work.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.538566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.538566Z digest=sha256:d94acaa3a467840e0f3561e33572637ef84abd01f414094539db786c12933c27

Observation 907589c4-ac8a-4711-86b1-82450819d6cd · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

A Survey on Video Temporal Grounding with Multimodal Large Language Model DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.543808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.543808Z digest=sha256:a0a369f78e786c1e0ac632090cdc34e1377301f107a6f65b4013fcd4eb2eea95

Observation 33dc5993-1d22-4d4d-ad40-f50ad8e7cc77 · outbound

This paper cites an unresolved cited work.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.550482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.550482Z digest=sha256:cb5b1c4f2890f8af8d4f0dc8a98ca76a87030f1ea59918e051efbeb92a54cd08

Observation f3e11a11-382d-4dc5-a13a-9dc3717a4ab0 · outbound

This paper cites an unresolved cited work.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.555595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.555595Z digest=sha256:78efe05ba7b88f56c429d586f6f57c79c6a940b2fbae8eedf44ab06d303ccdb2

Observation a2e85720-5584-4f24-bd07-71b9ebab44f8 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

A Survey on Video Temporal Grounding with Multimodal Large Language Model VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.561810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.561810Z digest=sha256:34799734ee1a58a8569628286bd4bd3403bad12d2e5d052706ff2ed61f4f60a8

Observation de74602a-762c-4ac9-9b46-cca957b72f45 · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

A Survey on Video Temporal Grounding with Multimodal Large Language Model PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.566703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.566703Z digest=sha256:830c7d84d327974241b9343a296200c6316ed11a2ccc3d3559387bc5ab039a79

Observation 44d77a73-e0ac-48f4-822b-474fab802c87 · outbound

This paper cites Huang, X.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Huang, X

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.571330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.571330Z digest=sha256:b57e7eb45635a9c30e4a629f5986ce09dfc7eeee4c7a683facc70e19e54c676f

Observation c8e0f96d-0b6c-408f-aede-f6307fd30b97 · outbound

This paper cites an unresolved cited work.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.576083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.576083Z digest=sha256:2266d2cb5be78aa5299ce541dba0dfd6a532172186ad9be1f06a86ba9eb88aea

Observation d330d962-caea-4a7e-bd1c-fbf39944b073 · outbound

This paper cites an unresolved cited work.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.580641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.580641Z digest=sha256:97bb4241ad9ac211d2ac5d940aa51318d13a8da7b5ffe11a63f5dee2253c4ffc

Observation 21ab5dd0-d1ef-4df7-b288-9b8eb5d86e05 · outbound

This paper cites an unresolved cited work.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.585383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.585383Z digest=sha256:05c30fab4d6e659b76c2914b68b39ce4b1e2702e62ef490c1c576cd418a722b0

Observation 8094023f-5c7d-4f0c-95cf-a041a5174776 · outbound

This paper cites Question-Answering Dense Video Events.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Question-Answering Dense Video Events

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-05T23:32:19.423928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-05T23:32:17.590085Z digest=sha256:ae8a71808b06b2f74d8cb1c280790e5eb0373ebc5159a5ec4f121923b45e90df

Observation 434ed86b-354c-41ac-bbfa-ed746c68d012 · outbound

This paper cites an unresolved cited work.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.594853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.594853Z digest=sha256:2346274e6129ef1e64e8d93ad8321b1cbd4ea0462e8a911bc79b9e78a61ca5d9

Observation 989f777a-e959-4819-8ad1-662748ee862c · outbound

This paper cites HawkEye: Training Video-Text LLMs for Grounding Text in Videos.

A Survey on Video Temporal Grounding with Multimodal Large Language Model HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.599520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.599520Z digest=sha256:f0210221cc40652096bf496ddb50606e8bf535cdd4a2bb8a21e0daa340212ee1

Observation a08790b8-dd72-424b-85b7-b1c9ae0dffc5 · outbound

This paper cites an unresolved cited work.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.604862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.604862Z digest=sha256:26fe86146920b58ee6f89326f85b65fef803cc7130ecd8e26bf8c1e8d35cbc0d

Observation 5ee27c7d-fb51-4241-8eaf-ef831dfba3f3 · outbound

This paper cites Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.609232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.609232Z digest=sha256:0344b3e7cf984728121ab5f6df0ec5bce4c2eca375d9d00372d562440cbb7a20

Observation 95e527c9-5938-419e-8dbf-f4abc8847158 · outbound

This paper cites an unresolved cited work.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.613707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.613707Z digest=sha256:a7a6c4c7428f63da41d68937f61dd27a292f7de21d28208e5bc875c679cd3739

Observation ef7c5fcc-0cc6-4ebd-b0cc-dc5daecb2176 · outbound

This paper cites Multimodal Foundation Models: From Specialists to General-Purpose Assistants.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Multimodal Foundation Models: From Specialists to General-Purpose Assistants

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.617742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.617742Z digest=sha256:974efbe5742b508b405260bab593f86c5f3aeab92a96b92331f75557298446d5

Observation a1115ff8-dc7e-4c5d-b44e-ff31a2438b41 · outbound

This paper cites A Comprehensive Study of Deep Video Action Recognition.

A Survey on Video Temporal Grounding with Multimodal Large Language Model A Comprehensive Study of Deep Video Action Recognition

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.622367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.622367Z digest=sha256:cd2ebe29bebf37de61793c05607e849c732ace2fdd9c636c38c8d74383fa024d

Observation 2975d617-022c-4f22-86bf-9aa15455e0a7 · outbound

This paper cites Abdar, M.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Abdar, M

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.627140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.627140Z digest=sha256:250e09bb04a45001e25706a1b2c72016df7fdbf795ad216cc93f0bc90f7e86f4

Observation 846e7eb6-fa64-48a2-b450-500334a692f1 · outbound

This paper cites an unresolved cited work.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.631444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.631444Z digest=sha256:7ebfac2189b148806a9f6aefab58a17c7eab0947786cd0962e89d2a1ab104a62

Observation d4881420-e7fa-4c70-b3e5-fa63564f43e9 · outbound

This paper cites Zhang, J.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Zhang, J

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.635989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.635989Z digest=sha256:4dda4ddc01d2f8ccdfd4ce198e935b6aebb58bee6b62280a4954264b67449081

Observation 014b09cb-8115-4e53-a693-aabaadb268e0 · outbound

This paper cites A Survey on Natural Language Video Localization.

A Survey on Video Temporal Grounding with Multimodal Large Language Model A Survey on Natural Language Video Localization

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.640437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.640437Z digest=sha256:088f36c0f1bd35a1d76b90d1bb3c436b1d5f694c39f960d40f7f0a7029767303

Observation 21fbba2c-03c7-447c-bdd3-e3ca6c267ddf · outbound

This paper cites an unresolved cited work.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.645031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.645031Z digest=sha256:ede925249c4056a1ac1ae33f892af12b0c885c29ead84617b3dd60a9e6b56cfa

Observation 55320261-df96-4821-9363-464649fb13b0 · outbound

This paper cites an unresolved cited work.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Unresolved cited work

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.649595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.649595Z digest=sha256:4b3655d263a42c45c10426b6b9ece7d530c8c2b168746a14e71fd854e813dabb

Observation 8ffb8ef7-7d1a-4ad8-8f33-38b013ebddb5 · outbound

This paper cites Zhang, A.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Zhang, A

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.654145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.654145Z digest=sha256:c67230db6b8c44373510526092dd1b11c3424274516895856fc0c33a5041f9ea

Observation d2d29c30-0c5b-47d8-9ec2-111db965fa1a · outbound

This paper cites Zhang, H.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Zhang, H

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.658647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.658647Z digest=sha256:29b9df3e75577f1a801831549845d50c2651bfeabaf08ae3f05043b37c90faaf

Observation 559c54d0-d5d1-4c3a-8da5-f6fbbc1c6f72 · outbound

This paper cites an unresolved cited work.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Unresolved cited work

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.663025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.663025Z digest=sha256:f3da2483ece8af4a40e723ba928c55fcf40a34d049b95a4239939f7c73b431e4

Observation a9dbd8d1-b3d8-4adb-be10-5b0c26411e1a · outbound

This paper cites an unresolved cited work.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Unresolved cited work

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.667535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.667535Z digest=sha256:40eda276fa760e7bbd2bb73ee05c30eb3741eb0490be221dd9b7a64fc912e857

Observation c93a3c89-3f11-48d1-9b9f-3b732bd20044 · outbound

This paper cites an unresolved cited work.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Unresolved cited work

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.672297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.672297Z digest=sha256:ce0c03210049eba6e3dea69070331760d2abac8f25f502a6995773ca5f4e0839

Observation 471c449e-de1f-4052-98be-129190e86b9a · outbound

This paper cites an unresolved cited work.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Unresolved cited work

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.676303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.676303Z digest=sha256:e69d0869ae466bd81f30a82eee33495ab6c270d98896fb0a06dcb1084085bd99

Observation 785bfcf4-c5d7-45b1-b340-75bc7c769e3f · outbound

This paper cites Zhang, A.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Zhang, A

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.680772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.680772Z digest=sha256:c4f25042e0ee8139bfb8d6ef8bb210e0ffb83ab3195dacabc1c4bf08ef533977

Observation f610a675-51b8-4e4d-ba0b-882b260d5c68 · outbound

This paper cites an unresolved cited work.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Unresolved cited work

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.684899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.684899Z digest=sha256:0d511e3c4acd01762aecb0a165b56c5efcba84dd60684b00251312b6c234548e

Observation 10c6805a-8eef-4283-ace3-ca14d249584a · outbound

This paper cites an unresolved cited work.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Unresolved cited work

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.689927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.689927Z digest=sha256:51e428ad3b40be9665993b260abd1cffea593781d500314627fba223036c06c4

Observation 544ec9ee-5746-496c-a53c-15a78c722964 · outbound

This paper cites Iashin and E.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Iashin and E

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.693951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.693951Z digest=sha256:39585ccb37ad7eb97bbca3790d253c54749ec31b300675a9de936be3a756871a

Observation 4c3be1a9-f883-4e45-9cc1-4acb42f792f8 · outbound

This paper cites an unresolved cited work.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Unresolved cited work

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.697949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.697949Z digest=sha256:f0ace6fb191b8b9221ba5ec21600b9a738254ff0784c77d22dbd5a56107ae4d7

Observation 19e47eb9-e53e-4ef6-b09f-014ff4bff5ee · outbound

This paper cites CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding.

A Survey on Video Temporal Grounding with Multimodal Large Language Model CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.702112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.702112Z digest=sha256:546344e7794b22bf7cf034dc8fb9101b160a02e9e61fa76a4c93f3ecc34f6da5

Observation eb7c405f-827e-40fe-9f4e-d9a7da9559bd · outbound

This paper cites an unresolved cited work.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Unresolved cited work

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.706691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.706691Z digest=sha256:d634163888e5d8f39e564dc3c897c3eea5624525cba55173975b4ba41487c814

Observation 3d3bdab4-1c93-4a54-b42b-ecc0b7b0bbff · outbound

This paper cites Zhang, X.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Zhang, X

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.711697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.711697Z digest=sha256:cf0cbde5210e5dd403ca0d8efabfcce8c09125c8131a573d60900fb266d7d6ac

Observation c7011462-3c0d-4498-bc56-d0c1e7897347 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.716937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.716937Z digest=sha256:f56ee022be000434e007802d59c7f7c9a6c41d0019e9f93e09248a4f43ae51b2

Observation 03c0c3d3-0db8-4f80-8ec7-9fd97c3db266 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

A Survey on Video Temporal Grounding with Multimodal Large Language Model VideoChat: Chat-Centric Video Understanding

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.721674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.721674Z digest=sha256:3017a4ad8fd3b3bf17a040a9f86d7b77594ef42fc965057b80c357504b619903

Observation 6b544616-83ca-4226-8f75-51a7f3224619 · outbound

This paper cites Dosovitskiy, L.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Dosovitskiy, L

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.726831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.726831Z digest=sha256:501ecd19bf5b8516dbc6846b78f1c9c3b64a4303e34a14169c1d9ce38e4a2193

Observation 3168c6c6-3f0d-4066-88d2-dc460577a500 · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

A Survey on Video Temporal Grounding with Multimodal Large Language Model EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.731677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.731677Z digest=sha256:6ea9d840db6a307de65dd1250521606bdab9743ea1fcaaec863ac3f826a88d75

Observation 47c5f693-624f-4adc-99b1-43066b9d9b89 · outbound

This paper cites an unresolved cited work.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Unresolved cited work

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.736248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.736248Z digest=sha256:b895e9903e01d711ce9051a7667879ddc75aa5a8b0de488e4d41f5483f4515b0

Observation d9769dd9-7495-450c-bd61-cbe0690111c7 · outbound

This paper cites Radford, J.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Radford, J

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.740487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.740487Z digest=sha256:35aa65eeef548374966ddb47757286c9cf3db954f9630594358c6006233d5532

Observation 1d6a3f28-cf46-4f13-be85-7c7b3ff04869 · outbound

This paper cites Feichtenhofer, Y.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Feichtenhofer, Y

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.745075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.745075Z digest=sha256:20584ca152ccbd3e9b65b33b73d18f13cbb54f1609fe64f920c9b1d7f9932428

Observation 1eb1a059-d6de-4815-a9a2-f124e60d9c9a · outbound

This paper cites an unresolved cited work.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Unresolved cited work

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.750415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.750415Z digest=sha256:d0de441e36a493b7a4f9e58d4be485535967b52abc3019185117d9b51e5149fe

Observation 8a9ae72e-bdc2-40b4-b00d-fb97646cdeec · outbound

This paper cites an unresolved cited work.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Unresolved cited work

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.755773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.755773Z digest=sha256:f11b2a0866b5dbf208e8bdaa549c36fbeea586b84fc55abfb4100fd94e5037e1

Observation e0f9fe1b-f90e-4d06-9af1-480b6a5ef781 · outbound

This paper cites an unresolved cited work.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Unresolved cited work

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.760224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.760224Z digest=sha256:b4f50bb1f5b5705f5329e6951765a911a3a18c4117452a176de6f19603ac841b

Observation 6d3917f6-fd6a-4437-b08e-e1c464182e5f · outbound

This paper cites an unresolved cited work.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Unresolved cited work

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.765061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.765061Z digest=sha256:ed2c33ff309ffbd007d34d5bbd8f58eaad161d7519355fe6a03096ea74239b85

Observation aecfe681-e828-44a6-94f9-67d088684a43 · outbound

This paper cites HierarQ: Task-Aware Hierarchical Q-Former for Enhanced Video Understanding.

A Survey on Video Temporal Grounding with Multimodal Large Language Model HierarQ: Task-Aware Hierarchical Q-Former for Enhanced Video Understanding

Reference 70

Resolution
verified exact
local_arxiv, observed 2026-08-05T23:32:19.163503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-05T23:32:17.769503Z digest=sha256:e31769feebef0f6521d6a6aa1605bf9a6666f0e87bc348a569856b35001a82e1

Observation 90b69520-e057-43d9-8cc0-42c77af96758 · outbound

This paper cites Alayrac, J.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Alayrac, J

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.775218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.775218Z digest=sha256:0eb3ca442444ce2448fd8897cdec393549d0c10f4bede67d7c93c907dfe4b0c3

Observation e5f12779-6a3e-4707-ad9e-4d98c238e7a1 · outbound

This paper cites an unresolved cited work.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Unresolved cited work

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.780085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.780085Z digest=sha256:fa3139734b718691806ad62168003e874b5748be3afaf13d8dca439d05cbf8c0

Observation af7cf594-9292-4f89-8d3e-4663faab3b93 · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

A Survey on Video Temporal Grounding with Multimodal Large Language Model mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.785286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.785286Z digest=sha256:3a46e7c96a6ce532a095bc1340c89d34e2c811790bf7645ccdecdb0f0ab7cda6

Observation 795bc96e-a601-40d7-982a-70f0bc170e33 · outbound

This paper cites Huang, L.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Huang, L

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.789902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.789902Z digest=sha256:c9b0e5199b205de17de20347ee2dc2c4a72a6364fffeedb1838b824c76adf3ff

Observation 9c32c71b-70ce-46b2-b471-906b0464179a · outbound

This paper cites an unresolved cited work.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Unresolved cited work

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.794850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.794850Z digest=sha256:8ea8bad054424dcd0a10b83ea95c526c6a49140b80ab754dbe6930ef9e9fb7cd

Observation 8d9f69a4-1136-4732-8936-b794417cb6bd · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

A Survey on Video Temporal Grounding with Multimodal Large Language Model LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.799225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.799225Z digest=sha256:a875d110291d3782fa932312f2e53893c61389e6528b4e2a4f0c19399276a37d

Observation 1949ae06-50bd-4c89-a926-48883de00917 · outbound

This paper cites mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models.

A Survey on Video Temporal Grounding with Multimodal Large Language Model mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.804125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.804125Z digest=sha256:3492facecd89404abacfe70d26be9a23732a404297704c606f8177358ecc5310

Observation 865f3073-7922-4280-8fb8-51adb5bb41d0 · outbound

This paper cites EVLM: An Efficient Vision-Language Model for Visual Understanding.

A Survey on Video Temporal Grounding with Multimodal Large Language Model EVLM: An Efficient Vision-Language Model for Visual Understanding

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.808793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.808793Z digest=sha256:d0648c598c5bcee49bb832ece58e99c5adc962cd675be70652a6cd06e012a8da

Observation 89728ef8-fc4a-4e38-ae24-691b01dab7f8 · outbound

This paper cites Slow-Fast Architecture for Video Multi-Modal Large Language Models.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Slow-Fast Architecture for Video Multi-Modal Large Language Models

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.814422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.814422Z digest=sha256:da7e59ec9fbb08b2c78a2b1c84588a7b3903fdeb11cf6b5ec1e1a6768070de51

Observation 819149ec-2a90-4398-bce9-4c9b07ecff66 · outbound

This paper cites Miech, D.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Miech, D

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.819028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.819028Z digest=sha256:d615fe7b6039ae0fc65ff0f75a4e8b0606c64e738d253e7dcf0bfc455973be45

Observation d9c08757-20ed-4c5c-bfb7-01ad525618cd · outbound

This paper cites an unresolved cited work.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Unresolved cited work

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.823530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.823530Z digest=sha256:bdbaa4d6bec7834d064c0d4605907e8da93076f73033ec94ede2172b5050344e

Observation f1de355d-a83b-4b4c-98ec-849371a1a58c · outbound

This paper cites Schuhmann, R.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Schuhmann, R

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.827672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.827672Z digest=sha256:16c86165a879bcdbd6608950de5ce40534093bd1afd396aaf2cb37fbe21a25f8

Observation 4d0bd7e3-10c6-4259-acb1-704fa9e30622 · outbound

This paper cites M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning.

A Survey on Video Temporal Grounding with Multimodal Large Language Model M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.831964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.831964Z digest=sha256:c69d13ab3c3d763ebb5565234299c2373df39c7c4a249a5485f24d5e0befcc93

Observation f1561613-df99-4fd2-96ae-a26e9ea1434e · outbound

This paper cites an unresolved cited work.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Unresolved cited work

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.836944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.836944Z digest=sha256:09a6140d844fd78de62f4d49c4b82795d20b6408b179cd18ab800ebbf01f6546

Observation bbd18ff9-bd60-41e7-8f7d-00d4155bd2e8 · outbound

This paper cites an unresolved cited work.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Unresolved cited work

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.842203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.842203Z digest=sha256:6847f26c1b93dccd729b0ed401cc419ed9961025dd4bad181ea11a63f53f2371

Observation cde0115a-451a-46d7-9c81-bc72ae14062d · outbound

This paper cites an unresolved cited work.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Unresolved cited work

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.846474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.846474Z digest=sha256:bb2cf85acf8cb5751664bd19eb3a83269bf06554d6e6085caa165b846ab324ec

Observation 45f422a6-bf94-4aba-a52a-2f60719e77aa · outbound

This paper cites Visual Prompting in Multimodal Large Language Models: A Survey.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Visual Prompting in Multimodal Large Language Models: A Survey

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.850843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.850843Z digest=sha256:9a1f820797f6968abebefe74f0e0a522e18de7fde718e00b90cd511bb90249b8

Observation 0db8b3c9-30bd-480e-a287-ede2fcbd48a3 · outbound

This paper cites an unresolved cited work.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Unresolved cited work

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.855145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.855145Z digest=sha256:f6bcdc47957addb9a7c6ac9e6700c6aa4d632e4f4f108116a986491bee4be2d1

Observation 8aff295d-f45c-49bf-8711-5fbe9bc37e66 · outbound

This paper cites an unresolved cited work.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Unresolved cited work

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.859811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.859811Z digest=sha256:db515c5f953de2fba9d31d93d88579a6d1dce0946f6400e6b6a042be558c2939

Observation 90112b56-42bf-4d27-b410-da5d77c56cbe · outbound

This paper cites Liu, C.-Y.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Liu, C.-Y

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.864113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.864113Z digest=sha256:14ecca41fbbc3fcc9f367e9f002cb9c211d9cb9e98f78cfee19750b887a36a17

Observation 7ffdd6e3-785b-4995-a9af-7dd2950cb6fb · outbound

This paper cites MLLM as Video Narrator: Mitigating Modality Imbalance in Video Moment Retrieval.

A Survey on Video Temporal Grounding with Multimodal Large Language Model MLLM as Video Narrator: Mitigating Modality Imbalance in Video Moment Retrieval

Reference 91

Resolution
verified exact
local_arxiv, observed 2026-08-05T23:32:19.030539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-05T23:32:17.868571Z digest=sha256:71dc5090acbe087ea6d2c5089a3548ea9f54d540930403bed15d332062797ea9

Observation 40ef3941-45ca-4495-967b-b3e6ee5cf6a0 · outbound

This paper cites Di and W.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Di and W

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.873383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.873383Z digest=sha256:037b7a8000e8140fc632bf9887fe526be6d91fccb14877748907bca7a86d34b3

Observation 156dfd67-c440-4efa-851d-09709eb8606e · outbound

This paper cites Zheng, X.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Zheng, X

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.878426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.878426Z digest=sha256:c6e21fd0693066626b59c91235aa6c54447898d6a61319bad3c49e83a3fc539d

Observation 5ddbcca4-14eb-4c7a-b53e-c010ce2d2124 · outbound

This paper cites Grounding-Prompter: Prompting LLM with Multimodal Information for Temporal Sentence Grounding in Long Videos.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Grounding-Prompter: Prompting LLM with Multimodal Information for Temporal Sentence Grounding in Long Videos

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.882834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.882834Z digest=sha256:f27a47e91d4638e7f5afd8c63d6cf329ea22f484fc73fd1d653d52d06f4f35a0

Observation 52300739-7cee-474e-977b-5d4e9eb60b6b · outbound

This paper cites an unresolved cited work.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Unresolved cited work

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.887713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.887713Z digest=sha256:1f4628fccd956d7376ff6bf8a8f2516ea8489c05a603a71e174134ef1b4a0436

Observation 96585e33-88e0-4671-a87b-258ad57956e1 · outbound

This paper cites VERIFIED: A Video Corpus Moment Retrieval Benchmark for Fine-Grained Video Understanding.

A Survey on Video Temporal Grounding with Multimodal Large Language Model VERIFIED: A Video Corpus Moment Retrieval Benchmark for Fine-Grained Video Understanding

Reference 96

Resolution
verified exact
local_arxiv, observed 2026-08-05T23:32:18.994201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-05T23:32:17.891994Z digest=sha256:5e127ad6cd40d5fff8e24bc97c06222719f169177f8248cef3115a2e677e3546

Observation 20f51e9d-5358-4905-ac83-e8e182b11593 · outbound

This paper cites Infusing Environmental Captions for Long-Form Video Language Grounding.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Infusing Environmental Captions for Long-Form Video Language Grounding

Reference 97

Resolution
verified exact
local_arxiv, observed 2026-08-05T23:32:18.972288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-05T23:32:17.896758Z digest=sha256:6672e4e6cecb451059116cb0b74889452fba0ba959ff2be9abad35597ffb5d10

Observation b960ac46-3284-4199-99bf-247be9d6cc91 · outbound

This paper cites an unresolved cited work.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Unresolved cited work

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.901602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.901602Z digest=sha256:3906dd5b441a8da027d84c5e41879bf45cc764d2f0c0421a54c9fb91f3d23047

Observation 73d51dbf-dc62-48a4-9c7e-12dc4c4c26af · outbound

This paper cites an unresolved cited work.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Unresolved cited work

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.906296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.906296Z digest=sha256:4d66aa5ce49bee5017198e080d5bf7d99287ed8c2dc6817e4a38e658ea64e933

Observation a55f8afc-6794-49d6-b2e8-f7c444f1765a · outbound

This paper cites Context-Enhanced Video Moment Retrieval with Large Language Models.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Context-Enhanced Video Moment Retrieval with Large Language Models

Reference 100

Resolution
verified exact
local_arxiv, observed 2026-08-05T23:32:18.882066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-05T23:32:17.911414Z digest=sha256:df7b31f2d875938935342c4946b85fdf2a9f12b54326e039736b50deaa50def2

Pith citing papers

Observation 88ca922f-510e-4504-9f08-5e73b4c37d34 · inbound

Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs cites this paper.

Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs A Survey on Video Temporal Grounding with Multimodal Large Language Model

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T13:27:58.123306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:27:58.123306Z digest=sha256:f9e0f84c3e913b2f83a2d3c275c8f080860dad46a4517625784ee4903d62bc71

Observation c62788a6-7f32-4a43-bd09-750b057d5d86 · inbound

Natural-Language Temporal Grounding in Hour-Long Videos is a Search Problem: A Benchmark and Empirical Decomposition cites this paper.

Natural-Language Temporal Grounding in Hour-Long Videos is a Search Problem: A Benchmark and Empirical Decomposition A Survey on Video Temporal Grounding with Multimodal Large Language Model

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:27:56.646763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-27T10:01:09.933880Z digest=sha256:1cfaa2db45a9dde84ece07dd21c935b84877e3b524291c64d49af7b49d4ad96e

Observation b4dd3e4b-6a31-4ccd-b5c9-eab64f69718f · inbound

NEST: Narrative Event Structures in Time for Long Video Understanding cites this paper.

NEST: Narrative Event Structures in Time for Long Video Understanding A Survey on Video Temporal Grounding with Multimodal Large Language Model

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:29:31.266362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-26T17:57:55.366051Z digest=sha256:855c85c8cf6c978fabc0ed458f86adcea0a3047c64c9e49ef6dfe343170c9d97