Pith. sign in

Paper Citation Record · LEDGER

Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 42 inbound Pith citation observations for arXiv:2410.03290.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.03290 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 42 of 42 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T19:49:43.592295Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:09:55.249321Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c8b08ce7-467b-493d-a67e-5a418909482d · inbound

Number it: Temporal Grounding Videos like Flipping Manga cites this paper.

Number it: Temporal Grounding Videos like Flipping Manga Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-12T19:49:43.592295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:49:43.592295Z digest=sha256:4d7488498c5685aab582805864ad219d6d7a74787c5ccea6ac9a21e8754368df

Observation bb49a29a-1bc5-44d1-a064-420709d161ba · inbound

Seq2Time: Sequential Knowledge Transfer for Video LLM Temporal Grounding cites this paper.

Seq2Time: Sequential Knowledge Transfer for Video LLM Temporal Grounding Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T12:47:21.279910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:47:21.279910Z digest=sha256:d8f8d1239fcfff9e23e5f5e2b0475fb2667a65d11f7516883a9469177097128b

Observation 87372a76-ae68-488f-8336-5c7927f086d5 · inbound

TimeRefine: Temporal Grounding with Time Refining Video LLM cites this paper.

TimeRefine: Temporal Grounding with Time Refining Video LLM Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:35.823573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:35.823573Z digest=sha256:5f95f8d0aa619c54655ca5b5f75ec0b2c0e0340440fa5f1521a07f558ab14c09

Observation d3958667-486e-4e66-b164-d7e9e50dd5bd · inbound

Task Preference Optimization: Improving Multimodal Large Language Models with Vision Task Alignment cites this paper.

Task Preference Optimization: Improving Multimodal Large Language Models with Vision Task Alignment Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-11T00:47:12.387207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:47:12.387207Z digest=sha256:cc7ba288ceecc70ac17016542cf10212539c468d4bcbd16a38d979b21d19b2ee

Observation d42da109-7df7-40f7-bff1-cb300db98040 · inbound

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs cites this paper.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:16.276918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:16.276918Z digest=sha256:3cf751db6e5c26da7700fb4fde93bf20d0ab3dfcbe04e9e4ff0c6029bc8d8bd3

Observation dcaf2cd8-1f50-401c-83f5-de3f9366e583 · inbound

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling cites this paper.

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:52:20.775068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-17T02:52:20.643070Z digest=sha256:1ce3d0e3157553ac385f2240f6fc4067839ef04c53ffdf3c938437eb6a29e6d0

Observation ce014ffa-038b-4798-a758-e213929b9342 · inbound

Temporal Preference Optimization for Long-Form Video Understanding cites this paper.

Temporal Preference Optimization for Long-Form Video Understanding Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T15:35:30.257411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:35:30.257411Z digest=sha256:7749e00eccab9c9944a1ad17cd99c58a2aef5b1334bfafc11f01bc04fd8fcd73

Observation 921a662e-7183-44a5-b656-983f70e2ff53 · inbound

Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency cites this paper.

Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:35:47.027169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:35:47.027169Z digest=sha256:0873642f13af6c73c49bb2b3f73b8789e1d5fd234f434316b9701d1bcdbdc965

Observation 67ca3aea-21d9-42f0-82e8-0d874e7f886f · inbound

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations cites this paper.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:53.529692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:53.529692Z digest=sha256:2af692de173c76ad587b2ee1e1d95fe7744e3ca6abe6e9412071ba50e99d2968

Observation 874431d9-8e44-4c9c-b80f-5fc8a6dcd538 · inbound

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering cites this paper.

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T23:29:24.133365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:29:24.133365Z digest=sha256:a351564a42aa6aedf22403c2301ea8e55fbe54559877ba02771a7b6b8f8c17d9

Observation 255843c6-4a7e-4c5c-892f-b71afcacd37b · inbound

IPFormer-VideoLLM: Enhancing Multi-modal Video Understanding for Multi-shot Scenes cites this paper.

IPFormer-VideoLLM: Enhancing Multi-modal Video Understanding for Multi-shot Scenes Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T22:37:38.924039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:37:38.924039Z digest=sha256:ac17b6f5c388aae3f0061c0e060d547d809c02ebf8a208b71a78e2ab0c4a1dd8

Observation e660d4f9-82e0-4d07-941d-532b62ab2a88 · inbound

Datasets and Recipes for Video Temporal Grounding via Reinforcement Learning cites this paper.

Datasets and Recipes for Video Temporal Grounding via Reinforcement Learning Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:28.022951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:44:28.022951Z digest=sha256:d70ff6ed1cd48f4cdd5e444913f3eaa7963c1fa87c3ddb23c1a4b2d0176e3118

Observation cb8d8afd-6d9a-4832-a6ef-ffdc5fd7d998 · inbound

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding cites this paper.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T05:30:36.525254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:30:36.525254Z digest=sha256:dcf5da13db288d6546309fd58d5e95ab6c5106ca15e9d23595277d9c88a4f68c

Observation cccf61d6-12aa-4df8-8839-8fd1fec4e665 · inbound

"Harmless to You, Hurtful to Me!": Investigating the Detection of Toxic Languages Grounded in the Perspective of Youth cites this paper.

"Harmless to You, Hurtful to Me!": Investigating the Detection of Toxic Languages Grounded in the Perspective of Youth Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-06T05:15:43.184734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:15:43.184734Z digest=sha256:b913710320b6fdd4537cba01a3d1e6af259e5ed5116d745e27b9846d6eff33ea

Observation 13f125b0-b270-4d4e-b194-a727aecc27fb · inbound

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models cites this paper.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:38.429768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:38.429768Z digest=sha256:dcc58ed4c48a9e51d89253e0461f5cca30e4244f83cddb6165ed0ab365199e51

Observation 5ee27c7d-fb51-4241-8eaf-ef831dfba3f3 · inbound

A Survey on Video Temporal Grounding with Multimodal Large Language Model cites this paper.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.609232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.609232Z digest=sha256:cf1455d3e2f4b447c01f41d63af21f68f1b3b029961b80cd0568c6d0f3d9f22a

Observation a9a7e9ff-10cd-4e27-b89f-d12e6896098a · inbound

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding cites this paper.

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T17:50:49.044330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:50:49.044330Z digest=sha256:8f90a05ae020e70851593b3740b09299be37936e10faf09f6f37699d625a8d5d

Observation 41ca1c2b-a3bf-4f21-bef5-9b4434fbc12a · inbound

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data cites this paper.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:17.747639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:17.747639Z digest=sha256:8016b7d627760c973c203830665303fe1de0cbfb90fcf7f49620bad451b3d750

Observation 551d36da-da25-4119-847e-9767d9931db7 · inbound

OneThinker: All-in-one Reasoning Model for Image and Video cites this paper.

OneThinker: All-in-one Reasoning Model for Image and Video Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:11:26.607886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-17T02:09:39.820651Z digest=sha256:94ed3d96d89f13879a47bc95fedf1c27b593670556aa32a15d19806d5eeecfec

Observation 19a5d37e-6486-4a8a-b153-7e6e0d905da5 · inbound

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding cites this paper.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:58:46.609124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:2b635c5fa6694e3bebe72c6fd69e191e80ac9cce7fedcc09e454ec5b43042c5b

Observation 70b78f38-3efd-4918-83b6-718938457278 · inbound

Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation cites this paper.

Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:40:46.375077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-16T08:38:49.075457Z digest=sha256:7298a82108415818c65f8abc36d5af64c383f60a5bf75bd9d888a237206b6f66

Observation d164cb4b-ad36-47f8-b0d4-75525aeb73c1 · inbound

Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation cites this paper.

Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T05:14:23.944562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:14:23.944562Z digest=sha256:b4771807b8e423a271c3a5e66373a72220145d076b93e5e7991afd498633d8cf

Observation a410e7bc-9a09-4f8b-8f2a-e680c85963b4 · inbound

A Multimodal Foundation Model of Spatial Transcriptomics and Histology for Biological Discovery and Clinical Prediction cites this paper.

A Multimodal Foundation Model of Spatial Transcriptomics and Histology for Biological Discovery and Clinical Prediction Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-13T12:44:39.273341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T12:44:39.273341Z digest=sha256:4d45338806d8ac0ba5c917d625ed775b48cafbb7641e84a8b903a5ac5d4ae49d

Observation d668af7f-8286-4138-881e-afcc6626edbf · inbound

Single-agent vs. Multi-agents for Automated Video Analysis of On-Screen Collaborative Learning Behaviors cites this paper.

Single-agent vs. Multi-agents for Automated Video Analysis of On-Screen Collaborative Learning Behaviors Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T17:33:02.311291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-13T17:32:06.256142Z digest=sha256:178ac43195d99e139b1d85fe4aea07d14665c117e72b431d111ce83fa76bee16

Observation 241d4e5e-d9e2-4581-9e32-7c01ee248be4 · inbound

How Should Video LLMs Output Time? An Analysis of Efficient Temporal Grounding Paradigms cites this paper.

How Should Video LLMs Output Time? An Analysis of Efficient Temporal Grounding Paradigms Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:05:58.925033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T16:50:48.559703Z digest=sha256:8ab0c1933bb0bbf0476bb1d63614642b04611abdebe7210fdab0b50a8533db4a

Observation 6dcc5857-4426-4d94-8fc3-e128035b778b · inbound

Towards Temporal Compositional Reasoning in Long-Form Sports Videos cites this paper.

Towards Temporal Compositional Reasoning in Long-Form Sports Videos Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-14T19:27:29.843866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T19:27:29.843866Z digest=sha256:a2f2436e5a8a7be2aec1a9bbe981e9a1d7259245e9f37de5da527478da4da8a8

Observation da6f3052-00a2-46b6-8175-55d420989eb9 · inbound

EvoGround: Self-Evolving Video Agents for Video Temporal Grounding cites this paper.

EvoGround: Self-Evolving Video Agents for Video Temporal Grounding Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:32:52.391468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-14T19:29:47.356665Z digest=sha256:0d341f9eb0b947c2951bbe6102495eedf6649ba61901f2a866fd1c2dae35838e

Observation ed7acdf5-df22-44a5-a1a3-a99d5d8beb02 · inbound

Video-Zero: Self-Evolution Video Understanding cites this paper.

Video-Zero: Self-Evolution Video Understanding Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-06-30T21:35:04.426284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T21:32:16.939563Z digest=sha256:3bb6d6683abedf0bee3fe6d3af3ccd0ce0a6396d5c73da46e4287aa4b0a5d9ff

Observation 8751f1f0-6f57-4225-ae09-27e7d4237a56 · inbound

MLLMs Know When Before Speaking: Revealing and Recovering Temporal Grounding via Attention Cues cites this paper.

MLLMs Know When Before Speaking: Revealing and Recovering Temporal Grounding via Attention Cues Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-22T07:14:42.418973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-22T07:13:43.716510Z digest=sha256:0134823e694f61a2ebea957090766361cc4a76015d7603a026a622c05f68454f

Observation e8e66c53-1aa6-462f-8eae-33e353ef4e39 · inbound

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence cites this paper.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:13:59.520753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:1bea79c86aa75ecddbe4ac6b094397f5236f20cd5663d2ceb22a7d7de8f50d1a

Observation ce3a092a-a93a-41bb-9264-0ca58af703fa · inbound

VidPrism: Heterogeneous Mixture of Experts for Image-to-Video Transfer cites this paper.

VidPrism: Heterogeneous Mixture of Experts for Image-to-Video Transfer Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:03:25.917929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-29T13:01:42.880738Z digest=sha256:84770543664a34e7931d164c13620dcd460fdce3f9400a14763bb71850a19ffc

Observation 43060d12-bce0-4242-89ee-08bd37db8ef6 · inbound

Watch, Remember, Reason: Human-View Video Understanding with MLLMs cites this paper.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:14.973292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:06598bca4188694e7ffc7a0472655becebb5ecb16b75724ed5cc0ee792d4286f

Observation 2f75dce3-b997-4a91-ab51-f7d37bfec2e4 · inbound

Counterfactual Reasoning for Fine-Grained Evidence Disentanglement in VideoQA cites this paper.

Counterfactual Reasoning for Fine-Grained Evidence Disentanglement in VideoQA Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:57:30.201488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-27T16:55:35.743040Z digest=sha256:74c5eb4768c9d0aa07bee30def5bef2a42ad0bed62e02687d702e67f7dca435d

Observation b68500d7-1bc9-4b0c-be33-e5e2f117903b · inbound

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning cites this paper.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models

Reference 135

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T10:48:03.119304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:6b1352492534024cd06c86fc6929014848c9a11c7b6746555691d7f7fdd997c6

Observation 56cffcb1-b355-4a21-8d17-f5d2b404f5d5 · inbound

Training LLMs with Reinforcement Learning over Digital Twin Representations for Reasoning-Intensive Surgical VideoQA cites this paper.

Training LLMs with Reinforcement Learning over Digital Twin Representations for Reasoning-Intensive Surgical VideoQA Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-03T17:48:46.270090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-27T03:40:26.565029Z digest=sha256:5d0f1f430f234ea483ed05051225f0ba1b3bee09c13874c1b1bbd619c3f99414

Observation f266c6e2-8f7f-4787-a57c-2c67421b3810 · inbound

NEST: Narrative Event Structures in Time for Long Video Understanding cites this paper.

NEST: Narrative Event Structures in Time for Long Video Understanding Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T03:29:31.221922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-26T17:57:55.366051Z digest=sha256:330a3840e50f2ea15c31dcd4bf1611ff54577edfd4969ac2c9ff84864dc34e39

Observation 069fe2d3-7ff4-4e73-b830-14eee3dafe06 · inbound

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models cites this paper.

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models

Reference 117

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:55.251092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T01:50:54.242508Z digest=sha256:2df556be42d54a9cab9f180df9db92dea684a49156ad4891274474c109d1dd1a

Observation dca9bb31-a4d2-417c-8449-830109fe3267 · inbound

TimeThink: Reasoning with Time for Video LLMs cites this paper.

TimeThink: Reasoning with Time for Video LLMs Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-07-11T08:59:46.244502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T08:59:46.244502Z digest=sha256:050406955d219461fb9194bdc5dfac5baeafaa0c55ea98c2f25bb91b360b1546

Observation 76950a0e-92bb-46e7-a5cc-8355dec280c7 · inbound

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding cites this paper.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models

Reference 110

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:50.709854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:50.709854Z digest=sha256:cd5c039c3662a95034346887b52ff0f4fe0417509e477bfd9af749239ba09c21

Observation 2c1b856a-8498-4fe4-aeb1-ca770c12920d · inbound

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs cites this paper.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:16.246667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:16.246667Z digest=sha256:50b9efe3c8d8f90a55bad8281fec7c9994c87f1938b9dee7c1f961a8999e38d3

Observation cc0bda32-8dbd-402d-9fe7-50afc2fb0e4b · inbound

TimePLE: Rethinking Temporal Representation for Video Temporal Grounding cites this paper.

TimePLE: Rethinking Temporal Representation for Video Temporal Grounding Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-31T23:32:53.699768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:32:53.699768Z digest=sha256:510a4780e77832ad2fa116875bbea2c3645a20462d6ea1f6cc647c5d06f002b2

Observation 82086325-bb4d-469d-89f3-a01a88db52af · inbound

FADE: From Passive Verification to Active Discovery in Counterfactual Video Understanding cites this paper.

FADE: From Passive Verification to Active Discovery in Counterfactual Video Understanding Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T17:56:48.437058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:56:48.437058Z digest=sha256:d8ffb077b735885a3210ecc2a7e19f076508ac580e88a300b4b4143afca417bc