Pith. sign in

Paper Citation Record · LEDGER

HawkEye: Training Video-Text LLMs for Grounding Text in Videos

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 28 inbound Pith citation observations for arXiv:2403.10228.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.10228 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 28 of 28 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:35:47.240790Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T20:38:56.195732Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 716f8125-2ca2-4693-9d43-eb94caa2de63 · inbound

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling cites this paper.

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T02:52:20.788080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T02:52:20.643070Z digest=sha256:2b33cde2ac122c127f1847a744b36f8e14a792994b4cfb651b0ad7a7819a4dad

Observation 85ce6aae-bad9-42a4-9987-1de562e8821c · inbound

Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency cites this paper.

Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:35:47.240790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:35:47.240790Z digest=sha256:e96ec9593b4b8de8a6d30d5b711ae325c9ed90e22cb1fc571b71a0663d7ecfd4

Observation b27e939a-30d3-47d2-b4aa-8417e26ffd4b · inbound

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering cites this paper.

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T23:29:26.972515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:29:26.972515Z digest=sha256:3c9dbb85340acfdd16baa352f51b01cb8c494bcc0070240bcb9d03747ae35b9e

Observation 87fcc106-ab07-4f0a-8792-1a6c1b708276 · inbound

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding cites this paper.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T05:30:36.705649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:30:36.705649Z digest=sha256:601314fa9eef6dee61e30a0cc2d18f546a803e630092407e85e43e2078627031

Observation 4a2f41c7-4695-441f-a94a-413cc4806014 · inbound

TAR: Temporal Anchor-Constrained Reasoning for Video Temporal Grounding cites this paper.

TAR: Temporal Anchor-Constrained Reasoning for Video Temporal Grounding HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T22:01:26.012600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:01:26.012600Z digest=sha256:7f7aff728295c3b316be028e0ada08d5f54e3a51bf4bf8287d7c9295bf802141

Observation 989f777a-e959-4819-8ad1-662748ee862c · inbound

A Survey on Video Temporal Grounding with Multimodal Large Language Model cites this paper.

A Survey on Video Temporal Grounding with Multimodal Large Language Model HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.599520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.599520Z digest=sha256:8d55f1b05c552855a1ca28436e42bf3779952ee14941387ca3bc96e67ac9fc81

Observation 8fa2f2a3-cd9b-42d7-ba04-97b5d38e51c1 · inbound

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding cites this paper.

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T07:23:07.344850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:23:07.344850Z digest=sha256:ac8056af9f7b60c6557d5429d5bd8653d5111d2763aa41d2d5c1c3c7026215b5

Observation c2633b38-8492-4fae-b0b6-62c18d264932 · inbound

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning cites this paper.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:18:52.343793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:dc9e9cc1062da701178cc45a7db245723892367dd5db837b0b440a6c44f05f25

Observation 040a74d2-2dee-4802-b188-0c456b63ffee · inbound

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding cites this paper.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:58:46.547114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:832e524950765048ee738b7c3202e95ae3975105d6bf2bb94833156e5e66e8a5

Observation a6aa77df-74c7-41c0-87d7-7282fe22140f · inbound

A Paradigm Shift: Fully End-to-End Training for Temporal Sentence Grounding in Videos cites this paper.

A Paradigm Shift: Fully End-to-End Training for Temporal Sentence Grounding in Videos HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-13T19:48:11.514531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T19:43:57.298338Z digest=sha256:7f69a380ff625dddd38852764c44981fe2410d13e6d718dc62c194f289226ff0

Observation f06b4652-9297-48bf-a0b1-87467b1b6445 · inbound

Bridging Time and Space: Decoupled Spatio-Temporal Alignment for Video Grounding cites this paper.

Bridging Time and Space: Decoupled Spatio-Temporal Alignment for Video Grounding HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:15:52.667017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:38:16.204012Z digest=sha256:ea0437286daa75bc90d45c4fdcb2ea54205e29771a683c7f60379959d3131e8e

Observation c35a53c6-e8b0-4b7f-9ab0-a1ec6b784bdb · inbound

UniversalVTG: A Universal and Lightweight Foundation Model for Video Temporal Grounding cites this paper.

UniversalVTG: A Universal and Lightweight Foundation Model for Video Temporal Grounding HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:51:02.282639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T17:55:56.385801Z digest=sha256:988cb9d81edacb3194799c3435e31238e4c003416f845495b110ded9a85fe2d1

Observation ba024df3-6600-4bda-8112-4382727cef7b · inbound

How Should Video LLMs Output Time? An Analysis of Efficient Temporal Grounding Paradigms cites this paper.

How Should Video LLMs Output Time? An Analysis of Efficient Temporal Grounding Paradigms HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:05:58.951662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T16:50:48.559703Z digest=sha256:bff3ff2b5a63c40edd715c6479c9ef418e5532fc9156ea1ef72e72c3d4ec72e6

Observation b05c671b-1fb7-4c2d-a597-34a0592d090b · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T08:30:57.115545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T16:36:33.264166Z digest=sha256:9f09f8f05e499f944dd3c6c2d4fc4e1c155f6a64c3fc5baf3735b58b4896a7d2

Observation 08fed6a5-2670-4c49-8ae0-e4892f7e1f78 · inbound

ViLL-E: Video LLM Embeddings for Retrieval cites this paper.

ViLL-E: Video LLM Embeddings for Retrieval HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:21:01.989634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:00:43.573409Z digest=sha256:40d7203c0c73460f21c1a097f58128020f9ff11dbf317e529332b0140d6b3212

Observation 2ef16935-95ca-4aef-a608-08692ba7edc3 · inbound

OmniVTG: A Large-Scale Dataset and Training Paradigm for Open-World Video Temporal Grounding cites this paper.

OmniVTG: A Large-Scale Dataset and Training Paradigm for Open-World Video Temporal Grounding HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:31:15.822999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-07T16:51:43.783133Z digest=sha256:4c3810057c2e70de0d325c6184b051feb93675792761bb030fef07c5f03b6a1f

Observation b0ba08a4-59c2-4a30-aaa1-0596f03ff238 · inbound

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding cites this paper.

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-12T00:46:13.699623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-07T14:10:27.416341Z digest=sha256:a1d357d39cdd70cc3237d9606c40279f8d47b3bab4a0069037cb2508f3efd7ba

Observation d46e3d6c-f44f-496b-9a3d-863043f454d6 · inbound

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding cites this paper.

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:45:34.762438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-01T08:41:31.700367Z digest=sha256:026d4ac5c29793e75c5a04240bd9fea6b95009910ee60841500fe2f383445b4f

Observation ea95495a-a242-433d-a6d2-ca7e8ac05a50 · inbound

EvoGround: Self-Evolving Video Agents for Video Temporal Grounding cites this paper.

EvoGround: Self-Evolving Video Agents for Video Temporal Grounding HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T19:32:52.481502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T19:29:47.356665Z digest=sha256:317328df796c912fd9d74fb4e8fba49fea59c38046071c667142708ae39e2ced

Observation d38aec51-379d-40fb-a3b8-4ae26b317623 · inbound

Foresee-to-Ground: From Predictive Temporal Perception to Evidence-Driven Reasoning for Video Temporal Grounding cites this paper.

Foresee-to-Ground: From Predictive Temporal Perception to Evidence-Driven Reasoning for Video Temporal Grounding HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-22T08:06:15.420230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-22T08:04:55.680668Z digest=sha256:f6e8768d82d47f8a370e770a500a31729ea197d4bf436d2be41041f1f6113d8d

Observation 25b9edde-1c18-4302-88f3-8f11480aea48 · inbound

EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding cites this paper.

EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:34:01.972853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T22:27:37.631870Z digest=sha256:4986f748a2791673fe4f8a0f128887a78692aa70cc06757937d867d2136eaa75

Observation b4438bce-39e3-4a81-88d8-8fcb2c7a5e47 · inbound

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models cites this paper.

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-07-02T11:46:55.672672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:db7eb99d48675a478d8bce438a19039fe7f37e23679d127356f42e0ee1200a85

Observation 00a19d4c-3a9f-4e3b-a7f2-bfc08c45edc1 · inbound

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding cites this paper.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 59

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T17:07:12.834984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:cf3b80cd8a5428f2d96585ad59a43974e22949004629a77a313ad45f5588c640

Observation 03120690-d260-4787-a437-a7550b942b44 · inbound

Temporal-Aware Reasoning Optimization for Video Temporal Grounding cites this paper.

Temporal-Aware Reasoning Optimization for Video Temporal Grounding HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 100

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:17:29.106952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T17:21:30.638275Z digest=sha256:965774009398e801dc78c5e0a0e163a0fcadbddd06ee489988887143f1f0cdbb

Observation 6046348f-f56f-46c0-93f6-ed10f481f627 · inbound

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning cites this paper.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 144

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:48:03.047929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:d64d31941202e99b7848bdbdc7b28107c7a7bdf7382801f07b9c7c0030b5c083

Observation 0ddc5890-898c-4e74-8d69-05728f1ca9a3 · inbound

Natural-Language Temporal Grounding in Hour-Long Videos is a Search Problem: A Benchmark and Empirical Decomposition cites this paper.

Natural-Language Temporal Grounding in Hour-Long Videos is a Search Problem: A Benchmark and Empirical Decomposition HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:27:56.644785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T10:01:09.933880Z digest=sha256:b017a09f59aef9c42b9b3de1063143f7b8fb4fb977626a5d6acc1bfa0700812d

Observation bc60d46a-72b6-452b-a97e-5da54f35f31d · inbound

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams cites this paper.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:38:56.197261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:2462c828db8a3f978d1a25baa837cea099963199348067bc7379f672d7fa5c9b

Observation 82ac2554-f4d3-4bf2-977d-3c49f7fa0cba · inbound

TimeThink: Reasoning with Time for Video LLMs cites this paper.

TimeThink: Reasoning with Time for Video LLMs HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 51

Resolution
unresolved
no resolver link, observed 2026-07-11T08:59:46.244502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T08:59:46.244502Z digest=sha256:07960f9fcdc0b720c9c171b4168b1dc94b044abc816576823802a6175fda8a38