Pith. sign in

Paper Citation Record · LEDGER

HawkEye: Training Video-Text LLMs for Grounding Text in Videos

As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 35 inbound Pith citation observations for arXiv:2403.10228.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.10228 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 35 of 35 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:17:40.261405Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T20:38:56.195732Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 5a20b3e7-ecab-4f1d-a804-bd3b2d5cc14e · inbound

Number it: Temporal Grounding Videos like Flipping Manga cites this paper.

Number it: Temporal Grounding Videos like Flipping Manga HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-12T19:49:43.614209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:49:43.614209Z digest=sha256:d75ba54d03a4075a824f45ec2e12bfa54034507bb8b2cdfbe58e6451fa468efc

Observation 8ad9f57c-8738-44b4-9e96-6089277393f9 · inbound

On the Consistency of Video Large Language Models in Temporal Comprehension cites this paper.

On the Consistency of Video Large Language Models in Temporal Comprehension HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T17:07:00.840663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:07:00.840663Z digest=sha256:2e8a8d8b75350d506d30b6d7a9e6d40fb2823d2ece4bc25a0a485930104b2ea5

Observation 09b2e3e7-88b7-464a-aabb-3bfc16f0a067 · inbound

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos cites this paper.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.683167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.683167Z digest=sha256:af20f511997ee40082a79dbadb307a93ed081c9d884d1e68c2473896a0c860e8

Observation 1a061070-dc74-439d-b113-4e22b65fc55e · inbound

Seq2Time: Sequential Knowledge Transfer for Video LLM Temporal Grounding cites this paper.

Seq2Time: Sequential Knowledge Transfer for Video LLM Temporal Grounding HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T12:47:21.283210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:47:21.283210Z digest=sha256:269d7d675e689bf3fb5877da431da7fe2c95048484f928d8bc21eaf4650f1488

Observation d1d2b970-e28e-4342-86a8-428ab4a2ecd2 · inbound

Task Preference Optimization: Improving Multimodal Large Language Models with Vision Task Alignment cites this paper.

Task Preference Optimization: Improving Multimodal Large Language Models with Vision Task Alignment HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-11T00:47:12.424055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:47:12.424055Z digest=sha256:1800b0bc9d3c66f8b1430d8a9ffc77a06d7e926a07952298a7edf4844c9313c8

Observation 716f8125-2ca2-4693-9d43-eb94caa2de63 · inbound

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling cites this paper.

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T02:52:20.788080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T02:52:20.643070Z digest=sha256:9e2f9f8fe0e089eda915ed50879f3ebf02a97c007e209c334ba2efd2b0868614

Observation 9946ca8e-25ea-4a73-be36-b9d15d7fc028 · inbound

TEMPURA: Temporal Event Masked Prediction and Understanding for Reasoning in Action cites this paper.

TEMPURA: Temporal Event Masked Prediction and Understanding for Reasoning in Action HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T04:17:40.261405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:17:40.261405Z digest=sha256:3d41aab4624f218731b55836763118b41b825d0aba4f174e253cc94faf8e79d7

Observation 85ce6aae-bad9-42a4-9987-1de562e8821c · inbound

Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency cites this paper.

Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:35:47.240790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:35:47.240790Z digest=sha256:830d1210482eb8635c622a4c731070d7af901834bdc5095e4a7dafadb171238f

Observation b27e939a-30d3-47d2-b4aa-8417e26ffd4b · inbound

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering cites this paper.

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T19:00:53.758638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:00:53.758638Z digest=sha256:5ac40c894107515947d9bb72c54b82031f62fa967ddc7fc63ed3ef7311b7b084

Observation 87fcc106-ab07-4f0a-8792-1a6c1b708276 · inbound

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding cites this paper.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T05:30:36.705649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:30:36.705649Z digest=sha256:fdce542ce1a10185137d571d90197edc308c8356302e6bf5f56d8dd3fd585a38

Observation 4a2f41c7-4695-441f-a94a-413cc4806014 · inbound

TAR: Temporal Anchor-Constrained Reasoning for Video Temporal Grounding cites this paper.

TAR: Temporal Anchor-Constrained Reasoning for Video Temporal Grounding HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T22:01:26.012600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:01:26.012600Z digest=sha256:b27588d6a3419afc9247c109405d0ce6e7608fa5bec2c9ff19885d5d312b3995

Observation 989f777a-e959-4819-8ad1-662748ee862c · inbound

A Survey on Video Temporal Grounding with Multimodal Large Language Model cites this paper.

A Survey on Video Temporal Grounding with Multimodal Large Language Model HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.599520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.599520Z digest=sha256:837e263ef4b58f7e3d042688dd789ba50aea44b571ac7100f202b373f90a0ff0

Observation 8fa2f2a3-cd9b-42d7-ba04-97b5d38e51c1 · inbound

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding cites this paper.

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T07:23:07.344850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:23:07.344850Z digest=sha256:2708bf7403ddeb2527842c6d7f5c9c13636a30e1e329a14a6ee0adc31dbc3765

Observation c2633b38-8492-4fae-b0b6-62c18d264932 · inbound

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning cites this paper.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:18:52.343793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:26ca16e93fd3a60f4257c5af8aca64e3c0789b84223c65dd92e91934abb8d834

Observation 040a74d2-2dee-4802-b188-0c456b63ffee · inbound

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding cites this paper.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:58:46.547114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:f798d7d59f04720aa390781932f0b18680799b80dfab756b225d732ccab2881c

Observation a6aa77df-74c7-41c0-87d7-7282fe22140f · inbound

A Paradigm Shift: Fully End-to-End Training for Temporal Sentence Grounding in Videos cites this paper.

A Paradigm Shift: Fully End-to-End Training for Temporal Sentence Grounding in Videos HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-13T19:48:11.514531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-13T19:43:57.298338Z digest=sha256:bc5626cc912feae641344a6cafc4b80cc00e57f2cd80382b8928c8adba5742d4

Observation f06b4652-9297-48bf-a0b1-87467b1b6445 · inbound

Bridging Time and Space: Decoupled Spatio-Temporal Alignment for Video Grounding cites this paper.

Bridging Time and Space: Decoupled Spatio-Temporal Alignment for Video Grounding HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:15:52.667017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T18:38:16.204012Z digest=sha256:d049e127c77d73714d51eed296d6bc4348d800212ded0e54d51d4d5398a72120

Observation c35a53c6-e8b0-4b7f-9ab0-a1ec6b784bdb · inbound

UniversalVTG: A Universal and Lightweight Foundation Model for Video Temporal Grounding cites this paper.

UniversalVTG: A Universal and Lightweight Foundation Model for Video Temporal Grounding HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:51:02.282639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T17:55:56.385801Z digest=sha256:597978fa3c3cf15a6f43dba863fa3c6e15de75f7c6744780fc0cf5cb7b8f4dc0

Observation ba024df3-6600-4bda-8112-4382727cef7b · inbound

How Should Video LLMs Output Time? An Analysis of Efficient Temporal Grounding Paradigms cites this paper.

How Should Video LLMs Output Time? An Analysis of Efficient Temporal Grounding Paradigms HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:05:58.951662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T16:50:48.559703Z digest=sha256:7241f36ac431466f4cc6346e79506cd8d698c97a4c8bb6296b3b95c24ed78e19

Observation b05c671b-1fb7-4c2d-a597-34a0592d090b · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T08:30:57.115545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T16:36:33.264166Z digest=sha256:2b5dd09495d0e34e501af7abfd9895f1d52a55b705e90986899105f1a6ff29c6

Observation 08fed6a5-2670-4c49-8ae0-e4892f7e1f78 · inbound

ViLL-E: Video LLM Embeddings for Retrieval cites this paper.

ViLL-E: Video LLM Embeddings for Retrieval HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:21:01.989634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T15:00:43.573409Z digest=sha256:1d27b63fa07d01d64341833160de5652b489ba2c0739fcee94c5b300ee173936

Observation 2ef16935-95ca-4aef-a608-08692ba7edc3 · inbound

OmniVTG: A Large-Scale Dataset and Training Paradigm for Open-World Video Temporal Grounding cites this paper.

OmniVTG: A Large-Scale Dataset and Training Paradigm for Open-World Video Temporal Grounding HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:31:15.822999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-07T16:51:43.783133Z digest=sha256:d2be412a8b07d429ba848a5d11af8426b95b2238b1a76ba00d23f664eef5a37b

Observation b0ba08a4-59c2-4a30-aaa1-0596f03ff238 · inbound

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding cites this paper.

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-12T00:46:13.699623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-07T14:10:27.416341Z digest=sha256:2d4efe7b8b32f005a438e1606e3e82d3fea5a75e759b8e307ebb5117f1d8cf17

Observation d46e3d6c-f44f-496b-9a3d-863043f454d6 · inbound

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding cites this paper.

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:45:34.762438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T08:41:31.700367Z digest=sha256:98b9cba28dfdfe461893584b28f931d92d372fbdb4cbf74f7ddd2355dedba0f3

Observation ea95495a-a242-433d-a6d2-ca7e8ac05a50 · inbound

EvoGround: Self-Evolving Video Agents for Video Temporal Grounding cites this paper.

EvoGround: Self-Evolving Video Agents for Video Temporal Grounding HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T19:32:52.481502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-14T19:29:47.356665Z digest=sha256:1e7c9e67728f6a44d9102e3b538ec686279f0b1a58ced4de1527e0c93813b3d6

Observation d38aec51-379d-40fb-a3b8-4ae26b317623 · inbound

Foresee-to-Ground: From Predictive Temporal Perception to Evidence-Driven Reasoning for Video Temporal Grounding cites this paper.

Foresee-to-Ground: From Predictive Temporal Perception to Evidence-Driven Reasoning for Video Temporal Grounding HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-22T08:06:15.420230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-22T08:04:55.680668Z digest=sha256:05dcaacf5bbaffe9dbb2b77db79448de7542c3d0f31ad0775936867110de0469

Observation 25b9edde-1c18-4302-88f3-8f11480aea48 · inbound

EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding cites this paper.

EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:34:01.972853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T22:27:37.631870Z digest=sha256:437c344461207ed6626fff79223b093047fa44b6b6a69b792adf92afc0513b7b

Observation b4438bce-39e3-4a81-88d8-8fcb2c7a5e47 · inbound

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models cites this paper.

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-07-02T11:46:55.672672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:09249ed99b2f6b4ec279061c1d2e4742d37a7673ba4e4ef43aba52338b6756c0

Observation 00a19d4c-3a9f-4e3b-a7f2-bfc08c45edc1 · inbound

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding cites this paper.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 59

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T17:07:12.834984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:839d82976fe3402f4501113886ec0626101011af1ccce17a6a5cf0e178d34375

Observation 03120690-d260-4787-a437-a7550b942b44 · inbound

Temporal-Aware Reasoning Optimization for Video Temporal Grounding cites this paper.

Temporal-Aware Reasoning Optimization for Video Temporal Grounding HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 100

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:17:29.106952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-27T17:21:30.638275Z digest=sha256:97dd87894b8db66f2960abe0ad602384d7cf786cd46d0cb1cccb5c63ef4e9640

Observation 6046348f-f56f-46c0-93f6-ed10f481f627 · inbound

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning cites this paper.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 144

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:48:03.047929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:082052c6f8368c356322504e58d061aaff8033a91d7ca41bcf89c875b2b8d044

Observation 0ddc5890-898c-4e74-8d69-05728f1ca9a3 · inbound

Natural-Language Temporal Grounding in Hour-Long Videos is a Search Problem: A Benchmark and Empirical Decomposition cites this paper.

Natural-Language Temporal Grounding in Hour-Long Videos is a Search Problem: A Benchmark and Empirical Decomposition HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:27:56.644785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-27T10:01:09.933880Z digest=sha256:33dfbbf9f3eb224a0129392d41bd2f3dadcf056a6e519e40d5ed0f2ee68e9744

Observation bc60d46a-72b6-452b-a97e-5da54f35f31d · inbound

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams cites this paper.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:38:56.197261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:81b1f62b40b99a12f12ceda38f2a4cdf90d7b652d47bab6cc0baccafb29c797b

Observation 82ac2554-f4d3-4bf2-977d-3c49f7fa0cba · inbound

TimeThink: Reasoning with Time for Video LLMs cites this paper.

TimeThink: Reasoning with Time for Video LLMs HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 51

Resolution
unresolved
no resolver link, observed 2026-07-11T08:59:46.244502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T08:59:46.244502Z digest=sha256:b9e4b941b33a6dfa69ffb0b06c95d2d85b6de793057385c50d401b4d71a7b41c

Observation 2f283da1-a739-4f44-bc42-ef5ab8393b3d · inbound

Your VLM Already Knows When: Training-Free Temporal Grounding by Asking Yes or No cites this paper.

Your VLM Already Knows When: Training-Free Temporal Grounding by Asking Yes or No HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T00:13:21.323847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:13:21.323847Z digest=sha256:ca710fa67e63e066c660db02b1731238457efbf3f5340855fd29d9a95e124291