Pith. sign in

Paper Citation Record · LEDGER

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking

As of 7 August 2026, this Paper Citation Record lists 74 of 74 outbound references and 2 inbound Pith citation observations for arXiv:2602.17555.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2602.17555 v3

Coverage vector

measured 74 of 74 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-15T20:48:44.933542Z

measured 76 of 76 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-11T12:10:52.100410Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-06-29T22:13:59.633393Z

Reference resolution

74 of 74 outbound references displayed

  • verified exact31
  • verified fuzzy42
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9f5fb39e-c0c8-4292-af5c-a5aa4623a85d · outbound

This paper cites The claude 3 model family: Opus, sonnet, haiku.Claude-3 Model Card, 1(1):4.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking The claude 3 model family: Opus, sonnet, haiku.Claude-3 Model Card, 1(1):4

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:50:17.912633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:d58d33a276e472999fd078f94fb60201bed458c1055bf07048cf65cb5f60c856

Observation 415449e0-40c5-4efb-b274-21a9fc0545b0 · outbound

This paper cites Qwen2.5-VL Technical Report.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking Qwen2.5-VL Technical Report

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-15T20:50:17.355161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:16730a892b70d37e42c3175047ff4f6023d2947377dcb7fe8d4bc4aafb477a7b

Observation f0f7b3eb-afa1-4583-88dd-bcfc239420b3 · outbound

This paper cites PerturboLLaVA: Reducing Multimodal Hallucinations with Perturbative Visual Training.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking PerturboLLaVA: Reducing Multimodal Hallucinations with Perturbative Visual Training

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:50:17.363150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:47a2f441a820a97e1c146f28bfe2a375b7e30bbb6cb5534b0d508b55349412db

Observation c8c6babe-6981-4bb2-a476-a7642d796942 · outbound

This paper cites Rextime: A benchmark suite for reasoning-across-time in videos.Advances in Neural In- formation Processing Systems, 37:28662–28673.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking Rextime: A benchmark suite for reasoning-across-time in videos.Advances in Neural In- formation Processing Systems, 37:28662–28673

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:50:17.903131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:e0ec7ed7bf41be1cd9f5db0f0eab0d9d1fa1822ca1fe539e514c2a4c121182bf

Observation 10bc9c55-0139-423a-be98-cd0f5bc24946 · outbound

This paper cites Sharegpt4video: Improving video understand- ing and generation with better captions.Advances in Neural Information Processing Systems, 37:19472–19495.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking Sharegpt4video: Improving video understand- ing and generation with better captions.Advances in Neural Information Processing Systems, 37:19472–19495

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:50:17.867739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:327ce1f927f03618ab82ddef3f22018c29e2f84be87df625abf6cd613338ee7e

Observation d019d1a7-b3ae-45af-aaa4-56e92d853f1c · outbound

This paper cites Scaling rl to long videos.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking Scaling rl to long videos

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:50:17.360259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:20a0d80ad3f3af0e21b095d54f476643721b63e38a4848f52d5c42aca929948c

Observation 3a46ea84-1d00-424c-9e5c-d0430f0c79a3 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-15T20:50:17.352986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:f7a3e3bca2a43b910f5804c1b30d69c258e4883fb5a92b0cae953f2993b3a9f0

Observation 4434c415-058f-4bc4-b1d9-d3381d95eb76 · outbound

This paper cites V-STaR: Benchmarking Video-LLMs on Video Spatio-Temporal Reasoning.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking V-STaR: Benchmarking Video-LLMs on Video Spatio-Temporal Reasoning

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:50:17.357781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:5eae8e0b5f27330e9aa5ae6988e280ad66190412bac2052e309160ebeb204bd0

Observation 7b3d732e-4471-4481-9e57-33ed2db58000 · outbound

This paper cites Spatial-temporal trans- former for dynamic scene graph generation.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking Spatial-temporal trans- former for dynamic scene graph generation

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:50:17.901120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:2805f721fc87aef23bae7a3478108cf532aada4616408b18ac615f63d7c5f59a

Observation d710b2ce-a93c-40f1-ab5d-6408887bcae1 · outbound

This paper cites Sophiavl-r1: Reinforcing mllms reasoning with thinking reward.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking Sophiavl-r1: Reinforcing mllms reasoning with thinking reward

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:50:17.326138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:c850a263a38b87217efa8e6fe325ac71c5d743e0b29bc51271b3f1e534db6202

Observation 1e388425-27dd-41c1-8682-95d614269836 · outbound

This paper cites Mmbench-video: A long-form multi-shot benchmark for holistic video under- standing.Advances in Neural Information Processing Sys- tems, 37:89098–89124.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking Mmbench-video: A long-form multi-shot benchmark for holistic video under- standing.Advances in Neural Information Processing Sys- tems, 37:89098–89124

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:50:17.917780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:b7ce3209d9e8569b0bc56d11bf16ad06321cf3a46dea853d85b079454cfdc7f7

Observation 0ae663e0-08ca-419d-8925-c9b064e5d3cd · outbound

This paper cites Video-of-thought: Step-by-step video reasoning from perception to cognition.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking Video-of-thought: Step-by-step video reasoning from perception to cognition

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:50:17.912993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:461e410cf5e270153f38e5b4b042219f9222462ed45937dd069489033420b8aa

Observation a4ffcf43-2769-47e2-88ed-92569688896c · outbound

This paper cites Video-R1: Reinforcing Video Reasoning in MLLMs.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking Video-R1: Reinforcing Video Reasoning in MLLMs

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-15T20:50:17.350762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:6d7f753a74fcfb44fd4d62c4ef0c15b5e400f8e73031e01f5cb48a8dd1c2a2c3

Observation fa7a02a4-645d-4a61-a3d5-1c8085a543b2 · outbound

This paper cites Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:50:17.907308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:6604bd0a0e3df3806d09d3fb67a34e7202e54bbadb5c5bb4f3466acf995cdf7a

Observation 2029d3f8-9b70-4a18-99c2-0fb321b3d092 · outbound

This paper cites Embodied AI Agents: Modeling the World.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking Embodied AI Agents: Modeling the World

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:50:17.337127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:90dc352489c5cbec79ff769b05a712c475cd15ae0e754837a6c5956168ad276f

Observation c45a468b-95a2-4c3b-977b-f9ea127e6646 · outbound

This paper cites Chain-of-Frames: Advancing Video Understanding in Multimodal LLMs via Frame-Aware Reasoning.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking Chain-of-Frames: Advancing Video Understanding in Multimodal LLMs via Frame-Aware Reasoning

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-15T20:50:17.331346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:31e774465c5939de9ad7adda7ddd28939751b01069a80314c9229f891ac9c128

Observation 2c65fb75-4f22-4c20-8e04-1431f1850f7b · outbound

This paper cites Real-time scene understanding for blind users: Enhancing vision-language models for accessibility.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking Real-time scene understanding for blind users: Enhancing vision-language models for accessibility

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:50:17.908979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:378a24e770bfad14ee46807d156b1cbb3ca6093e00e4cc826ddb279b1033ebff

Observation ebb49bc6-80ca-470d-b840-06d10b4dd78f · outbound

This paper cites Deepseek-r1 incentivizes reasoning in llms through reinforcement learning.Nature, 645(8081):633– 638.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking Deepseek-r1 incentivizes reasoning in llms through reinforcement learning.Nature, 645(8081):633– 638

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:50:17.909349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:04af7ae61f040554065ad988115a61bde498267091f7e90e4ff82cc658db726f

Observation 86ef4eb3-55aa-463f-95eb-5e51f1234ed0 · outbound

This paper cites TRACE: Temporal Grounding Video LLM via Causal Event Modeling.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking TRACE: Temporal Grounding Video LLM via Causal Event Modeling

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:50:17.305860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:88c75c29f005498c8b8f819643b4e82cc1a657c6522b7a3573402ef9ce2fa596

Observation 6d894c32-3ecb-45dd-a86a-c6762ad69e19 · outbound

This paper cites TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:50:17.342680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:9c5c4771ebab462bcb1cddbcedc551e1cdff826b798fc291139cbe2d935b51a3

Observation 50e746e7-4b5e-4bc7-993e-949296716072 · outbound

This paper cites Videoespresso: A large-scale chain-of-thought dataset for fine-grained video reasoning via core frame selection.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking Videoespresso: A large-scale chain-of-thought dataset for fine-grained video reasoning via core frame selection

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:50:17.897299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:70dd312db8ea2cb8a3ade63350825fbf3f27391d0d95606382c5762cc4e1be89

Observation 31cdb3d0-9a54-4998-a147-75949029aa1b · outbound

This paper cites To- wards open-vocabulary scene graph generation with prompt- based finetuning.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking To- wards open-vocabulary scene graph generation with prompt- based finetuning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:50:17.899558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:71c1ef5b4948a8a41da39b33a4707961c4771aa42fdb31bda9ee26221a2e278a

Observation 915d2a63-00a5-4147-a73e-01012cd22d7c · outbound

This paper cites Glm-4.1 v-thinking: Towards versatile multi- modal reasoning with scalable reinforcement learning.arXiv e-prints, pages arXiv–2507.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking Glm-4.1 v-thinking: Towards versatile multi- modal reasoning with scalable reinforcement learning.arXiv e-prints, pages arXiv–2507

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:50:17.919486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:2a82cecac5cd9f71d8d4c875513549a15ca1b13e445c4370be64b7bd1820aa3f

Observation 21156c8a-d0fd-423f-94f0-37dd52a0d419 · outbound

This paper cites Uncertainty-quantified Rollout Policy Adaptation for Unlabelled Cross-domain Temporal Grounding.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking Uncertainty-quantified Rollout Policy Adaptation for Unlabelled Cross-domain Temporal Grounding

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:50:17.331446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:c688785a5515066b893cac064adad32c2493a9adff1a3a1d0c1974ae9d57c829

Observation 5d5cdca8-a824-44a3-af0a-987e97696a04 · outbound

This paper cites Vtimellm: Empower llm to grasp video moments.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking Vtimellm: Empower llm to grasp video moments

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:50:17.897202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:02d36749f243b707797166abd45fb1fc252638009ea945688caae686756f0846

Observation 7e8bd92e-e70b-4914-88c3-f42ae7fee634 · outbound

This paper cites Lita: Language instructed temporal-localization assistant.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking Lita: Language instructed temporal-localization assistant

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:50:17.904883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:4eff9ac6d49287c1a092a7eece7260ec03533604d5d4982337d3ec2be7ba816c

Observation 7c28ffae-31e7-4afc-a5ce-581409848079 · outbound

This paper cites Building a mind palace: Structuring environment-grounded semantic graphs for ef- fective long video analysis with llms.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking Building a mind palace: Structuring environment-grounded semantic graphs for ef- fective long video analysis with llms

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:50:17.906816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:827c742be457f28ace612ec0ff926edbc6b15e4e173341114cbbe55964a5bcb4

Observation f99db999-b1f3-466b-a32d-73d1bafa01d6 · outbound

This paper cites Action genome: Actions as compositions of spatio- temporal scene graphs.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking Action genome: Actions as compositions of spatio- temporal scene graphs

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:50:17.916053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:447e612da979b08bb52067c34080fb08b789be900ff1999d26e2c6742eb9eb42

Observation 39b8bd07-c8a7-407e-adcd-ebcbe0fd6a97 · outbound

This paper cites Look again, think slowly: Enhancing visual reflection in vision-language models.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking Look again, think slowly: Enhancing visual reflection in vision-language models

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:50:17.355679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:d57cbfcd6f0942fe74cff468458dc52546d6d17be9f28a9940c67bbdd0f07014

Observation f4875134-737b-4d6d-82d2-0bcf503a9e06 · outbound

This paper cites Chat-univi: Unified visual representation em- powers large language models with image and video un- derstanding.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking Chat-univi: Unified visual representation em- powers large language models with image and video un- derstanding

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:50:17.914708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:3d2f5123d20bbfbb916ed4679dbe69d4c3e43627769728d2c6d8712b15e8a0ce

Observation f6e7061f-110a-4648-82be-f4b883ee572d · outbound

This paper cites Do you remember? dense video captioning with cross-modal memory retrieval.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking Do you remember? dense video captioning with cross-modal memory retrieval

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:50:17.924034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:8ed2c3dc2d8cce2c3cfccc3f90a2577dfdbd7982edbeba6e7ce05418253d3ab7

Observation 6cf79116-d81c-4194-bfaf-59bcefa6f1de · outbound

This paper cites Vidhalluc: Evaluating temporal hallucinations in multimodal large lan- guage models for video understanding.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking Vidhalluc: Evaluating temporal hallucinations in multimodal large lan- guage models for video understanding

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:50:17.929153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:8455e785254a63f71ac25c64ced5d0b8d7a1f993d9cb3f2175f3fee62e1bf265

Observation d62ebeaa-5f0d-4f4c-bb7e-d320650c0026 · outbound

This paper cites Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:50:17.320193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:8f7d8e098d70ddd051a59261e3bd11b7bf7a53e47c7e693d39dbc4dd89c517e2

Observation 3e96dc93-a236-49ae-93a4-c01ecf63556c · outbound

This paper cites Inten- tqa: Context-aware video intent reasoning.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking Inten- tqa: Context-aware video intent reasoning

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:50:17.891407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:7b3662b25886ad1aa93b77b62b464d71458d8870e3bfa9d2597031361eb21530

Observation 8230c58b-072f-4f6f-a24b-35f25ae5a698 · outbound

This paper cites Temporal Reasoning Transfer from Text to Video.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking Temporal Reasoning Transfer from Text to Video

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:50:17.317205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:b89ef1829915266d54b7ce28f810ca36b799cd026050baeadfaea7fafbf9594e

Observation 8c64cd1d-55e7-4260-9a2d-361a8ed12191 · outbound

This paper cites Embodied agent inter- face: Benchmarking llms for embodied decision making.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking Embodied agent inter- face: Benchmarking llms for embodied decision making

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:50:17.880624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:2182e3d837c90b777d64523eb8c2693498dda3897b854765d8b7f5e8d7af3026

Observation 200d6909-d5c2-49af-b057-b5f5a31ee70f · outbound

This paper cites From pixels to graphs: Open-vocabulary scene graph generation with vision-language models.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking From pixels to graphs: Open-vocabulary scene graph generation with vision-language models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:50:17.882447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:7759d08ff470177298ffbf38df47c956f01a3bc000de079239e7881b30ec4308

Observation 7b18476e-6052-409e-920d-660b4714a53b · outbound

This paper cites VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:56:07.931779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:9811bad08fd2ea832de9ca001476bb17b40efe221d071524075d07175e8670dc

Observation f8633f1c-9ec2-4432-b97c-13e15abbe6fc · outbound

This paper cites Factorizable net: an efficient subgraph-based framework for scene graph generation.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking Factorizable net: an efficient subgraph-based framework for scene graph generation

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:50:17.893095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:ce111759d7d3bb9a74e29232ed2f43860fab7d781bf4ca9955753025dd8a6c62

Observation e58adc24-469a-4f4d-ae59-0221806007ba · outbound

This paper cites Video-llava: Learning united visual repre- sentation by alignment before projection.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking Video-llava: Learning united visual repre- sentation by alignment before projection

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:50:17.871435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:95d4be67f7be1aad2e4c84c6d8415727ebf4c3f627353ef6767afde29ea60377

Observation 33aef3a7-863e-4a5a-8fe4-0a1613a0a488 · outbound

This paper cites Vila: On pre-training for vi- sual language models.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking Vila: On pre-training for vi- sual language models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:50:17.899383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:c371c7f625001f0bd96a364e5238dd5f7067f09abdb6f930fce876437600665f

Observation 6d7f2c74-79b1-47a0-b6d3-5cfe3a3be4b9 · outbound

This paper cites Univtg: Towards unified video- language temporal grounding.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking Univtg: Towards unified video- language temporal grounding

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:50:17.882299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:616de20fbcf143f4800c9561a042ee61d456ed75706a1117c3472c4d517bc28e

Observation 42004759-81a5-471a-9f69-a12430e11910 · outbound

This paper cites More Thinking, Less Seeing? Assessing Amplified Hallucination in Multimodal Reasoning Models.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking More Thinking, Less Seeing? Assessing Amplified Hallucination in Multimodal Reasoning Models

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T20:50:17.334014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:3161282a20355ccb5b99879c8c7626330e1d765d0456d97f041b1f9ddeba18ed

Observation 6b13fddb-5b7b-4527-826c-4a725636c01f · outbound

This paper cites MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-15T20:50:17.314324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:b6dfb1bf09fee1a136f1744449966e5cc5227eaefb3a1668be39671d4f7da457

Observation 9dfaf9cd-7642-4161-8223-70b988d07863 · outbound

This paper cites When thinking drifts: Evidential grounding for robust video reasoning.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking When thinking drifts: Evidential grounding for robust video reasoning

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:50:17.344715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:224cee3043f730e98e4c793cc5472b6947637ac0794b0e7db76779d369cdec6e

Observation cc03584b-6285-49d5-9323-ce1fb009052c · outbound

This paper cites Video-chatgpt: Towards detailed video un- derstanding via large vision and language models.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking Video-chatgpt: Towards detailed video un- derstanding via large vision and language models

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:50:17.884546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:a9a23ddcf3f12697c0101a6b3cc1a47658a0ce0eb0b2f26e3e495e6766305772

Observation 257d8de6-f7ab-479f-a1c0-9d9ceca21ecb · outbound

This paper cites Correlation-Guided Query-Dependency Calibration for Video Temporal Grounding.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking Correlation-Guided Query-Dependency Calibration for Video Temporal Grounding

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:50:17.328936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:b42097008719bdc60519c893eaafd043532a25788e758218293ff701d1ec4f29

Observation ef59aff4-b3a7-416c-a779-3371d284ebe1 · outbound

This paper cites Unbiased scene graph generation in videos.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking Unbiased scene graph generation in videos

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:50:17.865731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:4a3edb1ef0fd033da5aaa02f9650b84a51f26b8453fbc766da6d7c1974e3e87d

Observation 4fcf03e9-060d-4cd6-9dd5-99f900caba39 · outbound

This paper cites Hig: Hierarchical interlacement graph approach to scene graph generation in video understanding.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking Hig: Hierarchical interlacement graph approach to scene graph generation in video understanding

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:50:17.860904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:3b3b60560b27eb95c35115dca462878b01d75b31bd8dd8bfa356993f4fd83218

Observation 96258dd8-77d7-4a12-bad6-7a836e461633 · outbound

This paper cites Hyperglm: Hypergraph for video scene graph generation and anticipation.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking Hyperglm: Hypergraph for video scene graph generation and anticipation

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:50:17.921448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:84d49e85f0529515d70de3ce939c459507edd96169b74dfbc058a43fbb0555c3

Observation 4c634398-e264-4c3e-a689-5b5c796bb710 · outbound

This paper cites Gpt-4 technical report.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking Gpt-4 technical report

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:50:17.914392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:951341805152d81aedf85330435ff891488422d97d367eda1ead349fddd78894

Observation 20cc96f2-af3b-453c-9e46-92bf9ad55239 · outbound

This paper cites Gpt-4o system card.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking Gpt-4o system card

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:50:17.910864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:208ca4d470b864baaa840a1d22bc8c8f60b610bab77e8a7b38d825de30c484b4

Observation ac77e235-4cbe-41f3-a396-d25bf3c38437 · outbound

This paper cites ZoomV: Temporal Zoom-in for Efficient Long Video Understanding.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking ZoomV: Temporal Zoom-in for Efficient Long Video Understanding

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-08-06T02:01:23.268869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:2fe4ac365c5557b1d0f208e95b6f97c2c596486fddb818cfa948902b5818ede5

Observation e4b213bb-51c7-4153-a0e0-5a20dfa05fb9 · outbound

This paper cites Question- answering dense video events.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking Question- answering dense video events

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:50:17.859108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:970dd35f52c7fcc17a38d5b4a19277f33a36971aa47ca137087d38799f68783a

Observation b22e985c-0c3e-4a94-a3b4-34efded62655 · outbound

This paper cites Step: Enhancing video-llms’ com- positional reasoning by spatio-temporal graph-guided self- training.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking Step: Enhancing video-llms’ com- positional reasoning by spatio-temporal graph-guided self- training

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:50:17.925612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:9137b362a3834a662ac4b85183f73a93bed78280d1238a783a76ff782ada3be8

Observation 52eaa934-63d4-4398-9834-78c42842010b · outbound

This paper cites Timechat: A time-sensitive multimodal large lan- guage model for long video understanding.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking Timechat: A time-sensitive multimodal large lan- guage model for long video understanding

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:50:17.864409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:0335c38972787afe7bcef9a2065ea42278effc45d7be4e41cca0e80e4efa7db5

Observation 9ede4606-d07b-4c39-bc6f-5ba72652507a · outbound

This paper cites VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-05-15T20:50:17.290213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:fb9935ceec3ee3c03d286daec53480aab03ad3a68aa8b492d8b40fe5f28f8fae

Observation fb90480b-a89e-4a6f-a611-2bf5662a0432 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking Gemini: A Family of Highly Capable Multimodal Models

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-05-15T20:50:17.348363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:2b9885bc232a7773b1900e6558fb25ae7c3a71eb3a5f4d230fd6a70178b8ea5f

Observation e28eb4f0-7cfa-45e4-880c-b2695d6e4ce8 · outbound

This paper cites Reka Core, Flash, and Edge: A Series of Powerful Multimodal Language Models.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking Reka Core, Flash, and Edge: A Series of Powerful Multimodal Language Models

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:50:17.299732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:a5280617b511786fd2670edadeca7c38d8d9f3e3e94a5af84fd69d219e0b319a

Observation 511600ad-fa45-463a-b372-d1339b3197bd · outbound

This paper cites Causal ai scientist: Facilitating causal data science with large language models.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking Causal ai scientist: Facilitating causal data science with large language models

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:50:17.931383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:b081123d08f71ee57a4ddb130ce82489fac06476def2b1db799ca0111fe6849f

Observation e5003bc1-b6e4-4d96-84f5-f96518f6cdbd · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-05-15T20:50:17.295417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:d90ccd805102071a27475dd411794ef23fbae670cc0100ac0931f2cc37dcd5a8

Observation 4495bbd2-de89-44cc-a83e-b0b39d756171 · outbound

This paper cites Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:40:06.763282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:a20463b448cb5d4cbf474426f342ad06b5cf97574fe30417b106cd0ba007c0f7

Observation d43ac4bc-f95d-4211-8c3b-b48c6a578c23 · outbound

This paper cites Sportshhi: A dataset for human-human interaction detection in sports videos.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking Sportshhi: A dataset for human-human interaction detection in sports videos

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:50:17.927223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:f108f205fecd0ee842f2aeb067f1072530fd69d6654a27298c97e154eca14c71

Observation d476a100-c10d-4a3b-b461-6bea06375118 · outbound

This paper cites Visionary-r1: Mitigating shortcuts in visual reasoning with reinforcement learning.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking Visionary-r1: Mitigating shortcuts in visual reasoning with reinforcement learning

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:50:17.340125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:60e5af019f426d97fdf43d26da1d5f073270b70af73913e3f9d4b91d46996aca

Observation 08557017-0441-455c-9ffa-9aaec944151f · outbound

This paper cites Can i trust your answer? visually grounded video question answering.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking Can i trust your answer? visually grounded video question answering

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:50:17.918002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:0bc3527d4a7aa1958997e04617f98ad1b4cc3e22903eccce006398025436595b

Observation a077a12b-9366-4d12-bdf0-aa1653f2517d · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-05-15T20:50:17.342034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:ae104b94bb1df1fd5c6ea2557d5f5cc0772452e8680ca2b6387055bb5c1498b8

Observation 57132801-4c3f-436a-8365-f5cafe8907da · outbound

This paper cites Panoptic video scene graph generation.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking Panoptic video scene graph generation

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:50:17.916364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:b623f84a2c220b8577b96da3fe2bcadd8bfca6d085fa9ab6b0283a2daf71ffe0

Observation 44b0c5ee-6c92-4316-93b8-d39258c929d7 · outbound

This paper cites Thinking in space: How mul- timodal large language models see, remember, and recall spaces.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking Thinking in space: How mul- timodal large language models see, remember, and recall spaces

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:50:17.919647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:657f90156ee43a7b3b616b784d15c3ebec176e8c1fc63b9e2e72ce81ccfb6e61

Observation b05026ca-166f-4fb1-ba7a-1d870bce46dd · outbound

This paper cites VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

Reference 69

Resolution
verified exact
local_arxiv, observed 2026-05-15T20:50:17.353150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:99953effa775689b89448b612ed9692c0b7e38fdb0bcd8d9b3720859824680cb

Observation 8ef12da1-0c3f-4e51-aeab-d8f38fddd796 · outbound

This paper cites Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:50:17.308558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:39169c4ae3ecc8b361d8778ec1334a3e44fff65990cf2818faffe669307a366a

Observation d712bd22-8129-47e5-bdc7-8484d71ea0e1 · outbound

This paper cites Vtime- 11 cot: Thinking by drawing for video temporal grounding and reasoning.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking Vtime- 11 cot: Thinking by drawing for video temporal grounding and reasoning

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:50:17.905072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:4c95e90b283de2d6e22e3fc1a77700f1ace417dcbf4b5f5e8a7d25292ee0ba98

Observation 3c1f2cfd-c49a-4438-a077-9db5e8271014 · outbound

This paper cites Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:50:17.350435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:be4b230acd3314c528b90d15f39d54386b9a51e2b358f5e56709afbabc7408c3

Observation 05b6845a-ca89-4feb-9deb-3c6c34cc87a3 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 73

Resolution
verified exact
local_arxiv, observed 2026-05-15T20:50:17.302649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:0831cc5467fd03858953519a5247f79a2e81e69ce8d120a0a8f490f7e6243530

Observation 94d69096-4dc9-45a8-af32-517308399139 · outbound

This paper cites Towards video thinking test: A holistic benchmark for advanced video reasoning and understanding.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking Towards video thinking test: A holistic benchmark for advanced video reasoning and understanding

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:50:17.901551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:f8cfa5b0009e25b45958ae6986a4c6d3422ccdbf819a7f27982c8b68fc60c24f

Pith citing papers

Observation ca9de3d6-bab6-4863-899b-95acdc4c3104 · inbound

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence cites this paper.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:13:59.634880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:4653928e286ee5fa432aa6347320061310ce4295af86fc5696f184c34f0c891e

Observation 85dba891-455d-446f-8815-b7e93f4c9785 · inbound

EventCoT: Event-centric Video Chain-of-thought for Reasoning Temporal Localization cites this paper.

EventCoT: Event-centric Video Chain-of-thought for Reasoning Temporal Localization GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-11T12:10:52.100410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T12:10:52.100410Z digest=sha256:a5b4b8dbfb6527f10a1139cb6b8a4a064cf059cbfac156a98898c46d8fc2bd01