Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:10:50.329904Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 5 inbound Pith citation observations for arXiv:2505.22457.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:10:50.329904Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-03T09:30:02.227226Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T19:30:08.052620Z
65 of 65 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation a8b243e2-10bb-45c5-9874-d119d25b607b · outbound
Fostering Video Reasoning via Next-Event Prediction GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1344f05-fac6-4083-b108-4d82a83f50ae · outbound
Fostering Video Reasoning via Next-Event Prediction Qwen Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7b28bc2-524b-4bd9-828b-d89f18750afa · outbound
Fostering Video Reasoning via Next-Event Prediction Qwen2.5-VL Technical Report
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b97c9119-9168-41fa-acfe-81ee98f6fb38 · outbound
Fostering Video Reasoning via Next-Event Prediction Perspective—discovery within validation logic: Deliberately surfacing, complementing, and substituting abductive reasoning in hypothetico- deductive inquiry
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8c97c53f-09c4-49cc-a2be-eebe6f273efc · outbound
Fostering Video Reasoning via Next-Event Prediction TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ef549fe-3383-4fb6-8f03-86eaa348bf98 · outbound
Fostering Video Reasoning via Next-Event Prediction Exploring the Effect of Reinforcement Learning on Video Understanding: Insights from SEED-Bench-R1
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0791a0bd-f9d3-47f2-86ad-e93cedb888e3 · outbound
Fostering Video Reasoning via Next-Event Prediction Inductive or Deductive? Rethinking the Fundamental Reasoning Abilities of LLMs
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec8e0ee3-de3c-4682-8d6c-becc925384ba · outbound
Fostering Video Reasoning via Next-Event Prediction Lost in Time: A New Temporal Benchmark for VideoLLMs
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbbad63a-714e-4059-9099-f57c7e3879de · outbound
Fostering Video Reasoning via Next-Event Prediction Abduction
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1b26eba1-f86e-44d0-a8c9-12b66e821277 · outbound
Fostering Video Reasoning via Next-Event Prediction Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e760cb1-9cb8-4792-91bb-0b81ab3a281b · outbound
Fostering Video Reasoning via Next-Event Prediction Predicting the future: A jointly learnt model for action anticipation
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9b41be81-a1e1-4170-869d-90f1cb2964ec · outbound
Fostering Video Reasoning via Next-Event Prediction Inductive and deductive reasoning
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation da3b778a-9a18-401d-8f84-3217ffe933f1 · outbound
Fostering Video Reasoning via Next-Event Prediction DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96b992c6-8922-45e8-99df-b3e8f37d511e · outbound
Fostering Video Reasoning via Next-Event Prediction Large Language Models Are Reasoning Teachers
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b1b8a55-698f-4d8f-be06-2c98a3185948 · outbound
Fostering Video Reasoning via Next-Event Prediction GPT-4o System Card
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14476e3b-fe1c-4b0a-b790-b5bb495e1be7 · outbound
Fostering Video Reasoning via Next-Event Prediction A hierarchical representation for future action prediction
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 59de5145-1f66-4698-94f4-6ffbe66328ce · outbound
Fostering Video Reasoning via Next-Event Prediction LLaVA-OneVision: Easy Visual Task Transfer
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 481d3352-d8e1-4c05-97e5-2b773533b625 · outbound
Fostering Video Reasoning via Next-Event Prediction LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a60049e2-c685-48d0-a8f6-1cad7c51724f · outbound
Fostering Video Reasoning via Next-Event Prediction Mvbench: A comprehensive multi-modal video understanding benchmark
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 450bb94b-80e9-4bc3-ae25-9225c7105aa6 · outbound
Fostering Video Reasoning via Next-Event Prediction A survey of multimodel large language models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84733f08-cbc8-4014-8a9b-a32583d0674b · outbound
Fostering Video Reasoning via Next-Event Prediction Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0160fd17-3872-42e9-8f2a-753a0538dea6 · outbound
Fostering Video Reasoning via Next-Event Prediction Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc1f4c3c-6c0f-41b0-a710-4a3506c49205 · outbound
Fostering Video Reasoning via Next-Event Prediction TempCompass: Do Video LLMs Really Understand Videos?
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afe341c0-7073-4c1d-882c-34ab70420a31 · outbound
Fostering Video Reasoning via Next-Event Prediction Training language models to follow instructions with human feedback
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9610f74-06b6-42ca-8980-6a344a6c5767 · outbound
Fostering Video Reasoning via Next-Event Prediction Learning transferable visual models from natural language supervision
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 454627e6-d7b8-46a8-8f4f-5d3637e20f86 · outbound
Fostering Video Reasoning via Next-Event Prediction Video (language) modeling: a baseline for generative models of natural videos
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 633a548a-9fec-4e11-aa6c-8bd1ebd3fa99 · outbound
Fostering Video Reasoning via Next-Event Prediction DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9bf018a1-6fb0-448d-b95e-efa4b250609d · outbound
Fostering Video Reasoning via Next-Event Prediction HybridFlow: A Flexible and Efficient RLHF Framework
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f42e9495-07d5-43ff-a659-14d3113954da · outbound
Fostering Video Reasoning via Next-Event Prediction To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e931fd75-e0b7-4087-af4f-9c8819e9dbff · outbound
Fostering Video Reasoning via Next-Event Prediction The wisdom of crowds: Temporal progressive attention for early action prediction
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f8ba82e5-fe4e-461e-8d5e-52c2b8254de2 · outbound
Fostering Video Reasoning via Next-Event Prediction Video understanding with large language models: A survey
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5465fe51-8e45-4d3b-86f2-106cb8700161 · outbound
Fostering Video Reasoning via Next-Event Prediction Gemini: A Family of Highly Capable Multimodal Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02023953-c1a5-41ac-8ba0-e57c4663e488 · outbound
Fostering Video Reasoning via Next-Event Prediction Generating videos with scene dynamics
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 761182a1-f92a-4d9d-b1f1-15faf9eef1d0 · outbound
Fostering Video Reasoning via Next-Event Prediction Critique Fine-Tuning: Learning to Critique is More Effective than Learning to Imitate
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2aa31ecd-1b36-4c79-a23a-2fd80ed5d43f · outbound
Fostering Video Reasoning via Next-Event Prediction Chain-of-thought prompting elicits reasoning in large language models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d84633e9-bfe1-4859-8de3-ad1ba266b0c6 · outbound
Fostering Video Reasoning via Next-Event Prediction Longvideobench: A benchmark for long- context interleaved video-language understanding
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7786aaa2-42f6-4983-8025-236172866b37 · outbound
Fostering Video Reasoning via Next-Event Prediction Next-qa: Next phase of question- answering to explaining temporal actions
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4bdd91e-916d-48b7-8ae6-ba21a585a8d9 · outbound
Fostering Video Reasoning via Next-Event Prediction Tree of thoughts: Deliberate problem solving with large language models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f62cfd8-7bdb-4bc1-b40d-e42e352344dc · outbound
Fostering Video Reasoning via Next-Event Prediction ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 554e8628-ed79-4325-8914-01eb210ef96d · outbound
Fostering Video Reasoning via Next-Event Prediction Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 154b2e67-dfd9-49d3-8247-306637013fb2 · outbound
Fostering Video Reasoning via Next-Event Prediction LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5065b583-6453-4089-8ddf-b0e1f36b6f9d · outbound
Fostering Video Reasoning via Next-Event Prediction LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f835619-fbea-477d-a8f5-aa583d881a46 · outbound
Fostering Video Reasoning via Next-Event Prediction LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9a95e2c-c00f-4580-9b5d-6551957c4539 · outbound
Fostering Video Reasoning via Next-Event Prediction Achild…”},{Scene2:“Legsare…
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6208c882-a1e4-4b0e-bd89-8c07fdcb3b5c · outbound
Fostering Video Reasoning via Next-Event Prediction Unresolved cited work
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d76ef48b-4255-4107-a1a6-693015df19f4 · outbound
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation db8aad78-2cee-4390-97f6-b7b146a25760 · outbound
Fostering Video Reasoning via Next-Event Prediction Unresolved cited work
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9ab6bacb-9e18-42c4-8962-c6b42e505ed6 · outbound
Fostering Video Reasoning via Next-Event Prediction Unresolved cited work
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f597f086-9003-4464-ac15-afa20cba306b · outbound
Fostering Video Reasoning via Next-Event Prediction suitable
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 92fcad3b-2201-482f-a665-5f0f263ba338 · outbound
Fostering Video Reasoning via Next-Event Prediction d e s c r i p t i o n
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 65ad7b78-c82f-496e-b895-ac0d179512ba · outbound
Fostering Video Reasoning via Next-Event Prediction Unresolved cited work
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3dc27cc6-0011-43e3-9455-e1c67fec9fd4 · outbound
Fostering Video Reasoning via Next-Event Prediction d e s c r i p t i o n
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 84b3ed2f-467c-4790-8f1d-79d50c88cff6 · outbound
Fostering Video Reasoning via Next-Event Prediction d e s c r i p t i o n
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation eccaab00-adda-4cac-ae6f-5c2cb4d935ad · outbound
Fostering Video Reasoning via Next-Event Prediction Unresolved cited work
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1f1cdcd4-1d27-43a9-b28a-efcbd22a09db · outbound
Fostering Video Reasoning via Next-Event Prediction Unresolved cited work
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b511f957-3141-4348-bff4-5b1eab709b5e · outbound
Fostering Video Reasoning via Next-Event Prediction Unresolved cited work
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 22328854-7e91-457c-ab7c-b0555fe4e041 · outbound
Fostering Video Reasoning via Next-Event Prediction C o n c l u s i o n : right
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 627e4afa-4a32-46f5-b949-8c9971d99f08 · outbound
Fostering Video Reasoning via Next-Event Prediction Question
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ff0e6b85-545b-419d-ab95-d3c6aa46106c · outbound
Fostering Video Reasoning via Next-Event Prediction - Ensure only one correct answer and that the r em ai nin g three options are wrong
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 13c7557d-d7c2-4d67-b2ab-c3048e22dcae · outbound
Fostering Video Reasoning via Next-Event Prediction Question
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0a22f794-bfb4-45da-bb6b-cda04bf24d4c · outbound
Fostering Video Reasoning via Next-Event Prediction - Ensure only one correct answer and that the r em ai nin g three options are wrong
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3e9e1026-edaa-4cd2-b433-e990e86b4565 · outbound
Fostering Video Reasoning via Next-Event Prediction Question
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 02fa0c0c-327e-4877-86c2-325c0418a7f0 · outbound
Fostering Video Reasoning via Next-Event Prediction - Ensure only one correct answer and that the r em ai nin g three options are wrong
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation aaf58cc8-58b2-43d2-b779-18a2023b4d3d · outbound
Fostering Video Reasoning via Next-Event Prediction Question
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 26d5b6dc-02ae-4016-a9b7-aec1ff6e3ed3 · outbound
Fostering Video Reasoning via Next-Event Prediction - Answer options should be built upon the scenes after the observed scenes and before the last scene
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 999cd9e9-7450-41b6-81d1-bffa35dbba45 · inbound
FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs Fostering Video Reasoning via Next-Event Prediction
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a9dbaae-8460-4fa6-a39e-5f07afa41cf8 · inbound
EgoSelf: From Memory to Personalized Egocentric Assistant Fostering Video Reasoning via Next-Event Prediction
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d5dbfd85-61ed-4384-aae0-8589d7c3ab8a · inbound
EmoTrans: A Benchmark for Understanding, Reasoning, and Predicting Emotion Transitions in Multimodal LLMs Fostering Video Reasoning via Next-Event Prediction
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b5ace437-a666-4879-885c-3a35be18a4dc · inbound
Imagine Before You Predict: Interleaved Latent Visual Reasoning for Video Event Prediction Fostering Video Reasoning via Next-Event Prediction
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 416437f4-13df-4c83-977f-687485dc6d1b · inbound
FeVOS: Foresight Expression Video Object Segmentation Fostering Video Reasoning via Next-Event Prediction
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.