Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-12T09:09:15.248815Z
Paper Citation Record · LEDGER
As of 6 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 1 inbound Pith citation observation for arXiv:2607.02607.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-12T09:09:15.248815Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-01T17:47:13.597841Z
A source-named dated measurement, never combined with another source.
Source: cited_works
60 of 60 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 3bbd38c2-e143-40bf-afa6-1957de4191dc · outbound
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e56d3b84-e583-4396-ad0b-4e707daf6bb1 · outbound
Latent Visual Cache for Video Reasoning Vision language models in autonomous driving: A survey and outlook.IEEE Transactions on Intelligent Vehicles, 2024
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69fa301c-f4e0-4958-ab71-15c30a20c502 · outbound
Latent Visual Cache for Video Reasoning Vad: Vectorized scene representation for efficient autonomous driving
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9c9a082-004e-4b32-b35f-095feea72685 · outbound
Latent Visual Cache for Video Reasoning Video generation models as world simulators
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33cd0623-9126-41cc-96ac-b9f1090fdcad · outbound
Latent Visual Cache for Video Reasoning Wan: Open and advanced large-scale video generative models, 2025
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55c5adcd-9e7e-4785-8998-848a30f207f7 · outbound
Latent Visual Cache for Video Reasoning The past mistake is the future wisdom: Error-driven contrastive probability optimization for chinese spell checking
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b9b706f-a077-49ca-8152-1b253f3b6527 · outbound
Latent Visual Cache for Video Reasoning Let’s think with images efficiently! an interleaved-modal chain-of-thought reasoning framework with dynamic and precise visual thoughts
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29ec90d7-9c7e-417a-beb3-0147665c700a · outbound
Latent Visual Cache for Video Reasoning Youtu-llm: Unlocking the native agentic potential for lightweight large language models, 2026
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f86084fc-f191-4f9a-8c9f-8af9945009d7 · outbound
Latent Visual Cache for Video Reasoning Kimi-VL Technical Report
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 182386d5-da9d-4b38-a0a3-b32b0c21d0ee · outbound
Latent Visual Cache for Video Reasoning Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 609a7e00-84a1-4cac-9f6e-8560753367ed · outbound
Latent Visual Cache for Video Reasoning Openai gpt-5 system card, 2025
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f248c81-668c-4268-a7ca-286ff0f6de64 · outbound
Latent Visual Cache for Video Reasoning Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb15ed41-eb52-4cf7-83a7-3eeb282da656 · outbound
Latent Visual Cache for Video Reasoning Unresolved cited work
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6336a6b7-b930-4190-b7c5-1433443b967b · outbound
Latent Visual Cache for Video Reasoning Cognitive Mismatch in Multimodal Large Language Models for Discrete Symbol Understanding
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56ddba45-e26e-4da8-923f-16ebd976c2f1 · outbound
Latent Visual Cache for Video Reasoning Video-chatgpt: Towards detailed video understanding via large vision and language models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a243d13e-e109-48e7-9ced-a9e5d91bcc5a · outbound
Latent Visual Cache for Video Reasoning Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b591ea9b-7666-4e83-a32c-a2b3aaa26823 · outbound
Latent Visual Cache for Video Reasoning VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7be36d32-4b03-4d44-a5b4-61b288913e7d · outbound
Latent Visual Cache for Video Reasoning Internvl3.5: Advancing open-source multimodal models in versatility, reasoning, and efficiency, 2025
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0cce0f75-ac71-455d-a92b-7ee08cc8cd80 · outbound
Latent Visual Cache for Video Reasoning AutoCAP: Towards automatic cross-lingual alignment planning for zero-shot chain-of-thought
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0aa1d5c1-0320-439f-81aa-10102372b1f9 · outbound
Latent Visual Cache for Video Reasoning Qwen2-vl: Enhancing vision- language model’s perception of the world at any resolution, 2024
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 571f510e-56ee-49d4-be9a-24ffbcd8901a · outbound
Latent Visual Cache for Video Reasoning Moviechat: From dense token to 11 sparse memory for long video understanding
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f52c87b-d21b-4ee2-9e19-93137f97795c · outbound
Latent Visual Cache for Video Reasoning Wrong-of-thought: An integrated reasoning framework with multi-perspective verification and wrong information
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25bfd2fe-7451-4b06-96f4-2d50707e5b6a · outbound
Latent Visual Cache for Video Reasoning Qwen3-vl technical report, 2025
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3958ca1a-4f2a-4066-849f-719caf024697 · outbound
Latent Visual Cache for Video Reasoning Gemma 3 technical report, 2025
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 337f02a6-4e19-480c-9450-0f73058640a7 · outbound
Latent Visual Cache for Video Reasoning CCHall: A novel benchmark for joint cross-lingual and cross-modal hallucinations detection in large language models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4b504f6-641d-4a0a-b1b5-d38fc1bf8f93 · outbound
Latent Visual Cache for Video Reasoning Mitigating visual forgetting via take-along visual conditioning for multi-modal long cot reasoning
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ef79481-a29b-4e86-823e-f9a9a23e6f4f · outbound
Latent Visual Cache for Video Reasoning Seeing through the chain: Mitigate hallucination in multimodal reasoning models via cot compression and contrastive preference optimization.arXiv preprint arXiv:2602.03380, 2026
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 993d79d9-c14e-40eb-ab8c-d3bfc9ed7830 · outbound
Latent Visual Cache for Video Reasoning Context length alone hurts llm performance despite perfect retrieval.arXiv preprint arXiv:2510.05381, 2025
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c8c7acb-7923-4123-addc-45c73cecc36d · outbound
Latent Visual Cache for Video Reasoning Visual hallucinations of multi-modal large language models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 569710b0-7a25-44f3-86f0-3fa73b97db47 · outbound
Latent Visual Cache for Video Reasoning Vigc: Visual instruction generation and correction
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ac93fe6-1637-4287-a98c-47ed8dc3837b · outbound
Latent Visual Cache for Video Reasoning Hourvideo: 1-hour video- language understanding.Advances in Neural Information Processing Systems, 37:53168–53197, 2024
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 379a674e-ed8d-4f53-b14f-c3c70947bf0a · outbound
Latent Visual Cache for Video Reasoning Qwen3.5: Towards native multimodal agents, February 2026
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94fcc864-d632-488c-9bf1-ebc6a73093f7 · outbound
Latent Visual Cache for Video Reasoning TempCompass: Do Video LLMs Really Understand Videos?
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f761258-7789-40b7-a964-ca38a7dc3464 · outbound
Latent Visual Cache for Video Reasoning Mvbench: A comprehensive multi-modal video understanding benchmark
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4c2cb37-99d7-49b1-8cd2-1efa60cbe954 · outbound
Latent Visual Cache for Video Reasoning MMVU: Measuring Expert-Level Multi-Discipline Video Understanding
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79b7c0e5-ac12-46d6-bdf4-d8a4afaae848 · outbound
Latent Visual Cache for Video Reasoning Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60c3f555-64cd-49c8-ac60-787c95fbc898 · outbound
Latent Visual Cache for Video Reasoning Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7a51859-3c19-437f-8d27-b043a7d6ab6a · outbound
Latent Visual Cache for Video Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99b40fe0-9b62-4c20-81ed-7adfafccec25 · outbound
Latent Visual Cache for Video Reasoning Open-o3-video: Grounded video reasoning with explicit spatio-temporal evidence.arXiv preprint arXiv:2510.20579, 2025
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 159025a9-c049-4390-8a1c-323bf468d698 · outbound
Latent Visual Cache for Video Reasoning Video-R1: Reinforcing Video Reasoning in MLLMs
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35d8a661-3f92-4a48-94f2-fc65384c5157 · outbound
Latent Visual Cache for Video Reasoning LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1f6a88e-21b4-4933-a8f8-5e9fd83e794c · outbound
Latent Visual Cache for Video Reasoning VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40a3a8b0-f12c-481e-8a7c-ad113e85c7d1 · outbound
Latent Visual Cache for Video Reasoning Long Context Transfer from Language to Vision
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71df1acf-dc5e-426b-a1f6-e2a36be71d5f · outbound
Latent Visual Cache for Video Reasoning Vila: On pre-training for visual language models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ecd8c8d3-10c9-4324-8578-0a89ae81813b · outbound
Latent Visual Cache for Video Reasoning Unhackable Temporal Rewarding for Scalable Video MLLMs
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4336ba2-ee6e-4266-80e1-6c5feb0c7415 · outbound
Latent Visual Cache for Video Reasoning LLaVA-OneVision: Easy Visual Task Transfer
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01d453a6-8134-402f-ba94-496cc1c508df · outbound
Latent Visual Cache for Video Reasoning Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1307fab0-b333-4ece-a3b5-91afc10ea3b1 · outbound
Latent Visual Cache for Video Reasoning Chain-of-thought prompting elicits reasoning in large language models
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c582509d-ebf2-4e43-aa05-0c74b6d342ca · outbound
Latent Visual Cache for Video Reasoning Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b60f57ee-590d-4c42-9387-e709ebb22a73 · outbound
Latent Visual Cache for Video Reasoning Video-llava: Learning united visual representation by alignment before projection
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d96fd88-739f-4aa3-acbd-c2684148a4d4 · outbound
Latent Visual Cache for Video Reasoning Unresolved cited work
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b01da90-e26c-4395-a3ad-47823087f670 · outbound
Latent Visual Cache for Video Reasoning Thinking with videos: Multimodal tool-augmented reinforcement learning for long video reasoning
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e3ebc10-c568-4798-884f-7eb7dc5775ee · outbound
Latent Visual Cache for Video Reasoning Video-of-thought: step-by-step video reasoning from perception to cognition
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27d46470-37dc-4282-a15f-b80223777081 · outbound
Latent Visual Cache for Video Reasoning Vitcot: Video-text interleaved chain-of-thought for boosting video understanding in large language models
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69ea96d0-e42a-40c0-a0ef-0980032c5ac1 · outbound
Latent Visual Cache for Video Reasoning Training Large Language Models to Reason in a Continuous Latent Space
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 307331b4-7612-45cf-8db7-08264b22e4b5 · outbound
Latent Visual Cache for Video Reasoning SoftCoT: Soft chain-of-thought for efficient reasoning with LLMs
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2000a443-5da2-48a0-8880-1f649cca7e5d · outbound
Latent Visual Cache for Video Reasoning Hybrid latent reasoning via reinforcement learning.Advances in Neural Information Processing Systems, 38:5501–5530, 2026
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ccbdba9-ab49-4122-be42-b98628f36953 · outbound
Latent Visual Cache for Video Reasoning Reasoning beyond language: A comprehensive survey on latent chain-of-thought reasoning.arXiv preprint arXiv:2505.16782, 2025
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 152b8ca8-3380-46b6-b07c-c9a008e6cc01 · outbound
Latent Visual Cache for Video Reasoning Monet: Reasoning in latent visual space beyond image and language
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74f0c947-6424-4ac1-a4e3-74e3928e4042 · outbound
Latent Visual Cache for Video Reasoning Scaling up test-time compute with latent reasoning: A recurrent depth approach.Advances in Neural Information Processing Systems, 38:41340–41391, 2026
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30cf12d7-beed-4d12-9caa-a1ba8ba98eb2 · inbound
Thinking in Video: Can Video Generators Really Reason About the Real World? Latent Visual Cache for Video Reasoning
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.