Pith. sign in

Paper Citation Record · LEDGER

Latent Visual Cache for Video Reasoning

As of 6 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 1 inbound Pith citation observation for arXiv:2607.02607.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.02607 v1

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-12T09:09:15.248815Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T17:47:13.597841Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

60 of 60 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved59
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3bbd38c2-e143-40bf-afa6-1957de4191dc · outbound

This paper cites A Generalist Agent.

Latent Visual Cache for Video Reasoning A Generalist Agent

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-12T09:09:15.248815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:09:15.248815Z digest=sha256:983c4e7899c0566903981472d93f5fb36b997e9cffae3dde95b87b5f443edab6

Observation e56d3b84-e583-4396-ad0b-4e707daf6bb1 · outbound

This paper cites Vision language models in autonomous driving: A survey and outlook.IEEE Transactions on Intelligent Vehicles, 2024.

Latent Visual Cache for Video Reasoning Vision language models in autonomous driving: A survey and outlook.IEEE Transactions on Intelligent Vehicles, 2024

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-12T09:09:15.248815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:09:15.248815Z digest=sha256:9506adbcba8ae4ae1c03c8cc0a6efe7f8471a70ad50611c4ae47a9df25db46aa

Observation 69fa301c-f4e0-4958-ab71-15c30a20c502 · outbound

This paper cites Vad: Vectorized scene representation for efficient autonomous driving.

Latent Visual Cache for Video Reasoning Vad: Vectorized scene representation for efficient autonomous driving

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-12T09:09:15.248815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:09:15.248815Z digest=sha256:fcb331afa76a46bd1645e6582bb9480bae71da9111f68bf2560ab6d6a9893ce5

Observation e9c9a082-004e-4b32-b35f-095feea72685 · outbound

This paper cites Video generation models as world simulators.

Latent Visual Cache for Video Reasoning Video generation models as world simulators

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-12T09:09:15.248815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:09:15.248815Z digest=sha256:3860abedbeb61a660d015a55f1bed20add9f597105bfb088bc21318019d8e45c

Observation 33cd0623-9126-41cc-96ac-b9f1090fdcad · outbound

This paper cites Wan: Open and advanced large-scale video generative models, 2025.

Latent Visual Cache for Video Reasoning Wan: Open and advanced large-scale video generative models, 2025

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-12T09:09:15.248815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:09:15.248815Z digest=sha256:efb423fc13caa559e7e2d23b3b4ebdd7588fa0a2f1bf8eca5d7539ad118e7fac

Observation 55c5adcd-9e7e-4785-8998-848a30f207f7 · outbound

This paper cites The past mistake is the future wisdom: Error-driven contrastive probability optimization for chinese spell checking.

Latent Visual Cache for Video Reasoning The past mistake is the future wisdom: Error-driven contrastive probability optimization for chinese spell checking

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-12T09:09:15.248815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:09:15.248815Z digest=sha256:8a62a3a80a06c177a44d167aa00a04ffc2b620dd31b18efdc460b8a7b6af1037

Observation 2b9b706f-a077-49ca-8152-1b253f3b6527 · outbound

This paper cites Let’s think with images efficiently! an interleaved-modal chain-of-thought reasoning framework with dynamic and precise visual thoughts.

Latent Visual Cache for Video Reasoning Let’s think with images efficiently! an interleaved-modal chain-of-thought reasoning framework with dynamic and precise visual thoughts

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-12T09:09:15.248815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:09:15.248815Z digest=sha256:5b73e3cbd4aa1e94163e2b25a5954c03ea8958c6510720d2e11e3ba19e4a6be8

Observation 29ec90d7-9c7e-417a-beb3-0147665c700a · outbound

This paper cites Youtu-llm: Unlocking the native agentic potential for lightweight large language models, 2026.

Latent Visual Cache for Video Reasoning Youtu-llm: Unlocking the native agentic potential for lightweight large language models, 2026

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-12T09:09:15.248815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:09:15.248815Z digest=sha256:a5ea146ed30795d48a70abfbb13ad14ce9ac46f7f1724654b32dbaca481118b9

Observation f86084fc-f191-4f9a-8c9f-8af9945009d7 · outbound

This paper cites Kimi-VL Technical Report.

Latent Visual Cache for Video Reasoning Kimi-VL Technical Report

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-12T09:09:15.248815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:09:15.248815Z digest=sha256:3d1721c829cb91fc81ea1545b5bfda007463ad348c239d3c1ae1d3ed63ec6e62

Observation 182386d5-da9d-4b38-a0a3-b32b0c21d0ee · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

Latent Visual Cache for Video Reasoning Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-12T09:09:15.248815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:09:15.248815Z digest=sha256:7845d58ab1b9f2cd5b995ae4871a31810314a1a7a183f9194e856e3312394c98

Observation 609a7e00-84a1-4cac-9f6e-8560753367ed · outbound

This paper cites Openai gpt-5 system card, 2025.

Latent Visual Cache for Video Reasoning Openai gpt-5 system card, 2025

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-12T09:09:15.248815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:09:15.248815Z digest=sha256:1305c45c5acb8fe66e8435ac1fe43d56d4911cfabc4f2013047a94f11166b1cb

Observation 6f248c81-668c-4268-a7ca-286ff0f6de64 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Latent Visual Cache for Video Reasoning Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-12T09:09:15.248815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:09:15.248815Z digest=sha256:2322e4555d50fa9f4c3385f59e8fe7e6869fd2ef156ba7bdc6cec87098aaa978

Observation eb15ed41-eb52-4cf7-83a7-3eeb282da656 · outbound

This paper cites an unresolved cited work.

Latent Visual Cache for Video Reasoning Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-12T09:09:15.248815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:09:15.248815Z digest=sha256:0da738b664d97e54f1ad0c5d20423bba54645d1f7cf92e335d27232ba84baf20

Observation 6336a6b7-b930-4190-b7c5-1433443b967b · outbound

This paper cites Cognitive Mismatch in Multimodal Large Language Models for Discrete Symbol Understanding.

Latent Visual Cache for Video Reasoning Cognitive Mismatch in Multimodal Large Language Models for Discrete Symbol Understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-12T09:09:15.248815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:09:15.248815Z digest=sha256:112dd428fb538a37eaf3841826caf56a333b381d5761acee5f2fb37c770d4b28

Observation 56ddba45-e26e-4da8-923f-16ebd976c2f1 · outbound

This paper cites Video-chatgpt: Towards detailed video understanding via large vision and language models.

Latent Visual Cache for Video Reasoning Video-chatgpt: Towards detailed video understanding via large vision and language models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-12T09:09:15.248815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:09:15.248815Z digest=sha256:f11dcb3ddfe83af89cfccfd950456d95e678ff9ba0712b05b1696534cd0983fb

Observation a243d13e-e109-48e7-9ced-a9e5d91bcc5a · outbound

This paper cites Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis.

Latent Visual Cache for Video Reasoning Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-12T09:09:15.248815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:09:15.248815Z digest=sha256:3903884c0251bc504ad64d940a3b586304dcce93d892a1c99f09c0f10e9bd191

Observation b591ea9b-7666-4e83-a32c-a2b3aaa26823 · outbound

This paper cites VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding.

Latent Visual Cache for Video Reasoning VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-12T09:09:15.248815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:09:15.248815Z digest=sha256:1f0406071afe5b883b332ca100229d561363eaa3448ad120857b392763c44df8

Observation 7be36d32-4b03-4d44-a5b4-61b288913e7d · outbound

This paper cites Internvl3.5: Advancing open-source multimodal models in versatility, reasoning, and efficiency, 2025.

Latent Visual Cache for Video Reasoning Internvl3.5: Advancing open-source multimodal models in versatility, reasoning, and efficiency, 2025

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-12T09:09:15.248815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:09:15.248815Z digest=sha256:98c8b15a5489e0fca0bc25be9ec47c008f07124705e1b91c8798697c2155dcef

Observation 0cce0f75-ac71-455d-a92b-7ee08cc8cd80 · outbound

This paper cites AutoCAP: Towards automatic cross-lingual alignment planning for zero-shot chain-of-thought.

Latent Visual Cache for Video Reasoning AutoCAP: Towards automatic cross-lingual alignment planning for zero-shot chain-of-thought

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-12T09:09:15.248815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:09:15.248815Z digest=sha256:46dd1264f1241e717370dea69fe3f2f042860143035bc504ecd527c3e98ce7d8

Observation 0aa1d5c1-0320-439f-81aa-10102372b1f9 · outbound

This paper cites Qwen2-vl: Enhancing vision- language model’s perception of the world at any resolution, 2024.

Latent Visual Cache for Video Reasoning Qwen2-vl: Enhancing vision- language model’s perception of the world at any resolution, 2024

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-12T09:09:15.248815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:09:15.248815Z digest=sha256:4126f2bcde5010eaed91b6bf8a7fb39939736fa1f7b7bbbe43b2adf68d401a90

Observation 571f510e-56ee-49d4-be9a-24ffbcd8901a · outbound

This paper cites Moviechat: From dense token to 11 sparse memory for long video understanding.

Latent Visual Cache for Video Reasoning Moviechat: From dense token to 11 sparse memory for long video understanding

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-12T09:09:15.248815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:09:15.248815Z digest=sha256:4da60a32dd7fa25c168002bd8051a6ee76426d2f209457ab6f13e5febffbceeb

Observation 2f52c87b-d21b-4ee2-9e19-93137f97795c · outbound

This paper cites Wrong-of-thought: An integrated reasoning framework with multi-perspective verification and wrong information.

Latent Visual Cache for Video Reasoning Wrong-of-thought: An integrated reasoning framework with multi-perspective verification and wrong information

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-12T09:09:15.248815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:09:15.248815Z digest=sha256:80e33d4ebb216d7b9fc241bb6d68024cd2de4f006bb6a033d8898ebaaf2d8f31

Observation 25bfd2fe-7451-4b06-96f4-2d50707e5b6a · outbound

This paper cites Qwen3-vl technical report, 2025.

Latent Visual Cache for Video Reasoning Qwen3-vl technical report, 2025

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-12T09:09:15.248815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:09:15.248815Z digest=sha256:963278a4bb72625fb50506dc1a4288489e6d9bfab766c60b6105c65678d1e94f

Observation 3958ca1a-4f2a-4066-849f-719caf024697 · outbound

This paper cites Gemma 3 technical report, 2025.

Latent Visual Cache for Video Reasoning Gemma 3 technical report, 2025

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-12T09:09:15.248815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:09:15.248815Z digest=sha256:95999127e46e46bec4537fd431ee9f0db4c972815b6c143f043a762301511a75

Observation 337f02a6-4e19-480c-9450-0f73058640a7 · outbound

This paper cites CCHall: A novel benchmark for joint cross-lingual and cross-modal hallucinations detection in large language models.

Latent Visual Cache for Video Reasoning CCHall: A novel benchmark for joint cross-lingual and cross-modal hallucinations detection in large language models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-12T09:09:15.248815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:09:15.248815Z digest=sha256:9b9077dae3d940d0b1df91d255f689efcd45bdefbef1a68f22a22b37f51d2c66

Observation a4b504f6-641d-4a0a-b1b5-d38fc1bf8f93 · outbound

This paper cites Mitigating visual forgetting via take-along visual conditioning for multi-modal long cot reasoning.

Latent Visual Cache for Video Reasoning Mitigating visual forgetting via take-along visual conditioning for multi-modal long cot reasoning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-12T09:09:15.248815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:09:15.248815Z digest=sha256:7817da47d754bd29a0292767c3e5c58a8e5bc758e1bb35c21aae57f6ccc11689

Observation 7ef79481-a29b-4e86-823e-f9a9a23e6f4f · outbound

This paper cites Seeing through the chain: Mitigate hallucination in multimodal reasoning models via cot compression and contrastive preference optimization.arXiv preprint arXiv:2602.03380, 2026.

Latent Visual Cache for Video Reasoning Seeing through the chain: Mitigate hallucination in multimodal reasoning models via cot compression and contrastive preference optimization.arXiv preprint arXiv:2602.03380, 2026

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-12T09:09:15.248815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:09:15.248815Z digest=sha256:e6861d86c750ecdc573f0bafcc7bc3039e5f3c0f89c0a5c26ded521819d4f036

Observation 993d79d9-c14e-40eb-ab8c-d3bfc9ed7830 · outbound

This paper cites Context length alone hurts llm performance despite perfect retrieval.arXiv preprint arXiv:2510.05381, 2025.

Latent Visual Cache for Video Reasoning Context length alone hurts llm performance despite perfect retrieval.arXiv preprint arXiv:2510.05381, 2025

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-12T09:09:15.248815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:09:15.248815Z digest=sha256:1f50b104b0c3eb36756f5e03f1c6a11872e21ccc8d6ce06ede0b3d8c1c652d2e

Observation 3c8c7acb-7923-4123-addc-45c73cecc36d · outbound

This paper cites Visual hallucinations of multi-modal large language models.

Latent Visual Cache for Video Reasoning Visual hallucinations of multi-modal large language models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-12T09:09:15.248815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:09:15.248815Z digest=sha256:141bd805d9cee4f5e5d02783ed01791c95a4965ffe180c52d4917178eeae8a04

Observation 569710b0-7a25-44f3-86f0-3fa73b97db47 · outbound

This paper cites Vigc: Visual instruction generation and correction.

Latent Visual Cache for Video Reasoning Vigc: Visual instruction generation and correction

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-12T09:09:15.248815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:09:15.248815Z digest=sha256:76449c100d49819c0f588291737159c7f36147ffc60ca909f67310f631d3f1a8

Observation 2ac93fe6-1637-4287-a98c-47ed8dc3837b · outbound

This paper cites Hourvideo: 1-hour video- language understanding.Advances in Neural Information Processing Systems, 37:53168–53197, 2024.

Latent Visual Cache for Video Reasoning Hourvideo: 1-hour video- language understanding.Advances in Neural Information Processing Systems, 37:53168–53197, 2024

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-12T09:09:15.248815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:09:15.248815Z digest=sha256:9363237547053769a9449f38660cf40ef86bfe1bf7e53f56d48b5190c20b60e3

Observation 379a674e-ed8d-4f53-b14f-c3c70947bf0a · outbound

This paper cites Qwen3.5: Towards native multimodal agents, February 2026.

Latent Visual Cache for Video Reasoning Qwen3.5: Towards native multimodal agents, February 2026

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-12T09:09:15.248815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:09:15.248815Z digest=sha256:611f6f144172ccf33fcdede5d87367ae92724ecb0a48134d0c00c8b8ae9fde71

Observation 94fcc864-d632-488c-9bf1-ebc6a73093f7 · outbound

This paper cites TempCompass: Do Video LLMs Really Understand Videos?.

Latent Visual Cache for Video Reasoning TempCompass: Do Video LLMs Really Understand Videos?

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-12T09:09:15.248815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:09:15.248815Z digest=sha256:a6a412c049e25a72d29589d3b985a44ebcec45b3b64785ebeae420a2cb9056f9

Observation 3f761258-7789-40b7-a964-ca38a7dc3464 · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark.

Latent Visual Cache for Video Reasoning Mvbench: A comprehensive multi-modal video understanding benchmark

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-12T09:09:15.248815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:09:15.248815Z digest=sha256:1d1838a44b1af2e928e88c8216595f41f7095bc27c18c8e0b31396d1a0f1526e

Observation d4c2cb37-99d7-49b1-8cd2-1efa60cbe954 · outbound

This paper cites MMVU: Measuring Expert-Level Multi-Discipline Video Understanding.

Latent Visual Cache for Video Reasoning MMVU: Measuring Expert-Level Multi-Discipline Video Understanding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-12T09:09:15.248815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:09:15.248815Z digest=sha256:2581ccc7182a1c19d47ac2933fac00d7d63e1d0d1db7e249602bdcbb791eda87

Observation 79b7c0e5-ac12-46d6-bdf4-d8a4afaae848 · outbound

This paper cites Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos.

Latent Visual Cache for Video Reasoning Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-12T09:09:15.248815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:09:15.248815Z digest=sha256:01876d35ca339b46e75c2208b9143508c75e51fc36fbfb1324c89b61077c21be

Observation 60c3f555-64cd-49c8-ac60-787c95fbc898 · outbound

This paper cites Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces.

Latent Visual Cache for Video Reasoning Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-12T09:09:15.248815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:09:15.248815Z digest=sha256:283eb5b3756eae30c2a035813f9cff707730c995e582407e0b609d462b0e26e8

Observation d7a51859-3c19-437f-8d27-b043a7d6ab6a · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Latent Visual Cache for Video Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-12T09:09:15.248815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:09:15.248815Z digest=sha256:c7f1171fb6f752de32ee511de4a77907d0523fb380f42daf6934488dc5b16b6a

Observation 99b40fe0-9b62-4c20-81ed-7adfafccec25 · outbound

This paper cites Open-o3-video: Grounded video reasoning with explicit spatio-temporal evidence.arXiv preprint arXiv:2510.20579, 2025.

Latent Visual Cache for Video Reasoning Open-o3-video: Grounded video reasoning with explicit spatio-temporal evidence.arXiv preprint arXiv:2510.20579, 2025

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-12T09:09:15.248815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:09:15.248815Z digest=sha256:7f392e7e22d6fa95b02f64257dcf1a39cc96e39c1d8640445f6f4ce76512f9bb

Observation 159025a9-c049-4390-8a1c-323bf468d698 · outbound

This paper cites Video-R1: Reinforcing Video Reasoning in MLLMs.

Latent Visual Cache for Video Reasoning Video-R1: Reinforcing Video Reasoning in MLLMs

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-12T09:09:15.248815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:09:15.248815Z digest=sha256:6f86e04bd2eef557e328d36a725867191c6fbbe483bed777758e6718cf4cbe72

Observation 35d8a661-3f92-4a48-94f2-fc65384c5157 · outbound

This paper cites LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models.

Latent Visual Cache for Video Reasoning LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-12T09:09:15.248815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:09:15.248815Z digest=sha256:6153d9b57b7bfb78b2b798f468c9fdfb0ad5d83215d1d3230fb9a216977cead3

Observation f1f6a88e-21b4-4933-a8f8-5e9fd83e794c · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

Latent Visual Cache for Video Reasoning VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-12T09:09:15.248815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:09:15.248815Z digest=sha256:4e6da7e88b392a3ccce7e3f6d99b1cdf0440c1d220823694c0a89d95f3fc78d5

Observation 40a3a8b0-f12c-481e-8a7c-ad113e85c7d1 · outbound

This paper cites Long Context Transfer from Language to Vision.

Latent Visual Cache for Video Reasoning Long Context Transfer from Language to Vision

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-12T09:09:15.248815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:09:15.248815Z digest=sha256:8e92bc13a318f436086792e794aa902a3df16020f140e2ecfe441d643096aa3a

Observation 71df1acf-dc5e-426b-a1f6-e2a36be71d5f · outbound

This paper cites Vila: On pre-training for visual language models.

Latent Visual Cache for Video Reasoning Vila: On pre-training for visual language models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-12T09:09:15.248815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:09:15.248815Z digest=sha256:e09d1655494fabf42ba04befd3ca21e5a2436c419b4cbd95c88866d2667c5179

Observation ecd8c8d3-10c9-4324-8578-0a89ae81813b · outbound

This paper cites Unhackable Temporal Rewarding for Scalable Video MLLMs.

Latent Visual Cache for Video Reasoning Unhackable Temporal Rewarding for Scalable Video MLLMs

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-12T09:09:15.248815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:09:15.248815Z digest=sha256:67f051e715ed5381d9fe43ad106864cc9133674888becdebd7ce1a86d903f5d2

Observation e4336ba2-ee6e-4266-80e1-6c5feb0c7415 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Latent Visual Cache for Video Reasoning LLaVA-OneVision: Easy Visual Task Transfer

Reference 46

Resolution
unresolved
no resolver link, observed 2026-07-12T09:09:15.248815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:09:15.248815Z digest=sha256:2e8decbe77fb2d169e0e8602ffcad32168673090acc9921302f89778b3563973

Observation 01d453a6-8134-402f-ba94-496cc1c508df · outbound

This paper cites Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input.

Latent Visual Cache for Video Reasoning Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-12T09:09:15.248815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:09:15.248815Z digest=sha256:9a43bc18773e4a0e4f6e49728c90f0da15e5ba318a096860b4ae99e1fba51d59

Observation 1307fab0-b333-4ece-a3b5-91afc10ea3b1 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Latent Visual Cache for Video Reasoning Chain-of-thought prompting elicits reasoning in large language models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-12T09:09:15.248815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:09:15.248815Z digest=sha256:d49aae050c3c76c0709627cda3e5fc698c4297094cdb99504fafe6ad4100e2bf

Observation c582509d-ebf2-4e43-aa05-0c74b6d342ca · outbound

This paper cites Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022.

Latent Visual Cache for Video Reasoning Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-12T09:09:15.248815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:09:15.248815Z digest=sha256:a012243db0e2ac3d6d04db9beea6fedefec75380fbe5d9ba6db544725e736724

Observation b60f57ee-590d-4c42-9387-e709ebb22a73 · outbound

This paper cites Video-llava: Learning united visual representation by alignment before projection.

Latent Visual Cache for Video Reasoning Video-llava: Learning united visual representation by alignment before projection

Reference 50

Resolution
unresolved
no resolver link, observed 2026-07-12T09:09:15.248815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:09:15.248815Z digest=sha256:81e7dd1afa9c6ff81fc291472600b709a56dccb6f255fec11c580464f95161d6

Observation 4d96fd88-739f-4aa3-acbd-c2684148a4d4 · outbound

This paper cites an unresolved cited work.

Latent Visual Cache for Video Reasoning Unresolved cited work

Reference 51

Resolution
unresolved
no resolver link, observed 2026-07-12T09:09:15.248815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:09:15.248815Z digest=sha256:866c3d5bac26917e9a1925c3ce1876591675c6076edb2ec1b4d2217a067a01b9

Observation 6b01da90-e26c-4395-a3ad-47823087f670 · outbound

This paper cites Thinking with videos: Multimodal tool-augmented reinforcement learning for long video reasoning.

Latent Visual Cache for Video Reasoning Thinking with videos: Multimodal tool-augmented reinforcement learning for long video reasoning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-07-12T09:09:15.248815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:09:15.248815Z digest=sha256:d4dd774ce04e7ef807e8363fb36d7a04a1daf937badee9aed519fa3c1bc445f0

Observation 6e3ebc10-c568-4798-884f-7eb7dc5775ee · outbound

This paper cites Video-of-thought: step-by-step video reasoning from perception to cognition.

Latent Visual Cache for Video Reasoning Video-of-thought: step-by-step video reasoning from perception to cognition

Reference 53

Resolution
unresolved
no resolver link, observed 2026-07-12T09:09:15.248815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:09:15.248815Z digest=sha256:4672e51fa8dcd06f913343ddce1132f18fed5442424937685a1acaee83a03ee8

Observation 27d46470-37dc-4282-a15f-b80223777081 · outbound

This paper cites Vitcot: Video-text interleaved chain-of-thought for boosting video understanding in large language models.

Latent Visual Cache for Video Reasoning Vitcot: Video-text interleaved chain-of-thought for boosting video understanding in large language models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-07-12T09:09:15.248815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:09:15.248815Z digest=sha256:b4e08a8d4de1d5a99371247c7090d96d66d340a5feb548e9bb22b3af032dac43

Observation 69ea96d0-e42a-40c0-a0ef-0980032c5ac1 · outbound

This paper cites Training Large Language Models to Reason in a Continuous Latent Space.

Latent Visual Cache for Video Reasoning Training Large Language Models to Reason in a Continuous Latent Space

Reference 55

Resolution
unresolved
no resolver link, observed 2026-07-12T09:09:15.248815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:09:15.248815Z digest=sha256:eea8b0014995f98a1db8b37d84bae10b8de93999509888cc89f723963cdee017

Observation 307331b4-7612-45cf-8db7-08264b22e4b5 · outbound

This paper cites SoftCoT: Soft chain-of-thought for efficient reasoning with LLMs.

Latent Visual Cache for Video Reasoning SoftCoT: Soft chain-of-thought for efficient reasoning with LLMs

Reference 56

Resolution
unresolved
no resolver link, observed 2026-07-12T09:09:15.248815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:09:15.248815Z digest=sha256:8ab7294146ff72d0376f06495922363142d3532528cc433851ac5e637228bb23

Observation 2000a443-5da2-48a0-8880-1f649cca7e5d · outbound

This paper cites Hybrid latent reasoning via reinforcement learning.Advances in Neural Information Processing Systems, 38:5501–5530, 2026.

Latent Visual Cache for Video Reasoning Hybrid latent reasoning via reinforcement learning.Advances in Neural Information Processing Systems, 38:5501–5530, 2026

Reference 57

Resolution
unresolved
no resolver link, observed 2026-07-12T09:09:15.248815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:09:15.248815Z digest=sha256:b969cd840faec36b0567bfa9cc340cd71bbb686013e7f956dc070a9b68200e85

Observation 4ccbdba9-ab49-4122-be42-b98628f36953 · outbound

This paper cites Reasoning beyond language: A comprehensive survey on latent chain-of-thought reasoning.arXiv preprint arXiv:2505.16782, 2025.

Latent Visual Cache for Video Reasoning Reasoning beyond language: A comprehensive survey on latent chain-of-thought reasoning.arXiv preprint arXiv:2505.16782, 2025

Reference 58

Resolution
unresolved
no resolver link, observed 2026-07-12T09:09:15.248815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:09:15.248815Z digest=sha256:aeb1c20d841106d7328c2c73f6ff0f5e723f7a53644af2313a24be3cd402f61c

Observation 152b8ca8-3380-46b6-b07c-c9a008e6cc01 · outbound

This paper cites Monet: Reasoning in latent visual space beyond image and language.

Latent Visual Cache for Video Reasoning Monet: Reasoning in latent visual space beyond image and language

Reference 59

Resolution
unresolved
no resolver link, observed 2026-07-12T09:09:15.248815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:09:15.248815Z digest=sha256:d5df18af8ade67e9c9e654355320288b0ebb654fdbe94c4091ad94b42fffb4eb

Observation 74f0c947-6424-4ac1-a4e3-74e3928e4042 · outbound

This paper cites Scaling up test-time compute with latent reasoning: A recurrent depth approach.Advances in Neural Information Processing Systems, 38:41340–41391, 2026.

Latent Visual Cache for Video Reasoning Scaling up test-time compute with latent reasoning: A recurrent depth approach.Advances in Neural Information Processing Systems, 38:41340–41391, 2026

Reference 60

Resolution
malformed identifier
no resolver link, observed 2026-07-12T09:09:15.248815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:09:15.248815Z digest=sha256:3fcee636ddc415ba293b62ccc438970eabef266f3adfe352d267a0f573181866

Pith citing papers

Observation 30cf12d7-beed-4d12-9caa-a1ba8ba98eb2 · inbound

Thinking in Video: Can Video Generators Really Reason About the Real World? cites this paper.

Thinking in Video: Can Video Generators Really Reason About the Real World? Latent Visual Cache for Video Reasoning

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:13.597841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:13.597841Z digest=sha256:76c593d61f4d4a40346c079c1bf42d87890e31235b8d7939ae77a78e892b1071