Pith. sign in

Paper Citation Record · LEDGER

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames

As of 9 August 2026, this Paper Citation Record lists 74 of 74 outbound references and 9 inbound Pith citation observations for arXiv:2507.02001.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.02001 v1

Coverage vector

measured 74 of 74 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:06:26.838928Z

measured 83 of 83 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T05:58:12.475954Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T02:26:27.039487Z

Reference resolution

74 of 74 outbound references displayed

  • verified exact0
  • verified fuzzy32
  • unresolved42
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bcde3505-031a-47dc-8c7e-28b63fd461ce · outbound

This paper cites Gpt-4o mini.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Gpt-4o mini

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:06:28.175558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:06:26.538741Z digest=sha256:45a2f00dace19a116ef53b78bf65e49d25f2c71733a979270eda5372d155e117

Observation 387d2518-9558-4198-8746-035ed713012a · outbound

This paper cites GPT-4 Technical Report.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.543483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.543483Z digest=sha256:98a378d9a5240fde88c193fd4a31c38cd7601d9caba7650178c6d3b4d9ee0506

Observation df9f3100-4bd4-49f7-bc8a-19f485fb67d3 · outbound

This paper cites Gpt-4v(ision) system card.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Gpt-4v(ision) system card

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:06:28.162319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:06:26.548152Z digest=sha256:3aeb6391f11276cf27a550eb70d530461a26230c53a8bd2d68d6a27108cecebe

Observation a3fff511-2d54-4ba5-924c-34e764c4af5d · outbound

This paper cites The claude 3 model family: Opus, sonnet, haiku.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames The claude 3 model family: Opus, sonnet, haiku

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.552134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.552134Z digest=sha256:c652cb1cc570954f6409b82e5bcc22bd0578509fe7e301fe4104d4b50a1a3f27

Observation 1d6782b4-c2e0-431a-94e4-f55867eea3e2 · outbound

This paper cites Goldfish: Vision-language understanding of arbitrarily long videos.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Goldfish: Vision-language understanding of arbitrarily long videos

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:06:28.138085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:06:26.556279Z digest=sha256:b73a9b6e4deca474f847ea9139f5a001fac70158412c6ed6cfe898eb9ef3b309

Observation 1bb50a06-d606-498e-8032-3fba727674ad · outbound

This paper cites Qwen2.5-VL Technical Report.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Qwen2.5-VL Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.560339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.560339Z digest=sha256:34182bd5508f4b242943b4d56d444b9ed1b45e22d8ea9ca96348d52eaf82d6f9

Observation 2c33cb58-731a-4887-a63f-3562ed3586e6 · outbound

This paper cites Memory consolidation enables long-context video understanding.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Memory consolidation enables long-context video understanding

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:06:28.124247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:06:26.564549Z digest=sha256:e108fe4a156d25c9af8af89fbca992a3c83ce2f48118758ac3bcd416f0839b07

Observation 3bfd2c42-7b29-4306-b3ca-7b797d31df8b · outbound

This paper cites Token merging: Your vit but faster.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Token merging: Your vit but faster

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.568878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.568878Z digest=sha256:a60ac6fc9032445897762089231f67654047d33f6f818ffe69a3aae8df08e021

Observation 72cec564-9794-4885-b89a-fc12884bd531 · outbound

This paper cites Large Language Monkeys: Scaling Inference Compute with Repeated Sampling.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Large Language Monkeys: Scaling Inference Compute with Repeated Sampling

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.572809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.572809Z digest=sha256:dcdc9d5622ff3ffd1a5fd6a6864abc4041fa7d9e514cbf1a0a2c2438ad73db43

Observation 9eade738-2545-4b18-8bc3-c3da1305d5a0 · outbound

This paper cites Revisiting the" video" in video-language understanding.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Revisiting the" video" in video-language understanding

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:06:28.097441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:06:26.577040Z digest=sha256:e0f5306c43fda95e19f1440c8220585d02ec9879887bc6e5d8b3db524a0872ff

Observation 020fde4c-2d0c-472a-b5d5-1404fd8f37f7 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.581070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.581070Z digest=sha256:6d9ac0eefca11ec40086030387ccee7b81f4295f878a6a97e8c2556444aadb57

Observation ccc1dec2-8b47-4260-84de-ddc1f6e35162 · outbound

This paper cites Scannet: Richly-annotated 3d reconstructions of indoor scenes.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Scannet: Richly-annotated 3d reconstructions of indoor scenes

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:06:28.081757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:06:26.586086Z digest=sha256:9bed6a9768ce6d3fc6eeda043979c4b3db5534377f8f4e200d05946e91698703

Observation d9f97327-4c3b-4d82-befa-5acd394c998b · outbound

This paper cites Instructblip: Towards general-purpose vision-language models with instruction tuning.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Instructblip: Towards general-purpose vision-language models with instruction tuning

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:06:28.067867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:06:26.589942Z digest=sha256:e7dd7b9dcfd17d981830e6dbd3c402e2d62ae3802222a60127942242cbf211fc

Observation 5b0c2ca9-0df1-409b-b09e-da75570ff7f3 · outbound

This paper cites Structured information extraction from complex scientific text with fine-tuned large language models.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Structured information extraction from complex scientific text with fine-tuned large language models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.593731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.593731Z digest=sha256:b57163ee17cef5408b2b15be4de495480a18be6273faf9484784a441c82cdf7a

Observation 31d1563e-9bba-492c-af3d-4bf8f36f16ff · outbound

This paper cites Videoagent: A memory- augmented multimodal agent for video understanding.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Videoagent: A memory- augmented multimodal agent for video understanding

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:06:28.053783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:06:26.598035Z digest=sha256:d754ea09f52b424f72ab62b4f40cb15d87423ecfa89a3abd73658f2fe8b39554

Observation 292ead10-8221-4de8-baf0-223240bfdcd7 · outbound

This paper cites Video-of-thought: Step-by-step video reasoning from perception to cognition.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Video-of-thought: Step-by-step video reasoning from perception to cognition

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:06:28.039620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:06:26.601910Z digest=sha256:f780436556c8143105bbb789fb1485bd773dac1592ad683b52d62a7ffae0b90d

Observation 1a18afce-cbae-4a1e-9105-f474d547a1b7 · outbound

This paper cites Vertex api.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Vertex api

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:06:28.024304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:06:26.605942Z digest=sha256:a39e84191009ee23c7e9e7baf86955a805284797540c2c5280bbe9acb089fa78

Observation 890ce2db-a15c-4099-9972-913c0d04f47f · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Ego4d: Around the world in 3,000 hours of egocentric video

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.610165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.610165Z digest=sha256:50b4b06211ee428e7f7f402ef362097f82f5668df03802e3ac58a53e97082e0f

Observation b86b1bc8-177e-4ca4-989c-95d8b663f6bc · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.614422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.614422Z digest=sha256:e9a6e7b08b0ea1b25809b79391b280eb5b2a59eb7118ecd578137575b58f594d

Observation 55638bfe-9742-4f24-9d1a-33ba2c263953 · outbound

This paper cites VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.618589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.618589Z digest=sha256:67d8ea07104ee041a756df4469ca07714605e7edef70317cbb82a6efad9e1673

Observation 5c80fa2e-4634-4f6b-9b35-e3e13991eac9 · outbound

This paper cites Cogagent: A visual language model for gui agents.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Cogagent: A visual language model for gui agents

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.622743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.622743Z digest=sha256:76c5d86b2550d7a1ce887b8569909d044270f5839318de153c96d62053773d92

Observation 68a1f1d5-3093-4483-970f-36a8124c05eb · outbound

This paper cites RULER: What's the Real Context Size of Your Long-Context Language Models?.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames RULER: What's the Real Context Size of Your Long-Context Language Models?

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.626813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.626813Z digest=sha256:b830a4b6670e12d50ed3cb29a413d79c8feff735a0977252a27e5a11a6c66e38

Observation 8a4bb54f-4386-45f1-a7e6-378cef52d6ef · outbound

This paper cites Unsupervised dense information retrieval with contrastive learning.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Unsupervised dense information retrieval with contrastive learning

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:06:27.990130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:06:26.630935Z digest=sha256:5b0cfdbb35f32b0f0cb0c144f11d0bc8bcb803ac2f8c2f92572849b03f23da05

Observation 46c95784-9ad3-47d5-987f-fb2bc8b469a7 · outbound

This paper cites Perceiver IO: A general architecture for structured inputs & outputs.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Perceiver IO: A general architecture for structured inputs & outputs

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:06:27.975679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:06:26.634902Z digest=sha256:736acd571c03a109bc6dfa67a08b96a88f93be1b8e5b8e1c9499f99241634b79

Observation 2c94b909-942e-498a-bd5c-b7481be88167 · outbound

This paper cites Action genome: Actions as compositions of spatio-temporal scene graphs.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Action genome: Actions as compositions of spatio-temporal scene graphs

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:06:27.959177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:06:26.638949Z digest=sha256:e222cde753495fd37e938b03830232d4a984d038dae32b7b59c9244e3b5aa664

Observation e64ffbc9-b9c6-458c-9a1f-163cf154ef0e · outbound

This paper cites Billion-scale similarity search with gpus.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Billion-scale similarity search with gpus

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:06:27.945103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:06:26.643001Z digest=sha256:83dc7e07d7d827e5f3cc91b9cc7a02f079a3f5804ba92f36b5129b2abd7e38a6

Observation 378dba24-81bd-4297-9f4f-974a00cd3418 · outbound

This paper cites Language Repository for Long Video Understanding.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Language Repository for Long Video Understanding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.647384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.647384Z digest=sha256:90fa17a2b127ea9ca2405c93d56edb1ebc77e10e2d38bb0c0a097d9554cfacc3

Observation 08709ce0-90ed-46df-8326-bb7a61063c1b · outbound

This paper cites Large language models are zero-shot reasoners.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Large language models are zero-shot reasoners

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.651869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.651869Z digest=sha256:d81f4490f59b72722a11d55bb47d56bc3cbdbe5783ca376f23523ad6a00c4bc1

Observation 72dcd6ef-668b-4361-9d1a-ef3c4d38e50f · outbound

This paper cites Text-conditioned resampler for long form video understanding.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Text-conditioned resampler for long form video understanding

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:06:27.920453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:06:26.655620Z digest=sha256:b37e2a4e2bca099ff0e6e8f7d6c150c6e15b5196ab082d9f350f92809593dc19

Observation a86a8290-85a0-4b46-9027-cb17eeb6047d · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.659753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.659753Z digest=sha256:aeedbec09259f236fd8535cef616bce872c24e567fb8a33e7081c3bca55e7539

Observation b59e93c8-4d71-4d5e-bcb7-d983720ce84e · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:06:27.905959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:06:26.663923Z digest=sha256:e2e3b17e1cd8c7db0f789c9d4594b90c7eebff445e581a383699da324489b2cb

Observation 3aae1bfb-1b72-4d58-9f17-afe081a67c85 · outbound

This paper cites Invariant grounding for video question answering.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Invariant grounding for video question answering

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.667963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.667963Z digest=sha256:90b48ffcc2917e917b3ef010fc433d215f85f89fa920a5900189612584335046

Observation 72e9616c-1eb2-4419-b41e-984702c0ae8b · outbound

This paper cites Visual instruction tuning.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Visual instruction tuning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.672045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.672045Z digest=sha256:353d435bbe10be33e386e6288f55ac483568ec28f768f054439456fa8ad62dca

Observation 6f2ad1ef-5f6e-4272-9f7e-b501963dcd99 · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge, 2024.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Llava-next: Improved reasoning, ocr, and world knowledge, 2024

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.677167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.677167Z digest=sha256:d7e9ac2db6eecb7aee601038adabda7c5e0a291d4c07e06e8cfdbfb370e8c4f0

Observation c91d1091-05ce-4085-88e1-8358f049a61f · outbound

This paper cites Ring attention with blockwise transformers for near-infinite context.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Ring attention with blockwise transformers for near-infinite context

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:06:27.862733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:06:26.681952Z digest=sha256:fa87e3889164443cff7304944d2d8e9de392d1152ef3a87205ec16f2f9f98e7f

Observation 9a47afc2-3c1f-4dc4-9976-eb7ea133f318 · outbound

This paper cites Lost in the middle: How language models use long contexts.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Lost in the middle: How language models use long contexts

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.686250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.686250Z digest=sha256:3e5cf0954cc88619d5f9702425dd9a7220b68f3aba707cdfcc99d6a7f7e4efcb

Observation 0c6a7bba-51f7-4a1b-b18b-271944449f42 · outbound

This paper cites BOLT: Boost Large Vision-Language Model Without Training for Long-form Video Understanding.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames BOLT: Boost Large Vision-Language Model Without Training for Long-form Video Understanding

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.690026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.690026Z digest=sha256:32fe0474ab71e21c137320b96e3be49c786af61b3d6bd460307c4fb0cfdde0ea

Observation 0637e5cb-e5e8-4aef-8210-94ea4cecd6de · outbound

This paper cites Video-rag: Visually-aligned retrieval-augmented long video comprehension.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Video-rag: Visually-aligned retrieval-augmented long video comprehension

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.694115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.694115Z digest=sha256:2da6af52f5f9fcac12adcada1ab07e82212bb4ffe1493d67d2f39d5cd63b984a

Observation 936b0c68-3747-425b-9665-a59e445ecbc3 · outbound

This paper cites Openeqa: Embodied question answering in the era of foundation models.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Openeqa: Embodied question answering in the era of foundation models

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:06:27.840702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:06:26.698208Z digest=sha256:90e51ca4057acc9ff56e875999d37c572701acb769c37b03e8e2209ee59ae132

Observation f7641231-55bd-4424-92f1-d787c02454d1 · outbound

This paper cites Egoschema: A diagnostic benchmark for very long-form video language understanding.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Egoschema: A diagnostic benchmark for very long-form video language understanding

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.702136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.702136Z digest=sha256:c3c50942301a27a7c486d20b350aa8a7268c7efc41cc31dfd0558f5cd46e388e

Observation 65f12c65-6fcf-4e40-971e-26e749985c4d · outbound

This paper cites Morevqa: Exploring modular reasoning models for video question answering.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Morevqa: Exploring modular reasoning models for video question answering

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:06:27.816727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:06:26.705931Z digest=sha256:a2cf0c70aae514ea2459d36bfca27dc12a62cfef8cf649714210e73a014793fc

Observation ed9989cc-d9e3-4f85-a6fd-7b067971dd23 · outbound

This paper cites PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.709761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.709761Z digest=sha256:c4354a3a5c8a837d0c4d6d0eb4c359eb47401bd1cc7e8ef902bb6c123765083b

Observation f5c42bad-8fee-43c8-ba53-5c63d6515837 · outbound

This paper cites Training language models to follow instructions with human feedback.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Training language models to follow instructions with human feedback

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:06:27.802700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:06:26.713878Z digest=sha256:dccf4f74ad8d9f1c5e877f8bbf8dca9d11619692a7509fa1459ae3e190f58d7c

Observation 8ba5c474-171e-44c9-99e1-d673a68f9c1f · outbound

This paper cites Too many frames, not all useful: Efficient strategies for long-form video qa.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Too many frames, not all useful: Efficient strategies for long-form video qa

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.717830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.717830Z digest=sha256:b84e30d169760914f202f22b0a6cf08781097db01093c84a8c25e506616dcd09

Observation 179bd670-daa2-447a-976c-d5715b713e58 · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Robust speech recognition via large-scale weak supervision

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:06:27.788784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:06:26.721629Z digest=sha256:351f42031ba9f3878a623a8e2bb82c4126f80058dd8e93b1dfa1dbecc2142f9c

Observation b3cd413b-2e5d-4cde-99a6-9cd73acc26e5 · outbound

This paper cites Habitat-Matterport 3D Dataset (HM3D): 1000 Large-scale 3D Environments for Embodied AI.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Habitat-Matterport 3D Dataset (HM3D): 1000 Large-scale 3D Environments for Embodied AI

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.725687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.725687Z digest=sha256:ac78921728bb2c7292d1f98694c38f585d5b291517a2520d7be7c7f11b1919bf

Observation babb1833-a079-4dea-97c1-6dcba476389a · outbound

This paper cites Tokenlearner: What can 8 learned tokens do for images and videos? In NeurIPS, 2021.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Tokenlearner: What can 8 learned tokens do for images and videos? In NeurIPS, 2021

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:06:27.774633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:06:26.729922Z digest=sha256:50645a0e321a10facaf26cd3b4f978036eb78c3d867076371359ff7f1e1d0308

Observation 4a1aaef5-5c7c-4b4b-8d7b-7f3318b8541a · outbound

This paper cites LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.733908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.733908Z digest=sha256:86da7f1a4d6000339252fce25eb417d014b77bb4c355b224b0867c62c6e273cb

Observation cec287de-e943-4ef5-8f23-4a17d6b0da51 · outbound

This paper cites Aligning and prompting everything all at once for universal visual perception.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Aligning and prompting everything all at once for universal visual perception

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:06:27.758875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:06:26.737936Z digest=sha256:a019003fc4675da09fd8a45e88d40d35957842a8c49b6fc193220e8ff2012d69

Observation 0d671523-3eee-4c70-b7c1-dff89fd15863 · outbound

This paper cites Reflexion: Language agents with verbal reinforcement learning.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Reflexion: Language agents with verbal reinforcement learning

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:06:27.745028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:06:26.741639Z digest=sha256:8abff32c2c06a6e8cbfd03ed92f6801afc55e7c3a310f60b2259c72bcb4cc6d5

Observation bd5d0968-1d97-4b36-b179-595b46cab0f2 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.745808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.745808Z digest=sha256:32adcd602b1b022420f035facfabc62ee5c33305ef39e790fbec58646324cddd

Observation 4d1c2b4e-a74b-4f70-8c1c-f0b2b2932a61 · outbound

This paper cites EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.749785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.749785Z digest=sha256:1cde67b12352ec08fdcd3b1a777d5cc9a35a1cc7862f43aa1ca50ba07802e6e4

Observation 010f03f9-7700-4291-ba3d-1c67a8939b41 · outbound

This paper cites Vipergpt: Visual inference via python execution for reasoning.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Vipergpt: Visual inference via python execution for reasoning

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:06:27.730371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:06:26.753979Z digest=sha256:1f0722c6059c83f4af79ad0a71174ec58f00eac9ed01b439494594421b070701

Observation 5d0abb2d-0042-4ea3-8b9c-3996a522b0fa · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.757838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.757838Z digest=sha256:2a70786585b7e6d04913627df2de016b89b69c84b6bdcaf336c891233a4a9488

Observation 19e9e558-b6aa-4f0c-a164-001f97a492af · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.762347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.762347Z digest=sha256:35beb9706fdeb6a486bd90c7389d916bc6aabc9ddf16dee8b47de10e08e72d30

Observation 29234379-7dac-44aa-be2c-74cb9bb9350f · outbound

This paper cites Vamos: Versatile action models for video understanding.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Vamos: Versatile action models for video understanding

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:06:27.716273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:06:26.766293Z digest=sha256:c973e522862e6168999d72b4629e9bae677be2a3c99738517f6704947ac9ac08

Observation ec6bc11f-8871-478e-a1b9-f976c7cb2318 · outbound

This paper cites LVBench: An Extreme Long Video Understanding Benchmark.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames LVBench: An Extreme Long Video Understanding Benchmark

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.770136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.770136Z digest=sha256:9e0909f86fdea9ccfc61842fd0084648bd89ff916468ac55f67f64fdf34a5bc8

Observation 3e25099e-ea0a-45d7-84ab-2ca445722a01 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.774431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.774431Z digest=sha256:f1915232e58c18cfa155da2c51be8ac20081444477e8c184fbdc23a37a82ca52

Observation 0878dda8-ff1b-4e0b-abd1-b10f1735c29a · outbound

This paper cites Vila: Efficient video-language alignment for video question answering.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Vila: Efficient video-language alignment for video question answering

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:06:27.702596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:06:26.778822Z digest=sha256:14952009d035eeb2122ca40a8f9e6b16cd8ad6a196ddd4f3f0ed2613cafdce3e

Observation 9a207b0b-a775-4415-aac5-094f9018a8c6 · outbound

This paper cites VideoAgent: Long-form Video Understanding with Large Language Model as Agent.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames VideoAgent: Long-form Video Understanding with Large Language Model as Agent

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.782893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.782893Z digest=sha256:dc2091ee8ea3af495932351ea22a3361c65188e050c8d5dcd9092b0fbf7d2b74

Observation f1cc9b77-687e-4840-91f8-69b982c35dc5 · outbound

This paper cites VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.787020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.787020Z digest=sha256:fd3c5a2c6bdc8c3672b22e4cf53264443c97bf2d9c7938e8d48b020d3d67a39e

Observation 558c4423-1900-440e-8699-d995e8977c03 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Chain-of-thought prompting elicits reasoning in large language models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.791453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.791453Z digest=sha256:09acef32bbe0a5eb598fc47b22a5507cb58836aad36ed4053ef0c0ac554cf191

Observation 9d3bcf66-c07c-44de-9637-691bed3f7b76 · outbound

This paper cites Visual Haystacks: A Vision-Centric Needle-In-A-Haystack Benchmark.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Visual Haystacks: A Vision-Centric Needle-In-A-Haystack Benchmark

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.795465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.795465Z digest=sha256:a64f8839ec0c76cccdb0243e2322baa83031254fafafb54b0b1a1118e22ad250

Observation b404a63e-0ac2-403b-b925-c13a9b907b38 · outbound

This paper cites Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.799622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.799622Z digest=sha256:a75a59cf16c418fca5183dba0a597a6d390b9322d72ea94e52ff88ac4058e80b

Observation b60ec109-f39c-4120-9e14-83ded3c3a62a · outbound

This paper cites Next-qa: Next phase of question-answering to explaining temporal actions.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Next-qa: Next phase of question-answering to explaining temporal actions

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:06:27.677233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:06:26.803791Z digest=sha256:7dedd4cc62e23851db9320f10912cd10b9198943684058c48a2828150197ffa7

Observation 4dd48ba5-4468-46ad-9d09-6acab640dd89 · outbound

This paper cites Retrieval-based video language model for efficient long video question answering.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Retrieval-based video language model for efficient long video question answering

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.807519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.807519Z digest=sha256:834f4af3f0722f8268d533b21442df287c6240153f30a31f3c436c538b3c305f

Observation 408380f8-1539-406f-a9fb-cd5719957c8a · outbound

This paper cites Tree of thoughts: Deliberate problem solving with large language models.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Tree of thoughts: Deliberate problem solving with large language models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.811686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.811686Z digest=sha256:37226f1f679603d8f24f12912b83a06576df30e813d4ab4a78503deccd538b5a

Observation b09bf94d-d062-4b5e-8064-fafc3665cd94 · outbound

This paper cites Self-chained image-language model for video localization and question answering.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Self-chained image-language model for video localization and question answering

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:06:27.652196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:06:26.815384Z digest=sha256:f50c0368bf48ce45bdbefc6413116042c29e0124e2be40ad3e517ccec397534d

Observation c8a105ce-1fb0-4fc7-9ab4-58b260b8c5fd · outbound

This paper cites Lv-eval: A balanced long-context benchmark with 5 length levels up to 256k.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Lv-eval: A balanced long-context benchmark with 5 length levels up to 256k

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.819219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.819219Z digest=sha256:27813a46fba790b4a5eaa8c583b7425c2a117c85aa6c467aeeef0ffe39b1c0d9

Observation 3e450e21-1196-4122-802d-cf414fabd1bf · outbound

This paper cites Sigmoid loss for language image pre-training.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Sigmoid loss for language image pre-training

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.823206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.823206Z digest=sha256:5c2bc9885ec6a38beb7218eeeb4137aff52104de627f00349db9c7bc5713b75c

Observation 88277197-a309-436c-bdcd-3c159b1114e4 · outbound

This paper cites A simple llm framework for long-range video question-answering.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames A simple llm framework for long-range video question-answering

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:06:27.627536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:06:26.827083Z digest=sha256:c3b9ea9dcd4d100e9dafecbad94ecead70522bc11d49c1b27179406e82342c99

Observation 9359e176-943c-4861-9513-79e000d20a13 · outbound

This paper cites Learning video representations from large language models.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Learning video representations from large language models

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:06:27.612666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:06:26.830891Z digest=sha256:6909191c6c93c6ffbb8418158f4c34e3ec76b9c4db07cca30b1f534d0ccdfefc

Observation a798454b-76a3-45fa-8c2d-f29a9d1940b3 · outbound

This paper cites Needle In A Video Haystack: A Scalable Synthetic Evaluator for Video MLLMs.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Needle In A Video Haystack: A Scalable Synthetic Evaluator for Video MLLMs

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.834825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.834825Z digest=sha256:2712b4640e3bf066f1163a784f3c058dd0efacc23d23fa1e207f53a4e1054d1e

Observation 64cb19ff-1738-4904-bd56-a835154c65f1 · outbound

This paper cites temporal certificate.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames temporal certificate

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:06:27.597368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:06:26.838928Z digest=sha256:d7cde2db7ae90281e7cf92cb87aa3592ecc7d206ba6ecef8b020edb03dd7f076

Pith citing papers

Observation 370ffc45-bcfe-4ee6-88f1-589cbeb5489e · inbound

Internalized Reasoning for Long-Context Visual Document Understanding cites this paper.

Internalized Reasoning for Long-Context Visual Document Understanding Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:53:28.169730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T23:53:19.148407Z digest=sha256:278d27e5e75ce8c69711ed39fa4a1586caa419ea6eca9accbb2454ed1278a8a2

Observation 1ee4d66e-d02f-4c03-80d8-7cd159425f45 · inbound

Internalized Reasoning for Long-Context Visual Document Understanding cites this paper.

Internalized Reasoning for Long-Context Visual Document Understanding Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-13T15:50:49.083652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T15:50:49.083652Z digest=sha256:d984d8273e7b6129fb08aa3b5f4fb80bb22fcff6a8bf0c62b1533f5356fef0df

Observation 18b041e5-2a84-4125-a0b2-c384cbdce99d · inbound

Event-Causal RAG: A Retrieval-Augmented Generation Framework for Long Video Reasoning in Complex Scenarios cites this paper.

Event-Causal RAG: A Retrieval-Augmented Generation Framework for Long Video Reasoning in Complex Scenarios Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:06:13.284145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T10:15:15.129358Z digest=sha256:83a39f12f3e7b2728ca6bb4cb0a56b7c3697b9df105135ed1a6543217ed16389

Observation 3a28ed55-3902-4a5a-bbe6-ce055291ed38 · inbound

Personal Visual Context Learning in Large Multimodal Models cites this paper.

Personal Visual Context Learning in Large Multimodal Models Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:06:37.196047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T03:42:15.402131Z digest=sha256:8eab0cc5fc464665c8d3c78f8160556b1b8655f94db0467c948824a5edc32415

Observation 09716617-b19c-4d9a-b9a6-7492d3692ca2 · inbound

Swift Sampling: Selecting Temporal Surprises via Taylor Series cites this paper.

Swift Sampling: Selecting Temporal Surprises via Taylor Series Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:56:07.853544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T05:55:23.479344Z digest=sha256:4caf109503209f51bef16644bbc969d52eb2b03cd8ba3aaa112d8fe253e64010

Observation 30a53c02-f069-4af1-9e5b-240630cbf94d · inbound

Zero-Shot 3D Question Answering via Hierarchical View-to-Token Transportation cites this paper.

Zero-Shot 3D Question Answering via Hierarchical View-to-Token Transportation Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames

Reference 76

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T02:26:27.041380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T10:54:02.188634Z digest=sha256:89571c76344e758928baa0bf4764be70ddca749feb4c1606a18b7f852b231a93

Observation 18343c31-284e-4b4d-bcea-e216b7bff1d3 · inbound

MemLearner: Learning to Query Context memory for Video World Models cites this paper.

MemLearner: Learning to Query Context memory for Video World Models Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T10:25:41.855643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-01T05:30:56.140465Z digest=sha256:21f41dab8e039c8afe57f562dd78cb7aef0bd8a3ce9c6f522ea829ec1fb21bd5

Observation be4b30a9-dbd6-4ef6-80e1-276ccb3bb29e · inbound

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos cites this paper.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:42.836941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:42.836941Z digest=sha256:a75fb8cfe4f7dd1ca6eab66f9728eb237fc45acc8afbb66b07abe6e3b8fd8e9f

Observation b0d01f99-df2b-4f19-80e9-377357e4ccb3 · inbound

Beyond Frame Selection: Rethinking Long-Video Understanding with MLLMs cites this paper.

Beyond Frame Selection: Rethinking Long-Video Understanding with MLLMs Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-08T05:58:12.475954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:58:12.475954Z digest=sha256:9bfa63cd47725aa9b86712b889151d981eec5ae22c6387db5e93574b90ee021e