Pith. sign in

Paper Citation Record · LEDGER

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames

As of 7 August 2026, this Paper Citation Record lists 74 of 74 outbound references and 8 inbound Pith citation observations for arXiv:2507.02001.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.02001 v1

Coverage vector

measured 74 of 74 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:06:26.838928Z

measured 82 of 82 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T21:22:42.836941Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T02:26:27.039487Z

Reference resolution

74 of 74 outbound references displayed

  • verified exact0
  • verified fuzzy32
  • unresolved42
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bcde3505-031a-47dc-8c7e-28b63fd461ce · outbound

This paper cites Gpt-4o mini.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Gpt-4o mini

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:06:28.175558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:06:26.538741Z digest=sha256:036e73f097d7dc3ac50b4aa36d2b408cf606f244225a1e6004317321fc62b4c1

Observation 387d2518-9558-4198-8746-035ed713012a · outbound

This paper cites GPT-4 Technical Report.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.543483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.543483Z digest=sha256:a0bd55a26faf5ee029267edf04ba9dfeb2bfa7742a23b4ec32bf86979fc7d567

Observation df9f3100-4bd4-49f7-bc8a-19f485fb67d3 · outbound

This paper cites Gpt-4v(ision) system card.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Gpt-4v(ision) system card

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:06:28.162319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:06:26.548152Z digest=sha256:4e5ed8578f39081fae005561c736104bc8dd6d8700dd37de48ddbcbb87ad87a3

Observation a3fff511-2d54-4ba5-924c-34e764c4af5d · outbound

This paper cites The claude 3 model family: Opus, sonnet, haiku.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames The claude 3 model family: Opus, sonnet, haiku

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.552134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.552134Z digest=sha256:97e69a7691e9993d48a511844f4f9dcb4b40a988567b06279a65f46532107e33

Observation 1d6782b4-c2e0-431a-94e4-f55867eea3e2 · outbound

This paper cites Goldfish: Vision-language understanding of arbitrarily long videos.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Goldfish: Vision-language understanding of arbitrarily long videos

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:06:28.138085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:06:26.556279Z digest=sha256:8e35fe7e35774e1bfac709c6f5ad068229dedb78962fdcd0d55be3da8e01a51f

Observation 1bb50a06-d606-498e-8032-3fba727674ad · outbound

This paper cites Qwen2.5-VL Technical Report.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Qwen2.5-VL Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.560339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.560339Z digest=sha256:913238dabd7e8d3f9a93586ebd1ce462ab0e16857e6b16871342d4a713180249

Observation 2c33cb58-731a-4887-a63f-3562ed3586e6 · outbound

This paper cites Memory consolidation enables long-context video understanding.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Memory consolidation enables long-context video understanding

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:06:28.124247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:06:26.564549Z digest=sha256:70623ddfc1e17e27d03c4987a0143bdac9e99df9030d69f30fe37011930034e6

Observation 3bfd2c42-7b29-4306-b3ca-7b797d31df8b · outbound

This paper cites Token merging: Your vit but faster.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Token merging: Your vit but faster

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.568878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.568878Z digest=sha256:f058e16dab67769ffa384b5d8dd415055660ce6c12c5bb0417bdc138ee94e9c1

Observation 72cec564-9794-4885-b89a-fc12884bd531 · outbound

This paper cites Large Language Monkeys: Scaling Inference Compute with Repeated Sampling.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Large Language Monkeys: Scaling Inference Compute with Repeated Sampling

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.572809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.572809Z digest=sha256:b98e0703e2c3c2a923b149e16bf08f3166ba48c8efedee39c07f67f19a005da3

Observation 9eade738-2545-4b18-8bc3-c3da1305d5a0 · outbound

This paper cites Revisiting the" video" in video-language understanding.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Revisiting the" video" in video-language understanding

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:06:28.097441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:06:26.577040Z digest=sha256:9dd78c754f1983a4d66f42124ed4f971f9a028f4f784c6a254973edecf19e25e

Observation 020fde4c-2d0c-472a-b5d5-1404fd8f37f7 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.581070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.581070Z digest=sha256:2300763f308ef813741a72a51b4f8314ad5668e8c365f10a9ac09ff1b4cf45eb

Observation ccc1dec2-8b47-4260-84de-ddc1f6e35162 · outbound

This paper cites Scannet: Richly-annotated 3d reconstructions of indoor scenes.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Scannet: Richly-annotated 3d reconstructions of indoor scenes

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:06:28.081757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:06:26.586086Z digest=sha256:d7b6ce02e5a6ee4ee8f4bbb923b8919c10481c6380a3edbaee7834356a50b420

Observation d9f97327-4c3b-4d82-befa-5acd394c998b · outbound

This paper cites Instructblip: Towards general-purpose vision-language models with instruction tuning.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Instructblip: Towards general-purpose vision-language models with instruction tuning

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:06:28.067867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:06:26.589942Z digest=sha256:b57beabd3ab074d60f3a70b079d638dfef14705fca35ae016162a36427b406e0

Observation 5b0c2ca9-0df1-409b-b09e-da75570ff7f3 · outbound

This paper cites Structured information extraction from complex scientific text with fine-tuned large language models.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Structured information extraction from complex scientific text with fine-tuned large language models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.593731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.593731Z digest=sha256:ea141952d98c8e67c32b8adbdff9c385fec2555bf94dfca3e2afe1116afaa631

Observation 31d1563e-9bba-492c-af3d-4bf8f36f16ff · outbound

This paper cites Videoagent: A memory- augmented multimodal agent for video understanding.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Videoagent: A memory- augmented multimodal agent for video understanding

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:06:28.053783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:06:26.598035Z digest=sha256:e4d65aeda94c287ff82d86891b58f07962b8dbcaa2ed15e78446f9cab1243081

Observation 292ead10-8221-4de8-baf0-223240bfdcd7 · outbound

This paper cites Video-of-thought: Step-by-step video reasoning from perception to cognition.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Video-of-thought: Step-by-step video reasoning from perception to cognition

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:06:28.039620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:06:26.601910Z digest=sha256:43159bc3d75db58c02f23e3597cb1783a5703d9ff8d45e8ee047c633c331b48d

Observation 1a18afce-cbae-4a1e-9105-f474d547a1b7 · outbound

This paper cites Vertex api.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Vertex api

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:06:28.024304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:06:26.605942Z digest=sha256:98babd856c7465b4df5523bf5c9d79ffce842c01b307cac9920cd042f67fffaf

Observation 890ce2db-a15c-4099-9972-913c0d04f47f · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Ego4d: Around the world in 3,000 hours of egocentric video

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.610165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.610165Z digest=sha256:bb3044a8f886e809efbc03afc440b6409c67c6c40d811f9f25af069eea34c883

Observation b86b1bc8-177e-4ca4-989c-95d8b663f6bc · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.614422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.614422Z digest=sha256:7f336a46782e35e4330c289971f941c49e85dafd2e06ce736a1f2787333b5247

Observation 55638bfe-9742-4f24-9d1a-33ba2c263953 · outbound

This paper cites VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.618589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.618589Z digest=sha256:16cad8170daf01744835d20ce7ae625839907365aa8f60c3f1d320c224f3f2ab

Observation 5c80fa2e-4634-4f6b-9b35-e3e13991eac9 · outbound

This paper cites Cogagent: A visual language model for gui agents.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Cogagent: A visual language model for gui agents

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.622743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.622743Z digest=sha256:a9db0d63c773a31e5f48e22d3fb5cec1ad1ad3ac9cbe2ba3bd3793c8bfe3b371

Observation 68a1f1d5-3093-4483-970f-36a8124c05eb · outbound

This paper cites RULER: What's the Real Context Size of Your Long-Context Language Models?.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames RULER: What's the Real Context Size of Your Long-Context Language Models?

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.626813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.626813Z digest=sha256:a37c8f80e0e9788fb410630d93222f1b9cf95bf14673b4a35cc39016416f3b61

Observation 8a4bb54f-4386-45f1-a7e6-378cef52d6ef · outbound

This paper cites Unsupervised dense information retrieval with contrastive learning.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Unsupervised dense information retrieval with contrastive learning

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:06:27.990130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:06:26.630935Z digest=sha256:0cbb026bef8d3c94e8e2afa95b6a925de6c35974d8bf7c53bbbea2e3bf72f5d2

Observation 46c95784-9ad3-47d5-987f-fb2bc8b469a7 · outbound

This paper cites Perceiver IO: A general architecture for structured inputs & outputs.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Perceiver IO: A general architecture for structured inputs & outputs

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:06:27.975679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:06:26.634902Z digest=sha256:1348db4f50713277ff37ef9c8ddfc212ad3fdb24da47b26d72e6b2601b66a72a

Observation 2c94b909-942e-498a-bd5c-b7481be88167 · outbound

This paper cites Action genome: Actions as compositions of spatio-temporal scene graphs.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Action genome: Actions as compositions of spatio-temporal scene graphs

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:06:27.959177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:06:26.638949Z digest=sha256:f140b7b5b550bd8a0b4fa928fe4043759510d3bac446a460e1b7f9b4334a6c34

Observation e64ffbc9-b9c6-458c-9a1f-163cf154ef0e · outbound

This paper cites Billion-scale similarity search with gpus.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Billion-scale similarity search with gpus

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:06:27.945103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:06:26.643001Z digest=sha256:6178ea4d33619041b6efa2b3d232f96dc8f4dd72f45499c75c6b9eb802996817

Observation 378dba24-81bd-4297-9f4f-974a00cd3418 · outbound

This paper cites Language Repository for Long Video Understanding.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Language Repository for Long Video Understanding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.647384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.647384Z digest=sha256:a77c5d4c3f2af8723b87d9156e4d57b2ea3b5b34ca338aa40b1c09d0f4a643f3

Observation 08709ce0-90ed-46df-8326-bb7a61063c1b · outbound

This paper cites Large language models are zero-shot reasoners.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Large language models are zero-shot reasoners

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.651869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.651869Z digest=sha256:7484ac0ddb90bede0f678ee2d33bc1c9068c86a7f69724269c86baad67601b30

Observation 72dcd6ef-668b-4361-9d1a-ef3c4d38e50f · outbound

This paper cites Text-conditioned resampler for long form video understanding.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Text-conditioned resampler for long form video understanding

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:06:27.920453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:06:26.655620Z digest=sha256:597d871bda8b50ba9d4b3a89bc2e711c0e4b1e4150a2150a052115d0cb7b7523

Observation a86a8290-85a0-4b46-9027-cb17eeb6047d · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.659753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.659753Z digest=sha256:3b42691176ae2f961c718c867708eca047cc332762cfcce17a0e8ff7d41cd540

Observation b59e93c8-4d71-4d5e-bcb7-d983720ce84e · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:06:27.905959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:06:26.663923Z digest=sha256:cf84344bae0c80bc77c02930da61db84214d9de3af072ce21ef27ddff09ce912

Observation 3aae1bfb-1b72-4d58-9f17-afe081a67c85 · outbound

This paper cites Invariant grounding for video question answering.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Invariant grounding for video question answering

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.667963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.667963Z digest=sha256:4ac0137ccf667dd7b812c46b27c9bfc2616b17947c1b4548d1c4555237bf3d97

Observation 72e9616c-1eb2-4419-b41e-984702c0ae8b · outbound

This paper cites Visual instruction tuning.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Visual instruction tuning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.672045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.672045Z digest=sha256:e0c9001bcff9b84770f5bf6b1173a99159b82b80f8d4cce2f6f9cd652dda6cd8

Observation 6f2ad1ef-5f6e-4272-9f7e-b501963dcd99 · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge, 2024.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Llava-next: Improved reasoning, ocr, and world knowledge, 2024

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.677167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.677167Z digest=sha256:7c50764c9e5a292970d493f9bbfbcde0484ac09e23559b61fe9aace0cf1bceb4

Observation c91d1091-05ce-4085-88e1-8358f049a61f · outbound

This paper cites Ring attention with blockwise transformers for near-infinite context.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Ring attention with blockwise transformers for near-infinite context

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:06:27.862733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:06:26.681952Z digest=sha256:363f8df1611f01c7098f0b5b1148ade3674341dc5b542c44512f8d853b851dde

Observation 9a47afc2-3c1f-4dc4-9976-eb7ea133f318 · outbound

This paper cites Lost in the middle: How language models use long contexts.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Lost in the middle: How language models use long contexts

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.686250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.686250Z digest=sha256:d72b13142be7b48f786118e4890d38f9bca775bbd1fc72f0fa93fbb515712d24

Observation 0c6a7bba-51f7-4a1b-b18b-271944449f42 · outbound

This paper cites BOLT: Boost Large Vision-Language Model Without Training for Long-form Video Understanding.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames BOLT: Boost Large Vision-Language Model Without Training for Long-form Video Understanding

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.690026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.690026Z digest=sha256:9981f120b97148321701f2ebbbb022cd152e79e88696aaa45e1ae652f1b562d7

Observation 0637e5cb-e5e8-4aef-8210-94ea4cecd6de · outbound

This paper cites Video-rag: Visually-aligned retrieval-augmented long video comprehension.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Video-rag: Visually-aligned retrieval-augmented long video comprehension

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.694115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.694115Z digest=sha256:c91bca98d0bab68183c2d9f24e028d9f29e5a4779d6958864c2859d947862114

Observation 936b0c68-3747-425b-9665-a59e445ecbc3 · outbound

This paper cites Openeqa: Embodied question answering in the era of foundation models.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Openeqa: Embodied question answering in the era of foundation models

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:06:27.840702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:06:26.698208Z digest=sha256:4fc8c6f3e9fcc778cc6b7163701dc64d59e514339770b216d97870f9aa5ff82a

Observation f7641231-55bd-4424-92f1-d787c02454d1 · outbound

This paper cites Egoschema: A diagnostic benchmark for very long-form video language understanding.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Egoschema: A diagnostic benchmark for very long-form video language understanding

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.702136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.702136Z digest=sha256:a5b09d2b9e5630625c69d0390b741d6dc0772d593428cfacec169a854caac323

Observation 65f12c65-6fcf-4e40-971e-26e749985c4d · outbound

This paper cites Morevqa: Exploring modular reasoning models for video question answering.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Morevqa: Exploring modular reasoning models for video question answering

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:06:27.816727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:06:26.705931Z digest=sha256:39c702c5739fd708872853e0227d1c1f7504814fcfd5a1501817b4991f797ac2

Observation ed9989cc-d9e3-4f85-a6fd-7b067971dd23 · outbound

This paper cites PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.709761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.709761Z digest=sha256:6aa23ed82bfbb7c5bc978dba999adb9301ac7ea65428b5af0070a0e055beb12f

Observation f5c42bad-8fee-43c8-ba53-5c63d6515837 · outbound

This paper cites Training language models to follow instructions with human feedback.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Training language models to follow instructions with human feedback

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:06:27.802700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:06:26.713878Z digest=sha256:910af02d3aa910bd7dcc4f72849071a81440798bd8a29c29f5a14a96e990371e

Observation 8ba5c474-171e-44c9-99e1-d673a68f9c1f · outbound

This paper cites Too many frames, not all useful: Efficient strategies for long-form video qa.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Too many frames, not all useful: Efficient strategies for long-form video qa

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.717830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.717830Z digest=sha256:34f0d556d75dc26bc56e009c633cc31a11e6b76b6ff91b27a2b1d0155f202624

Observation 179bd670-daa2-447a-976c-d5715b713e58 · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Robust speech recognition via large-scale weak supervision

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:06:27.788784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:06:26.721629Z digest=sha256:59239cb711353a4550dd505bf60dc58ac63a5941030b6596e820ec40a960ca39

Observation b3cd413b-2e5d-4cde-99a6-9cd73acc26e5 · outbound

This paper cites Habitat-Matterport 3D Dataset (HM3D): 1000 Large-scale 3D Environments for Embodied AI.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Habitat-Matterport 3D Dataset (HM3D): 1000 Large-scale 3D Environments for Embodied AI

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.725687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.725687Z digest=sha256:6770fc6e9f33721ea90b43792f4938100d71f15998b3aaaf6fef63fed565ee5b

Observation babb1833-a079-4dea-97c1-6dcba476389a · outbound

This paper cites Tokenlearner: What can 8 learned tokens do for images and videos? In NeurIPS, 2021.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Tokenlearner: What can 8 learned tokens do for images and videos? In NeurIPS, 2021

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:06:27.774633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:06:26.729922Z digest=sha256:ef3822ebfbc41e27431b47e2dee86939ad26e89ae3901fcc095ec7942a4e8391

Observation 4a1aaef5-5c7c-4b4b-8d7b-7f3318b8541a · outbound

This paper cites LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.733908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.733908Z digest=sha256:8a3c6ba1e6057a86fbdbf38a716ae4efcd7f2d9981761a9dfd6c2e4d7d29897d

Observation cec287de-e943-4ef5-8f23-4a17d6b0da51 · outbound

This paper cites Aligning and prompting everything all at once for universal visual perception.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Aligning and prompting everything all at once for universal visual perception

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:06:27.758875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:06:26.737936Z digest=sha256:1b09c0ebab7bc1376a4de45cc3319b1c9d767b90277e5979174395cf0af5f3b0

Observation 0d671523-3eee-4c70-b7c1-dff89fd15863 · outbound

This paper cites Reflexion: Language agents with verbal reinforcement learning.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Reflexion: Language agents with verbal reinforcement learning

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:06:27.745028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:06:26.741639Z digest=sha256:a80406d9a922af460a915c28c721419747bc694c7fe6d08f9d36047a93c54ecd

Observation bd5d0968-1d97-4b36-b179-595b46cab0f2 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.745808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.745808Z digest=sha256:bb75aeb8b00a06fea6663cd7d281dfdb4a004de0b3be0fe02411f77a60328b81

Observation 4d1c2b4e-a74b-4f70-8c1c-f0b2b2932a61 · outbound

This paper cites EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.749785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.749785Z digest=sha256:de8ef6223674f4fb70497477cf3ce96238767d5cc01d147f3b853d1be82a88d5

Observation 010f03f9-7700-4291-ba3d-1c67a8939b41 · outbound

This paper cites Vipergpt: Visual inference via python execution for reasoning.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Vipergpt: Visual inference via python execution for reasoning

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:06:27.730371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:06:26.753979Z digest=sha256:dd2158c23779c953b072e0f520d7ec75d79c343f1acb4f551215e6a35eceb823

Observation 5d0abb2d-0042-4ea3-8b9c-3996a522b0fa · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.757838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.757838Z digest=sha256:e810c9e9727a2d535bf4cecbe9982811535b76f8177ae90dabc70b932375e8e2

Observation 19e9e558-b6aa-4f0c-a164-001f97a492af · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.762347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.762347Z digest=sha256:3cebc1ed1557e413b26a40575a43b74f74baa912a52681d056eff9805343d9b9

Observation 29234379-7dac-44aa-be2c-74cb9bb9350f · outbound

This paper cites Vamos: Versatile action models for video understanding.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Vamos: Versatile action models for video understanding

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:06:27.716273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:06:26.766293Z digest=sha256:7f79080a97e70651fe0d7984d1394a266accc939f54c2721cd6c300c7e7dbca3

Observation ec6bc11f-8871-478e-a1b9-f976c7cb2318 · outbound

This paper cites LVBench: An Extreme Long Video Understanding Benchmark.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames LVBench: An Extreme Long Video Understanding Benchmark

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.770136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.770136Z digest=sha256:e4086ee860318e94a60a40497776440f475fadd4d984f29c94466e09a29615b5

Observation 3e25099e-ea0a-45d7-84ab-2ca445722a01 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.774431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.774431Z digest=sha256:89009d090a6063c264e51929ae0937d3546678d62b3506c7a0c424e4beffd7f4

Observation 0878dda8-ff1b-4e0b-abd1-b10f1735c29a · outbound

This paper cites Vila: Efficient video-language alignment for video question answering.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Vila: Efficient video-language alignment for video question answering

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:06:27.702596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:06:26.778822Z digest=sha256:7b8995d62cb3dde7b6e73610c5409942eec21a7ade0e081472e01366067fdea7

Observation 9a207b0b-a775-4415-aac5-094f9018a8c6 · outbound

This paper cites VideoAgent: Long-form Video Understanding with Large Language Model as Agent.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames VideoAgent: Long-form Video Understanding with Large Language Model as Agent

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.782893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.782893Z digest=sha256:47c1b8283aba67140efbaa85ab449df81796a2e9579dee012101366e6e5dde3e

Observation f1cc9b77-687e-4840-91f8-69b982c35dc5 · outbound

This paper cites VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.787020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.787020Z digest=sha256:a4b61760af90e60b61919c2029f989cb2b702d3e251e21eea40bad22a6591393

Observation 558c4423-1900-440e-8699-d995e8977c03 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Chain-of-thought prompting elicits reasoning in large language models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.791453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.791453Z digest=sha256:95c87fa7993b6ffa2d484866356a080220aa97f459e5da35dd59fb4e6acebdfe

Observation 9d3bcf66-c07c-44de-9637-691bed3f7b76 · outbound

This paper cites Visual Haystacks: A Vision-Centric Needle-In-A-Haystack Benchmark.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Visual Haystacks: A Vision-Centric Needle-In-A-Haystack Benchmark

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.795465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.795465Z digest=sha256:ac5cc8782b909a509e14ee4eee74d64b0cee88a2a5fff231a795713c38b1e624

Observation b404a63e-0ac2-403b-b925-c13a9b907b38 · outbound

This paper cites Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.799622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.799622Z digest=sha256:2a2c9d403319256b8285ac43cb528095be6dd49a576c29e9211571efa90ffd7e

Observation b60ec109-f39c-4120-9e14-83ded3c3a62a · outbound

This paper cites Next-qa: Next phase of question-answering to explaining temporal actions.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Next-qa: Next phase of question-answering to explaining temporal actions

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:06:27.677233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:06:26.803791Z digest=sha256:84069b92756c1bf804730d52beba4872d37005ef9907cc47ab06beae30b9a5c3

Observation 4dd48ba5-4468-46ad-9d09-6acab640dd89 · outbound

This paper cites Retrieval-based video language model for efficient long video question answering.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Retrieval-based video language model for efficient long video question answering

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.807519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.807519Z digest=sha256:5a4af9ec95241c68516f302844ed6e89993b14f01bb7f071aaee964122148382

Observation 408380f8-1539-406f-a9fb-cd5719957c8a · outbound

This paper cites Tree of thoughts: Deliberate problem solving with large language models.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Tree of thoughts: Deliberate problem solving with large language models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.811686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.811686Z digest=sha256:0cbbfa988d01e592a2d26bd47cf692c6c9b8b991c67d796fc165388985cc54bb

Observation b09bf94d-d062-4b5e-8064-fafc3665cd94 · outbound

This paper cites Self-chained image-language model for video localization and question answering.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Self-chained image-language model for video localization and question answering

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:06:27.652196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:06:26.815384Z digest=sha256:19ab4c7f0d5be16a1c68569398fd7b4638b4cba1f088ed15cbde2af4b47370fa

Observation c8a105ce-1fb0-4fc7-9ab4-58b260b8c5fd · outbound

This paper cites Lv-eval: A balanced long-context benchmark with 5 length levels up to 256k.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Lv-eval: A balanced long-context benchmark with 5 length levels up to 256k

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.819219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.819219Z digest=sha256:17241d42588b8810b741726b4b374added331549354dc4e358fab2d62d9dbec7

Observation 3e450e21-1196-4122-802d-cf414fabd1bf · outbound

This paper cites Sigmoid loss for language image pre-training.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Sigmoid loss for language image pre-training

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.823206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.823206Z digest=sha256:95fdbee2fee1d27ee90a3cf22fda710c4a9ef0c7a459dc1534ee714c680cce5c

Observation 88277197-a309-436c-bdcd-3c159b1114e4 · outbound

This paper cites A simple llm framework for long-range video question-answering.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames A simple llm framework for long-range video question-answering

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:06:27.627536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:06:26.827083Z digest=sha256:942004ead0878c5b6fd760921ab809b13cb39f1fff1ff31a873f69d84a1ba69f

Observation 9359e176-943c-4861-9513-79e000d20a13 · outbound

This paper cites Learning video representations from large language models.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Learning video representations from large language models

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:06:27.612666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:06:26.830891Z digest=sha256:b1e1a22f5a49befdfc1087b9f35f975ed71883d0332350d2c4b615eeb5bb8dd4

Observation a798454b-76a3-45fa-8c2d-f29a9d1940b3 · outbound

This paper cites Needle In A Video Haystack: A Scalable Synthetic Evaluator for Video MLLMs.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Needle In A Video Haystack: A Scalable Synthetic Evaluator for Video MLLMs

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.834825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.834825Z digest=sha256:861bced45a851d50cbd0af5dcb842b6752555555e31f12121017c9126590be42

Observation 64cb19ff-1738-4904-bd56-a835154c65f1 · outbound

This paper cites temporal certificate.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames temporal certificate

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:06:27.597368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:06:26.838928Z digest=sha256:30cf2baa1af125183e463464a1e461d456de70b04bb6413f1a10f2810ad3295f

Pith citing papers

Observation 370ffc45-bcfe-4ee6-88f1-589cbeb5489e · inbound

Internalized Reasoning for Long-Context Visual Document Understanding cites this paper.

Internalized Reasoning for Long-Context Visual Document Understanding Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:53:28.169730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T23:53:19.148407Z digest=sha256:62149bac8c209f46fdf19341398d61b8499798ebff6db19fe19dcb191ce62d0b

Observation 1ee4d66e-d02f-4c03-80d8-7cd159425f45 · inbound

Internalized Reasoning for Long-Context Visual Document Understanding cites this paper.

Internalized Reasoning for Long-Context Visual Document Understanding Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-13T15:50:49.083652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T15:50:49.083652Z digest=sha256:694f306c1b0232af5ab406b4a1e32ce69ef92a3cb509f10cac8281bb7dee89ac

Observation 18b041e5-2a84-4125-a0b2-c384cbdce99d · inbound

Event-Causal RAG: A Retrieval-Augmented Generation Framework for Long Video Reasoning in Complex Scenarios cites this paper.

Event-Causal RAG: A Retrieval-Augmented Generation Framework for Long Video Reasoning in Complex Scenarios Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:06:13.284145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T10:15:15.129358Z digest=sha256:ea46fa968760b243eef56cb8469e7e9ea3b2be1cbc7f5c8791bc57e1140b180f

Observation 3a28ed55-3902-4a5a-bbe6-ce055291ed38 · inbound

Personal Visual Context Learning in Large Multimodal Models cites this paper.

Personal Visual Context Learning in Large Multimodal Models Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:06:37.196047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T03:42:15.402131Z digest=sha256:f8752ffb481ed271188619e13a237a581fc0359f863fd55bdc7825797a9c9414

Observation 09716617-b19c-4d9a-b9a6-7492d3692ca2 · inbound

Swift Sampling: Selecting Temporal Surprises via Taylor Series cites this paper.

Swift Sampling: Selecting Temporal Surprises via Taylor Series Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:56:07.853544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T05:55:23.479344Z digest=sha256:cf3cb7638b267e5c9a5043941da037e0b3b033b012b85a4afad0c3c3ef0db1d0

Observation 30a53c02-f069-4af1-9e5b-240630cbf94d · inbound

Zero-Shot 3D Question Answering via Hierarchical View-to-Token Transportation cites this paper.

Zero-Shot 3D Question Answering via Hierarchical View-to-Token Transportation Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames

Reference 76

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T02:26:27.041380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-28T10:54:02.188634Z digest=sha256:678944fc19520484efee10c7ade55bb46df8ab7c75da2162dd6f97e120007db6

Observation 18343c31-284e-4b4d-bcea-e216b7bff1d3 · inbound

MemLearner: Learning to Query Context memory for Video World Models cites this paper.

MemLearner: Learning to Query Context memory for Video World Models Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T10:25:41.855643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-01T05:30:56.140465Z digest=sha256:a5bfe29783ce525308af7cb2dca4fcfe5103de4487939a17c507a02fb5c74d9e

Observation be4b30a9-dbd6-4ef6-80e1-276ccb3bb29e · inbound

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos cites this paper.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:42.836941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:42.836941Z digest=sha256:ecce1653477cca5a170c5b1b15835a9d36d930dce2aa5e25f50b2e5db4936482