Pith. sign in

Paper Citation Record · LEDGER

Do Thought Streams Matter? Evaluating Reasoning in Gemini Vision-Language Models for Video Scene Understanding

As of 4 August 2026, this Paper Citation Record lists 11 of 11 outbound references and 0 inbound Pith citation observations for arXiv:2604.11177.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.11177 v1

Coverage vector

measured 11 of 11 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T16:05:46.467102Z

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

11 of 11 outbound references displayed

  • verified exact5
  • verified fuzzy6
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 184a4c4a-f035-4b59-a0fb-ea6f0f766382 · outbound

This paper cites ActivityNet: A large-scale video benchmark for human activity understanding.

Do Thought Streams Matter? Evaluating Reasoning in Gemini Vision-Language Models for Video Scene Understanding ActivityNet: A large-scale video benchmark for human activity understanding

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T17:01:59.495645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:05:46.467102Z digest=sha256:78b72060b2052919c7a7cd54674ecc68ada1a7338b5e709809d9754efe6b96bc

Observation 9ea846ef-fd0b-49df-bc55-3f3136ae3b8d · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Do Thought Streams Matter? Evaluating Reasoning in Gemini Vision-Language Models for Video Scene Understanding Evaluating Large Language Models Trained on Code

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-11T09:21:00.049963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:05:46.467102Z digest=sha256:1993bb6bdf3f5130bfefdb482e39ea574373594dfcdb2012296dd3f9194fccf9

Observation 12623f48-c162-468b-ae0e-d22fd43ce55c · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

Do Thought Streams Matter? Evaluating Reasoning in Gemini Vision-Language Models for Video Scene Understanding Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-11T09:21:00.028188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:05:46.467102Z digest=sha256:91bb280f80da38303c28e41a529a9afd52c72f1443c89ebb3820644972d16507

Observation d6180581-857c-4979-9ce9-996ea3826343 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Do Thought Streams Matter? Evaluating Reasoning in Gemini Vision-Language Models for Video Scene Understanding Gemini: A Family of Highly Capable Multimodal Models

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-11T09:21:00.044696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:05:46.467102Z digest=sha256:a8a2d976ffc580e60bd1ab2750ff0eeea8c210914b1b52785ee41be636b30a45

Observation 74f86ccc-4312-489a-bd07-8b948d636de4 · outbound

This paper cites Ego4D: Around the world in 3,000 hours of egocentric video.

Do Thought Streams Matter? Evaluating Reasoning in Gemini Vision-Language Models for Video Scene Understanding Ego4D: Around the world in 3,000 hours of egocentric video

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T17:01:59.482617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:05:46.467102Z digest=sha256:a7ed4b0b285b7270b3074c4ebcd0dac25604c9cdae48fc450493717aba8c84b1

Observation 0c654cf5-93d0-49ec-b169-92825527aa05 · outbound

This paper cites Measuring massive multitask language understanding.

Do Thought Streams Matter? Evaluating Reasoning in Gemini Vision-Language Models for Video Scene Understanding Measuring massive multitask language understanding

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T17:01:59.478851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:05:46.467102Z digest=sha256:8e5be7b3e6b5626f1ffd9e732b61753630c3b5f8c72e2c3e6ad228dea3f8656b

Observation ed5db929-2208-4d51-a68c-dbe79ed601de · outbound

This paper cites GPT-4 Technical Report.

Do Thought Streams Matter? Evaluating Reasoning in Gemini Vision-Language Models for Video Scene Understanding GPT-4 Technical Report

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-11T09:21:00.031299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:05:46.467102Z digest=sha256:4b9bf999dbb4ea13ffa4d9b4a840ab8d83be25c03cdd91151dc08e986e14d0a9

Observation e8f66c76-3cf0-4b62-9add-3b97e0757e59 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Do Thought Streams Matter? Evaluating Reasoning in Gemini Vision-Language Models for Video Scene Understanding Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-11T09:21:00.037094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:05:46.467102Z digest=sha256:1e226a421b70db240a9e8a26288e688d4f0ced0371ebdaa5b258a4487d9450bd

Observation fc4373d3-ec42-4a9d-a61b-3e2a3b9914f1 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Do Thought Streams Matter? Evaluating Reasoning in Gemini Vision-Language Models for Video Scene Understanding Chain-of-thought prompting elicits reasoning in large language models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T17:01:59.486427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:05:46.467102Z digest=sha256:06516cba3adb824512057c605bb6cb6b60b543ba36dfdb6c7593461921f812c8

Observation 09e4eda8-0908-4689-abce-7ec6be1ce3c9 · outbound

This paper cites HellaSwag: Can a ma- chine really finish your sentence? InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics (ACL), pages 4791–4800.

Do Thought Streams Matter? Evaluating Reasoning in Gemini Vision-Language Models for Video Scene Understanding HellaSwag: Can a ma- chine really finish your sentence? InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics (ACL), pages 4791–4800

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T17:01:59.490380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:05:46.467102Z digest=sha256:3d6a0960ca8e5e1a4fe0200e4b67925270e0afc0b3d06fa0373c87559389807e

Observation 8b66a822-f6ee-4374-a7b9-590412a570e5 · outbound

This paper cites Xing, et al.

Do Thought Streams Matter? Evaluating Reasoning in Gemini Vision-Language Models for Video Scene Understanding Xing, et al

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T17:01:59.474637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:05:46.467102Z digest=sha256:90ec0eb8b31adda51c7d20b4def63343ebe319ca025962ca808309b5459e45d3

Pith citing papers

No inbound Pith citation observations are available.