Pith. sign in

Paper Citation Record · LEDGER

Team of One: Cracking Complex Video QA with Model Synergy

As of 8 August 2026, this Paper Citation Record lists 9 of 9 outbound references and 0 inbound Pith citation observations for arXiv:2507.13820.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.13820 v1

Coverage vector

measured 9 of 9 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T16:19:20.572832Z

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

9 of 9 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7942b469-8135-4613-b723-615543de5f56 · outbound

This paper cites write newline.

Team of One: Cracking Complex Video QA with Model Synergy write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T16:19:19.594085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:19:19.594085Z digest=sha256:a507cbd49fe46f0f1a8a761a5a5719ed275f17cc9631881e5dea116e36f4be04

Observation a26eab60-698e-4f3f-a062-4c57e6fe0641 · outbound

This paper cites Gemini 2.5 Pro Preview (2025-03-25) , 2025.

Team of One: Cracking Complex Video QA with Model Synergy Gemini 2.5 Pro Preview (2025-03-25) , 2025

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:19:21.159092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T16:19:19.680262Z digest=sha256:43c2bdc6d58f591f01e990c79b0a57c1645953a622678fdb19ec0bbc08388c29

Observation 51d3faf4-ad0f-42bd-afbe-a5773eadb2e6 · outbound

This paper cites How Good is my Video LMM? Complex Video Reasoning and Robustness Evaluation Suite for Video-LMMs.

Team of One: Cracking Complex Video QA with Model Synergy How Good is my Video LMM? Complex Video Reasoning and Robustness Evaluation Suite for Video-LMMs

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T16:19:19.770396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:19:19.770396Z digest=sha256:f207aa95184ac68b8c0a83fe0bf5f0eb82b8694eaca3abdf5e4836c46013ee69

Observation fb465385-0d6d-433f-ad89-d4a14ff65ae1 · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models.

Team of One: Cracking Complex Video QA with Model Synergy Llama-vid: An image is worth 2 tokens in large language models

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:19:20.890382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T16:19:19.916856Z digest=sha256:051e2ae7a932873eaca2117e77be242b6c3acbc3a5411c6ad7d3782b4bc189f9

Observation c4b533c1-c4c3-466b-99f3-ac902afaac00 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

Team of One: Cracking Complex Video QA with Model Synergy Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T16:19:20.053425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:19:20.053425Z digest=sha256:c5ae470d7e8b1125f23a239563124557f219c424f400f081489bc12ecb2cc7a4

Observation 7c45b3d7-ddaa-431b-adf1-0af8a14fca00 · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

Team of One: Cracking Complex Video QA with Model Synergy Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T16:19:20.204460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:19:20.204460Z digest=sha256:634dfa808305973674ada85f4012383479a73ae0fd99af9bcc82c44a17762539

Observation 3f310098-559d-497b-aa17-2aa12af0fde3 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Team of One: Cracking Complex Video QA with Model Synergy Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T16:19:20.329167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:19:20.329167Z digest=sha256:b78684ad06ce986a20f41afd3b82c1c9862ebea6a6b1ce166ccf564b3970a408

Observation df85bc52-675d-40b2-9061-7dfa2486f749 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Team of One: Cracking Complex Video QA with Model Synergy Chain-of-thought prompting elicits reasoning in large language models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T16:19:20.430057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:19:20.430057Z digest=sha256:e6892fbbc0d48045b708ccb56e9c78e0ce2b541755822de2c75c358ea6259c29

Observation 1b7c2315-9f7b-48f5-9ee5-19ff8e9f5dee · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

Team of One: Cracking Complex Video QA with Model Synergy Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T16:19:20.572832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:19:20.572832Z digest=sha256:01645a4a2ab02f0daf9d02dfd9a8fde195d9d25e006bc017f17f7e25df4298ac

Pith citing papers

No inbound Pith citation observations are available.