Pith. sign in

Paper Citation Record · LEDGER

Revealing Single Frame Bias for Video-and-Language Learning

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2206.03428.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2206.03428 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:24:08.054008Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-19T12:57:17.859675Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 3b609533-6cb3-4a1a-95c7-23944a314510 · inbound

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation cites this paper.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Revealing Single Frame Bias for Video-and-Language Learning

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:30:22.675262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:c07973bac2a2d4cb3da1fffb42c92ea7101886c0dc0181ba037159bd1a94b1de

Observation 4a5d9929-c04c-4477-8cce-51421cbc6261 · inbound

LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment cites this paper.

LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment Revealing Single Frame Bias for Video-and-Language Learning

Reference 182

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T03:27:59.157987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-17T03:27:58.952076Z digest=sha256:2d4053f88575e7773638feca2bdb836d8314b47c41aee45f09d49158fb8e3630

Observation b6a91d0e-1fbf-4fa8-9f31-d31fde10f0c8 · inbound

TempCompass: Do Video LLMs Really Understand Videos? cites this paper.

TempCompass: Do Video LLMs Really Understand Videos? Revealing Single Frame Bias for Video-and-Language Learning

Reference 93

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T02:46:16.881605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-17T02:46:16.632743Z digest=sha256:cea5292cdce890de9cff96a03bbe64ad24bcc8aff5dc772cb782fc66e09d782b

Observation 5db36479-bdbe-403c-96c4-d7aeddf531c1 · inbound

LLaVA-Video: Video Instruction Tuning With Synthetic Data cites this paper.

LLaVA-Video: Video Instruction Tuning With Synthetic Data Revealing Single Frame Bias for Video-and-Language Learning

Reference 151

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:20:33.094705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:d6ead0a1f233809ebc290a6488cf1573561c1f850af164279e0e261724704317

Observation 4e4bce62-752d-4272-ab3f-e477d18fe4a4 · inbound

RTime-QA: A Benchmark for Atomic Temporal Event Understanding in Large Multi-modal Models cites this paper.

RTime-QA: A Benchmark for Atomic Temporal Event Understanding in Large Multi-modal Models Revealing Single Frame Bias for Video-and-Language Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:08.054008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:24:08.054008Z digest=sha256:695e4c749b35b66ad2b2a48cb3a48621115ecc67aecf42b8595a7363c135887a

Observation 31890dce-20eb-4158-b682-ef76c8a461d4 · inbound

One Trajectory, One Token: Grounded Video Tokenization via Panoptic Sub-object Trajectory cites this paper.

One Trajectory, One Token: Grounded Video Tokenization via Panoptic Sub-object Trajectory Revealing Single Frame Bias for Video-and-Language Learning

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-19T12:57:17.861503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T12:54:31.765909Z digest=sha256:9d1d19d23620f873d96c421ceb1530aeaaede46b272b555846d4bfb91f3254d8

Observation 942b5cc6-c1e0-4267-903d-465920c8a26b · inbound

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering cites this paper.

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Revealing Single Frame Bias for Video-and-Language Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T23:29:20.635934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:29:20.635934Z digest=sha256:9904ba7eda3e50dd1923d71e1ab4caef948a154a93a8037329ad98dd427a0cad

Observation d810de34-9890-46c5-9a87-0269ca7b868c · inbound

Adapting MLLMs for Nuanced Video Retrieval cites this paper.

Adapting MLLMs for Nuanced Video Retrieval Revealing Single Frame Bias for Video-and-Language Learning

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:21:18.850444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T22:20:09.051957Z digest=sha256:1cf26d3844cedb8c6d562e94b264ea28a47037b832f65aa44710e7286dcb8a23