Pith. sign in

Paper Citation Record · LEDGER

Revealing Single Frame Bias for Video-and-Language Learning

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2206.03428.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2206.03428 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:24:08.054008Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-19T12:57:17.859675Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 3b609533-6cb3-4a1a-95c7-23944a314510 · inbound

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation cites this paper.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Revealing Single Frame Bias for Video-and-Language Learning

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:30:22.675262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:0db21d5134e091e27cf50cf670dcbdcc408913144976473a79a32f2ec70a346a

Observation 4a5d9929-c04c-4477-8cce-51421cbc6261 · inbound

LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment cites this paper.

LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment Revealing Single Frame Bias for Video-and-Language Learning

Reference 182

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T03:27:59.157987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-17T03:27:58.952076Z digest=sha256:476969a758cdb7d23387ca5fba0d0fc7e919a17cf3c8b0d4d67055d33f27f401

Observation b6a91d0e-1fbf-4fa8-9f31-d31fde10f0c8 · inbound

TempCompass: Do Video LLMs Really Understand Videos? cites this paper.

TempCompass: Do Video LLMs Really Understand Videos? Revealing Single Frame Bias for Video-and-Language Learning

Reference 93

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T02:46:16.881605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-17T02:46:16.632743Z digest=sha256:b97bac0eb0f145c2bf30fd180219875050f6d3a46ab410e15fafedd89fa4331f

Observation 5db36479-bdbe-403c-96c4-d7aeddf531c1 · inbound

LLaVA-Video: Video Instruction Tuning With Synthetic Data cites this paper.

LLaVA-Video: Video Instruction Tuning With Synthetic Data Revealing Single Frame Bias for Video-and-Language Learning

Reference 151

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:20:33.094705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:3f471f3ba0a4f804e136812110115c5872a5db8659798ca0eb9d797fea746148

Observation 4e4bce62-752d-4272-ab3f-e477d18fe4a4 · inbound

RTime-QA: A Benchmark for Atomic Temporal Event Understanding in Large Multi-modal Models cites this paper.

RTime-QA: A Benchmark for Atomic Temporal Event Understanding in Large Multi-modal Models Revealing Single Frame Bias for Video-and-Language Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:08.054008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:24:08.054008Z digest=sha256:695e4c749b35b66ad2b2a48cb3a48621115ecc67aecf42b8595a7363c135887a

Observation 31890dce-20eb-4158-b682-ef76c8a461d4 · inbound

One Trajectory, One Token: Grounded Video Tokenization via Panoptic Sub-object Trajectory cites this paper.

One Trajectory, One Token: Grounded Video Tokenization via Panoptic Sub-object Trajectory Revealing Single Frame Bias for Video-and-Language Learning

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-19T12:57:17.861503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T12:54:31.765909Z digest=sha256:650adcbb0604b6a66b788768e9cb407d8f2436f30b6572045d0809cfaa35e165

Observation 942b5cc6-c1e0-4267-903d-465920c8a26b · inbound

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering cites this paper.

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Revealing Single Frame Bias for Video-and-Language Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T23:29:20.635934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:29:20.635934Z digest=sha256:9904ba7eda3e50dd1923d71e1ab4caef948a154a93a8037329ad98dd427a0cad

Observation d810de34-9890-46c5-9a87-0269ca7b868c · inbound

Adapting MLLMs for Nuanced Video Retrieval cites this paper.

Adapting MLLMs for Nuanced Video Retrieval Revealing Single Frame Bias for Video-and-Language Learning

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:21:18.850444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T22:20:09.051957Z digest=sha256:44562cf3faf55260cfdd4ef27c78cb623cb4c9f45bc31c7d21e97a79be60a3ed