Pith. sign in

Paper Citation Record · LEDGER

Perception Test: A Diagnostic Benchmark for Multimodal Video Models

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2305.13786.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2305.13786 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:34:12.843409Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-24T05:03:55.426403Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 282c0876-504b-4e4d-aacd-3e1f4b66717f · inbound

Gemini: A Family of Highly Capable Multimodal Models cites this paper.

Gemini: A Family of Highly Capable Multimodal Models Perception Test: A Diagnostic Benchmark for Multimodal Video Models

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-24T05:03:55.429529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-24T05:00:28.453838Z digest=sha256:d11a6a2822c50a3e5aea0872c7af662c601ce34ae190c9e2a333c46a8ce16907

Observation 996a85db-de6c-4d09-933c-9919afc93569 · inbound

TempCompass: Do Video LLMs Really Understand Videos? cites this paper.

TempCompass: Do Video LLMs Really Understand Videos? Perception Test: A Diagnostic Benchmark for Multimodal Video Models

Reference 113

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T02:46:16.794430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-17T02:46:16.632743Z digest=sha256:f6984cd1966d27a43374e1e895dd0c3cb0a4a1683981a019d84db0273b3f1b87

Observation 52d23f46-1c7e-4ab5-ada8-1cdc7367bec6 · inbound

MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale cites this paper.

MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale Perception Test: A Diagnostic Benchmark for Multimodal Video Models

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-11T20:56:42.326943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:56:42.326943Z digest=sha256:9e7fdd6ae78b5e9ec33f22533793b249e1c370612f457f65ea64219bb49edf89

Observation a9b6ac0b-271d-4707-9b02-efa9b456f11e · inbound

Movie2Story: A framework for understanding videos and telling stories in the form of novel text cites this paper.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text Perception Test: A Diagnostic Benchmark for Multimodal Video Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.204032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.204032Z digest=sha256:77fa0f7d721f8ce84bf8483bd36e5337b0a6b54088767544c345967383fc4870

Observation 3630f9cd-e051-4639-8f7b-db90f12e0eb9 · inbound

J-EDI QA: Benchmark for deep-sea organism-specific multimodal LLM cites this paper.

J-EDI QA: Benchmark for deep-sea organism-specific multimodal LLM Perception Test: A Diagnostic Benchmark for Multimodal Video Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T11:19:55.636847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:19:55.636847Z digest=sha256:e53b11a8ec88dd7550c967f98d619cc8cba66167de94923c4834c1932308f8ba

Observation 14ccbf6e-e7e4-41ea-8b8d-4b0709e9d7bf · inbound

Correspondence of high-dimensional emotion structures elicited by video clips between humans and Multimodal LLMs cites this paper.

Correspondence of high-dimensional emotion structures elicited by video clips between humans and Multimodal LLMs Perception Test: A Diagnostic Benchmark for Multimodal Video Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T20:34:12.843409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:34:12.843409Z digest=sha256:bbb72c94b78c3ae8dab1314241ea4f26c248e1dbafbc5f689186ccb573bbdd35

Observation 4a4ee60c-492e-4d88-a047-99ce02a09dfd · inbound

CausalVQA: A Physically Grounded Causal Reasoning Benchmark for Video Models cites this paper.

CausalVQA: A Physically Grounded Causal Reasoning Benchmark for Video Models Perception Test: A Diagnostic Benchmark for Multimodal Video Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T04:45:01.031903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:45:01.031903Z digest=sha256:33d49d5ae88bfedadcc0d91ef041140514ac48fa4093c3bf03fa78fc38e420c3

Observation bdd09533-6b82-46f0-921e-33d77498e756 · inbound

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models cites this paper.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Perception Test: A Diagnostic Benchmark for Multimodal Video Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:30.481104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:30.481104Z digest=sha256:c4e4ae68ef571c5ad9b3fa69627f55d01ac917bd242742a40e950f3936ea3ab4

Observation 66a69bad-3da5-4e39-a2cc-3e7d8433c4d8 · inbound

ReGATE: Learning Faster and Better with Fewer Tokens in MLLMs cites this paper.

ReGATE: Learning Faster and Better with Fewer Tokens in MLLMs Perception Test: A Diagnostic Benchmark for Multimodal Video Models

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-19T03:22:01.209863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-19T03:18:11.993413Z digest=sha256:931aead1ca41b9097fc1506e6be42017bcebe09ee90d06426274a16cb3696351

Observation 7f192378-6656-46d9-82d4-cce3ea3690ab · inbound

Seeing More, Saying More: Lightweight Language Experts are Dynamic Video Token Compressors cites this paper.

Seeing More, Saying More: Lightweight Language Experts are Dynamic Video Token Compressors Perception Test: A Diagnostic Benchmark for Multimodal Video Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T13:05:55.211311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:05:55.211311Z digest=sha256:1e203b037ba0c74d97bfa01d927ace0bf06116959a1a78abd4b7b05e2c5ce609

Observation 3c68a8dd-bae2-424e-9ef6-57b524686eb7 · inbound

HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understanding cites this paper.

HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understanding Perception Test: A Diagnostic Benchmark for Multimodal Video Models

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:57:53.852870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T12:55:04.564442Z digest=sha256:7a4d6ee97328b2fc75183d8153049f8dc779b7f59a43868160e89ed3b1ae5614