Pith. sign in

Paper Citation Record · LEDGER

Perception Test: A Diagnostic Benchmark for Multimodal Video Models

As of 16 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2305.13786.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2305.13786 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:34:12.843409Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-24T05:03:55.426403Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 282c0876-504b-4e4d-aacd-3e1f4b66717f · inbound

Gemini: A Family of Highly Capable Multimodal Models cites this paper.

Gemini: A Family of Highly Capable Multimodal Models Perception Test: A Diagnostic Benchmark for Multimodal Video Models

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-24T05:03:55.429529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-24T05:00:28.453838Z digest=sha256:9571915e583d27ded296a17c651bdce8e7257719bf71aa7643e7d0ab81811194

Observation 996a85db-de6c-4d09-933c-9919afc93569 · inbound

TempCompass: Do Video LLMs Really Understand Videos? cites this paper.

TempCompass: Do Video LLMs Really Understand Videos? Perception Test: A Diagnostic Benchmark for Multimodal Video Models

Reference 113

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T02:46:16.794430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-17T02:46:16.632743Z digest=sha256:3421bccebe5a2de8576fb3dbfcdc4c73d78bed67e3e80cf88e65c2269fc77f40

Observation 52d23f46-1c7e-4ab5-ada8-1cdc7367bec6 · inbound

MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale cites this paper.

MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale Perception Test: A Diagnostic Benchmark for Multimodal Video Models

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-11T20:56:42.326943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:56:42.326943Z digest=sha256:ddf956b009f417a058a3416f5d59ef93cbbb3602953d90cb37b7f5926fbd5756

Observation a9b6ac0b-271d-4707-9b02-efa9b456f11e · inbound

Movie2Story: A framework for understanding videos and telling stories in the form of novel text cites this paper.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text Perception Test: A Diagnostic Benchmark for Multimodal Video Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.204032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.204032Z digest=sha256:bea7911b209f410cef87ccdb2ddd913c6715f87056fad3bde9004b4df6b4a269

Observation 3630f9cd-e051-4639-8f7b-db90f12e0eb9 · inbound

J-EDI QA: Benchmark for deep-sea organism-specific multimodal LLM cites this paper.

J-EDI QA: Benchmark for deep-sea organism-specific multimodal LLM Perception Test: A Diagnostic Benchmark for Multimodal Video Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T11:19:55.636847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:19:55.636847Z digest=sha256:0b0b60dc6053c6dc4a744efdae06b74504c4d69ef01d0b0b3e702cd18427e7e0

Observation 14ccbf6e-e7e4-41ea-8b8d-4b0709e9d7bf · inbound

Correspondence of high-dimensional emotion structures elicited by video clips between humans and Multimodal LLMs cites this paper.

Correspondence of high-dimensional emotion structures elicited by video clips between humans and Multimodal LLMs Perception Test: A Diagnostic Benchmark for Multimodal Video Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T20:34:12.843409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:34:12.843409Z digest=sha256:0273ce0a44c10eae64c0c69398590f77b67c0c43922dbdf72bac6231f0912a6d

Observation 4a4ee60c-492e-4d88-a047-99ce02a09dfd · inbound

CausalVQA: A Physically Grounded Causal Reasoning Benchmark for Video Models cites this paper.

CausalVQA: A Physically Grounded Causal Reasoning Benchmark for Video Models Perception Test: A Diagnostic Benchmark for Multimodal Video Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T04:45:01.031903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:45:01.031903Z digest=sha256:a9efedce30e44a95674cbd1d9f8dd1f331ab0c8ef477ef088e6b800d930667c5

Observation bdd09533-6b82-46f0-921e-33d77498e756 · inbound

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models cites this paper.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Perception Test: A Diagnostic Benchmark for Multimodal Video Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:30.481104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:30.481104Z digest=sha256:e0da872ade0a0a1505fc74c92243de79f58146eea68eb706cbc2b1dbcfbcc399

Observation 66a69bad-3da5-4e39-a2cc-3e7d8433c4d8 · inbound

ReGATE: Learning Faster and Better with Fewer Tokens in MLLMs cites this paper.

ReGATE: Learning Faster and Better with Fewer Tokens in MLLMs Perception Test: A Diagnostic Benchmark for Multimodal Video Models

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-19T03:22:01.209863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-19T03:18:11.993413Z digest=sha256:bad3300d9567db31bb69e84f21c7ce9e3c57692b3a828a368fbb149152942df8

Observation 7f192378-6656-46d9-82d4-cce3ea3690ab · inbound

Seeing More, Saying More: Lightweight Language Experts are Dynamic Video Token Compressors cites this paper.

Seeing More, Saying More: Lightweight Language Experts are Dynamic Video Token Compressors Perception Test: A Diagnostic Benchmark for Multimodal Video Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T13:05:55.211311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:05:55.211311Z digest=sha256:360ad91dba615cb5be60deb4bc74faec4f94258a2dd892b32e03921471fdd06b

Observation 3c68a8dd-bae2-424e-9ef6-57b524686eb7 · inbound

HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understanding cites this paper.

HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understanding Perception Test: A Diagnostic Benchmark for Multimodal Video Models

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:57:53.852870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-16T12:55:04.564442Z digest=sha256:679946f71c9113d98f9364a7427ed88b47b1672e21080f002c274bd4eea98483