Pith. sign in

Paper Citation Record · LEDGER

ViLMA: A Zero-Shot Benchmark for Linguistic and Temporal Grounding in Video-Language Models

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2311.07022.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.07022 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:59:33.726301Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T17:58:29.104450Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b78e71a2-12a7-4abb-a751-b8e93525728f · inbound

VLM-AD: End-to-End Autonomous Driving through Vision-Language Model Supervision cites this paper.

VLM-AD: End-to-End Autonomous Driving through Vision-Language Model Supervision ViLMA: A Zero-Shot Benchmark for Linguistic and Temporal Grounding in Video-Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T12:17:54.231007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:17:54.231007Z digest=sha256:5b84c3c8d6fbde9f5605e592437813a1e483cceb8712a5a5c29ef4e79b882cde

Observation d78f15c0-b7e3-4db3-8dc0-380def64992e · inbound

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models cites this paper.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models ViLMA: A Zero-Shot Benchmark for Linguistic and Temporal Grounding in Video-Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.726301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.726301Z digest=sha256:6f936b31a4d7c9b2f511226e1ea9a02d4d2a8c7c4dfd701749cc14654bc2d33a

Observation ae2016cf-432c-4a34-a33f-3e6feffeb33d · inbound

EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World? cites this paper.

EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World? ViLMA: A Zero-Shot Benchmark for Linguistic and Temporal Grounding in Video-Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:10.258668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:10.258668Z digest=sha256:8d2cf8d37f18194a4e44ba7139ca90820ccff8a17985a602e0175a5087d34a23

Observation 2ea50030-f356-41cf-8e3b-dffd43350ee0 · inbound

Argus Inspection: Do Multimodal Large Language Models Possess the Eye of Panoptes? cites this paper.

Argus Inspection: Do Multimodal Large Language Models Possess the Eye of Panoptes? ViLMA: A Zero-Shot Benchmark for Linguistic and Temporal Grounding in Video-Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:18:59.175905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:18:59.175905Z digest=sha256:7bda3b492a87d7774028a72846f7b888ab9aeecdf007b67406f98475f4094727

Observation 5a7676cf-84d0-49ea-9aba-f61972b1194c · inbound

GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them? cites this paper.

GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them? ViLMA: A Zero-Shot Benchmark for Linguistic and Temporal Grounding in Video-Language Models

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T17:58:29.172776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:58:27.448450Z digest=sha256:00dc9a6c922fe802077ae18f7ec0e728e2b47cbf95fa9be6d62872629ea9ce44