Pith. sign in

Paper Citation Record · LEDGER

Vibe-Eval: A hard evaluation suite for measuring progress of multimodal language models

As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2405.02287.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.02287 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T14:31:37.593033Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T19:07:17.481277Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f44115e1-e065-4567-bb46-09df096acc19 · inbound

VideoPhy: Evaluating Physical Commonsense for Video Generation cites this paper.

VideoPhy: Evaluating Physical Commonsense for Video Generation Vibe-Eval: A hard evaluation suite for measuring progress of multimodal language models

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:34:37.784037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T11:34:37.599691Z digest=sha256:df3fb607d2954bf42f38e1148f11f26702a82dc94e1dc62b33325c7a404c1375

Observation 124d0e77-9ba9-4cad-b295-463f0b2ff2f8 · inbound

LLaVA-OneVision: Easy Visual Task Transfer cites this paper.

LLaVA-OneVision: Easy Visual Task Transfer Vibe-Eval: A hard evaluation suite for measuring progress of multimodal language models

Reference 111

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:23:49.482250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:52992abcf34bfc778bc854e344bc68a5780d285c938431758855818ce817fea9

Observation 8b235945-633a-46b1-92a5-50dcdeb2758d · inbound

MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs cites this paper.

MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs Vibe-Eval: A hard evaluation suite for measuring progress of multimodal language models

Reference 216

Resolution
unresolved
no resolver link, observed 2026-08-12T14:31:37.593033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:31:37.593033Z digest=sha256:1b436ca590f9379eb4dbb300219742b30474594decc51325c0c7dc293d36768c

Observation 0b8ce77c-ad5b-4b39-82a5-5caa1005f66e · inbound

VL-RewardBench: A Challenging Benchmark for Vision-Language Generative Reward Models cites this paper.

VL-RewardBench: A Challenging Benchmark for Vision-Language Generative Reward Models Vibe-Eval: A hard evaluation suite for measuring progress of multimodal language models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T12:12:30.714912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:12:30.714912Z digest=sha256:d9e74a8573237b54b140f93b8ca323637adb174138c06a219c30097eb3af5e29

Observation d3f7a737-f74a-40f8-9644-b349ecab8e7f · inbound

ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models cites this paper.

ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models Vibe-Eval: A hard evaluation suite for measuring progress of multimodal language models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T20:55:05.905698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:55:05.905698Z digest=sha256:8d96ea6523987af4ec7dbc390c86e922336d0e8b3649808d4017338d1b58b747

Observation f0e75ab4-d008-43ec-b0e7-367f2f3232d2 · inbound

Seeing is Believing, but How Much? A Comprehensive Analysis of Verbalized Calibration in Vision-Language Models cites this paper.

Seeing is Believing, but How Much? A Comprehensive Analysis of Verbalized Calibration in Vision-Language Models Vibe-Eval: A hard evaluation suite for measuring progress of multimodal language models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:24.600283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:01:24.600283Z digest=sha256:ec9bdda9e8611dc337e62a5235c4d09de4fd4a760d34332fe5f25b6ff2b2b2f3

Observation 393c8537-f874-4d32-8226-f091b85810bc · inbound

MiMo-VL Technical Report cites this paper.

MiMo-VL Technical Report Vibe-Eval: A hard evaluation suite for measuring progress of multimodal language models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.186518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.186518Z digest=sha256:5f4d2d624464393bbf4f68c3f23c43ba68d384c4a6f3bc1f0d37f8ca066defa5

Observation ff721b82-d98c-462e-96e4-5b1561ae09f6 · inbound

Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities cites this paper.

Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities Vibe-Eval: A hard evaluation suite for measuring progress of multimodal language models

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-19T05:52:07.817113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-19T05:48:02.828938Z digest=sha256:068ed9bf3db315caf9dff1c1f15cb3b8f53de0035419d17b706784d947bf70ce

Observation 1bed5e4a-32df-43e2-84af-9274d5c77ca4 · inbound

Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity cites this paper.

Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity Vibe-Eval: A hard evaluation suite for measuring progress of multimodal language models

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:07:17.482823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-02T18:57:46.841456Z digest=sha256:569d1f73278ad52e5efa160aa46bc637c2d0c1e08960735bb844c74e90b7f038