Pith. sign in

Paper Citation Record · LEDGER

UniBench: Visual Reasoning Requires Rethinking Vision-Language Beyond Scaling

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2408.04810.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2408.04810 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:29:00.908466Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-19T05:57:08.155113Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 17e6553f-9f94-4f44-aac6-1cd611c55cae · inbound

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens cites this paper.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens UniBench: Visual Reasoning Requires Rethinking Vision-Language Beyond Scaling

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:23.786991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:23.786991Z digest=sha256:4082b915460f56511d6935068b166b9476431a23a65d0140cb7e85b0e9565a0c

Observation c1ee6603-8a4c-402e-a9eb-f70f557932e4 · inbound

AdvDreamer Unveils: Are Vision-Language Models Truly Ready for Real-World 3D Variations? cites this paper.

AdvDreamer Unveils: Are Vision-Language Models Truly Ready for Real-World 3D Variations? UniBench: Visual Reasoning Requires Rethinking Vision-Language Beyond Scaling

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T22:56:44.434584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:56:44.434584Z digest=sha256:1cf7adcaea3ec3314ace4b2b60328384338bc626a3e2beb053f8dd656b6f0bc0

Observation 03e8894b-d264-443c-920a-ab6fd4bb7e97 · inbound

Generative Physical AI in Vision: A Survey cites this paper.

Generative Physical AI in Vision: A Survey UniBench: Visual Reasoning Requires Rethinking Vision-Language Beyond Scaling

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-10T18:53:00.226717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:53:00.226717Z digest=sha256:531898170ca00365675122867a5a9ae0e826e87e511e0e810ef9564e81f4ac35

Observation 6499a395-ef09-4816-90d8-3637fc2a5575 · inbound

CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting cites this paper.

CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting UniBench: Visual Reasoning Requires Rethinking Vision-Language Beyond Scaling

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T11:29:00.908466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:29:00.908466Z digest=sha256:b3d827870df72277082020bc6aab663a55d922f5abd587c0f15695ff8708f89f

Observation 9771b082-b0d9-4e98-a987-f8f79d71e0b5 · inbound

BYO-Eval: Build Your Own Dataset for Fine-Grained Visual Assessment of Multimodal Language Models cites this paper.

BYO-Eval: Build Your Own Dataset for Fine-Grained Visual Assessment of Multimodal Language Models UniBench: Visual Reasoning Requires Rethinking Vision-Language Beyond Scaling

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T10:35:40.558440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:35:40.558440Z digest=sha256:8aa371f30321759438d2915b1a1b1558018f45a4d62a11e2bd785296e92b2a02

Observation cc4905d7-0b8d-4292-9cb9-a399ab984a99 · inbound

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks cites this paper.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks UniBench: Visual Reasoning Requires Rethinking Vision-Language Beyond Scaling

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-19T05:57:08.159043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:9b40be9564a5bf2fdbecc4756a1f0d5514cde206021270e6aa44955656a3d332

Observation 5d1b1a9b-21ad-4476-b3a4-bd2a3f721114 · inbound

EvoLMM: Self-Evolving Large Multimodal Models with Continuous Rewards cites this paper.

EvoLMM: Self-Evolving Large Multimodal Models with Continuous Rewards UniBench: Visual Reasoning Requires Rethinking Vision-Language Beyond Scaling

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T21:09:21.007755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:09:21.007755Z digest=sha256:6d849642b363fa9058436f36d6132f8c483ac5425b234e4a3e273b723b9a03af

Observation 4de9475d-178e-4730-b397-5851b843b4a1 · inbound

FeynmanBench: Benchmarking Multimodal LLMs on Diagrammatic Physics Reasoning cites this paper.

FeynmanBench: Benchmarking Multimodal LLMs on Diagrammatic Physics Reasoning UniBench: Visual Reasoning Requires Rethinking Vision-Language Beyond Scaling

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-13T16:52:59.939322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-13T16:51:48.705876Z digest=sha256:6b7541acf83b08ff457cd6b2b11b4ba5978c26ea84331245fd81e1eb13a5de34

Observation 332e3cb4-fafd-4284-aa76-18dfda44f7e1 · inbound

FeynmanBench: Benchmarking Multimodal LLMs on Diagrammatic Physics Reasoning cites this paper.

FeynmanBench: Benchmarking Multimodal LLMs on Diagrammatic Physics Reasoning UniBench: Visual Reasoning Requires Rethinking Vision-Language Beyond Scaling

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-13T12:10:53.720348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T12:10:53.720348Z digest=sha256:7271faf02b39ebba2d300df61f7426e53d42b5607c4469483c269f63905e49a6