Pith. sign in

Paper Citation Record · LEDGER

VisIT-Bench: A Benchmark for Vision-Language Instruction Following Inspired by Real-World Use

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:2308.06595.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2308.06595 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:47:21.063096Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T08:49:42.520344Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 645e2664-f561-446e-b5f4-cb4fb8d7985d · inbound

LLM Evaluators Recognize and Favor Their Own Generations cites this paper.

LLM Evaluators Recognize and Favor Their Own Generations VisIT-Bench: A Benchmark for Vision-Language Instruction Following Inspired by Real-World Use

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-22T18:44:28.800274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-22T18:44:28.766639Z digest=sha256:8e3786a382edf2a13c6367cf859b1da042f7c88cffafce39a629a76b2866f6be

Observation 97bf3576-2483-41ec-bf3b-e25af6f00177 · inbound

Cracking the Code of Juxtaposition: Can AI Models Understand the Humorous Contradictions cites this paper.

Cracking the Code of Juxtaposition: Can AI Models Understand the Humorous Contradictions VisIT-Bench: A Benchmark for Vision-Language Instruction Following Inspired by Real-World Use

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-24T00:53:40.644813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-24T00:52:52.056076Z digest=sha256:8ae5bd03d81b905f652c5d512dbf81cabe58f1fb87061a1e070e0e731544283d

Observation 6324c99f-25a7-4c9b-b78d-73f539624710 · inbound

VideoPhy: Evaluating Physical Commonsense for Video Generation cites this paper.

VideoPhy: Evaluating Physical Commonsense for Video Generation VisIT-Bench: A Benchmark for Vision-Language Instruction Following Inspired by Real-World Use

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:34:37.649327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T11:34:37.599691Z digest=sha256:589e24d37fb1e3db0d63d61f66519c7909a9d17e0eba98fdf57525e485ae2c95

Observation 151a6662-a5c1-4d3e-b2f4-2db3112de5ee · inbound

MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans? cites this paper.

MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans? VisIT-Bench: A Benchmark for Vision-Language Instruction Following Inspired by Real-World Use

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:59:32.821120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-16T07:59:32.638758Z digest=sha256:086129237bfef43779218b928c7c9c9f49623dfd84090459a70606e76d85c2ca

Observation 6c52f880-66dc-44bc-b6d0-94ef03a84f3a · inbound

Polymath: A Challenging Multi-modal Mathematical Reasoning Benchmark cites this paper.

Polymath: A Challenging Multi-modal Mathematical Reasoning Benchmark VisIT-Bench: A Benchmark for Vision-Language Instruction Following Inspired by Real-World Use

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:05:47.842622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-23T20:03:38.336841Z digest=sha256:1d6f410b7a341823656a4132a46e958e4b4895b05128c521a81f5d2126362681

Observation 448baf20-ff7e-400e-9967-8ca766215267 · inbound

MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs cites this paper.

MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs VisIT-Bench: A Benchmark for Vision-Language Instruction Following Inspired by Real-World Use

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T14:31:36.808308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:31:36.808308Z digest=sha256:70416978660c624d78008c56228dff78da64b1d98eac46755a9eb072ad46a47b

Observation 9d71215c-51e5-4115-a899-f61f15b75b00 · inbound

VladVA: Discriminative Fine-tuning of LVLMs cites this paper.

VladVA: Discriminative Fine-tuning of LVLMs VisIT-Bench: A Benchmark for Vision-Language Instruction Following Inspired by Real-World Use

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T21:29:26.274862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:29:26.274862Z digest=sha256:52eb7272e9b41db89527b6df8c5fbb20ef6cc1b78bbee191c3051d26a3b4ce61

Observation b6300708-95dd-4364-98a6-de3baf853e30 · inbound

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment cites this paper.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment VisIT-Bench: A Benchmark for Vision-Language Instruction Following Inspired by Real-World Use

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:49.564656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:49.564656Z digest=sha256:eb7a1c931895fc827acc3460ffdc1b9081d3a307d81b3e78dba6cfd2987314c0

Observation 6b1111fa-762f-4928-9212-46b65acf39e4 · inbound

When 'YES' Meets 'BUT': Can Large Models Comprehend Contradictory Humor Through Comparative Reasoning? cites this paper.

When 'YES' Meets 'BUT': Can Large Models Comprehend Contradictory Humor Through Comparative Reasoning? VisIT-Bench: A Benchmark for Vision-Language Instruction Following Inspired by Real-World Use

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-22T22:42:13.662482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-22T22:38:35.969273Z digest=sha256:1d469cb9179d992fe9456fae62f8e4d4d2eb63ddc1dcb75cf225b38d0dc47a6d

Observation cd280595-afce-46ff-96b7-565d47a44a47 · inbound

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping cites this paper.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping VisIT-Bench: A Benchmark for Vision-Language Instruction Following Inspired by Real-World Use

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:05.026455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:05.026455Z digest=sha256:aae9667619bfdba0ccac459a14d4d2dcb95a215fce51f66d019d64fa57860a74

Observation 8a537809-b3ad-4d46-a0ed-39b94ad498ba · inbound

Breaking the Mirror: Activation-Based Mitigation of Self-Preference in LLM Evaluators cites this paper.

Breaking the Mirror: Activation-Based Mitigation of Self-Preference in LLM Evaluators VisIT-Bench: A Benchmark for Vision-Language Instruction Following Inspired by Real-World Use

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-05T10:51:44.475907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:51:44.475907Z digest=sha256:560ea53247413167ece5ecb92e431739b7a0b4e1f084f95ed04adc79147bf954

Observation ba44526c-1df7-4e12-9fab-bb70b0da45ab · inbound

Inverse IFEval: Can LLMs Unlearn Stubborn Training Conventions to Follow Real Instructions? cites this paper.

Inverse IFEval: Can LLMs Unlearn Stubborn Training Conventions to Follow Real Instructions? VisIT-Bench: A Benchmark for Vision-Language Instruction Following Inspired by Real-World Use

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T10:16:51.066071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:16:51.066071Z digest=sha256:9ba9a8277339e32f0ae5caa2b750a2455f1efd8840b7a8a09f30d57481c87ffd

Observation 9b10dda8-45ad-4bc0-8edf-8be54046ef0e · inbound

VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes cites this paper.

VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes VisIT-Bench: A Benchmark for Vision-Language Instruction Following Inspired by Real-World Use

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T15:47:21.063096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:47:21.063096Z digest=sha256:1c125d30dae1fcb4922e5f04402b5338ef384633dcc61501a4e9cc449f19fbb7

Observation 07d9bdea-0282-4803-a703-bddc4dd1e19f · inbound

LLM-as-Judge Framework for Evaluating Tone-Induced Hallucination in Vision-Language Models cites this paper.

LLM-as-Judge Framework for Evaluating Tone-Induced Hallucination in Vision-Language Models VisIT-Bench: A Benchmark for Vision-Language Instruction Following Inspired by Real-World Use

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:38:43.153061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T05:11:01.039309Z digest=sha256:e20b79f8bb246d432db691ade0fa96b8d6493bd0cee21141242cb08cfadccc14

Observation bdf34ac1-4e87-4854-b000-b80797822dcd · inbound

DisasterBench: A Multimodal Benchmark for UAV-Based Disaster Response in Complex Environments cites this paper.

DisasterBench: A Multimodal Benchmark for UAV-Based Disaster Response in Complex Environments VisIT-Bench: A Benchmark for Vision-Language Instruction Following Inspired by Real-World Use

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-07-02T11:56:55.937243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-28T02:36:11.721143Z digest=sha256:7cff97707a41e560b2f899b6ee7e488ab3fc429a97d2a5665d68cfcef84a587d

Observation 07333a90-e63f-471b-9bdb-05a47d8ed0c3 · inbound

CVSBench: A Comprehensive Benchmark for Cross-view Spatial Reasoning and Dreaming cites this paper.

CVSBench: A Comprehensive Benchmark for Cross-view Spatial Reasoning and Dreaming VisIT-Bench: A Benchmark for Vision-Language Instruction Following Inspired by Real-World Use

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T08:49:42.521679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-26T10:54:49.774490Z digest=sha256:bd88ee0dbbe6abacd9eb3d3e23ee7faf99dbd771c0ae4a4f65324448a0f89376

Observation c6619a0e-fa7a-4650-b7a9-652ab7d0f249 · inbound

Poller: Are LLMs Suitable for Evaluating the Poetry Understanding Task? cites this paper.

Poller: Are LLMs Suitable for Evaluating the Poetry Understanding Task? VisIT-Bench: A Benchmark for Vision-Language Instruction Following Inspired by Real-World Use

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-06-30T06:04:21.562417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-30T05:59:58.183264Z digest=sha256:b0bdf56c5d374efc449afa801c541ce2ca704a1ecca5dfa9199f1ca1df0e597d