Pith. sign in

Paper Citation Record · LEDGER

TouchStone: Evaluating Vision-Language Models by Language Models

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2308.16890.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2308.16890 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T18:23:49.558553Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T17:18:44.036310Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 504ac152-8454-4025-b9ba-0b8070a74063 · inbound

TempCompass: Do Video LLMs Really Understand Videos? cites this paper.

TempCompass: Do Video LLMs Really Understand Videos? TouchStone: Evaluating Vision-Language Models by Language Models

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:46:16.764225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-17T02:46:16.632743Z digest=sha256:af26e9e2fe25026e8b817d50ab37bb748d390367657224b343b9df56d499ad16

Observation 5c644f71-cdfb-41ec-a154-953b0256558e · inbound

MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans? cites this paper.

MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans? TouchStone: Evaluating Vision-Language Models by Language Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:59:32.814596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T07:59:32.638758Z digest=sha256:8b820118067be9c785422d6fe9493c8eaae3971bb919ca3773936bf3f1b99355

Observation 2dcc7400-3811-4bfa-9905-b9f83fd47c1f · inbound

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment cites this paper.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment TouchStone: Evaluating Vision-Language Models by Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:49.558553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:49.558553Z digest=sha256:ea4ed6045a21634e6c3fe72672a16ebbdfb87392236be5544b2a273cd67901da

Observation f66346ef-ac07-4053-9dfd-1d66b8b07ad0 · inbound

P2P: Automated Paper-to-Poster Generation and Fine-Grained Benchmark cites this paper.

P2P: Automated Paper-to-Poster Generation and Fine-Grained Benchmark TouchStone: Evaluating Vision-Language Models by Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T15:23:20.236436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:23:20.236436Z digest=sha256:b3429e7e836012967ed77c65d7b8e323296d6bd97bb6245f81341f84c68d462c

Observation cc05ce37-688a-45b6-a477-0b980139db4f · inbound

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping cites this paper.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping TouchStone: Evaluating Vision-Language Models by Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:04.798457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:04.798457Z digest=sha256:afe11cf2b5fef08a97b72ee61d7464809c9dc224a176215be8c24dc3366eae4b

Observation 8208d6aa-a0b3-4ca8-8279-123b8ef4fb3c · inbound

From Image Captioning to Visual Storytelling cites this paper.

From Image Captioning to Visual Storytelling TouchStone: Evaluating Vision-Language Models by Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T10:27:39.160499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T10:27:39.160499Z digest=sha256:8155430835184e4fa3c33628c37572309c7ef309f5e2742e4f852fc91b82f7d0

Observation 767be77d-c9a4-44d1-87ef-859039bb6f60 · inbound

OxyEcomBench: Benchmarking Multimodal Foundation Models across E-Commerce Ecosystems cites this paper.

OxyEcomBench: Benchmarking Multimodal Foundation Models across E-Commerce Ecosystems TouchStone: Evaluating Vision-Language Models by Language Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-14T02:18:37.571735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-14T02:13:53.218695Z digest=sha256:4cb3f716b33f5d4b849794dc87d92ada234971b5f0e7450866d867161408e2a5

Observation 2cd4afe7-caa7-4bbb-97ad-3e8574ee3dcc · inbound

DisasterBench: A Multimodal Benchmark for UAV-Based Disaster Response in Complex Environments cites this paper.

DisasterBench: A Multimodal Benchmark for UAV-Based Disaster Response in Complex Environments TouchStone: Evaluating Vision-Language Models by Language Models

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-07-02T11:56:55.940571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T02:36:11.721143Z digest=sha256:3ae6adb15ee4e3041b74e5fc715508bddd03856f34e21bb379fc85bee12b443c

Observation d0d7280d-7980-4aac-9987-55276a82775b · inbound

Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation cites this paper.

Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation TouchStone: Evaluating Vision-Language Models by Language Models

Reference 94

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T17:18:44.037704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T04:19:26.332718Z digest=sha256:9a78082dd026c851dbf8a0ab0651e80846e1c6505dd41747ed0e1c24f9863955

Observation eb30ef8c-b841-49a5-8ddb-3203d059bf3a · inbound

Qwen-Audio-VAE Technical Report cites this paper.

Qwen-Audio-VAE Technical Report TouchStone: Evaluating Vision-Language Models by Language Models

Reference 109

Resolution
unresolved
no resolver link, observed 2026-07-14T03:31:19.309532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T03:31:19.309532Z digest=sha256:2580f40bb4e13901a88cba823b7ce6a74970a4c873f3898818cd3241273f9799