Pith. sign in

Paper Citation Record · LEDGER

VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 26 inbound Pith citation observations for arXiv:2407.11691.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.11691 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 26 of 26 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:49:27.762903Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

3
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 9b2b6afd-e249-426a-9884-5d899c6e2245 · inbound

Number it: Temporal Grounding Videos like Flipping Manga cites this paper.

Number it: Temporal Grounding Videos like Flipping Manga VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T19:49:43.334478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:49:43.334478Z digest=sha256:b1b78481c08cf822f7e8dfc82cce94ffcdf068e4e6d1c335f69f1474cae41428

Observation bcafcba6-8f48-41d5-9d96-1e189e1fea68 · inbound

MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs cites this paper.

MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models

Reference 212

Resolution
unresolved
no resolver link, observed 2026-08-12T14:31:37.574912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:31:37.574912Z digest=sha256:924fe4feedaf8d2b6cbbc8f8119f532a9a2226eba79c313ade79b61aa044d2d7

Observation 5714edc0-0fb7-401a-aa5b-ba5e6dfac8b5 · inbound

VARCO-VISION: Expanding Frontiers in Korean Vision-Language Models cites this paper.

VARCO-VISION: Expanding Frontiers in Korean Vision-Language Models VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T10:36:35.890601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:36:35.890601Z digest=sha256:324a44719b172f71cc675b0b212912c0ffc76bdae4d5a5ce48164b4565e58304

Observation 988de421-57cb-44d4-a950-f2f4a2a42171 · inbound

Advancing Myopia To Holism: Fully Contrastive Language-Image Pre-training cites this paper.

Advancing Myopia To Holism: Fully Contrastive Language-Image Pre-training VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T05:27:14.659539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:27:14.659539Z digest=sha256:17d315365765f6a71392282ab78d77b2816f54b25704af436f4fa28921f9132b

Observation 4fada578-696f-4cd4-88b8-77005655e502 · inbound

CompCap: Improving Multimodal Large Language Models with Composite Captions cites this paper.

CompCap: Improving Multimodal Large Language Models with Composite Captions VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models

Reference 2023

Resolution
malformed identifier
no resolver link, observed 2026-08-11T20:52:40.367394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:52:40.367394Z digest=sha256:b9c81370c3b91b6d263cc1654ee3aeb6f47d171354dbfc8a59169178adfb8255

Observation a8af7fbc-f283-488f-937e-0cb0f76f67e7 · inbound

POINTS1.5: Building a Vision-Language Model towards Real World Applications cites this paper.

POINTS1.5: Building a Vision-Language Model towards Real World Applications VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T17:53:35.627431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:53:35.627431Z digest=sha256:09b6fac810d6b9f76bd4b6ecba36059d6a446e40fd1a9ff3a06915400214d983

Observation 302ada15-5c9f-4496-a6b7-5971401e8e94 · inbound

Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage cites this paper.

Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T11:26:27.644916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:26:27.644916Z digest=sha256:5b98cde737075a79ea91f8e030f725a58860fa82ee7f27faf3b7efecd90257ce

Observation 61be594d-ee87-474d-8a2f-6031032e4763 · inbound

MVTamperBench: Evaluating Robustness of Vision-Language Models cites this paper.

MVTamperBench: Evaluating Robustness of Vision-Language Models VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T23:55:17.550346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:55:17.550346Z digest=sha256:a8c5efefc92a5c8648171535018fafb67103d8092f4d5e89b1749d1d74cbd504

Observation 2638719a-6663-4c36-8752-37b15a7b659b · inbound

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training cites this paper.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.762903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.762903Z digest=sha256:3e824ea20b6078b3826c5023881a02d81ae10e7fb380f3acec1c10985c479af2

Observation d21366e8-78ce-4120-92db-4a02f47b2928 · inbound

Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models cites this paper.

Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-07T02:19:36.652705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T14:19:34.622854Z digest=sha256:77e7a2b4adba8a8f1797c6db7be5b518383efb77712c52f345347c5f6c670e90

Observation 37d8ce2c-fe48-4922-b4cb-40e1f2b309c9 · inbound

Argus Inspection: Do Multimodal Large Language Models Possess the Eye of Panoptes? cites this paper.

Argus Inspection: Do Multimodal Large Language Models Possess the Eye of Panoptes? VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T11:18:56.940137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:18:56.940137Z digest=sha256:7eb45e48ff1d4ea4f35233cd00349b4d6fa43d6a0e8100466964d4ee75b779b6

Observation ac4ca1f2-65ff-41c0-a482-281972f3c65e · inbound

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models cites this paper.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models

Reference 132

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:08.783284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:08.783284Z digest=sha256:72e374735a7db46a03a4601d34b2cece03b8ea0a54c4b1f512e970f0fd9683c2

Observation 2bc60fe3-88be-415e-803c-f886f360bf8e · inbound

In-context Learning of Vision Language Models for Detection of Physical and Digital Attacks against Face Recognition Systems cites this paper.

In-context Learning of Vision Language Models for Detection of Physical and Digital Attacks against Face Recognition Systems VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-06T15:43:16.074287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:43:16.074287Z digest=sha256:6194f5cda065154c14fe54bd90d4df2b2ac4d3c4bf7c259ece8a7877dc9a61dd

Observation 42bc6ae6-ce22-4d6e-9970-e4b71f5ef03c · inbound

Improving Large Vision and Language Models by Learning from a Panel of Peers cites this paper.

Improving Large Vision and Language Models by Learning from a Panel of Peers VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.332321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.332321Z digest=sha256:4a24f763ce512e98e6ab6775289b80c0f750f99112cbeec679050d5218bd59ea

Observation 2483dcad-774a-4fce-8d5c-ef8546d96383 · inbound

LaV-CoT: Language-Aware Visual CoT with Multi-Aspect Reward Optimization for Real-World Multilingual VQA cites this paper.

LaV-CoT: Language-Aware Visual CoT with Multi-Aspect Reward Optimization for Real-World Multilingual VQA VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-07T02:19:36.652705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T17:53:08.136677Z digest=sha256:590e5e79604226c367d05fff44e3cff2476e17e60997e8197849eb384e13e243

Observation 51089e0d-fd39-4343-8c03-d24a391be2b8 · inbound

ChartVerse: Scaling Chart Reasoning via Reliable Programmatic Synthesis from Scratch cites this paper.

ChartVerse: Scaling Chart Reasoning via Reliable Programmatic Synthesis from Scratch VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-07T02:19:36.652705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T13:12:01.889341Z digest=sha256:33b09d3829c601233fb49d955026230d1373ca90dbe8f0bc8c7efaa1dff3d1d0

Observation fce6002d-6acf-4e5a-8652-57a1bf04b315 · inbound

S2H-DPO: Hardness-Aware Preference Optimization for Vision-Language Models cites this paper.

S2H-DPO: Hardness-Aware Preference Optimization for Vision-Language Models VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models

Reference 82

Resolution
metadata mismatch
arxiv_id, observed 2026-07-07T02:19:36.652705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-10T05:38:01.208136Z digest=sha256:57be5d3efe4beb88336ab344a6b16bef9a6a6229a011cf05ac8bb7c906237c26

Observation 7b4a8009-1c28-46da-85b4-e173e73760d8 · inbound

GeoSym127K: Scalable Symbolically-verifiable Synthesis for Multimodal Geometric Reasoning cites this paper.

GeoSym127K: Scalable Symbolically-verifiable Synthesis for Multimodal Geometric Reasoning VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-07T02:19:36.652705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T22:42:23.066277Z digest=sha256:f1d72d78c5002f5376ec0e25c7a4a716f583d9e230a4f4be83ecf267a64dcbf7

Observation 9fb821af-4b1f-4227-9b44-71bb0591aa1a · inbound

Beyond Mode Collapse: Distribution Matching for Diverse Reasoning cites this paper.

Beyond Mode Collapse: Distribution Matching for Diverse Reasoning VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-07T02:19:36.652705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T05:30:37.685873Z digest=sha256:5c1d3b311c53c0d87bb2144a7527411ea383eefc6741bc930e5d5282dd42cc0f

Observation 14906927-c213-4dfe-a9f6-9bdad2c30e1f · inbound

DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack cites this paper.

DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-07T02:19:36.652705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T01:28:59.297634Z digest=sha256:c780bf6e7a863fb727d8838d9e8271f760000e7c0a1716d22a7715191375b9e6

Observation 0e29ed91-b731-4575-acf7-6014584ac405 · inbound

LLM-Based Examination of Eligibility Criteria from Securities Prospectuses at the German Central Bank cites this paper.

LLM-Based Examination of Eligibility Criteria from Securities Prospectuses at the German Central Bank VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-07-07T02:19:36.652705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-26T03:56:28.271760Z digest=sha256:3ae4bee941771c3446dcacc2f307dfba301cdcc910251f7315b15ba97e0d9ad1

Observation 2caa173d-f753-475c-a264-790ef9896e3c · inbound

Causal Connections: Leveraging Multilingual Fine-Tuning for Financial QA@FinCausal 2026 cites this paper.

Causal Connections: Leveraging Multilingual Fine-Tuning for Financial QA@FinCausal 2026 VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-07-07T02:19:36.652705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-29T02:17:26.872854Z digest=sha256:c24f3dd419767430e605a0904c8e5bd6fd1dc40d1bb649e5451422f1d60a2e8a

Observation 76c71cff-bfa0-4cdb-bdd5-16f4e584dbb8 · inbound

MotionAtlas: Detailed Region Captioning for Motion-Centric Videos cites this paper.

MotionAtlas: Detailed Region Captioning for Motion-Centric Videos VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-07T02:19:36.652705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T07:21:46.783970Z digest=sha256:0be2d512670e82b36122194cbcc1a2b35a916e8137b27575b38f283d1e81c784

Observation ce412b0f-15e9-4e5c-b6d0-536b2bf5d278 · inbound

Nemotron-Labs-Diffusion: A Tri-Mode Language Model Unifying Autoregressive, Diffusion, and Self-Speculation Decoding cites this paper.

Nemotron-Labs-Diffusion: A Tri-Mode Language Model Unifying Autoregressive, Diffusion, and Self-Speculation Decoding VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-07-11T03:07:51.040962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-11T03:04:12.500342Z digest=sha256:018db3b26045a037787b96e0323471391793d179c946b098727f6803f5bc68a9

Observation 15cb102d-67ec-4f8c-9f8c-69894e76adbb · inbound

SLAPBench: Benchmarking Multimodal Large Language Models for Four-Finger SLAP Fingerprint Verification cites this paper.

SLAPBench: Benchmarking Multimodal Large Language Models for Four-Finger SLAP Fingerprint Verification VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T23:09:19.992647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:09:19.992647Z digest=sha256:1ae75c8cf279ae2f9f9585039abe9833220bf18620cdfcd8e34878934ac1ccf9

Observation 75b98791-32eb-4264-a0f5-e0a1d862d3ca · inbound

Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs cites this paper.

Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T00:13:47.883840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:13:47.883840Z digest=sha256:a8390b379cdd3a3019f402a36a09a6f1159f694b99a596c856a9aa714d64e578