Pith. sign in

Paper Citation Record · LEDGER

MM-Vet v2: A Challenging Benchmark to Evaluate Large Multimodal Models for Integrated Capabilities

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 19 inbound Pith citation observations for arXiv:2408.00765.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2408.00765 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:47:21.254143Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T11:37:03.174107Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0daf9697-f1ba-4749-a4da-a5d8c05b3a84 · inbound

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling cites this paper.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling MM-Vet v2: A Challenging Benchmark to Evaluate Large Multimodal Models for Integrated Capabilities

Reference 285

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:23:58.269046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:70d1c0a7d6cba305fb8a6437198b4032a2c82b2583330ac353c58987b69f4d79

Observation 2911ac93-87c4-4598-af85-422fa575980a · inbound

Visual question answering: from early developments to recent advances -- a survey cites this paper.

Visual question answering: from early developments to recent advances -- a survey MM-Vet v2: A Challenging Benchmark to Evaluate Large Multimodal Models for Integrated Capabilities

Reference 194

Resolution
unresolved
no resolver link, observed 2026-08-10T21:46:29.005825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:46:29.005825Z digest=sha256:05eb8eb9371bf01ccaaf335401931477a66e069c98ae2c0b60e08f1f04cf6966

Observation 113a5431-9327-4409-8ebf-729f311fce2e · inbound

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models cites this paper.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models MM-Vet v2: A Challenging Benchmark to Evaluate Large Multimodal Models for Integrated Capabilities

Reference 140

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:41:08.159522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:57e138ba68378205139343bc22be413e1255189bf418b1a079e615726af83fae

Observation e2d2eab0-142b-433a-bce1-7e25feec12ab · inbound

Generative RLHF-V: Learning Principles from Multi-modal Human Preference cites this paper.

Generative RLHF-V: Learning Principles from Multi-modal Human Preference MM-Vet v2: A Challenging Benchmark to Evaluate Large Multimodal Models for Integrated Capabilities

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:56.100138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:56.100138Z digest=sha256:80628d883cd37a396f3ce35aecdf71276f06e70e817b0cbe6f3f8719ba34fb19

Observation bc2251a8-4877-40a5-b6ef-654dd56118e6 · inbound

Point-RFT: Improving Multimodal Reasoning with Visually Grounded Reinforcement Finetuning cites this paper.

Point-RFT: Improving Multimodal Reasoning with Visually Grounded Reinforcement Finetuning MM-Vet v2: A Challenging Benchmark to Evaluate Large Multimodal Models for Integrated Capabilities

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:12:06.428682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:12:06.428682Z digest=sha256:5feb8fce8639c92d6ff409aadf328a111c803b8f684725efeebe0737e88af20f

Observation bcbbe6df-eef7-4337-be5b-744b37dbf1b9 · inbound

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning cites this paper.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning MM-Vet v2: A Challenging Benchmark to Evaluate Large Multimodal Models for Integrated Capabilities

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:47.364932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:47.364932Z digest=sha256:ff3a9103d1ff4602310e8848b3fb7e3037d985f6c6ce49581b231a554fa69dc6

Observation 5c2d2d6d-1592-4354-9563-59a96f206908 · inbound

FlagEvalMM: A Flexible Framework for Comprehensive Multimodal Model Evaluation cites this paper.

FlagEvalMM: A Flexible Framework for Comprehensive Multimodal Model Evaluation MM-Vet v2: A Challenging Benchmark to Evaluate Large Multimodal Models for Integrated Capabilities

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T05:14:44.700185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:14:44.700185Z digest=sha256:bfdec72b7fe48023f69fc420188f462fb6998fdd4148c104ba148201457c52d1

Observation 1764ca9d-63de-47ee-a84e-1487f08812a7 · inbound

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs cites this paper.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs MM-Vet v2: A Challenging Benchmark to Evaluate Large Multimodal Models for Integrated Capabilities

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:12.782437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:12.782437Z digest=sha256:824590e5373ca0e967124ae40a5b22287042a60a5153b02af91a221b65523660

Observation ca1d0d46-b4d3-48d6-ac3b-4260a4c5b49b · inbound

GenRecal: Generation after Recalibration from Large to Small Vision-Language Models cites this paper.

GenRecal: Generation after Recalibration from Large to Small Vision-Language Models MM-Vet v2: A Challenging Benchmark to Evaluate Large Multimodal Models for Integrated Capabilities

Reference 116

Resolution
unresolved
no resolver link, observed 2026-08-06T23:57:28.057021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:57:28.057021Z digest=sha256:3a768ef216bf227b18962c3e2d4aa606eba6bdfa6e85ea3d6fc7ddb880ed5323

Observation 518c8d90-2136-40d2-9882-b7fbcb30a6cc · inbound

Towards Interpretable Renal Health Decline Forecasting via Multi-LMM Collaborative Reasoning Framework cites this paper.

Towards Interpretable Renal Health Decline Forecasting via Multi-LMM Collaborative Reasoning Framework MM-Vet v2: A Challenging Benchmark to Evaluate Large Multimodal Models for Integrated Capabilities

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T11:44:11.416912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:44:11.416912Z digest=sha256:365258282fbbd3e4885839bc8fe8ab6e31869c39445435277ff9c2b10fcdf89d

Observation 8ea29b2a-bb5a-48db-890c-1e3531eb80c0 · inbound

The First Differentiable Transfer-Based Algorithm for Discrete MicroLED Repair cites this paper.

The First Differentiable Transfer-Based Algorithm for Discrete MicroLED Repair MM-Vet v2: A Challenging Benchmark to Evaluate Large Multimodal Models for Integrated Capabilities

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T22:21:02.854526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:21:02.854526Z digest=sha256:a7cab6efc1c44f87561b5044f04d62c28dfc9f7f5f7cd0279730082510d2019e

Observation 25e01ba0-e45f-4211-bc4b-8fefb87b8b50 · inbound

Towards Better Dental AI: A Multimodal Benchmark and Instruction Dataset for Panoramic X-ray Analysis cites this paper.

Towards Better Dental AI: A Multimodal Benchmark and Instruction Dataset for Panoramic X-ray Analysis MM-Vet v2: A Challenging Benchmark to Evaluate Large Multimodal Models for Integrated Capabilities

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-04T19:27:40.862949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:27:40.862949Z digest=sha256:e0d2e7e1020b56365f5cc288886a2aa04546b700f6cce63d24355257c89e3f38

Observation 10f1514b-204d-408f-af04-91d59147c537 · inbound

VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes cites this paper.

VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes MM-Vet v2: A Challenging Benchmark to Evaluate Large Multimodal Models for Integrated Capabilities

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T15:47:21.254143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:47:21.254143Z digest=sha256:8f8dbba0c0259077b76b1019e12ed53c9f8f4ca4116d0616704e4dfe4a5fd68d

Observation c9fbb055-9dac-4648-8cc1-e1cba22082b6 · inbound

LEC: Linear Expectation Constraints for Selection-Conditioned Risk Control in Selective Prediction and Routing Systems cites this paper.

LEC: Linear Expectation Constraints for Selection-Conditioned Risk Control in Selective Prediction and Routing Systems MM-Vet v2: A Challenging Benchmark to Evaluate Large Multimodal Models for Integrated Capabilities

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T19:18:19.182308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:18:19.182308Z digest=sha256:261bedb74a371b49715ca5f9a6efa327b76fabe8c2693a873dc3603e0f385aa1

Observation f3f8a07e-7f06-4bcd-b9a3-5dc82e882719 · inbound

Dictionary-Aligned Concept Control for Safeguarding Multimodal LLMs cites this paper.

Dictionary-Aligned Concept Control for Safeguarding Multimodal LLMs MM-Vet v2: A Challenging Benchmark to Evaluate Large Multimodal Models for Integrated Capabilities

Reference 121

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:35:57.379952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T18:04:05.157103Z digest=sha256:e8062ab986a1d22a881092297562022fbb8650477841c33d7ed98561a9a014f6

Observation bdcc2ca5-1010-447c-9d10-9b0773729c07 · inbound

Not Every Rubric Teaches Equally: Policy-Aware Rubric Rewards for RLVR cites this paper.

Not Every Rubric Teaches Equally: Policy-Aware Rubric Rewards for RLVR MM-Vet v2: A Challenging Benchmark to Evaluate Large Multimodal Models for Integrated Capabilities

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:03:03.481345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T05:02:35.271960Z digest=sha256:5c3f39dade339ab6200f85309798a460ae63a22156402b06b530b665e5dbe4f9

Observation 4f5d4185-a4ba-4e09-9bec-ba8339e0200e · inbound

Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients cites this paper.

Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients MM-Vet v2: A Challenging Benchmark to Evaluate Large Multimodal Models for Integrated Capabilities

Reference 139

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:48:56.003766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T01:08:52.981296Z digest=sha256:17e5382d3ddbc681c4f5d84a112baf1a7de1d1087c92394d743c2d585760674e

Observation b05bc6b7-2b9d-4a53-a7f8-3f4109dd9083 · inbound

Dive Into the Implicit Biases of Low-rank Vision-language Alignment cites this paper.

Dive Into the Implicit Biases of Low-rank Vision-language Alignment MM-Vet v2: A Challenging Benchmark to Evaluate Large Multimodal Models for Integrated Capabilities

Reference 53

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T11:37:03.175748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-10T11:34:56.843122Z digest=sha256:aacb00b9e966081ff480e3fbd853abf76e4b2735eb6a43631c3183f541ba164f

Observation 399d2775-a449-43cb-9752-8b1a32528212 · inbound

Knowing the Self, Understanding the World: A Dual-Cognition Benchmark for UAV Spatio-temporal Reasoning with MLLMs cites this paper.

Knowing the Self, Understanding the World: A Dual-Cognition Benchmark for UAV Spatio-temporal Reasoning with MLLMs MM-Vet v2: A Challenging Benchmark to Evaluate Large Multimodal Models for Integrated Capabilities

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-01T21:08:22.268714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:08:22.268714Z digest=sha256:387dd4182cfcf424ead4c2a4d0a10a35ba3709fa4b77533d1583c850900ef6e3