Pith. sign in

Paper Citation Record · LEDGER

MANBench: Is Your Multimodal Model Smarter than Human?

As of 7 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 0 inbound Pith citation observations for arXiv:2506.11080.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.11080 v1

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:04:36.266807Z

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

36 of 36 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved36
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9cbbc84f-f1cd-4c8e-94de-025e13f9b210 · outbound

This paper cites MathQA: Towards Interpretable Math Word Problem Solving with Operation-Based Formalisms.

MANBench: Is Your Multimodal Model Smarter than Human? MathQA: Towards Interpretable Math Word Problem Solving with Operation-Based Formalisms

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:33.019433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:33.019433Z digest=sha256:3b016cc39eedc8acb3ccabb4f012546b424302d56713a93c9cd10ace5f5a1a37

Observation 6fc41757-0bb0-4ca5-9a50-5ec21542d75b · outbound

This paper cites an unresolved cited work.

MANBench: Is Your Multimodal Model Smarter than Human? Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:33.100287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:33.100287Z digest=sha256:4a9264f61831edad47b5d837d31859604e268bc45ce3c4359262efbee3c16c38

Observation c4c0966f-75f9-43c1-94ef-c43269d723e0 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

MANBench: Is Your Multimodal Model Smarter than Human? Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:33.203398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:33.203398Z digest=sha256:edf73d873b2d4f26e297f1abf3090de5b10a6183b1463c6ec4ba0f50450a3cf6

Observation faf179a8-d2c9-4552-931e-acdd0bd8466a · outbound

This paper cites GMAI-MMBench: A Comprehensive Multimodal Evaluation Benchmark Towards General Medical AI.

MANBench: Is Your Multimodal Model Smarter than Human? GMAI-MMBench: A Comprehensive Multimodal Evaluation Benchmark Towards General Medical AI

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:33.269754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:33.269754Z digest=sha256:b8efcda70c2036746f70456bba28f1f913d7872167af209456c70d5bd6390484

Observation e79d540c-5bd5-4e1e-b525-ad946d81fda6 · outbound

This paper cites an unresolved cited work.

MANBench: Is Your Multimodal Model Smarter than Human? Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:04:38.122814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:04:33.342134Z digest=sha256:ca8e65e9925ac2cd4502bebc01ecfbc29347213c1ef17fbb6cc8771332fa7311

Observation 46b97ec5-2d1c-424f-aaac-7a0c7fe1c17b · outbound

This paper cites an unresolved cited work.

MANBench: Is Your Multimodal Model Smarter than Human? Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:04:37.922036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:04:33.413796Z digest=sha256:ba58fdcefb26ef1a3b2b90ba2def4b59726e8c33763ad407ee64a57933241c06

Observation 45ff32d9-7ea2-4adc-885d-acb2aed40042 · outbound

This paper cites The Llama 3 Herd of Models.

MANBench: Is Your Multimodal Model Smarter than Human? The Llama 3 Herd of Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:33.473651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:33.473651Z digest=sha256:9d4ea684b21267a05a7eb99e5c7d04e4e6f9b37ea38b04befd331effe44f18ba

Observation 6840f7bc-b91c-4ab6-8362-746e39cd4238 · outbound

This paper cites BLINK: Multimodal Large Language Models Can See but Not Perceive.

MANBench: Is Your Multimodal Model Smarter than Human? BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:33.576525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:33.576525Z digest=sha256:90c1aec66738c5577ec13ef2094d81104a9efb15c19482019b48c6b0c343d95e

Observation a7295d2f-a5da-461c-bb3a-9925fd622e13 · outbound

This paper cites OpenAI o1 System Card.

MANBench: Is Your Multimodal Model Smarter than Human? OpenAI o1 System Card

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:33.681836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:33.681836Z digest=sha256:ae68de0952b4983eb92e0222d1598eac4c720ed0f76b4a215e2b256c26c0172d

Observation 79938242-1aaa-48e2-8de3-877f771a804b · outbound

This paper cites DeepSeek-V3 Technical Report.

MANBench: Is Your Multimodal Model Smarter than Human? DeepSeek-V3 Technical Report

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:33.753861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:33.753861Z digest=sha256:f6e0e30cf05dd8f8d4ae2b3c0e5739bd0986eb3564d176060c5fa8a7217eff7b

Observation a1dfed1f-bfa6-46e5-a649-b08ab28235d8 · outbound

This paper cites an unresolved cited work.

MANBench: Is Your Multimodal Model Smarter than Human? Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:33.852870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:33.852870Z digest=sha256:33f535b754052ea63ae88b3e3fe56bf9d37e27ce95d690ab00f71418989b9eee

Observation f3cceeb6-bdc8-40de-8901-36b3e3619c04 · outbound

This paper cites an unresolved cited work.

MANBench: Is Your Multimodal Model Smarter than Human? Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:04:37.770114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:04:33.952899Z digest=sha256:973be546dcef8240c4536eac350fa9851a54cb19bf5ff192ee1f97cf5a0f26b2

Observation ced02137-89da-4243-82e2-44d681229640 · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

MANBench: Is Your Multimodal Model Smarter than Human? MMBench: Is Your Multi-modal Model an All-around Player?

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:34.025517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:34.025517Z digest=sha256:7099162d11217fffd540b12e8e5b2d84057f2159c625afbb0bad58546c4470ea

Observation c2780544-60bf-41f8-ae47-dc9314b71fe6 · outbound

This paper cites an unresolved cited work.

MANBench: Is Your Multimodal Model Smarter than Human? Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:34.099423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:34.099423Z digest=sha256:2f3f3014ae1edbb8d567f3ea29fc80a9ce345df8c638e28d73f31773768a35d0

Observation d94c870a-ba9f-4acc-8d8c-bd44a87f2ac6 · outbound

This paper cites an unresolved cited work.

MANBench: Is Your Multimodal Model Smarter than Human? Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:34.194730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:34.194730Z digest=sha256:f19185a7d8651bf57ffdd1e4f6b2b11357bd6e9788604fcf77de9c56a14987c3

Observation de2927ee-074a-4a77-9fda-d5f316e2214b · outbound

This paper cites MMIU: Multimodal Multi-image Understanding for Evaluating Large Vision-Language Models.

MANBench: Is Your Multimodal Model Smarter than Human? MMIU: Multimodal Multi-image Understanding for Evaluating Large Vision-Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:34.275314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:34.275314Z digest=sha256:b54c5fb2aeaf72b1fa14082ca3fa7e72546e85282a0b66577e22c4fdea3657d1

Observation 3b4515d8-05fb-458c-aa7c-fd9cf9c1bd68 · outbound

This paper cites an unresolved cited work.

MANBench: Is Your Multimodal Model Smarter than Human? Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:34.353892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:34.353892Z digest=sha256:313005e3e5d3c56576ebe9b7d708f1e56c683ea9879bfec08a5863d01b9bbd65

Observation b38614fb-0d8e-44cd-9e28-7398a53e3386 · outbound

This paper cites an unresolved cited work.

MANBench: Is Your Multimodal Model Smarter than Human? Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:04:37.549855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:04:34.481908Z digest=sha256:4e15c822bbbef3722c29c735c8d5cb742867ba090291527269319cc75fe74a6b

Observation 6a764844-966c-4a55-adf4-7d89ad29ec5f · outbound

This paper cites an unresolved cited work.

MANBench: Is Your Multimodal Model Smarter than Human? Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:34.555811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:34.555811Z digest=sha256:62955327590e69abc07acb6214c39c47af925b8d95eb51d06fb8975d45dd2c1c

Observation 84ea90b2-4610-4482-9b1e-57435bfd7e6f · outbound

This paper cites an unresolved cited work.

MANBench: Is Your Multimodal Model Smarter than Human? Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:34.627505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:34.627505Z digest=sha256:486bb93cb114da1380e434fb6d1ab814d0ef54a1c38ca55e4b2c80e2e95c7a6a

Observation ce03bda5-2ee7-4d7d-b199-22debc836638 · outbound

This paper cites ImageNet Large Scale Visual Recognition Challenge.

MANBench: Is Your Multimodal Model Smarter than Human? ImageNet Large Scale Visual Recognition Challenge

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:34.736985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:34.736985Z digest=sha256:9f75e3cfda73b3334735c24c861ce941c9c74ee4cd4f3ff18c87083ba9f9b8f1

Observation aa645b74-9800-47bf-8a4a-9f9697541a3d · outbound

This paper cites an unresolved cited work.

MANBench: Is Your Multimodal Model Smarter than Human? Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:04:37.398739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:04:34.811098Z digest=sha256:b94ca230819c822ed0888c5fcdba51939b2c90e94da745c068576dc4a0b46d42

Observation 10dcd35f-3ad8-4473-9d3b-2533e9f33ab9 · outbound

This paper cites an unresolved cited work.

MANBench: Is Your Multimodal Model Smarter than Human? Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:04:37.245133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:04:34.906468Z digest=sha256:ffbea3de0675bf5c853504c8824f03ee1f5d3be84f1479a7c859f752d42f665f

Observation 91cdb192-d7c1-4429-8c88-834a38467284 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

MANBench: Is Your Multimodal Model Smarter than Human? Gemini: A Family of Highly Capable Multimodal Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:34.957896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:34.957896Z digest=sha256:bf282655452f2f03a51964ed6701fddbe17f834f1af65081493f6ab2e3116d48

Observation 178da340-62c3-46c5-b605-1794e37e26ea · outbound

This paper cites an unresolved cited work.

MANBench: Is Your Multimodal Model Smarter than Human? Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:04:37.077215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:04:35.053070Z digest=sha256:9da9f391808227290e9ea66d4b1866b1d40050010aad08427ff534dc791639e4

Observation 1ff12e9b-7eac-4590-aa1a-a924fec77b41 · outbound

This paper cites an unresolved cited work.

MANBench: Is Your Multimodal Model Smarter than Human? Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:35.122944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:35.122944Z digest=sha256:f78ab4cdcf87158b0a9513a10b04dff7c13b0df99d2aa15e475a1e25839be7db

Observation b143a180-e90c-4d7c-8a66-0230a099cedf · outbound

This paper cites MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding.

MANBench: Is Your Multimodal Model Smarter than Human? MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:35.321699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:35.321699Z digest=sha256:89f8b1fc2084d1cd12a9adaa19586243f17b13a6208ca73533273e8570fb71aa

Observation cbe520ed-f158-447a-9e94-ed8fce53cbd6 · outbound

This paper cites Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset.

MANBench: Is Your Multimodal Model Smarter than Human? Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:35.401166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:35.401166Z digest=sha256:72c487b296f7b0d72ddcd5d5bdaf7b8839c75d937e61425c437c8dc04931c7c5

Observation 0a02c4e2-b065-46fe-8300-7b4fba8ddcc1 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

MANBench: Is Your Multimodal Model Smarter than Human? Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:35.585470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:35.585470Z digest=sha256:ba435470faed3cb9301e068b9f6473772d3677bb6989df5ed9cdca209133ed12

Observation 9697a96f-5196-47ad-9dad-7c8174cf9069 · outbound

This paper cites Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization.

MANBench: Is Your Multimodal Model Smarter than Human? Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:35.660401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:35.660401Z digest=sha256:fd481262ed695fa7d81e78bfc0000798f465957b0aacbf36a8a64d9e1cb452da

Observation 6f69844b-b07b-468b-9e95-fa2590d3990c · outbound

This paper cites LiveBench: A Challenging, Contamination-Limited LLM Benchmark.

MANBench: Is Your Multimodal Model Smarter than Human? LiveBench: A Challenging, Contamination-Limited LLM Benchmark

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:35.770013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:35.770013Z digest=sha256:6279bab34c5056f8f79366fbc836c19e01a44ad69b5fa92d64ec77ff6892d641

Observation 85842938-03b8-4905-9646-413aee53cda0 · outbound

This paper cites DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding.

MANBench: Is Your Multimodal Model Smarter than Human? DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:35.843690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:35.843690Z digest=sha256:6c608359327a1f961757e792c4494f299bdcfdcf4d2d6005c55496f0ea6bc2b3

Observation bc046f9a-de8e-4942-814f-7be1158cc679 · outbound

This paper cites an unresolved cited work.

MANBench: Is Your Multimodal Model Smarter than Human? Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:35.920353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:35.920353Z digest=sha256:ca97cf43ac6a4e734194bf44922fa23944d470fbed257e795de93e2179997f0a

Observation 9a2fe13c-2714-4032-b3f4-6bdb30c1446e · outbound

This paper cites A Careful Examination of Large Language Model Performance on Grade School Arithmetic.

MANBench: Is Your Multimodal Model Smarter than Human? A Careful Examination of Large Language Model Performance on Grade School Arithmetic

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:36.030597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:36.030597Z digest=sha256:9fdaaea96a9eca86fbbd7c9b13c9656b09eb7eab499763d0d23f105bdcf52021

Observation 57dceeea-435a-43b1-87b3-602d38865fa7 · outbound

This paper cites an unresolved cited work.

MANBench: Is Your Multimodal Model Smarter than Human? Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:04:36.917406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:04:36.122616Z digest=sha256:1495723913cf0c31c0e2055742bcf92929e18fa8133bca5200b7f737903ac8c3

Observation 94a18d1e-2c73-4c44-8f0f-5ff4448a9ece · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

MANBench: Is Your Multimodal Model Smarter than Human? MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:36.266807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:36.266807Z digest=sha256:aaf2f36114adf608883b5530505dd71dcd1f02b7fab636398cd8eb15e84bedd7

Pith citing papers

No inbound Pith citation observations are available.