Pith. sign in

Paper Citation Record · LEDGER

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans

As of 22 August 2026, this Paper Citation Record lists 100 of 121 outbound references and 0 inbound Pith citation observations for arXiv:2505.11141.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.11141 v2

Coverage vector

measured 100 of 121 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:00:59.555618Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 121 outbound references displayed

  • verified exact0
  • verified fuzzy15
  • unresolved85
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9b01b0bd-e42f-4612-bb51-3ab8bf8db8a2 · outbound

This paper cites Reasoning, problem solving, and intelligence.Handbook of human intelligence, pages 225–307, 1982.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Reasoning, problem solving, and intelligence.Handbook of human intelligence, pages 225–307, 1982

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.143833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.143833Z digest=sha256:9dce22bb7a99d4e4a9090dcaa13aa277767f509a384e495f30a3d1587410bee8

Observation 6fbd7297-af31-4bd5-ae5b-677c292a6edf · outbound

This paper cites Intelligence and reasoning.The Cambridge handbook of intelligence, pages 419–441, 2011.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Intelligence and reasoning.The Cambridge handbook of intelligence, pages 419–441, 2011

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.148090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.148090Z digest=sha256:c2f848c0d1a09083575b77529b5ec233adad6e89cd2adc5ebff5dd119436ca2a

Observation d465d40c-c445-45cb-83de-38e55385ded4 · outbound

This paper cites Levels of agi: Operationalizing progress on the path to agi.arXiv preprint arXiv:2311.02462, 2023.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Levels of agi: Operationalizing progress on the path to agi.arXiv preprint arXiv:2311.02462, 2023

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.151841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.151841Z digest=sha256:5c835fad9384043ac4bb7787986407351a8133cc7f78d99c637f711ace72eea6

Observation 942b823d-7c6d-445a-ad55-f40fbab378a9 · outbound

This paper cites LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.155573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.155573Z digest=sha256:156c04707faa5589a16024c45a82cb0d6a45fa647f3c6455459581890c3cfb61

Observation 38a7a480-6e43-4b31-a541-0d7fa169fc2f · outbound

This paper cites Qwen2.5 Technical Report.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Qwen2.5 Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.159550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.159550Z digest=sha256:820b253eeb44d667168122042599336fab8505eda508b7f98fba81a32aae0233

Observation 2ccad829-7663-4913-8d64-402cf4fadae7 · outbound

This paper cites ChatGLM-Math: Improving Math Problem-Solving in Large Language Models with a Self-Critique Pipeline.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans ChatGLM-Math: Improving Math Problem-Solving in Large Language Models with a Self-Critique Pipeline

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.163223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.163223Z digest=sha256:447974bae3c0514491319f6a2ee6ed5c2ca622a8a0394f8177e9a2511ea3c326

Observation b072b084-d996-444d-87d9-724d6cd8fb77 · outbound

This paper cites MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.167402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.167402Z digest=sha256:c050545996365215b6dc436af2d0c42363dc1dff1013cce35be530ffe37c6302

Observation 8f164bf5-6f56-4ca9-8c72-3483d5e86bb7 · outbound

This paper cites LogicAsker: Evaluating and Improving the Logical Reasoning Ability of Large Language Models.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans LogicAsker: Evaluating and Improving the Logical Reasoning Ability of Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.171247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.171247Z digest=sha256:3ce967c19e38e26458789980089bedbef6edbca44e9690c29d17128ed0004f40

Observation 111fb5e2-45aa-442c-9c4c-26b696a16f56 · outbound

This paper cites Symbol-LLM: Towards Foundational Symbol-centric Interface For Large Language Models.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Symbol-LLM: Towards Foundational Symbol-centric Interface For Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.175386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.175386Z digest=sha256:4963217ccc1792b44de7affb986283958f3e99572930d9f98f1503d27b62fddb

Observation d79de149-d067-4563-bb0b-5c2652268e76 · outbound

This paper cites Language Models can be Logical Solvers.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Language Models can be Logical Solvers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.179401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.179401Z digest=sha256:9d684ccb99ad580481761b86a63b6fcb03fd9c32895a831d2991ebb451e9f1b6

Observation deeab4a9-3e0a-417e-9d34-c8aa23cceab1 · outbound

This paper cites LogiCoT: Logical Chain-of-Thought Instruction-Tuning.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans LogiCoT: Logical Chain-of-Thought Instruction-Tuning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.183818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.183818Z digest=sha256:d31123c1661dee9748a622ee13f2e25eabb7c8c8072cbe377f756196b79f0f70

Observation d27b262c-4f30-40c8-b17c-685928490a0a · outbound

This paper cites OpenCodeReasoning: Advancing Data Distillation for Competitive Coding.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans OpenCodeReasoning: Advancing Data Distillation for Competitive Coding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.187803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.187803Z digest=sha256:71c1c17135675230c11c92a4377a7529dbd4e7f7dd2b070bafe018445acf3f61

Observation 5cbd7a7d-d1a6-45fa-b735-850e25339e1e · outbound

This paper cites Self-planning code generation with large language models.ACM Transactions on Software Engineering and Methodology, 33(7):1–30, 2024.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Self-planning code generation with large language models.ACM Transactions on Software Engineering and Methodology, 33(7):1–30, 2024

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.192248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.192248Z digest=sha256:9b57eff3d8b8a6a98f1211395277ffa6cf20ed8d6c6df4a3d382e70457beb30b

Observation 590fdc54-8eee-402c-bac4-0ab10656dbf9 · outbound

This paper cites Structured chain-of-thought prompting for code generation.ACM Transactions on Software Engineering and Methodology, 34(2):1–23, 2025.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Structured chain-of-thought prompting for code generation.ACM Transactions on Software Engineering and Methodology, 34(2):1–23, 2025

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.196631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.196631Z digest=sha256:4cafaf3a5d46bad6d9c2502352d8a40da4c55096e0c51c843e8f65b2f7555909

Observation 74d08069-4dd6-4a02-81bd-180aaa5140b3 · outbound

This paper cites CodeCoT: Tackling Code Syntax Errors in CoT Reasoning for Code Generation.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans CodeCoT: Tackling Code Syntax Errors in CoT Reasoning for Code Generation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.200419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.200419Z digest=sha256:a581a097cd6975dfbb716051e9bbade6722d7ab975009e2669129935153424e3

Observation f2f304f6-4926-4a90-bb0f-3a28489982e6 · outbound

This paper cites OpenAI o1 System Card.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans OpenAI o1 System Card

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.204340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.204340Z digest=sha256:f2197a9337fe9917dc6564e5a5e66219d5045d69213045e574bee585ee7ee5e8

Observation 7aa369a8-f284-403c-be99-edb45601d507 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.208247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.208247Z digest=sha256:d09f398d5705a33d6de808dda876283abeee9d3df30475f3ba57ef4c1a6957e2

Observation 2b1c60bd-76c2-446f-9331-519a992666e7 · outbound

This paper cites Springer, 2007.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Springer, 2007

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.212160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.212160Z digest=sha256:881c0530b68e05cf67d3bfacb6c0188b29e193674e8b47e8efccb21b1420e32f

Observation aa0867be-5301-4ea6-9526-6a22d5d3ff55 · outbound

This paper cites Exploring the Reasoning Abilities of Multimodal Large Language Models (MLLMs): A Comprehensive Survey on Emerging Trends in Multimodal Reasoning.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Exploring the Reasoning Abilities of Multimodal Large Language Models (MLLMs): A Comprehensive Survey on Emerging Trends in Multimodal Reasoning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.216711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.216711Z digest=sha256:431219e6a3b98eebb66bc8fa1c605e3580aa5bc776f47ee64e5c19ba2baf9fa6

Observation 4b5d843b-68ae-4bba-8048-27346506a819 · outbound

This paper cites R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.220726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.220726Z digest=sha256:6dc88343985271019ecf9c8daafb21f3ce164afc4be75824219a24993683d4c7

Observation 126ca753-9ff0-4f06-8aab-cf7da700f9e2 · outbound

This paper cites R1-v: Reinforcing super generalization ability in vision-language models with less than $3.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans R1-v: Reinforcing super generalization ability in vision-language models with less than $3

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.224784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.224784Z digest=sha256:cd99d3bcdc4084b10fe3b6e36df7ca2c39b584ff11fd8e542116accf69d1d5b9

Observation 0c668467-9f5f-4a06-8aad-64b234a0a74d · outbound

This paper cites Visual-RFT: Visual Reinforcement Fine-Tuning.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Visual-RFT: Visual Reinforcement Fine-Tuning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.228526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.228526Z digest=sha256:4d5da47377e273bed112503151a11d7dffa53215b92354ee858930ea9ff281fe

Observation 946cd72b-23ce-41bc-8664-4c5aff50a98f · outbound

This paper cites VisualPRM: An Effective Process Reward Model for Multimodal Reasoning.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans VisualPRM: An Effective Process Reward Model for Multimodal Reasoning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.232625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.232625Z digest=sha256:1aa824b7accbe13f522cbe937c5658bca612f2e626dbd3306929d2fce332fbd6

Observation 45b70ba3-813b-405d-ab7d-7248588f551b · outbound

This paper cites OThink-MR1: Stimulating multimodal generalized reasoning capabilities via dynamic reinforcement learning.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans OThink-MR1: Stimulating multimodal generalized reasoning capabilities via dynamic reinforcement learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.236677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.236677Z digest=sha256:e90c38bcee72a009d32e9c42ac633af9f6a003441b7e28ff1373cc7e61573d71

Observation 87dcb263-bd00-413b-823f-3fb3bab311c9 · outbound

This paper cites Vlm-r1: A stable and generalizable r1-style large vision-language model.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Vlm-r1: A stable and generalizable r1-style large vision-language model

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.240685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.240685Z digest=sha256:fe259156c7d459af0d1331aa979150acd16b1fed1129151c9d2b09e9347bea28

Observation 6fcabbe8-6ade-4025-b8db-731faac2ffe0 · outbound

This paper cites Open-r1-video.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Open-r1-video

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.244061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.244061Z digest=sha256:301f277ef0829ecfc9269592a8adb745acf7daa06ae532b75e5b5cedf4d4e910

Observation 3e1fca7e-9e72-4ef5-bbf0-8dbbbf594ba0 · outbound

This paper cites VisualPuzzles: Decoupling Multimodal Reasoning Evaluation from Domain Knowledge.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans VisualPuzzles: Decoupling Multimodal Reasoning Evaluation from Domain Knowledge

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.247883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.247883Z digest=sha256:d319a6df66473eb50f78ce7ad7cad3595f88913ddef18cd33931c6cab7dd8c64

Observation 23aa7b98-9633-401e-b355-e834fa88e42e · outbound

This paper cites MM-IQ: Benchmarking Human-Like Abstraction and Reasoning in Multimodal Models.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans MM-IQ: Benchmarking Human-Like Abstraction and Reasoning in Multimodal Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.251643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.251643Z digest=sha256:044a861e897758017c68980c84c0b282dd7e559d57538e19c144a9c1d06ac266

Observation b82014cd-bf82-458d-87fb-1ca793ecf2d1 · outbound

This paper cites VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.255612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.255612Z digest=sha256:f2dc4ab1cf1df7af2a2b64285657ccc1b5f81547feab01f4ff554abde78027c4

Observation 96dc97c7-5308-4420-86ed-34a2111bd7c1 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.259756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.259756Z digest=sha256:888fc92afe3ee166ea03199821ad05917b3d0a6b901b80b47250b6cb847cbede

Observation 266301ef-630b-4fc3-bd6a-df4c2a398974 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.263536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.263536Z digest=sha256:b7c4140565f895e01f8d37d8d0bf18fb2f0defee4bef3a08fb8257cb2d47466c

Observation e3c30b03-3d84-4f58-87d4-78358e7ebfa7 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.Advances in neural information processing systems, 35:23716–23736, 2022.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Flamingo: a visual language model for few-shot learning.Advances in neural information processing systems, 35:23716–23736, 2022

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.267404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.267404Z digest=sha256:6980dfbe35fb0e4bd4c6501c6903f0e6aead319212a42db4f2af4fd1210bf76d

Observation 0da66503-5c32-4921-a30e-1d2a527a2a1e · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.271391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.271391Z digest=sha256:19152531f1a0c4af4875d6184451ba70c8a20cce167375ad61f74c1d4cec5368

Observation 5aa9be15-7c32-4625-8aee-827af55d7df4 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36, 2024.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Visual instruction tuning.Advances in neural information processing systems, 36, 2024

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.275344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.275344Z digest=sha256:f369e32f247abff6e875cefa67a811a5146c79c5e29175e63b706cfdd7a2f64a

Observation 06c8bf22-2dbc-4557-9cc3-559ab2af888d · outbound

This paper cites Minigpt-4: Enhancing vision- language understanding with advanced large language models.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Minigpt-4: Enhancing vision- language understanding with advanced large language models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.279006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.279006Z digest=sha256:fc646487134a47bd90116f4ad3bb97b97277189d24ee094d39d034588f40edfb

Observation 31cf828d-8819-4179-9fe6-098ae51117d0 · outbound

This paper cites Gpt-4o: A multimodal language model, 2024.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Gpt-4o: A multimodal language model, 2024

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.282932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.282932Z digest=sha256:5c226ec033d9454edb97beaa02cbdb50b89030b55a39a507b76d413752b115f4

Observation af6bb968-af5e-45ad-8ac9-70924b9cce31 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Gemini: A Family of Highly Capable Multimodal Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.286731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.286731Z digest=sha256:98cacbb06d7d31e10fb306cc2c0ffcd6b47ad98291afd5679b8b293ac11d2190

Observation b7141138-c70c-4db1-8660-cbebf4c6f3a9 · outbound

This paper cites Claude: A conversational ai assistant, 2024.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Claude: A conversational ai assistant, 2024

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.290609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.290609Z digest=sha256:77798abb258ee311bb3c307553b25ad854ac1c8c72ed95b6ee2ffa7f93ce71fe

Observation 2d0f7752-44c5-4f92-b603-c9c2f81f0e40 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.294831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.294831Z digest=sha256:bf91d3f55d6dd3bdacff3b2f71df8b970a967b76ac93ab3616749ba0474484ea

Observation 2694a045-83cc-4cd6-939d-ac4a826f4422 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.299292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.299292Z digest=sha256:69514a5c2bbab77da779902675421638d3076d6e777d9eb27a24b6123079a9fc

Observation f8274c36-01d8-4b3e-a817-b0f3cbc43459 · outbound

This paper cites Qwen2.5-VL Technical Report.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Qwen2.5-VL Technical Report

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.303953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.303953Z digest=sha256:f9d3313234c4888a366d147e798da95ff65a51c9d173bd0fbc56d443d78ed8f7

Observation 93733ea4-3ca0-455a-85c4-6f1abd24f5b3 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.307992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.307992Z digest=sha256:0c4ff62c15904b440e03d380cadcbc2cd17ceece0ef25949e779d918fb9b3b6c

Observation 2c1a9576-a1dc-4ac7-80c5-7252db5066c8 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.312073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.312073Z digest=sha256:53702ae034c21f11799a4eca5b9797dd5f03d38751bdcf3e05573bc26574d682

Observation 2d672641-e81d-4906-ac7d-46a8c4d7f37e · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.316681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.316681Z digest=sha256:99018c4b5ef600733aa5cf91081c75c7ca2c6d2ffb2a3a6f217de1893b609a1f

Observation fe0b1d75-14b5-4aae-af7a-f5ae797f4393 · outbound

This paper cites Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.321375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.321375Z digest=sha256:fceac6dd41ef81ebf7b1e1f6a495512f92b99f18a067197487adc5ab0e7eab3a

Observation da2870a5-ca2b-4d6d-9c79-30c6cb2ad0d3 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.326590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.326590Z digest=sha256:3415360113cb31136248cfdaa7cbdac66842a44faf358e50cc44375449211d93

Observation 2c4efc12-8a41-4da1-abbc-af479be850e7 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans LLaVA-OneVision: Easy Visual Task Transfer

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.331240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.331240Z digest=sha256:1223231c024beccff3807f698915ab74ea91542bf450d64ded0e8296d32cdd84

Observation 0f548ba1-bc33-40ad-b000-52b6df849d2a · outbound

This paper cites Pangea: A fully open multilingual multimodal LLM for 39 languages.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Pangea: A fully open multilingual multimodal LLM for 39 languages

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.335135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.335135Z digest=sha256:7802ce6273cc71be11db4d29b18952079f64574e434065d2941185dfba7049c6

Observation 51f20111-f6b5-49f0-9ace-c65bd5bc8055 · outbound

This paper cites Harnessing Webpage UIs for Text-Rich Visual Understanding.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Harnessing Webpage UIs for Text-Rich Visual Understanding

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.339748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.339748Z digest=sha256:806d2777bbb9be68c46b87114ebfb83bf1c125d59374cef4cc156de7158cc507

Observation ffe5cd88-b90e-43c6-9847-ac75d913a313 · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.348610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.348610Z digest=sha256:4f85189e5927cb36811ad9cd2278d6658b434e60492127f6c04628891ec8730e

Observation c3000e9c-444e-44fe-a609-934d48f0fe48 · outbound

This paper cites The Llama 3 Herd of Models.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans The Llama 3 Herd of Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.353067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.353067Z digest=sha256:3749396a430a807a69918253d79b2bd08d3dde1096c9262af344e22d16425ed8

Observation f486d850-d21c-4f7f-b595-9f2197a97a44 · outbound

This paper cites Math-llava: Bootstrapping mathematical reasoning for multimodal large language models.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Math-llava: Bootstrapping mathematical reasoning for multimodal large language models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.356822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.356822Z digest=sha256:8a438f586ceaaa1d857daafa0e28b74121db569597ca02aa6993679915fc093e

Observation be250d21-e842-49a3-8f98-6d9065a73682 · outbound

This paper cites MultiMath: Bridging Visual and Mathematical Reasoning for Large Language Models.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans MultiMath: Bridging Visual and Mathematical Reasoning for Large Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.361895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.361895Z digest=sha256:280c5f133e84efd68ec032e81b6d01b7b66c3478dd02bae1a6863eb5a2052a0c

Observation 183a1482-dbf6-4aba-99b6-5aeba392488f · outbound

This paper cites Med-flamingo: a multimodal medical few-shot learner.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Med-flamingo: a multimodal medical few-shot learner

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.365921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.365921Z digest=sha256:e088ee592b722b7bf77c641c292cda65ac21f3c18c036bec38880d22dc66bbce

Observation 0fade284-7987-467d-91d4-02c561ee2ba6 · outbound

This paper cites Med-moe: Mixture of domain-specific experts for lightweight medical vision-language models.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Med-moe: Mixture of domain-specific experts for lightweight medical vision-language models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.369418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.369418Z digest=sha256:a6ceee0d0204176d4cc9c901cf49036112ea2b898a23e3bc33908b25660b839c

Observation 027bb3a7-7c97-4254-b7bf-2601e24db763 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.Advances in neural information processing systems, 35:24824–24837, 2022.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Chain-of-thought prompting elicits reasoning in large language models.Advances in neural information processing systems, 35:24824–24837, 2022

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.373609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.373609Z digest=sha256:fbe63aae277a044da0058e06171260de6f9fb786a9a544cdbd61b30d8d53ac43

Observation 52b3d24c-7bcc-479b-a14d-a47501771163 · outbound

This paper cites DPO Meets PPO: Reinforced Token Optimization for RLHF.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.377307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.377307Z digest=sha256:1a277bf83cfe19a162d510ec843aa39347d5c41fc672ca66d86608d6d4f4fb31

Observation 9462d6c9-bf94-4a4f-b89c-11a99540f6a6 · outbound

This paper cites Star: Self-taught reasoner bootstrapping reasoning with reasoning.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Star: Self-taught reasoner bootstrapping reasoning with reasoning

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.381458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.381458Z digest=sha256:1c333d7e679a32cd6798148f0db1aecfc1c0dfcfc60513a875c557ccb8d19c9e

Observation a3b024cb-163f-4f28-b6c9-6fb4b0c3c7c4 · outbound

This paper cites Quiet-STaR: Language Models Can Teach Themselves to Think Before Speaking.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Quiet-STaR: Language Models Can Teach Themselves to Think Before Speaking

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.385419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.385419Z digest=sha256:854ada8e4c5d06141f3dc3d85c5fdc6abec459584a9a0d291cc3d822d9f6a1d7

Observation 3c5b7f1a-73be-4a5b-aaa6-ab32ee6562e6 · outbound

This paper cites Qvq: To see the world with wisdom, December 2024.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Qvq: To see the world with wisdom, December 2024

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.390247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.390247Z digest=sha256:f6ee37f65bd608b43b780ceab16a895460734e78c3fd632f2973fadd8a2f7c10

Observation b7214e0e-bc38-4712-b2b5-b9eb1e42e5e6 · outbound

This paper cites Ocrbench: on the hidden mystery of ocr in large multimodal models.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Ocrbench: on the hidden mystery of ocr in large multimodal models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.394676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.394676Z digest=sha256:2a4329dbedc6b9e01915865709e5bdd6aa365172f6a343d6f4661d82ada39f0c

Observation b353542f-2d19-417b-902e-718beac5dd0a · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.398663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.398663Z digest=sha256:2c319fcb08f750e185e6fd92a77b36ece9cd36a2ada50af68fc845f221c33b42

Observation f22818f8-29b5-4380-ae4d-e55d70ec3e73 · outbound

This paper cites Docvqa: A dataset for vqa on document images.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Docvqa: A dataset for vqa on document images

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.402696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.402696Z digest=sha256:d29eccbf14cbee6b1006c30b4c231e176d9de640d323af6220c77443e0a033f8

Observation 2fb5dd0d-43da-4279-8def-12dcf65cfdeb · outbound

This paper cites AgentBench: Evaluating LLMs as Agents.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans AgentBench: Evaluating LLMs as Agents

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.406856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.406856Z digest=sha256:3ba274029844d7962df51fe82f840335cf341c63b189374fe4bafb76b953cce9

Observation 9a358f92-c552-45e3-b3fc-e11d25d47c04 · outbound

This paper cites ToolEyes: Fine-Grained Evaluation for Tool Learning Capabilities of Large Language Models in Real-world Scenarios.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans ToolEyes: Fine-Grained Evaluation for Tool Learning Capabilities of Large Language Models in Real-world Scenarios

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.411738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.411738Z digest=sha256:1fbd5e3c34a504f1bf174e5994b94d4f51cf00575772a90d1b7a0bd3083e96e2

Observation f58b9c21-638b-419b-ac58-047211f86d55 · outbound

This paper cites Egoschema: A diagnostic benchmark for very long-form video language understanding.Advances in Neural Information Processing Systems, 36:46212–46244, 2023.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Egoschema: A diagnostic benchmark for very long-form video language understanding.Advances in Neural Information Processing Systems, 36:46212–46244, 2023

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.415729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.415729Z digest=sha256:5c0defc751a7290d6c1dbac8ab5f841ac4f09294b86cbaaf887ce7ab09e29776

Observation d96fa136-b120-4857-bad7-20002bed72bc · outbound

This paper cites Egothink: Evaluating first-person perspective thinking capability of vision-language models.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Egothink: Evaluating first-person perspective thinking capability of vision-language models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.419749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.419749Z digest=sha256:4021eaf1609962181f1f5cdcacd24e0a94a0dfccc0b22e99ad26efa4136922a4

Observation 7e346f55-3341-4de5-ac5d-a1398baa9cb5 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.423494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.423494Z digest=sha256:3c4f58753a452609fd82882db4c79b800015148d89174aacb47844f8e26cf68e

Observation 80f72f2f-d8d7-412b-9d1c-02ee2aff2602 · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge.2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 3190–3199, 2019.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Ok-vqa: A visual question answering benchmark requiring external knowledge.2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 3190–3199, 2019

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.427615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.427615Z digest=sha256:2c467e72d4958086f426648d596f63f42ce4e37aedee20c01d56242f4d6477c0

Observation 5c1320a2-db16-4cbb-8926-6e2a3acea4cd · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pages 216–233.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pages 216–233

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.431764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.431764Z digest=sha256:cf42f2adee6aec1fff452dd6ba680ca5febd82f79e2ee1f98b89f8cda4314570

Observation 0b293c01-c3a9-43d5-bdff-79e2cb2ae5a1 · outbound

This paper cites MMMU: A massive multi-discipline multimodal understanding and reasoning benchmark for expert AGI.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans MMMU: A massive multi-discipline multimodal understanding and reasoning benchmark for expert AGI

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.435780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.435780Z digest=sha256:578c1316e73767b82996d356f11a7cbfe80ef8bec38f5cb0b6f8071d199ce8d2

Observation f461b015-adbb-412c-b320-b48ef8260221 · outbound

This paper cites Humanity's Last Exam.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Humanity's Last Exam

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.440333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.440333Z digest=sha256:6d7141497cb41986a6dfcdb9885f610652183d4237e7b4f6c2cc11b96352b5e9

Observation 3a1a2842-322c-4fec-b604-b4b16cc20081 · outbound

This paper cites Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? InEuropean Conference on Computer Vision, pages 169–186.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? InEuropean Conference on Computer Vision, pages 169–186

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.445137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.445137Z digest=sha256:e2422607bcb234d5afab19531d12a4dfa8d7f5f4283daabd32c97993750247b4

Observation f2b90fb1-d75a-44d1-b90a-cd0412ccae4e · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Training Verifiers to Solve Math Word Problems

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.448758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.448758Z digest=sha256:b66e5c5e689272d5f99c4f3d0fb92e6216036275e8c2f26abccac83db08ad881

Observation 86c09a99-d31d-471a-b009-cbacb4f554da · outbound

This paper cites EXAMS-V: A Multi-Discipline Multilingual Multimodal Exam Benchmark for Evaluating Vision Language Models.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans EXAMS-V: A Multi-Discipline Multilingual Multimodal Exam Benchmark for Evaluating Vision Language Models

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.455724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.455724Z digest=sha256:76823e77d7eeb5a86b6368c651088970f49d2c30ee76dbd24544ae00d92088fb

Observation 644e8399-f47e-4bb0-bd86-dc018eccb77a · outbound

This paper cites Scienceqa: A novel resource for question answering on scholarly articles.International Journal on Digital Libraries, 23 (3):289–301, 2022.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Scienceqa: A novel resource for question answering on scholarly articles.International Journal on Digital Libraries, 23 (3):289–301, 2022

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.460568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.460568Z digest=sha256:7345642d210410b8d5db69e7de1a9880fe4994d69ab92e383fcbd39914e1af70

Observation 1cb4f7d5-cc86-485c-9451-bee23e149ba5 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.465245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.465245Z digest=sha256:d8d33ba2321f491e420430b72205c0013070499a3b4475e7bdea91b891a9cdc2

Observation e611daba-748b-438a-bfc9-a85845645edc · outbound

This paper cites MMIU: Multimodal Multi-image Understanding for Evaluating Large Vision-Language Models.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans MMIU: Multimodal Multi-image Understanding for Evaluating Large Vision-Language Models

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.469347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.469347Z digest=sha256:b4c6e36402a06691f84d364824a4201b9bcbf8ac7e921ea3ac6ac412015dd97b

Observation 310d25b8-954a-49e2-b126-04bf88a79a6b · outbound

This paper cites Introducing openai o3 and o4-mini, 2025.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Introducing openai o3 and o4-mini, 2025

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.473292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.473292Z digest=sha256:df66828d29f5bcb1327ccc7c5ba48f0efc84e04222e0956975d2735510bc35d8

Observation 11e372b3-f8ff-40fb-a60c-0b6369ad6da2 · outbound

This paper cites Gemini 2.5: Our most intelligent ai model, 2025.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Gemini 2.5: Our most intelligent ai model, 2025

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.477768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.477768Z digest=sha256:c97f8d60420715eb070693e3c410e85b306545f01ffcfdd89c928ca93d02f4f7

Observation 97a17e25-d4a2-46e6-b471-675d3e17dc6c · outbound

This paper cites black + white = black.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans black + white = black

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.481666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.481666Z digest=sha256:05301a9bce417e7ed65f13a7a6b1d9008debb3776cdb0f7444372682692d8d7c

Observation 94aed1f4-904d-4adb-bb14-b01d495ddbba · outbound

This paper cites an unresolved cited work.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Unresolved cited work

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.486202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.486202Z digest=sha256:48b84d49fa3f4e1cd779e072a164d3460ee6a6b51eca896c5cc3c6aba7aef4fd

Observation a459a31a-7b38-460e-b1ec-232ef6e82ec0 · outbound

This paper cites an unresolved cited work.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Unresolved cited work

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.489974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.489974Z digest=sha256:db5fe015ee99eeec4e2e3f0509d65257df5d58a294fd7084c46c864569e7060f

Observation 0b2db3b0-2f37-453c-af85-afe20722b713 · outbound

This paper cites an unresolved cited work.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Unresolved cited work

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.493984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.493984Z digest=sha256:5187288e841d253b4f18684cb88ec7b0d3332eba4827c1129ba7ac381f60e632

Observation 20fa0037-5474-491d-9ef6-93cbd1a252fd · outbound

This paper cites during working hours,.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans during working hours,

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.497957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.497957Z digest=sha256:270c89c94b642ac37d66fcbfea1fd69388700e19e9d989923fc4845890c2c1a9

Observation eab776ae-9e6f-4d7b-ae9d-85ffb12cc433 · outbound

This paper cites for profit,.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans for profit,

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:00.815190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:00:59.501620Z digest=sha256:2d999d4fae4ee251e46d574a7593d37d09f5506e56324b16e0c441a6915323f1

Observation 683eea3b-4027-4f69-9e98-e25e7e30638c · outbound

This paper cites must," "main,.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans must," "main,

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:00.802838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:00:59.505286Z digest=sha256:8e671c84dec0f6d04a3654cf281c2a8505282014b00b2c28d1f8b59e2e805949

Observation eb89377d-292e-4eb6-950b-35520842aeba · outbound

This paper cites deconstruct.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans deconstruct

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:00.788369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:00:59.509218Z digest=sha256:6da443e2e46425a3e63ebc8c77987b2e2ca225be89544985a27715a9eefc755d

Observation 86a5383b-d01a-4a37-8db7-7935bd02cb1b · outbound

This paper cites - Step 5: Strictly and systematically compare the option’s information with the defini- tion’s core elements.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans - Step 5: Strictly and systematically compare the option’s information with the defini- tion’s core elements

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:00.775910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:00:59.512934Z digest=sha256:4bff98005935669729267ad89261f443ee9ca93fd10f484c57550f3bd60f9dda

Observation f6a62f66-96b9-4614-a9ba-16adf5cfe426 · outbound

This paper cites belongs to.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans belongs to

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:00.761912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:00:59.516889Z digest=sha256:8c30a595f45f4050388eb441ac62ac142c6000f82b038c0aa65b69e79ef97411

Observation e373f18c-daec-4418-9e31-9ba2bd6ef162 · outbound

This paper cites Choose the one that best matches or mismatches the definition.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Choose the one that best matches or mismatches the definition

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:00.749396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:00:59.520513Z digest=sha256:59058f6defa202e4bdf800ff6922f16ac45204f151da594042719934ccb4bbb3

Observation bc9867c2-1799-4ee0-b4a2-88ba3af3ddbe · outbound

This paper cites Always take the definition as the sole criterion.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Always take the definition as the sole criterion

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:00.731799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:00:59.524434Z digest=sha256:fb3e10138b490afcd446b4f90c99c3d0591f8044ca46c61527b301275f387628

Observation 8b4285e3-d13e-48cb-be54-65feff1d63f6 · outbound

This paper cites must," "main,.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans must," "main,

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:00.719458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:00:59.528151Z digest=sha256:bf978b5eba96e02c71e4c4e94b1bb2b892bb5aad4dfd20691fac68925458d78d

Observation 76020b2c-8134-4877-9893-e22fcaff5be9 · outbound

This paper cites justifiable defense.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans justifiable defense

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:00.706600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:00:59.532020Z digest=sha256:68921fd4934304cbfcdbff5f988a69a9f38613fee9571246a521cfd007ca2c27

Observation 2243a021-4918-4245-9872-16c964de8b97 · outbound

This paper cites Or" vs.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Or" vs

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:00.693791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:00:59.535732Z digest=sha256:6020b88e171a7569120d1a1d5b71a388990dd23516fc06afa93ed464f5f51e3f

Observation a7c078ef-b16d-4833-8f29-b69fdf4fd7c2 · outbound

This paper cites belongs to.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans belongs to

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:00.680001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:00:59.539713Z digest=sha256:ccfbc631c59be29fd314d20d57cb412c50555b8bc819f7c75d7d432a48d74136

Observation 18f4c6f3-10d2-442a-a5a1-6f967807db32 · outbound

This paper cites Word-Picking.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Word-Picking

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:00.661637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:00:59.543611Z digest=sha256:f0f239d0a1317d6257c23526249b1a225b2c5197c8a097a6ba12533ecd8f9c2d

Observation 906585cc-ec6b-4475-aeb8-b101b60e459e · outbound

This paper cites one-sentence summary.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans one-sentence summary

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:00.646462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:00:59.547332Z digest=sha256:cf3b41575033bd84e47b025d77235d2b08dcdbdab5f3511e2a9db959ac201173

Observation 24447a73-3b72-4239-b906-795e206198f7 · outbound

This paper cites 在工作时间”、“未经许可.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans 在工作时间”、“未经许可

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:00.632714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:00:59.551152Z digest=sha256:1ee168f70d29bc4cd623a34d84dfc73c280c9fb5dfc81c60d56d714b24238751

Observation 6ad78094-1698-4cab-a37c-7a55d69c3d76 · outbound

This paper cites 必须”、“主要”、“故意.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans 必须”、“主要”、“故意

Reference 101

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:00.617875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:00:59.555618Z digest=sha256:654ec390fd41393c4cecd9f916f8e53ab921d7f3a085d8cf9125d9eb86a431e4

Pith citing papers

No inbound Pith citation observations are available.