Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T21:00:59.555618Z
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 100 of 121 outbound references and 0 inbound Pith citation observations for arXiv:2505.11141.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T21:00:59.555618Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
100 of 121 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 9b01b0bd-e42f-4612-bb51-3ab8bf8db8a2 · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Reasoning, problem solving, and intelligence.Handbook of human intelligence, pages 225–307, 1982
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6fbd7297-af31-4bd5-ae5b-677c292a6edf · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Intelligence and reasoning.The Cambridge handbook of intelligence, pages 419–441, 2011
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d465d40c-c445-45cb-83de-38e55385ded4 · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Levels of agi: Operationalizing progress on the path to agi.arXiv preprint arXiv:2311.02462, 2023
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 942b823d-7c6d-445a-ad55-f40fbab378a9 · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38a7a480-6e43-4b31-a541-0d7fa169fc2f · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Qwen2.5 Technical Report
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ccad829-7663-4913-8d64-402cf4fadae7 · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans ChatGLM-Math: Improving Math Problem-Solving in Large Language Models with a Self-Critique Pipeline
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b072b084-d996-444d-87d9-724d6cd8fb77 · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f164bf5-6f56-4ca9-8c72-3483d5e86bb7 · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans LogicAsker: Evaluating and Improving the Logical Reasoning Ability of Large Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 111fb5e2-45aa-442c-9c4c-26b696a16f56 · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Symbol-LLM: Towards Foundational Symbol-centric Interface For Large Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d79de149-d067-4563-bb0b-5c2652268e76 · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Language Models can be Logical Solvers
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation deeab4a9-3e0a-417e-9d34-c8aa23cceab1 · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans LogiCoT: Logical Chain-of-Thought Instruction-Tuning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d27b262c-4f30-40c8-b17c-685928490a0a · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans OpenCodeReasoning: Advancing Data Distillation for Competitive Coding
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5cbd7a7d-d1a6-45fa-b735-850e25339e1e · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Self-planning code generation with large language models.ACM Transactions on Software Engineering and Methodology, 33(7):1–30, 2024
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 590fdc54-8eee-402c-bac4-0ab10656dbf9 · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Structured chain-of-thought prompting for code generation.ACM Transactions on Software Engineering and Methodology, 34(2):1–23, 2025
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74d08069-4dd6-4a02-81bd-180aaa5140b3 · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans CodeCoT: Tackling Code Syntax Errors in CoT Reasoning for Code Generation
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2f304f6-4926-4a90-bb0f-3a28489982e6 · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans OpenAI o1 System Card
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7aa369a8-f284-403c-be99-edb45601d507 · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b1c60bd-76c2-446f-9331-519a992666e7 · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Springer, 2007
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa0867be-5301-4ea6-9526-6a22d5d3ff55 · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Exploring the Reasoning Abilities of Multimodal Large Language Models (MLLMs): A Comprehensive Survey on Emerging Trends in Multimodal Reasoning
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b5d843b-68ae-4bba-8048-27346506a819 · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 126ca753-9ff0-4f06-8aab-cf7da700f9e2 · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans R1-v: Reinforcing super generalization ability in vision-language models with less than $3
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c668467-9f5f-4a06-8aad-64b234a0a74d · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Visual-RFT: Visual Reinforcement Fine-Tuning
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 946cd72b-23ce-41bc-8664-4c5aff50a98f · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans VisualPRM: An Effective Process Reward Model for Multimodal Reasoning
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45b70ba3-813b-405d-ab7d-7248588f551b · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans OThink-MR1: Stimulating multimodal generalized reasoning capabilities via dynamic reinforcement learning
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87dcb263-bd00-413b-823f-3fb3bab311c9 · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Vlm-r1: A stable and generalizable r1-style large vision-language model
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6fcabbe8-6ade-4025-b8db-731faac2ffe0 · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Open-r1-video
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e1fca7e-9e72-4ef5-bbf0-8dbbbf594ba0 · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans VisualPuzzles: Decoupling Multimodal Reasoning Evaluation from Domain Knowledge
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23aa7b98-9633-401e-b355-e834fa88e42e · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans MM-IQ: Benchmarking Human-Like Abstraction and Reasoning in Multimodal Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b82014cd-bf82-458d-87fb-1ca793ecf2d1 · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96dc97c7-5308-4420-86ed-34a2111bd7c1 · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 266301ef-630b-4fc3-bd6a-df4c2a398974 · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3c30b03-3d84-4f58-87d4-78358e7ebfa7 · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Flamingo: a visual language model for few-shot learning.Advances in neural information processing systems, 35:23716–23736, 2022
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0da66503-5c32-4921-a30e-1d2a527a2a1e · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5aa9be15-7c32-4625-8aee-827af55d7df4 · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Visual instruction tuning.Advances in neural information processing systems, 36, 2024
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06c8bf22-2dbc-4557-9cc3-559ab2af888d · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Minigpt-4: Enhancing vision- language understanding with advanced large language models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31cf828d-8819-4179-9fe6-098ae51117d0 · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Gpt-4o: A multimodal language model, 2024
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af6bb968-af5e-45ad-8ac9-70924b9cce31 · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Gemini: A Family of Highly Capable Multimodal Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7141138-c70c-4db1-8660-cbebf4c6f3a9 · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Claude: A conversational ai assistant, 2024
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d0f7752-44c5-4f92-b603-c9c2f81f0e40 · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2694a045-83cc-4cd6-939d-ac4a826f4422 · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8274c36-01d8-4b3e-a817-b0f3cbc43459 · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Qwen2.5-VL Technical Report
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93733ea4-3ca0-455a-85c4-6f1abd24f5b3 · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c1a9576-a1dc-4ac7-80c5-7252db5066c8 · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d672641-e81d-4906-ac7d-46a8c4d7f37e · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe0b1d75-14b5-4aae-af7a-f5ae797f4393 · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da2870a5-ca2b-4d6d-9c79-30c6cb2ad0d3 · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c4efc12-8a41-4da1-abbc-af479be850e7 · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans LLaVA-OneVision: Easy Visual Task Transfer
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f548ba1-bc33-40ad-b000-52b6df849d2a · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Pangea: A fully open multilingual multimodal LLM for 39 languages
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51f20111-f6b5-49f0-9ace-c65bd5bc8055 · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Harnessing Webpage UIs for Text-Rich Visual Understanding
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ffe5cd88-b90e-43c6-9847-ac75d913a313 · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3000e9c-444e-44fe-a609-934d48f0fe48 · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans The Llama 3 Herd of Models
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f486d850-d21c-4f7f-b595-9f2197a97a44 · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Math-llava: Bootstrapping mathematical reasoning for multimodal large language models
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be250d21-e842-49a3-8f98-6d9065a73682 · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans MultiMath: Bridging Visual and Mathematical Reasoning for Large Language Models
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 183a1482-dbf6-4aba-99b6-5aeba392488f · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Med-flamingo: a multimodal medical few-shot learner
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0fade284-7987-467d-91d4-02c561ee2ba6 · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Med-moe: Mixture of domain-specific experts for lightweight medical vision-language models
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 027bb3a7-7c97-4254-b7bf-2601e24db763 · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Chain-of-thought prompting elicits reasoning in large language models.Advances in neural information processing systems, 35:24824–24837, 2022
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52b3d24c-7bcc-479b-a14d-a47501771163 · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans DPO Meets PPO: Reinforced Token Optimization for RLHF
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9462d6c9-bf94-4a4f-b89c-11a99540f6a6 · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Star: Self-taught reasoner bootstrapping reasoning with reasoning
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3b024cb-163f-4f28-b6c9-6fb4b0c3c7c4 · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Quiet-STaR: Language Models Can Teach Themselves to Think Before Speaking
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c5b7f1a-73be-4a5b-aaa6-ab32ee6562e6 · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Qvq: To see the world with wisdom, December 2024
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7214e0e-bc38-4712-b2b5-b9eb1e42e5e6 · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Ocrbench: on the hidden mystery of ocr in large multimodal models
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b353542f-2d19-417b-902e-718beac5dd0a · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f22818f8-29b5-4380-ae4d-e55d70ec3e73 · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Docvqa: A dataset for vqa on document images
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fb5dd0d-43da-4279-8def-12dcf65cfdeb · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans AgentBench: Evaluating LLMs as Agents
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a358f92-c552-45e3-b3fc-e11d25d47c04 · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans ToolEyes: Fine-Grained Evaluation for Tool Learning Capabilities of Large Language Models in Real-world Scenarios
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f58b9c21-638b-419b-ac58-047211f86d55 · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Egoschema: A diagnostic benchmark for very long-form video language understanding.Advances in Neural Information Processing Systems, 36:46212–46244, 2023
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d96fa136-b120-4857-bad7-20002bed72bc · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Egothink: Evaluating first-person perspective thinking capability of vision-language models
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e346f55-3341-4de5-ac5d-a1398baa9cb5 · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80f72f2f-d8d7-412b-9d1c-02ee2aff2602 · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Ok-vqa: A visual question answering benchmark requiring external knowledge.2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 3190–3199, 2019
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c1320a2-db16-4cbb-8926-6e2a3acea4cd · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pages 216–233
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b293c01-c3a9-43d5-bdff-79e2cb2ae5a1 · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans MMMU: A massive multi-discipline multimodal understanding and reasoning benchmark for expert AGI
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f461b015-adbb-412c-b320-b48ef8260221 · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Humanity's Last Exam
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a1a2842-322c-4fec-b604-b4b16cc20081 · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? InEuropean Conference on Computer Vision, pages 169–186
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2b90fb1-d75a-44d1-b90a-cd0412ccae4e · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Training Verifiers to Solve Math Word Problems
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86c09a99-d31d-471a-b009-cbacb4f554da · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans EXAMS-V: A Multi-Discipline Multilingual Multimodal Exam Benchmark for Evaluating Vision Language Models
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 644e8399-f47e-4bb0-bd86-dc018eccb77a · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Scienceqa: A novel resource for question answering on scholarly articles.International Journal on Digital Libraries, 23 (3):289–301, 2022
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1cb4f7d5-cc86-485c-9451-bee23e149ba5 · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e611daba-748b-438a-bfc9-a85845645edc · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans MMIU: Multimodal Multi-image Understanding for Evaluating Large Vision-Language Models
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 310d25b8-954a-49e2-b126-04bf88a79a6b · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Introducing openai o3 and o4-mini, 2025
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11e372b3-f8ff-40fb-a60c-0b6369ad6da2 · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Gemini 2.5: Our most intelligent ai model, 2025
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97a17e25-d4a2-46e6-b471-675d3e17dc6c · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans black + white = black
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94aed1f4-904d-4adb-bb14-b01d495ddbba · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Unresolved cited work
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a459a31a-7b38-460e-b1ec-232ef6e82ec0 · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Unresolved cited work
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b2db3b0-2f37-453c-af85-afe20722b713 · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Unresolved cited work
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20fa0037-5474-491d-9ef6-93cbd1a252fd · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans during working hours,
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eab776ae-9e6f-4d7b-ae9d-85ffb12cc433 · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans for profit,
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 683eea3b-4027-4f69-9e98-e25e7e30638c · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans must," "main,
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation eb89377d-292e-4eb6-950b-35520842aeba · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans deconstruct
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 86a5383b-d01a-4a37-8db7-7935bd02cb1b · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans - Step 5: Strictly and systematically compare the option’s information with the defini- tion’s core elements
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation f6a62f66-96b9-4614-a9ba-16adf5cfe426 · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans belongs to
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation e373f18c-daec-4418-9e31-9ba2bd6ef162 · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Choose the one that best matches or mismatches the definition
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation bc9867c2-1799-4ee0-b4a2-88ba3af3ddbe · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Always take the definition as the sole criterion
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 8b4285e3-d13e-48cb-be54-65feff1d63f6 · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans must," "main,
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 76020b2c-8134-4877-9893-e22fcaff5be9 · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans justifiable defense
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 2243a021-4918-4245-9872-16c964de8b97 · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Or" vs
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation a7c078ef-b16d-4833-8f29-b69fdf4fd7c2 · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans belongs to
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 18f4c6f3-10d2-442a-a5a1-6f967807db32 · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans Word-Picking
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 906585cc-ec6b-4475-aeb8-b101b60e459e · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans one-sentence summary
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 24447a73-3b72-4239-b906-795e206198f7 · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans 在工作时间”、“未经许可
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 6ad78094-1698-4cab-a37c-7a55d69c3d76 · outbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans 必须”、“主要”、“故意
Reference 101
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
No inbound Pith citation observations are available.