Pith. sign in

Paper Citation Record · LEDGER

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset

As of 7 August 2026, this Paper Citation Record lists 93 of 93 outbound references and 5 inbound Pith citation observations for arXiv:2507.03483.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.03483 v2

Coverage vector

measured 93 of 93 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:15:52.926857Z

measured 98 of 98 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:02:43.598214Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T20:48:56.092401Z

Reference resolution

93 of 93 outbound references displayed

  • verified exact2
  • verified fuzzy10
  • unresolved81
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0196b675-2020-4ad9-afd8-792cc86d3ac2 · outbound

This paper cites Qwen2.5-VL Technical Report.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.545119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.545119Z digest=sha256:49cfa4f6188a522e447068b9f0f237fc698070abb8a2dc305e6b3d7568ef7df9

Observation c001efd2-9060-47e5-8216-628980107bda · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.549797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.549797Z digest=sha256:9f3dc2ff74a1d442244b8fbd6b7eff0d0b42a3dc1ad31b86c0591214bc95e10a

Observation 4e837ba1-77d1-4473-a9aa-2d5a7ee160d0 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.553874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.553874Z digest=sha256:352d534359c517441d832e684ed7ce95431041c4b580632321dde88c32673f63

Observation 5e4c053a-6351-49d2-97dc-cef948232513 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.557584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.557584Z digest=sha256:108555a9a3b8050bd9f67cc798dbf80724cac60ae0cbb0a03db9bb2d805280a3

Observation d0d01ff7-c1da-42f9-92ba-1fe13bdeb5ba · outbound

This paper cites MMMU: A Massive Multi-discipline Multimodal Understanding and Reason- ing Benchmark for Expert AGI, November 2023.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset MMMU: A Massive Multi-discipline Multimodal Understanding and Reason- ing Benchmark for Expert AGI, November 2023

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.561578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.561578Z digest=sha256:f470bca2c242f6c2e69c6321b082b3ee22e1ce1950e2cc35e3f0baeeab246b29

Observation 6dddc08f-5c55-4e0b-b062-4568f8cfaf57 · outbound

This paper cites Sci- enceqa: A novel resource for question answering on scholarly articles.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Sci- enceqa: A novel resource for question answering on scholarly articles

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.565003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.565003Z digest=sha256:52e2c22525dee5274f4ce046db00bc66371cec82aeb0fc200cdc82cba4be403d

Observation 461f4d28-97c3-4e46-83f2-50140d5ed506 · outbound

This paper cites What can large language models do in chemistry? a comprehensive benchmark on eight tasks.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset What can large language models do in chemistry? a comprehensive benchmark on eight tasks

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.568687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.568687Z digest=sha256:d5ae91e45437d467f986150b9f73c6ee1580fc323a68825ae7cd58dfbc9adf9e

Observation eb3368ba-232b-4847-9056-24d3778465a7 · outbound

This paper cites GPT-4o System Card.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset GPT-4o System Card

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.572152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.572152Z digest=sha256:711a65da287310c24cf840fbead7b63991ebb5741988fe677b32021486f6f772

Observation d73a7183-14f4-4193-b473-f8edbb384ad2 · outbound

This paper cites OpenAI o1 System Card.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset OpenAI o1 System Card

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.576901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.576901Z digest=sha256:291c56ecd6776a0c903c7775c4f94411cea156e927cabed7bfbb735b41fb6367

Observation 11ad014d-ed3b-4848-b195-3087ee8f5572 · outbound

This paper cites Introducing openai o3 and o4-mini.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Introducing openai o3 and o4-mini

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.581244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.581244Z digest=sha256:5efc678157a66311d616be60809199f182e934f9100a5a9a3925d9cc27427179

Observation 7a4441b2-d93b-4398-98a4-55d5c15d4457 · outbound

This paper cites Claude 3.7 sonnet.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Claude 3.7 sonnet

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.585255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.585255Z digest=sha256:c8a29ac5596a2e4e11d8775eb9bd7ddf2cb694fcb00ec9f5400ae64eb4d09a1c

Observation 7d4f9ae3-d4ac-478e-8621-9658363a762f · outbound

This paper cites SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.588780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.588780Z digest=sha256:f0fc728676f3b2a19f5bdd2f4a3e5e2354117ed4269ef7c2f253a4fdd6cdad4f

Observation c54c5915-1fc4-4630-8e7d-3e80d278ea9f · outbound

This paper cites P-mmeval: A parallel multilingual multitask benchmark for consistent evaluation of llms, 2024.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset P-mmeval: A parallel multilingual multitask benchmark for consistent evaluation of llms, 2024

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.592208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.592208Z digest=sha256:6014ae16f20eb8f700229a763f628eb8623fcf9f839e843580b9848eb91ea095

Observation a43d3464-a496-4d92-8aa2-7d56843564cb · outbound

This paper cites Mmlu-pro: A more robust and challenging multi-task language understanding benchmark.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Mmlu-pro: A more robust and challenging multi-task language understanding benchmark

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.596012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.596012Z digest=sha256:e17072309805222050d458db55d5fc1ba28228aa7291c48966ea5ff6431c3b54

Observation 1e0f301f-09f3-41a4-9887-c4b61bbd5631 · outbound

This paper cites Gpqa: A graduate-level google-proof q&a benchmark.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Gpqa: A graduate-level google-proof q&a benchmark

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.600800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.600800Z digest=sha256:e0c5bf26a32631cbd87066b27281f5be7c7ea5f42a9b161de5b4c679770ff7b0

Observation 61b533d6-b753-402f-a077-0458583cd244 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Measuring Mathematical Problem Solving With the MATH Dataset

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.604524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.604524Z digest=sha256:023c39b4a94dac2f42500c19d574ab34b6535b4578152a71413add1e50477975

Observation 11a9397a-b159-4a05-a6f6-70fb31f82779 · outbound

This paper cites Cumulative Reasoning with Large Language Models.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Cumulative Reasoning with Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.608615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.608615Z digest=sha256:0c1ad1959a36b947dce2733e168f8b70c1990b0e735760252cb98ed9890b3154

Observation c048c313-80d1-4d7a-aaa5-70a70998b260 · outbound

This paper cites Multimodal ArXiv: A dataset for improving scientific comprehension of large vision-language models.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Multimodal ArXiv: A dataset for improving scientific comprehension of large vision-language models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.617862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.617862Z digest=sha256:0eb75539c04c23ea4f54ec4a818b483aed5afb734c80eac7cce9708be5a4c4e0

Observation 2771ed89-b22c-48f2-9a6b-6cae50a9822b · outbound

This paper cites International standard classification of education.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset International standard classification of education

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.622384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.622384Z digest=sha256:bfad3e06f669cd0f9ff2b2802fc304e4ac845f644e359061187cf42d8e76dda9

Observation 19dbde4b-59d3-44ae-a3a5-bae7614a5a98 · outbound

This paper cites Mind with Eyes: from Language Reasoning to Multimodal Reasoning.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Mind with Eyes: from Language Reasoning to Multimodal Reasoning

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:15:53.550336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:15:52.625803Z digest=sha256:38c0457106fd568cc255105c5b56f00b410e2c8a69c80cc640199b3b359585c3

Observation b944415a-66da-46c3-85df-3faed605457d · outbound

This paper cites Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.629640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.629640Z digest=sha256:d59e84a49edd8f86652c4efea6380393befb5a4b82341e232dcccc553ae2ad88

Observation 814105ab-5d6c-4ee3-92ab-bf4575a160b2 · outbound

This paper cites Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.633699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.633699Z digest=sha256:2ae3c898d28796fef62e8ecca08ee6e9dee0619770388c221417736255ef256a

Observation 31c21de8-0702-4dd8-a981-a1f2b467ffae · outbound

This paper cites Generative Verifiers: Reward Modeling as Next-Token Prediction.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Generative Verifiers: Reward Modeling as Next-Token Prediction

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.637922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.637922Z digest=sha256:aa5fb63fda01b388f0f9f9d41a67df178f3793236b3676ae9e5d1633ca6f3095

Observation b76387e4-3dff-4ef0-8212-d05f5bca5e0f · outbound

This paper cites Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.641896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.641896Z digest=sha256:f88deef52a59a8ca64c03142bea57c1edab875ffbb7b7ba748d5743bc8222324

Observation e6bbca02-a527-4cb8-a7e2-05fd2c2b8596 · outbound

This paper cites VisualPRM: An Effective Process Reward Model for Multimodal Reasoning.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset VisualPRM: An Effective Process Reward Model for Multimodal Reasoning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.646946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.646946Z digest=sha256:58f8a374f9c5e07c515a15b338543e48bc83933dc3ae91335b5bd529fb2afa0e

Observation 14180049-22fc-4abd-be0a-90ce1b3dd6bb · outbound

This paper cites Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.651200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.651200Z digest=sha256:6749c678e2707bb509cde455b0e0b9a7c8420f124af7e46ac9f0660883d89a8f

Observation f5a3d15d-6335-4ea6-8e7c-e02618a8b1c4 · outbound

This paper cites Missing Premise exacerbates Overthinking: Are Reasoning Models losing Critical Thinking Skill?.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Missing Premise exacerbates Overthinking: Are Reasoning Models losing Critical Thinking Skill?

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.655037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.655037Z digest=sha256:019ee82cef18d52d9fb6d3ae07fb17eb402139f42890a39ee6c2512c23d6e064

Observation 9c59b816-d211-4455-b145-30517c69511a · outbound

This paper cites HallE-Control: Controlling Object Hallucination in Large Multimodal Models.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset HallE-Control: Controlling Object Hallucination in Large Multimodal Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.659099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.659099Z digest=sha256:2db22f525ed26a6315bee835eb8a006e4ea52171bf43c215eece6429def224f1

Observation 5fd9cdcc-245b-4092-bd2c-4ce3568ad1ef · outbound

This paper cites Hal-eval: A universal and fine-grained hallucination evaluation framework for large vision language models.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Hal-eval: A universal and fine-grained hallucination evaluation framework for large vision language models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.663171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.663171Z digest=sha256:47dfbcaa2cb80715bbbe6f0e8c062bdc4344bbcb645893bae5dd65bf4eae003b

Observation dff49deb-9081-4e57-a825-2d4fe27a794e · outbound

This paper cites The second half.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset The second half

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.666627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.666627Z digest=sha256:36e80d394a947e7dd6630b4be886c2f375890f145bc277610d043b6717e92b75

Observation 2fb3b996-4df1-46d4-aa02-6ed8afb6f6d9 · outbound

This paper cites Elevater: A benchmark and toolkit for evaluating language-augmented visual models.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Elevater: A benchmark and toolkit for evaluating language-augmented visual models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.670814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.670814Z digest=sha256:7f6bf880753d7950067257d4fd4a4b8184e7bd2a7d0e78fab9b9b4613ea20b17

Observation 59a54d08-71cd-4773-a16a-f1ba0da0294a · outbound

This paper cites Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.674522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.674522Z digest=sha256:dab7d5b0c455e57855c34ac1026034273a2c8e855e6b0328e29889a204b47d66

Observation 17c5c31d-b620-4493-b796-a8b6dd59cdd3 · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.678109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.678109Z digest=sha256:f59c67ed5df70450e3c2d0921a01b3867ae6a6e6b2da35cc8e991d0366c31adb

Observation eec4027b-0b94-4579-b9ae-994cd9e300f1 · outbound

This paper cites MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.681595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.681595Z digest=sha256:67eb3a91c05a316342392aa32597797c5e10a582d4de9a5f6f4341c20a1d37eb

Observation d88e2028-b93b-4ba4-aead-19819e32635f · outbound

This paper cites Gemini 2.5.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Gemini 2.5

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.685339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.685339Z digest=sha256:57099a4cef4804b55c74822b0702493c28988a0c57d80e16f4c7e2c66ce46179

Observation a4ae9a29-feaa-4155-a224-c9d95132001b · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.689598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.689598Z digest=sha256:fce8942c1c71636a181b682532a84711f4fac29b9f33f4a6db4a59aceac80592

Observation 57d9a831-24c3-49d7-b08c-97ae6260c925 · outbound

This paper cites MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.693320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.693320Z digest=sha256:1c33d547d2238ab6adecc3fa1b04fa53cf69c5a0a83384d4a1e419fb25ec2c2f

Observation 717384d5-dc58-4e87-b348-d36840ac1aa1 · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.697483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.697483Z digest=sha256:441bb90149de62f5cfbd9c69fee494e047531ccf1456007f35d1ce1855497503

Observation e509b6e9-ec44-41bd-865e-d1e0d170bfc1 · outbound

This paper cites DynaMath: A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language Models.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset DynaMath: A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.701699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.701699Z digest=sha256:e44b6e2a2189d3d3729ceb4f2e2bd06b81aa04fef56defb0860a93eeced428dc

Observation 59c29797-7406-4cc9-8f12-44a690f1425f · outbound

This paper cites CMMU: A Benchmark for Chinese Multi-modal Multi-type Question Understanding and Reasoning.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset CMMU: A Benchmark for Chinese Multi-modal Multi-type Question Understanding and Reasoning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.706978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.706978Z digest=sha256:b6fecfebe3a7fdd6ee04ea86c826834797749be5cd8295ec6670974695f5e131

Observation 8bfe9a92-6741-43ec-9626-7cc730452232 · outbound

This paper cites R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.710933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.710933Z digest=sha256:48c8e7df093f75513d9586fbc38e34b32fbeceb30dd332ef8e5fe14e4654f5ae

Observation c983e567-740d-4f0e-b249-3590d7cef571 · outbound

This paper cites Mmsci: A multimodal multi-discipline dataset for phd-level scientific comprehension.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Mmsci: A multimodal multi-discipline dataset for phd-level scientific comprehension

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:15:53.934789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:15:52.714872Z digest=sha256:fe89b28e9a8352f9d2abe062bb26589182348e07110b5d9ced9623851f26e8aa

Observation 1a482397-f3b1-4315-9fa2-f3cf176d83b3 · outbound

This paper cites WorldSimBench: Towards Video Generation Models as World Simulators.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset WorldSimBench: Towards Video Generation Models as World Simulators

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.718735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.718735Z digest=sha256:d75a92e8eb86caa99935665eee989e62da656e19493a175fc39dffd2c094ea56

Observation 415e36ab-9b69-4de4-a2ac-106cb43d10e3 · outbound

This paper cites Spatialvlm: Endowing vision-language models with spatial reasoning capabilities.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Spatialvlm: Endowing vision-language models with spatial reasoning capabilities

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.724401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.724401Z digest=sha256:58bccc9b908c3fc79f4c5d04426f1ce3dff0a0d0d9f766682c5b7af55d9a9651

Observation c03f60c5-0ec2-4492-94af-b179664f1177 · outbound

This paper cites MM-Spatial: Exploring 3D Spatial Understanding in Multimodal LLMs.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset MM-Spatial: Exploring 3D Spatial Understanding in Multimodal LLMs

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.729171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.729171Z digest=sha256:3b4330dca1e6207ce367905db222d7838b820547babb47a29ed3db7c7b9dd0fe

Observation c569df7c-8152-4629-8373-70c74568eca3 · outbound

This paper cites SpatialBot: Precise Spatial Understanding with Vision Language Models.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.733114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.733114Z digest=sha256:af9ea8e27793dbdab323b60903ae4c8888405a1489f3fb7bf0171dc0e226d23b

Observation f5e5e55d-0c93-42bc-8c6d-ac9082eba907 · outbound

This paper cites LLaVA-CoT: Let Vision Language Models Reason Step-by-Step.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset LLaVA-CoT: Let Vision Language Models Reason Step-by-Step

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.737246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.737246Z digest=sha256:3119262549543c80b61486362113a0cfcc31518277bc12a380dfb3038e0320d0

Observation 4d4fd369-42ff-48dd-aa16-8a3b2a588772 · outbound

This paper cites MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.742706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.742706Z digest=sha256:bc6a8f22fc7923987966cf962eceb3dc0fefa3b2e4517f9433b9e07e353593c0

Observation 8f0d0dbc-bb94-4e32-af2c-afdeb0ac4b09 · outbound

This paper cites MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.747130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.747130Z digest=sha256:c05b650c77d9f886ce377168f31da7dc3bb8a7eab35d714c4aecd3972859cb4d

Observation ab0e998a-2ab1-49a5-a659-96c0d5b83aea · outbound

This paper cites Let’s verify step by step.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Let’s verify step by step

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.752298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.752298Z digest=sha256:9d4c013d37928022b1198f88a53c9e316cee27e74c28aa0a5d5d14d6296cc533

Observation a8355300-10c5-4c64-bd8c-6442ff837d21 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Training Verifiers to Solve Math Word Problems

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.756735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.756735Z digest=sha256:d2b3ffc62ae05e4c95a97ce8185ba472ff5937a1a2838debf21f519c6e6b628b

Observation 849056ab-8732-4605-8bed-aa8161486bed · outbound

This paper cites Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.760921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.760921Z digest=sha256:2263289ab796abe80cd4985c21c94326e81810046021be03f9abfed760515f74

Observation 7fae0959-5729-4815-b117-045361ba19bc · outbound

This paper cites LLM Critics Help Catch LLM Bugs.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset LLM Critics Help Catch LLM Bugs

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.764833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.764833Z digest=sha256:f5a175967c03c9059ef347fd53fc664552552bda09d158a8c2fe161765581a3f

Observation 081ce2bf-beb7-4cee-80f6-57e90bf31711 · outbound

This paper cites Improve Mathematical Reasoning in Language Models by Automated Process Supervision.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Improve Mathematical Reasoning in Language Models by Automated Process Supervision

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.768877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.768877Z digest=sha256:bafe78e0d940cf68c423e03c3e661818977acc4643b9879072052d9e1ed9e3ab

Observation 29c5a000-9b78-4855-8090-ad9185c86536 · outbound

This paper cites Better Process Supervision with Bi-directional Rewarding Signals.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Better Process Supervision with Bi-directional Rewarding Signals

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.772660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.772660Z digest=sha256:5cc4820f261a2daf5f7acddf1c3df11c09471952d477222a293a9b913bf2f8ce

Observation 51cc0d4b-bd68-4089-99b2-3295328529ae · outbound

This paper cites Bandit based monte-carlo planning.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Bandit based monte-carlo planning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.776986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.776986Z digest=sha256:762a80e61caff0bdfa6c82faa9fa517747a234ec9e620e9be0bd23e2917ec40a

Observation b15b2ea2-9412-4577-b323-58c9e367dc85 · outbound

This paper cites Efficient selectivity and backup operators in monte-carlo tree search.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Efficient selectivity and backup operators in monte-carlo tree search

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:15:53.899442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:15:52.780751Z digest=sha256:f71af93318c0321a5dc2c474c2e319c743267eae78679f850dfd0c7f6fcb6001

Observation d6a07ceb-aa06-48d0-9f29-7c11db37d6ad · outbound

This paper cites OVM, Outcome-supervised Value Models for Planning in Mathematical Reasoning.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset OVM, Outcome-supervised Value Models for Planning in Mathematical Reasoning

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.784527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.784527Z digest=sha256:b291c8c5f60e6fee9d853f9c92bf4ee70b36fab63c37e893738e571b70489c7b

Observation 8e28e85a-a52d-48c4-81a8-bcdde58098ea · outbound

This paper cites Process Reward Model with Q-Value Rankings.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Process Reward Model with Q-Value Rankings

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.788301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.788301Z digest=sha256:4839c69f77df6bc419092558f68f3a1ff1a83b6697aadc0a658a14ebaaaf358f

Observation e9b66aa5-718f-4c68-a563-11dcb260dfc1 · outbound

This paper cites Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.792261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.792261Z digest=sha256:986350cf8b2142f2113ca850df33da069c7545a696dd0037320a7342b8fd5f2d

Observation e31829f5-60f1-4b74-87bc-5d1de299266b · outbound

This paper cites SALMON: Self-Alignment with Instructable Reward Models.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset SALMON: Self-Alignment with Instructable Reward Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.796856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.796856Z digest=sha256:3badca77205547c887bc22dd2e18813886a9b8d6d82050bf1b6940482917fc5a

Observation 1339cbfa-27d2-4f3c-b29e-0928a7dfcc23 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Constitutional AI: Harmlessness from AI Feedback

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.800618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.800618Z digest=sha256:3c7f625ab1e376cfe0da641b35af78b0bc39d135e9be416a8ecc1dd934f3dc70

Observation 97a9a0db-f0a4-4fdb-88dd-c2ff8fc7abf6 · outbound

This paper cites Improving Reinforcement Learning from Human Feedback Using Contrastive Rewards.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Improving Reinforcement Learning from Human Feedback Using Contrastive Rewards

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.803990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.803990Z digest=sha256:5ddfe93e7d114cadbaafbba28d93cd66a2f2982996324e5708a4b50c4b2d2a0f

Observation d92fb749-f394-4f1a-b187-4fc6ffdfd77d · outbound

This paper cites Examining false positives under inference scaling for mathematical reasoning.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Examining false positives under inference scaling for mathematical reasoning

Reference 65

Resolution
verified exact
doi, observed 2026-08-06T20:15:53.023616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:15:52.808246Z digest=sha256:c3bd30260b1216052a1aa45a3f5cda28a2c5295ccb8cc61b38cf386e3662d88b

Observation e3725d2b-e79d-425a-8ce4-15a9d4f1ad1a · outbound

This paper cites LLM Reasoners: New Evaluation, Library, and Analysis of Step-by-Step Reasoning with Large Language Models.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset LLM Reasoners: New Evaluation, Library, and Analysis of Step-by-Step Reasoning with Large Language Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.814532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.814532Z digest=sha256:7d8ce9432dffe02424bb8b4d1855ee4407ff88d213bee127d14fa333d36398be

Observation 820699b0-56dc-44d2-a1ca-65509caa0638 · outbound

This paper cites When benchmarks are targets: Revealing the sen- sitivity of large language model leaderboards.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset When benchmarks are targets: Revealing the sen- sitivity of large language model leaderboards

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:15:53.887422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:15:52.818291Z digest=sha256:265df1d374affa7f9cd8a31c0a843de8a2e156183fee403a6b2250c29179f1fe

Observation 881975a4-3c79-4864-95e0-367ca9a90bc4 · outbound

This paper cites LLMs may perform MCQA by selecting the least incorrect option.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset LLMs may perform MCQA by selecting the least incorrect option

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:15:53.874770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:15:52.827005Z digest=sha256:5c209b3417d7337308c345dee8992c90c43f35847d0eeb42fe391457e6037b60

Observation 29b595b6-4567-4e63-84df-20843a84b03d · outbound

This paper cites Llm-evaluation tropes: Perspectives on the validity of llm-evaluations, 2025.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Llm-evaluation tropes: Perspectives on the validity of llm-evaluations, 2025

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.830477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.830477Z digest=sha256:b329e2e6129ca77d2d3d431892f59039789998d01b74427a3018d7f00a98ce1b

Observation ac13ee45-6f5e-4d40-8333-138f75f524fc · outbound

This paper cites Right Answer, Wrong Score: Uncovering the Inconsistencies of LLM Evaluation in Multiple-Choice Question Answering.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Right Answer, Wrong Score: Uncovering the Inconsistencies of LLM Evaluation in Multiple-Choice Question Answering

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.834798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.834798Z digest=sha256:73410b9dc18c2a641d6f126472af99af9f9a402e9cad2287806bf8ea92006a12

Observation 2cbaac74-cec4-44ea-946f-42cf9c594d4a · outbound

This paper cites xfinder: Large language models as automated evaluators for reliable evaluation.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset xfinder: Large language models as automated evaluators for reliable evaluation

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:15:53.863526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:15:52.838908Z digest=sha256:71e0ba87edb4f9417967e951d1503e618d2e195a26a17643041e16762b889b3d

Observation 0f38035f-23bd-4724-9d29-014052b7e437 · outbound

This paper cites Gemini 2.5 flash.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Gemini 2.5 flash

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:15:53.851010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:15:52.842692Z digest=sha256:b5791e17d3da0c22e9ddca9a91cf9fbf88f1e4868704708e16b609b7b09eb12d

Observation 7d9dcc68-1485-4200-93be-794bf78da20d · outbound

This paper cites Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.846633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.846633Z digest=sha256:089726d1594f1c982ee9dcdf0fed4735de85c543229e4bc72479ffc64f5d0451

Observation 855707f7-c8b0-46b7-b564-0ce9163f59c1 · outbound

This paper cites Internvl3: Advancing open-source multimodal models with native mul- timodal pretraining.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Internvl3: Advancing open-source multimodal models with native mul- timodal pretraining

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:15:53.839160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:15:52.850191Z digest=sha256:b2daea0203cbea4da82ab322f84a706900d04fd80ff143da8f22f32b827cc175

Observation 56920f09-1579-4f22-8f5a-91d6b6520a27 · outbound

This paper cites Qvq: To see the world with wisdom, December 2024.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Qvq: To see the world with wisdom, December 2024

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.854871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.854871Z digest=sha256:066a2a5aaf65c346d1c9a57ab2d4aac33a9d32f13537d628d401e5f114cd4da1

Observation 1bca3adc-e7b8-4092-a8be-ad9def86939f · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.858559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.858559Z digest=sha256:3aeed36a377044c775d5ffe23b3fc018de8ef20f141c060ab0c66734deddcf58

Observation c39b1df0-7116-454c-8268-c435a71761fb · outbound

This paper cites Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.863191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.863191Z digest=sha256:a324b6c97af1e1fa6e04db62718067f0baffbdc1e610d81499c412949fbe9b5b

Observation 333414fa-1cd1-44c1-88f2-de0b5dab8950 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset LLaVA-OneVision: Easy Visual Task Transfer

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.867152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.867152Z digest=sha256:94ff41f8a4f7ef8d8fb9076cb5a2034817478c7351cd3701c9cbfe70919c4cb3

Observation 86989673-95a9-4fe3-87f8-8f610e120782 · outbound

This paper cites L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.871522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.871522Z digest=sha256:804b7636038303f894895a62df16849ed274f9148f8747bf334880328ba3e989

Observation c581f00f-528b-4a66-8d94-fa23c2970249 · outbound

This paper cites s1: Simple test-time scaling.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset s1: Simple test-time scaling

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.876328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.876328Z digest=sha256:6d203af65ceef0ba31e28ee732f4f81be61ec69ec8a2da869878d5b6ef8722fa

Observation 57807e59-0a43-4f93-b5c0-d06b67cc9349 · outbound

This paper cites Qwen3 Technical Report.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Qwen3 Technical Report

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.880163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.880163Z digest=sha256:cbbd06cc0245036445b1075283b6147937f99997eb260b99b70a08c03af37127

Observation 6928dccb-5ae1-499a-a476-29a0ae3c78a7 · outbound

This paper cites DeepSeek-V3 Technical Report.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset DeepSeek-V3 Technical Report

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.884044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.884044Z digest=sha256:99103d4879c86e0ab59bd0b886a6c6486ab38db58029f1c9d752e9676534ec86

Observation 28faa975-416c-4b4a-abf9-c8e60afd8a80 · outbound

This paper cites ReFT: Reasoning with Reinforced Fine-Tuning.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset ReFT: Reasoning with Reinforced Fine-Tuning

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.887684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.887684Z digest=sha256:f9dc7b5411404cc565f936cfeb49f523c9de1f47004ad5c99a1a6fae5bff1642

Observation 498d7beb-586b-413f-aca6-bce5709b0fc6 · outbound

This paper cites LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.891650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.891650Z digest=sha256:bec32031235ba3014064630dcaac95e2a291542670b04474475faf6b5b763004

Observation ca9e36a5-7d26-4505-aa6e-fda642f5da38 · outbound

This paper cites Swift:a scal- able lightweight infrastructure for fine-tuning, 2024.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Swift:a scal- able lightweight infrastructure for fine-tuning, 2024

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.896058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.896058Z digest=sha256:e97c99667862ed6d4c594eca16421c7872c07649b8920a63658194d10613e2e5

Observation 430102e6-d2be-48dc-87ed-cf6aa7da7169 · outbound

This paper cites fact verification.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset fact verification

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:15:53.812240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:15:52.900027Z digest=sha256:d4af7f9bde0a7ebd5f262375c6ae892dbbd3e3e5e07055a9fc900dfec18b6d92

Observation ecfd68ab-fee9-431c-a6db-9fa838cf3f6b · outbound

This paper cites an unresolved cited work.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Unresolved cited work

Reference 88

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:15:53.800449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:15:52.904012Z digest=sha256:5bc00b8c2aef4278273fef1a58165529a4b526c7480cf062acef42d72fd69180

Observation 554efa4c-683c-441d-bc93-e10c29d1600e · outbound

This paper cites an unresolved cited work.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Unresolved cited work

Reference 89

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:15:53.789297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:15:52.908029Z digest=sha256:dcb61cfd9104f340a1f63badf69f2e6057633878fcde0691b8895030b9543d6e

Observation 42294d97-2459-4b35-8f87-923266aed433 · outbound

This paper cites the southeastern part has rich forest resources.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset the southeastern part has rich forest resources

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:15:53.775845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:15:52.911579Z digest=sha256:f8b82a9a604c8cf831467d5b655d9300007a88b86df1f6e4cb48c2a8fd7b29a6

Observation e3b84485-b1fb-4c67-8ed9-aa64db3d9c3d · outbound

This paper cites an unresolved cited work.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Unresolved cited work

Reference 91

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:15:53.763119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:15:52.915564Z digest=sha256:6c8b70a5f09c4f995a91d84d379a4a1a32985c216efc807b416c4f34de290e5e

Observation 0912e1eb-79a9-48a8-a6d6-8a2c30e51ce7 · outbound

This paper cites an unresolved cited work.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Unresolved cited work

Reference 92

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:15:53.751893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:15:52.919130Z digest=sha256:42a9ce8fe9586f337b3ce746d48f279158f773f7a57938fa09e1255758702e88

Observation 929a71a1-9ac5-4fdc-b706-9a5234b8099d · outbound

This paper cites an unresolved cited work.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Unresolved cited work

Reference 93

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:15:53.740663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:15:52.923169Z digest=sha256:3ba260012c53174c6443c05c818f6272da353ddd9e0aec364843dd9041a115fc

Observation 19e66c48-fdad-40ae-bc07-676440fc883d · outbound

This paper cites F Limitations and Broader Impact BMMR is a dataset that focus on multidisciplinary reasoning for multimodal models.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset F Limitations and Broader Impact BMMR is a dataset that focus on multidisciplinary reasoning for multimodal models

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:15:53.728981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:15:52.926857Z digest=sha256:d3b4819e4793755112c7f3c4aee9fe98275ee95812c80153234324f48416d7d5

Observation 85ad6546-f4a4-4d92-9676-795774282688 · outbound

This paper cites doi: 10.18653/v1/2024.acl-long.744.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset doi: 10.18653/v1/2024.acl-long.744

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.822666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.822666Z digest=sha256:908323f9b56cc3e5d7c18a0bc61c808960acd462658046c730039783a4bf57b1

Pith citing papers

Observation 19eb5644-2edf-4276-bd43-6bfc7ba260a4 · inbound

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency cites this paper.

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset

Reference 153

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:58:58.884793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T11:58:58.660564Z digest=sha256:bd4915e1fbfffa7ddd100b2223ffeb5e3481d43be230389897f8a01fb009bf03

Observation 043f1a4b-b753-4a0b-9e06-afe4d804ee16 · inbound

SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents cites this paper.

SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T23:41:45.067017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:41:45.067017Z digest=sha256:9af45cd53d83bfa358e8a10ca1ab60d3474e029d2fff4d73c5e6929cd9050994

Observation ded2f978-9444-4286-af2a-5cfe152474db · inbound

Towards Characterizing Scientific Image Utility and Upgradability cites this paper.

Towards Characterizing Scientific Image Utility and Upgradability BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:06:27.854608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T11:10:37.254428Z digest=sha256:1813a1daca3060406907af594113933b6072bf593ce04ddadcf968cc4b4a656d

Observation 6e2299a8-0fe7-4042-a0f1-785578fb7fe6 · inbound

Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients cites this paper.

Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset

Reference 111

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:48:56.094168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T01:08:52.981296Z digest=sha256:7ecda78902292bd8bda0eddfd704c4e9d93556138289469db8d13c04ed5607a9

Observation 182f00bc-48d2-4ba2-9971-1b52146dab92 · inbound

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning cites this paper.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:43.598214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:43.598214Z digest=sha256:e13721a88619b8690a4058780b94e540a6bc15145a8613efef86c3761fd4aec2