Pith. sign in

Paper Citation Record · LEDGER

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation

As of 5 August 2026, this Paper Citation Record lists 78 of 78 outbound references and 1 inbound Pith citation observation for arXiv:2603.00546.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2603.00546 v2

Coverage vector

measured 78 of 78 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T19:56:35.246936Z

measured 79 of 79 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-29T11:53:35.315405Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T12:03:24.478700Z

Reference resolution

78 of 78 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved78
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b95ece5a-b998-493a-afb4-1960e4c6ae80 · outbound

This paper cites Qwen2.5-VL Technical Report.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:27.636015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:27.636015Z digest=sha256:b0e24a3d7f596700c2e27b0b1759fb6321b3e82dbbb7b09717ce49baeb836e96

Observation ba55f520-b18f-42e9-a081-06f8bc823306 · outbound

This paper cites MLLM-as-a-judge: Assessing multimodal LLM-as-a-judge with vision-language benchmark.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation MLLM-as-a-judge: Assessing multimodal LLM-as-a-judge with vision-language benchmark

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:27.702214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:27.702214Z digest=sha256:cb8ea498591e3c12cb9f2a24e63882f5a5361d2bc5ac8ac76bc93cc1d9e9918b

Observation 8fcd1560-87f6-4f87-900e-1b0d3c5986c8 · outbound

This paper cites Are we on the right way for eval- uating large vision-language models? InAdvances in Neural Information Processing Systems, pages 27056–27087.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Are we on the right way for eval- uating large vision-language models? InAdvances in Neural Information Processing Systems, pages 27056–27087

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:27.862650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:27.862650Z digest=sha256:641883b03eba5ab7937e819789f5e153e3654afa3ef9d538d0885c3529353082

Observation ef5384c5-836c-40fe-8259-79616f60e36b · outbound

This paper cites M 3CoT: A novel benchmark for multi- domain multi-step multi-modal chain-of-thought.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation M 3CoT: A novel benchmark for multi- domain multi-step multi-modal chain-of-thought

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:27.933953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:27.933953Z digest=sha256:9c9e5b36350c17e05586b86205293349219a3e36bb80e86b2b683ee7f65a5b13

Observation 84d50ffe-005a-44d1-9271-ab0a73a543e3 · outbound

This paper cites Every Sample Matters: Leveraging Mixture-of-Experts and High-Quality Data for Efficient and Accurate Code LLM.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Every Sample Matters: Leveraging Mixture-of-Experts and High-Quality Data for Efficient and Accurate Code LLM

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:27.990155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:27.990155Z digest=sha256:ba464c7213f5128b4ead71af65dcc26a67507bc84cfc17cee8a7608a196c6b18

Observation 0941feac-e08e-45b6-a2d3-e47d321952a4 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:28.086458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:28.086458Z digest=sha256:2a524bf75812c82be6f36d4144680fb7be473fb8cfcc861962fd4177dd660b3e

Observation 0cd671ac-3da1-430b-83e5-ca62c916d4f5 · outbound

This paper cites Efficient selectivity and backup operators in monte-carlo tree search.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Efficient selectivity and backup operators in monte-carlo tree search

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:28.281455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:28.281455Z digest=sha256:764d07588780a18f5bb06e53e9b93a96cd54855bfb8bd0a326481f336acd68cc

Observation 2bb1915a-a348-4929-842a-7dd66ba6281e · outbound

This paper cites Mm-ifengine: Towards multimodal instruction following.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Mm-ifengine: Towards multimodal instruction following

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:28.333590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:28.333590Z digest=sha256:81bccf6910283bde980f7bdd163fc9159ef0ceecdb0220e4ceac1113132ab5fa

Observation 1ca2f837-1ad3-464c-977a-95820fb8ab4f · outbound

This paper cites MM-PRM: Enhancing Multimodal Mathematical Reasoning with Scalable Step-Level Supervision.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation MM-PRM: Enhancing Multimodal Mathematical Reasoning with Scalable Step-Level Supervision

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:28.466278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:28.466278Z digest=sha256:3de68ce4cb2382da8642dd6f5f746c2a288cfc6f1f0a30b1707f3336db1db131

Observation 335b71fd-9a91-4702-bd6d-6032358880bd · outbound

This paper cites Seed1.5-VL Technical Report.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Seed1.5-VL Technical Report

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:28.590628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:28.590628Z digest=sha256:e850ccf476a89a906cce57d69450356d6d3322e0c2b719b00c5833e406d62fb4

Observation 1d0c79e5-8689-4bae-998d-69272bda73bc · outbound

This paper cites GPT-4o System Card.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation GPT-4o System Card

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:28.712586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:28.712586Z digest=sha256:feb997547aff1b3b96f112f90c8dc50606932e2e8b0965b33321e1bddcd567c5

Observation 39b7f9c5-7f9a-4158-9cc7-c8e0108ca516 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Gonzalez, Hao Zhang, and Ion Stoica

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:28.870577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:28.870577Z digest=sha256:e2252b9e5af9be1dcffec8cddb973597eba92d28cd9a6c3e82e051873984b415

Observation 3aee8528-579a-4fed-8205-11dc1dea540c · outbound

This paper cites Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:29.025781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:29.025781Z digest=sha256:7abaa27f3e90bcdff00557e869410b9104c09985683b61c1a974fd24968b2da7

Observation 87760e01-565b-4123-b07d-57f1790e7454 · outbound

This paper cites From generation to judgment: Op- portunities and challenges of LLM-as-a-judge.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation From generation to judgment: Op- portunities and challenges of LLM-as-a-judge

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:29.234607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:29.234607Z digest=sha256:ae6f1462755230ababd087d5ce764caeabe5dc2fae2178af2213f7ae571eb59a

Observation 71cc89be-fb92-444b-aeaf-a86bf525ddd8 · outbound

This paper cites Vl- rewardbench: A challenging benchmark for vision-language generative reward models.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Vl- rewardbench: A challenging benchmark for vision-language generative reward models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:29.459785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:29.459785Z digest=sha256:f5877b3c3c0ca03593e2801454feb4bb1636f3ecc0c89cb7610ea85f107315cd

Observation 7ee98173-9fc8-4dbf-8b53-fa26734bad75 · outbound

This paper cites Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:29.545585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:29.545585Z digest=sha256:bd6c892d8d7599d2c7363b177404703a6e06c95e79696d3e23acafd92296e099

Observation 58616d15-a35e-4aaf-a966-0a5b2c48cabd · outbound

This paper cites Improved baselines with visual instruction tuning.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Improved baselines with visual instruction tuning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:29.603365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:29.603365Z digest=sha256:459cb5752a4314e6260bce7b6b56ed405f9aa5b58df85b17e17916380b84ae5b

Observation 6a18d149-1aea-4896-837c-e464ed7cdb4a · outbound

This paper cites MIA-DPO: Multi-image augmented di- rect preference optimization for large vision-language mod- els.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation MIA-DPO: Multi-image augmented di- rect preference optimization for large vision-language mod- els

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:29.695853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:29.695853Z digest=sha256:c329720d8780861dfdda887ff778bb3dad6d0d2f94f6c2ea15e5734ebf138bc1

Observation 208fa9a9-3acd-4f28-8dec-c003cd5ceac1 · outbound

This paper cites Improve Mathematical Reasoning in Language Models by Automated Process Supervision.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Improve Mathematical Reasoning in Language Models by Automated Process Supervision

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:29.797926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:29.797926Z digest=sha256:310cf71cdc42aea43b956e2f2a3f2e643817fa4f3d07eb5de94ca2d61dbb95f5

Observation f61d1d53-b576-40ce-9acb-973822cea5b2 · outbound

This paper cites an unresolved cited work.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:29.895022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:29.895022Z digest=sha256:a3b4b175509490560eb65f336797b1a33f592187fbc9ebfb3378e675be70e064

Observation 2d4e6dec-6a8a-418e-8d02-bb3c703b7692 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Direct preference optimization: Your language model is secretly a reward model

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:30.102690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:30.102690Z digest=sha256:4b6808fa5cab797a2bd5a3452f2f3513ca4838c773c26ef7925106b9134aed9a

Observation c12b52a3-8990-4b27-9ced-1a516fc9929e · outbound

This paper cites GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:30.300126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:30.300126Z digest=sha256:32e5abb2b06c3799ddc8c3d60b70d0580a4217764043579258d424b0d6126bc7

Observation 43515f9f-132e-49c3-b733-22885f3f06ae · outbound

This paper cites Qwen3 technical report, 2025.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Qwen3 technical report, 2025

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:30.487065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:30.487065Z digest=sha256:f68768d0000ad6a81361826b20ce6920016f592a590e3f9623f17e9e23e6651d

Observation 9a820de4-a910-4a8d-a76c-6440165bd6fd · outbound

This paper cites MiMo-VL Technical Report.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation MiMo-VL Technical Report

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:30.621170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:30.621170Z digest=sha256:c403bbb7506dc66a2db9e0792552ff5e0474579f96d1dbc9108a5640000c3177

Observation 0d27393f-7520-41ec-b445-bf42ce9cb706 · outbound

This paper cites Mea- suring multimodal mathematical reasoning with math-vision dataset.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Mea- suring multimodal mathematical reasoning with math-vision dataset

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:30.762220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:30.762220Z digest=sha256:aabb5be1e23426ce84d2892406ba16f2f7be94fd8bc0eea0524390b897ad6eaf

Observation c956993e-8033-4bb3-b88b-75403199d96d · outbound

This paper cites Self-Taught Evaluators.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Self-Taught Evaluators

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:30.964243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:30.964243Z digest=sha256:ace4ae10e51c2f9b873b6c655baf7f5788797792cf1d8316fd0dea047a0a7a6e

Observation a626effc-9642-4931-824b-5e833831e6c4 · outbound

This paper cites Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:31.100419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:31.100419Z digest=sha256:0f83609d9e725e47240d07e0fa250b3c88323ea5ba1d41d537161e1e0b28a2c5

Observation ce050607-bab2-4464-920e-7d6e9f9d0c97 · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:31.186279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:31.186279Z digest=sha256:51ccbc2713e280e2ae06775873a17416e91b371864bd1bd352d72318296f8a1e

Observation c2f28a64-5501-45db-9814-687086437bd8 · outbound

This paper cites SoTA with less: MCTS-guided sample selection for data-efficient visual reasoning self-improvement.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation SoTA with less: MCTS-guided sample selection for data-efficient visual reasoning self-improvement

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:31.281051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:31.281051Z digest=sha256:897fc45926448394f6487f50ddb6ffa5f56d738a481dd6f283f83a4a175b00a3

Observation ad8e691b-00d2-4218-b3d7-c3be76064525 · outbound

This paper cites Unified multimodal chain-of-thought reward model through reinforcement fine- tuning.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Unified multimodal chain-of-thought reward model through reinforcement fine- tuning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:31.484261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:31.484261Z digest=sha256:ce9c059be079c1d6a611e66e577a90f549a1ac9b806a899144b2d436b4548954

Observation 31c2c557-fc52-4cda-b980-d82f73e7fe6d · outbound

This paper cites Unified Reward Model for Multimodal Understanding and Generation.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Unified Reward Model for Multimodal Understanding and Generation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:31.618718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:31.618718Z digest=sha256:b15961c2fdfe35cb3c8c3db0da274ebb8ce88454a270e37177332a2239f13001

Observation b91f6f3c-613a-4879-8461-3f7d70e5a749 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large lan- guage models.Advances in neural information processing systems, 35:24824–24837, 2022.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Chain-of-thought prompting elicits reasoning in large lan- guage models.Advances in neural information processing systems, 35:24824–24837, 2022

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:31.796208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:31.796208Z digest=sha256:a78f5fdd3c80647c148a1b095a40f6db79b70601ae4ad9b904efeb749060d7fd

Observation aa3fc127-7b93-4a47-adbb-472fafd0fd03 · outbound

This paper cites J1: Incentiviz- ing thinking in llm-as-a-judge via reinforcement learning.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation J1: Incentiviz- ing thinking in llm-as-a-judge via reinforcement learning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:31.858453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:31.858453Z digest=sha256:6e0c18e47334ca766cdc01f94c076bb569a94f3a5cac885ec8b0f6cdb8234106

Observation 9bd999c1-d447-4212-867a-446d82d60593 · outbound

This paper cites Multimodal Preference Data Synthetic Alignment with Re- ward Model, 2024.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Multimodal Preference Data Synthetic Alignment with Re- ward Model, 2024

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:31.956457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:31.956457Z digest=sha256:b1a46ceb2b8d12ac08b762b5baf49b9aade42cef89d43b3cec0832b4847966ef

Observation 0e781a24-5391-49b4-8b7b-e7aaef3e4f5f · outbound

This paper cites an unresolved cited work.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:32.055970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:32.055970Z digest=sha256:1773c6d25445e6fa38c51d0e76f2a4d7f59730d44edce08633e99d962ea52fa9

Observation 4663dc23-d0e0-44b4-9820-d6eee7906aa3 · outbound

This paper cites Monte carlo tree search boosts reasoning via iterative prefer- ence learning.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Monte carlo tree search boosts reasoning via iterative prefer- ence learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:32.196304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:32.196304Z digest=sha256:aeea417c30ba62646104191df91b7732107980c704cd913df97acc2e451c7daf

Observation b5119312-8e6d-43c8-b357-5e3047303c21 · outbound

This paper cites Llava- critic: Learning to evaluate multimodal models.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Llava- critic: Learning to evaluate multimodal models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:32.330762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:32.330762Z digest=sha256:c09369873a3a3893bf9c1fe98bb0d4e8f21719ac82a4d148e7d814c64df803a3

Observation 344db1ac-eb17-4a89-9d55-ceccdefb5236 · outbound

This paper cites Chen, Wenzheng Liu, Wei Zhang, Wenjie Zeng, Xikun Zhang, Jingyi Zhang, YuXin Song, Wenhao Wu, and Dacheng Tao.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Chen, Wenzheng Liu, Wei Zhang, Wenjie Zeng, Xikun Zhang, Jingyi Zhang, YuXin Song, Wenhao Wu, and Dacheng Tao

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:32.399862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:32.399862Z digest=sha256:0d28e90259705d53b0f46d23ce158d92c45dfbfaace495975f0168fdb1f8d3ed

Observation fd9a0d36-c308-4e71-96c4-c1ac633e6e16 · outbound

This paper cites Mulberry: Em- powering MLLM with o1-like reasoning and reflection via collective monte carlo tree search.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Mulberry: Em- powering MLLM with o1-like reasoning and reflection via collective monte carlo tree search

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:32.469544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:32.469544Z digest=sha256:0cb40c90a5e0806967ebb435d5db4ca2afca0bfee0b2a39575c85b35eef70442

Observation 9f2c75f4-4926-474b-a789-d457d18ac3e6 · outbound

This paper cites Multimodal RewardBench: Holistic Evaluation of Reward Models for Vision Language Models.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Multimodal RewardBench: Holistic Evaluation of Reward Models for Vision Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:32.626573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:32.626573Z digest=sha256:8e02df314412fae99ddce703db3d5ec6e14f3ab9aef7e4298905260d9790d17a

Observation 3cf4578e-7a22-4c43-abf0-b9ca285ac19e · outbound

This paper cites A survey on multimodal large language models.National Science Review, 11(12): nwae403, 2024.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation A survey on multimodal large language models.National Science Review, 11(12): nwae403, 2024

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:32.755599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:32.755599Z digest=sha256:b364b41ba1af6f1d70c9bc2be55d84e34edc66a2ccf79580036e493b7cbeef16

Observation d11f5705-9a13-488c-a8c2-2f9dd3c1a2e3 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:32.882052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:32.882052Z digest=sha256:84157f737b3023b1528f5b72b4c9f62ea90e270822f53b9755752cebc1d52891

Observation 9c4fd817-26f6-4cc6-b80e-2a2cb135ec94 · outbound

This paper cites Rlaif-v: Open-source ai feed- back leads to super gpt-4v trustworthiness.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Rlaif-v: Open-source ai feed- back leads to super gpt-4v trustworthiness

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:32.951359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:32.951359Z digest=sha256:809863c429e9cc238ae287a4d3bd18c0d1e1751d9248a90953c79f401c0cbd71

Observation 155dead4-f115-4b73-9b19-9b1e359ebcef · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understand- ing and reasoning benchmark for expert agi.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Mmmu: A massive multi-discipline multimodal understand- ing and reasoning benchmark for expert agi

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:33.005768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:33.005768Z digest=sha256:19f1b9875fed45b755c267e72fd9b9276a9566ab4d316a992821f29a32dfd89f

Observation 4a90b085-ec11-4f09-8888-9cca6fa2976b · outbound

This paper cites MMMU-pro: A more robust multi-discipline multi- modal understanding benchmark.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation MMMU-pro: A more robust multi-discipline multi- modal understanding benchmark

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:33.062176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:33.062176Z digest=sha256:d93d676d2005a15e78a4bfc2950815785ef501cc0f06da5e2d230e6597240f1e

Observation 03d601ef-cc7b-48f2-bd53-025463245bb5 · outbound

This paper cites InternLM-XComposer2.5-reward: A simple yet effective multi-modal reward model.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation InternLM-XComposer2.5-reward: A simple yet effective multi-modal reward model

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:33.095853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:33.095853Z digest=sha256:7d3b1675f8fd6d265a29d017fabe7e50096719f196c3929672548469c5c320b0

Observation b478a422-45dc-47ad-adaf-8b3ec038477c · outbound

This paper cites Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? InComputer Vision – ECCV 2024, pages 169–186, Cham, 2025.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? InComputer Vision – ECCV 2024, pages 169–186, Cham, 2025

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:33.160515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:33.160515Z digest=sha256:873867b139abc5b3cabf9ce1a1e30973f62db06c5b2b4dc054107878ef3c6d1e

Observation e328df1f-2149-4c82-8992-e059970c3bd9 · outbound

This paper cites GPT-4V(ision) as a Generalist Evaluator for Vision-Language Tasks.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation GPT-4V(ision) as a Generalist Evaluator for Vision-Language Tasks

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:33.229753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:33.229753Z digest=sha256:b813f32d8b6de2495336294dc0633e1cf29c47c295d9ba9174621e560708471c

Observation 0e55930e-bd20-4101-a527-91a1f656848d · outbound

This paper cites R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:33.297932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:33.297932Z digest=sha256:51efeae5628e84cb87379b4b25bb6bb383396f6ba81cec1420460b116e35b2c4

Observation e46844e7-e775-4c67-9b1d-6f524c4336b3 · outbound

This paper cites Llamafac- tory: Unified efficient fine-tuning of 100+ language mod- els.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Llamafac- tory: Unified efficient fine-tuning of 100+ language mod- els

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:33.389349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:33.389349Z digest=sha256:f7729ee98838916b1e834b10311c37ae430757973ccbacddd7402314358b5119

Observation 0634c388-8415-4f9c-bcc1-e76a24a1cd03 · outbound

This paper cites Easyr1: An efficient, scalable, multi-modality rl training framework, 2025.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Easyr1: An efficient, scalable, multi-modality rl training framework, 2025

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:33.492213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:33.492213Z digest=sha256:3f1256f5f78476e02db05e23142d757099aaeb7086847f2cb094b19bd0cedd97

Observation 30b42ec6-0547-43d4-b6b9-6efe512f3a97 · outbound

This paper cites Aligning Modalities in Vision Large Language Models via Preference Fine-tuning.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Aligning Modalities in Vision Large Language Models via Preference Fine-tuning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:33.552171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:33.552171Z digest=sha256:2198f675209c72fd7098c803bf2f4cdafa36dd4d8b931981d3f48c09aec2cba6

Observation e10bb0df-8bd6-4365-93ea-d163e706ac3b · outbound

This paper cites Capability-Oriented Evaluation Framework.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Capability-Oriented Evaluation Framework

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:33.610346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:33.610346Z digest=sha256:4e46c1a00a5f1db950de02f01b5f504b6fd24a68244c77b0e3e44940189417b6

Observation b0bfeda5-4260-452d-9455-42e018e144f3 · outbound

This paper cites Open-Source Training Data Collection.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Open-Source Training Data Collection

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:33.646776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:33.646776Z digest=sha256:89973a964a9e028d43ba43dcdbe63f4edd38cd6ea463116651183f4dbc0b9b1a

Observation 2569f6e2-ddd5-42d2-84e7-128362fddc86 · outbound

This paper cites Experimental Setup.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Experimental Setup

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:33.698935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:33.698935Z digest=sha256:c0c62ceee71d73a673e28822d021a8f8123971e7a39f009e59f9c072b21b5367

Observation 4a291b2e-12d8-4054-81af-e02c462b0813 · outbound

This paper cites Benchmark Examples 3 A.1.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Benchmark Examples 3 A.1

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:33.759584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:33.759584Z digest=sha256:1d077d5294bac9c187ed7cb6f96722f34855d5a2cc82cdb616a2b4d20b1e8b59

Observation 07341223-79eb-4b64-9143-15eec125faf3 · outbound

This paper cites **Gross operating surplus**: Net operating surplus + Con- sumption of fixed capital = 240,000 + 110,000 = 350,000 Rm; 3.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation **Gross operating surplus**: Net operating surplus + Con- sumption of fixed capital = 240,000 + 110,000 = 350,000 Rm; 3

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:33.832523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:33.832523Z digest=sha256:b4b21a82f7de74aaddd65a4f0a65cbd10af5976917f965731534e4c928d16c34

Observation 6284a8dd-0a6d-492e-a905-422ba40fa4c3 · outbound

This paper cites **Net operating surplus**: 240,000 Rm; 3.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation **Net operating surplus**: 240,000 Rm; 3

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:33.940890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:33.940890Z digest=sha256:b4443d643eee2d63a2a87dcf5efa6743c71d237cfe89729ed5be27cb0d5f3afc

Observation 362b74d5-ae8d-46c4-b058-ff5a35b0ea33 · outbound

This paper cites an unresolved cited work.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Unresolved cited work

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:34.012350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:34.012350Z digest=sha256:6775e58f4c4a6e5951abb37f2d38d1311a6e4c4793718971bafe9b2fcf305406

Observation 8396c4eb-63c1-4346-b453-3922c691afe6 · outbound

This paper cites **Sum of numbers in triangles: ** 4 + 11 + 18 = 33; **Sum of numbers in circles: ** 5 + 12 + 16 = 33.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation **Sum of numbers in triangles: ** 4 + 11 + 18 = 33; **Sum of numbers in circles: ** 5 + 12 + 16 = 33

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:34.045268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:34.045268Z digest=sha256:2cc03e3f1a367edc100da4ca071c60148003b53da8bfc01c9028377d1a84a939

Observation 6367bd4e-400b-41d6-98a2-0d8c14f1fc27 · outbound

This paper cites We can apply this rule to the squares.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation We can apply this rule to the squares

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:34.111025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:34.111025Z digest=sha256:305f6b669d10a4855198bf9904da3c1e01b801cfa2f69b5173b47194954d60eb

Observation 1613519c-70a6-40cb-a12f-3594159d77c2 · outbound

This paper cites Key Observations.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Key Observations

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:34.142051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:34.142051Z digest=sha256:bfaab803085cadb9a13b631d7132abe4d9ab7cedea1e33020cd11a8adfbe1215

Observation 473dd069-4e65-4f6c-ba93-10ee49400532 · outbound

This paper cites an unresolved cited work.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Unresolved cited work

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:34.206791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:34.206791Z digest=sha256:bd8c21e9002bac08f8dd8e90a0c679d9325bc251a11059bc7db4bfb948c1d08b

Observation 0bb5a77d-dc22-40ed-9233-9c4746f41ebf · outbound

This paper cites an unresolved cited work.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Unresolved cited work

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:34.244048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:34.244048Z digest=sha256:9bada4dad8cf2351e5d6e07d8aac3e3b6f25bd62e039b6908e89af44a6cba43a

Observation a9e52dd4-6919-48a0-b722-62c693ff2d28 · outbound

This paper cites an unresolved cited work.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Unresolved cited work

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:34.310999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:34.310999Z digest=sha256:de9101d7f3e5288f794762888dd0395c5fa7dfaee0d471ef2dd067f5e99ddde0

Observation 281b89df-2018-40eb-a9d5-47ae0071c3b3 · outbound

This paper cites an unresolved cited work.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Unresolved cited work

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:34.371808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:34.371808Z digest=sha256:3885b29e40ab89209682b70ab4ee4aeb2185334723a50ff481eab200cadd0413

Observation 9f1f8935-3028-4716-b6cd-a0407b82339d · outbound

This paper cites an unresolved cited work.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Unresolved cited work

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:34.464640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:34.464640Z digest=sha256:ca96683704979d6ee275a19dff48e1943e37f9155bc239c4460b7a47421af694

Observation e5b0b1b4-e301-432c-b732-fe556ad5b261 · outbound

This paper cites an unresolved cited work.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Unresolved cited work

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:34.535306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:34.535306Z digest=sha256:df9006dc9287214eae6eac90bfef169db7d7636c5030d44bac9673dac564cf52

Observation 5745fa83-121b-4620-93c9-36b8e4490c53 · outbound

This paper cites an unresolved cited work.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Unresolved cited work

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:34.568170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:34.568170Z digest=sha256:111d64d87d57d78d3c16587efc917b9e17526c14cfc4053648b2c99f6bc552fb

Observation 5fd8554f-5120-4673-a33d-06e2b89c8ab2 · outbound

This paper cites an unresolved cited work.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Unresolved cited work

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:34.635972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:34.635972Z digest=sha256:b4966dd02e9a8281b127a5b0000f0cef1272ad0a26b6eb622081a4aa786831b1

Observation 1465f68a-7383-4e3c-819f-26928d009ee8 · outbound

This paper cites an unresolved cited work.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Unresolved cited work

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:34.745434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:34.745434Z digest=sha256:b58f8574176f37df93823b0ed1d5de4f45b6fb2462511aac9afa61d06d3106c9

Observation 59d4d96f-b5a0-4ba1-bffc-3d8534be58a2 · outbound

This paper cites Modify the given correct solution text by introducing subtle, hard-to-detect errors.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Modify the given correct solution text by introducing subtle, hard-to-detect errors

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:34.814626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:34.814626Z digest=sha256:a183331a0ab3d6f19fe810e8b7a11f0a2bb012d91697a863cb8affac72d880cb

Observation 9d0f57b1-70f9-4258-8890-eae572b8befa · outbound

This paper cites Prioritize factual correctness above all other factors.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Prioritize factual correctness above all other factors

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:34.951668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:34.951668Z digest=sha256:587e8d4be61a4b9463ca87c5c609c73417a788e84dead9cfae7cc251d50d3946

Observation 12fabe4f-29c8-4cc2-b9ec-245d1b9af5bd · outbound

This paper cites Identify any reasoning fallacies, invalid inferences, or irrelevant logic chains that might affect reliability.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Identify any reasoning fallacies, invalid inferences, or irrelevant logic chains that might affect reliability

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:35.022760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:35.022760Z digest=sha256:5f5f7d7eff04e7a3b3f754745971e0017693b68e855f860badec54599172585c

Observation 2432c871-2dd6-4796-ac40-3f56498da1ea · outbound

This paper cites an unresolved cited work.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Unresolved cited work

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:35.076387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:35.076387Z digest=sha256:8f551c5dbd64a576f326636088908beecb36440bc4d06ab8632c835287708a31

Observation 06447396-6fab-46ae-8381-46821f8fe1bd · outbound

This paper cites The response should neither omit essential reasoning nor over-elaborate with redundant or misleading content.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation The response should neither omit essential reasoning nor over-elaborate with redundant or misleading content

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:35.154932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:35.154932Z digest=sha256:36d10f7547eb3f39e9c2249e63c5330c864a1c3b6f62ea5ec4c57463350125bf

Observation 386b1fba-a7de-4a2a-a35c-1eaf6c815f51 · outbound

This paper cites The decision should rely on reasoning quality, not stylistic fluency or writing pref- erence.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation The decision should rely on reasoning quality, not stylistic fluency or writing pref- erence

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:35.246936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:35.246936Z digest=sha256:517e4ef13a71a4ae9ebc84ef840cb0abc22d37704780cc752b143a73055160e4

Observation c91a1c5d-cc07-4537-81d6-5419be646f88 · outbound

This paper cites an unresolved cited work.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Unresolved cited work

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:29.346396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:29.346396Z digest=sha256:2067a8cbbd581fb7097e5e0986cdf60d7e10b62b8d274b9c9506d5c3cf701bcc

Pith citing papers

Observation ffe76620-05ea-44fc-a9dd-555a06543955 · inbound

MUSE: Benchmarking Manufacturable, Functional, and Assemblable Text-to-CAD Generation cites this paper.

MUSE: Benchmarking Manufacturable, Functional, and Assemblable Text-to-CAD Generation Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-16T02:22:33.558625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T11:53:35.315405Z digest=sha256:ba079d16bc327e1077978fb3f132799130b2d22bb411c75fa299cd5c4ef2d25d