Pith. sign in

Paper Citation Record · LEDGER

StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models

As of 23 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 0 inbound Pith citation observations for arXiv:2508.05383.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.05383 v1

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T23:29:17.102230Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

42 of 42 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved42
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 80ad6a69-6d28-470a-a1e7-4609fbb6fd38 · outbound

This paper cites Seed1.5-VL Technical Report.

StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models Seed1.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T23:29:13.766583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:29:13.766583Z digest=sha256:e5a104a5f2961ae34a9414906842b8ad6cdd95ea175f8f6bcca24738eb238636

Observation 66f41efe-fd1e-40ee-b23d-52eda21325e7 · outbound

This paper cites Learning to reason with llms, 2024.

StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models Learning to reason with llms, 2024

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T23:29:13.808513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:29:13.808513Z digest=sha256:849b15ea1d5620d7cdba9b87e822dcd5914208f53b7204194783ae977c725a36

Observation f7a78e0a-2303-4ced-8e0a-f587fe6416ab · outbound

This paper cites Gemini 2.5: Our most intelligent ai model, 2025.

StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models Gemini 2.5: Our most intelligent ai model, 2025

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T23:29:13.911172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:29:13.911172Z digest=sha256:d75ca7a6e0cf526a11bf9e2d7ab4bb12e04130e153d3089c1828c0acd52e854e

Observation e7ee3099-9567-4146-a212-0661e622810a · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T23:29:14.091099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:29:14.091099Z digest=sha256:6930eb2943b867d7f2cb6ceb3b172b030ce53fbf8f166eefea02b8b2fb970940

Observation 773c28a8-45e9-49ff-ab11-7031579c1417 · outbound

This paper cites Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In European Conference on Computer Vision, pages 169–186.

StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In European Conference on Computer Vision, pages 169–186

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T23:29:14.256851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:29:14.256851Z digest=sha256:82bdea32906424797b06a5f7597e94eb0d079404e865026c2a999f4af0127779

Observation fd97cdbd-84ef-44a5-8e05-c975d9995c0a · outbound

This paper cites SOLIDGEO: Measuring Multimodal Spatial Math Reasoning in Solid Geometry.

StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models SOLIDGEO: Measuring Multimodal Spatial Math Reasoning in Solid Geometry

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T23:29:14.367560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:29:14.367560Z digest=sha256:8def35a60654a34ad4b976163add49a70527c08220d5b3af96c7591bf3ec959d

Observation 17d1e031-80e0-44bd-a927-b968a3401fa6 · outbound

This paper cites We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?.

StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T23:29:14.487310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:29:14.487310Z digest=sha256:a9ee3eb5dc1141e95132734cc9b8e8655a1369550985cafd55f5a6055f83d992

Observation cf2837e5-30b2-4a4f-84b8-8c7912e0264c · outbound

This paper cites Polymath: A Challenging Multi-modal Mathematical Reasoning Benchmark.

StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models Polymath: A Challenging Multi-modal Mathematical Reasoning Benchmark

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T23:29:14.682753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:29:14.682753Z digest=sha256:d4971ec4f20ecbfadc6b77dbc931e98c24bae830eb5de445edbdbe3a82f2eeda

Observation 65d639d0-9eae-4eb5-a164-98cfa1390041 · outbound

This paper cites Mathscape: Evaluating mllms in multimodal math scenarios through a hierarchical benchmark.

StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models Mathscape: Evaluating mllms in multimodal math scenarios through a hierarchical benchmark

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T23:29:14.868582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:29:14.868582Z digest=sha256:eae90eef00b199fba147076870eb7aad89803c86b0ea9c84eab185e74bbe84fc

Observation 18cd1628-2a4b-4354-868e-acbceb063b63 · outbound

This paper cites R-Bench: Are your Large Multimodal Model Robust to Real-world Corruptions?.

StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models R-Bench: Are your Large Multimodal Model Robust to Real-world Corruptions?

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T23:29:15.024275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:29:15.024275Z digest=sha256:18667186ddea4f7685a5de45e0f5d7fe829e416d2001d6bf777cbf648e982f98

Observation 9231b1b3-8ec4-4a54-9baa-29eb6e09b9dc · outbound

This paper cites Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems.

StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T23:29:15.137929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:29:15.137929Z digest=sha256:f32688beebda95b406210907bb646b2f833a39f3578c9d24261ebea44fa4fa0f

Observation ad0ade1a-1244-40c8-b05e-4bc7fdd1d6a3 · outbound

This paper cites Scibench: Evaluating college-level scientific problem-solving abilities of large language models.

StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models Scibench: Evaluating college-level scientific problem-solving abilities of large language models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T23:29:15.264012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:29:15.264012Z digest=sha256:300f52b3ba04f96f632627e3c97e60efb7951e0e92492d7f97a05f183ce2ba7f

Observation 4bbd3274-5a24-4402-afc5-6da961b84d39 · outbound

This paper cites Scemqa: A scientific college entrance level multimodal question answering benchmark.

StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models Scemqa: A scientific college entrance level multimodal question answering benchmark

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T23:29:15.392526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:29:15.392526Z digest=sha256:602c460f3a44f6c89e35dbf4f151623083874e3707e59b1911612ee3ef3a69cb

Observation 29b6a9d1-578d-4c3e-a2fa-cfa6d2e16e32 · outbound

This paper cites VNHSGE: VietNamese High School Graduation Examination Dataset for Large Language Models.

StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models VNHSGE: VietNamese High School Graduation Examination Dataset for Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T23:29:15.552942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:29:15.552942Z digest=sha256:4c8f98f26c09c73a61fc9aee18d91753d3e06c07da1829b37dbb79cc64b5ffde

Observation 242b9482-5532-4ca9-b0a8-f34ff3e88156 · outbound

This paper cites Chemvlm: Exploring the power of multimodal large language models in chemistry area.

StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models Chemvlm: Exploring the power of multimodal large language models in chemistry area

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T23:29:15.664277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:29:15.664277Z digest=sha256:d0a3c37306ae61667468b5aef3736f3bcfd21ece08f400c0d3a9d4671f77599d

Observation 2c71bf73-9339-4b45-8e0b-b473c37bc94d · outbound

This paper cites ChemQuests: A Curated Chemistry Question-Answer Database Extracted from ChemRxiv papers.

StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models ChemQuests: A Curated Chemistry Question-Answer Database Extracted from ChemRxiv papers

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T23:29:15.787610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:29:15.787610Z digest=sha256:5c0815c42d3aee2db226b4cbadfb66341f24c8965c56a16246b5d597fedf73c9

Observation 2d9b7562-70e3-4b43-8bd4-31a4e10348c3 · outbound

This paper cites Evaluating the symbol binding ability of large language models for multiple-choice questions in vietnamese general education.

StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models Evaluating the symbol binding ability of large language models for multiple-choice questions in vietnamese general education

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T23:29:15.903931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:29:15.903931Z digest=sha256:b035e6e13ccf516c58c1f31e8537d8182e1dbe30b95a8c3ebe9c496b1bda08e2

Observation 48e8fdc2-78e4-493b-8bf5-b8d8d4134587 · outbound

This paper cites R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization.

StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T23:29:16.036443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:29:16.036443Z digest=sha256:deb67d4f398a44639b4958056952911a6a56325d15a4f2ef06b071086e8ed46b

Observation fc2651ec-3676-420e-bd47-10df60088fe7 · outbound

This paper cites Reason-rft: Reinforcement fine-tuning for visual reasoning.arXiv preprint arXiv:2503.20752, 2025.

StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models Reason-rft: Reinforcement fine-tuning for visual reasoning.arXiv preprint arXiv:2503.20752, 2025

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T23:29:16.106388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:29:16.106388Z digest=sha256:5adb379e5de526894e76b058e5a7345e7ab7bd1dbb3baf406e51a3e27c51228d

Observation dc0e6dd3-3e5a-460d-9af1-a39a376e33f0 · outbound

This paper cites VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model.

StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T23:29:16.170654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:29:16.170654Z digest=sha256:c927b9922582c0a1fc14f7a1c08653f2d3a583fd6371070d38d8b761518e7ae7

Observation 1e7e4c2b-afef-4f04-8e41-48b0bdee1539 · outbound

This paper cites Visual-RFT: Visual Reinforcement Fine-Tuning.

StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models Visual-RFT: Visual Reinforcement Fine-Tuning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T23:29:16.338406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:29:16.338406Z digest=sha256:72db617b7c7cc2bb32b970212c05abe93f016caddabd21c45a018e7a6e4e2374

Observation ca86a00a-d62e-4365-a261-53a173b04924 · outbound

This paper cites Perception-R1: Pioneering Perception Policy with Reinforcement Learning.

StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models Perception-R1: Pioneering Perception Policy with Reinforcement Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T23:29:16.461188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:29:16.461188Z digest=sha256:2638353116bbea024bffea19f2d021b5cb0794ad6aaa06b6178e56a9f1bffc0c

Observation e138b8bc-4358-41a5-ad59-8ba1c35ef888 · outbound

This paper cites R1-Zero's "Aha Moment" in Visual Reasoning on a 2B Non-SFT Model.

StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models R1-Zero's "Aha Moment" in Visual Reasoning on a 2B Non-SFT Model

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T23:29:16.503129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:29:16.503129Z digest=sha256:571481e63b37865c19206cce645faf57b395b1d4251d406e139a00ae0c945129

Observation f123ce55-4bbd-4388-82a6-b03a540296aa · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T23:29:16.575673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:29:16.575673Z digest=sha256:a46efbd0d93158edaa4449b51c867a40e90e479cfd097d0d6b383e6d091cce7c

Observation d551d9dd-76e3-45e6-81fb-cf68ba2e104f · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T23:29:16.646941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:29:16.646941Z digest=sha256:11627c7f6de56eb5dcbf1f249f8cd363e09b7dbe313fa93e0f7e8b888d731aac

Observation 56d5df1a-0304-408c-9ae2-c3c583c00b7b · outbound

This paper cites Demystifying Long Chain-of-Thought Reasoning in LLMs.

StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models Demystifying Long Chain-of-Thought Reasoning in LLMs

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T23:29:16.703242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:29:16.703242Z digest=sha256:0a657626a41f6f5fd001dd1b12fcfebf7b7adc7db0841a21a7dddcdd3f24ab94

Observation fd7b64ca-ab64-45e6-ad91-8132ecccf183 · outbound

This paper cites Open r1: A fully open reproduction of deepseek-r1, january 2025.

StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models Open r1: A fully open reproduction of deepseek-r1, january 2025

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T23:29:16.707048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:29:16.707048Z digest=sha256:f81f606e9f0af4856e9b4bcabc370b39fe78603820f45cd7bf32eb8f5ca27bfd

Observation a00dfc3d-725a-4f49-ba37-4f81fabc7e24 · outbound

This paper cites Rethinking RL Scaling for Vision Language Models: A Transparent, From-Scratch Framework and Comprehensive Evaluation Scheme.

StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models Rethinking RL Scaling for Vision Language Models: A Transparent, From-Scratch Framework and Comprehensive Evaluation Scheme

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T23:29:16.745968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:29:16.745968Z digest=sha256:e629c72f288a0ddaa4e090a94e84e03e14f2f9f1cecbad2bd47dde6a88552da7

Observation fdd19e39-4745-4c4a-990f-a61fd25109b0 · outbound

This paper cites VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks.

StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T23:29:16.807447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:29:16.807447Z digest=sha256:c268f4270544694460309f8eac06f7266ba74abb4b66810230200e6216f8df57

Observation da953aed-ea43-48d5-a0f7-abd7a5884468 · outbound

This paper cites VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning.

StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T23:29:16.856632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:29:16.856632Z digest=sha256:b1a851ee30de8704a08c171a117f0e29eef5a3ae8db0f74e3c02aed6eb2347ee

Observation 0cda74db-115e-4711-927d-1919d9f1a0fb · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T23:29:16.968225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:29:16.968225Z digest=sha256:4a356d979e403670f74bbf99f0178ca2b92a1a30348113a8f79d7732b026e36b

Observation 8969b9a4-7393-42e5-95c1-4d046f3c9338 · outbound

This paper cites RLPR: Extrapolating RLVR to General Domains without Verifiers.

StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models RLPR: Extrapolating RLVR to General Domains without Verifiers

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T23:29:17.022564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:29:17.022564Z digest=sha256:b45c648fcbaa4bc3ae4bf18abd171991db334c975de28329966e81a167dbb5e4

Observation 9b1af818-1681-4458-bc65-be417c1fb148 · outbound

This paper cites Grpo-lead: A difficulty-aware reinforcement learning approach for concise mathematical reasoning in language models.arXiv preprint arXiv:2504.09696, 2025.

StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models Grpo-lead: A difficulty-aware reinforcement learning approach for concise mathematical reasoning in language models.arXiv preprint arXiv:2504.09696, 2025

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T23:29:17.076702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:29:17.076702Z digest=sha256:c80dd094d96e612467f53365dd4ea6fe8613cd2ea608b034883bee17a18c0409

Observation 28beb4d5-350f-44e3-a2c8-bcf574f38a30 · outbound

This paper cites Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations.

StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T23:29:17.079340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:29:17.079340Z digest=sha256:df3247a9860b99aa6ff6cf94156c63f35c4ca05b88dfea13665ca6ee7332236e

Observation 9f30538b-8588-4cba-add5-7cc3256747de · outbound

This paper cites rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking.

StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T23:29:17.082217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:29:17.082217Z digest=sha256:884de9cb44b8bfa009ae7df32a1d6fe3cc141c42c2498ab6161f38a3b8e871fb

Observation 718a23d0-7989-4f4c-9022-a75a6ae2d9d8 · outbound

This paper cites R-PRM: Reasoning-Driven Process Reward Modeling.

StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models R-PRM: Reasoning-Driven Process Reward Modeling

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T23:29:17.085079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:29:17.085079Z digest=sha256:0e0ebbb75e50b175ccf506910a9ba8ddc8d2f8cb517f9586182bb95c775d150f

Observation 875fc884-4447-4b56-a5d6-a8046ecc9574 · outbound

This paper cites MM-PRM: Enhancing Multimodal Mathematical Reasoning with Scalable Step-Level Supervision.

StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models MM-PRM: Enhancing Multimodal Mathematical Reasoning with Scalable Step-Level Supervision

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T23:29:17.087874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:29:17.087874Z digest=sha256:7887bd91067aa25f724371dcc55bc29e57a7d93584e05ed8b48e7ff7ec133e2b

Observation 47e5b366-9243-453c-a533-9142e13422e2 · outbound

This paper cites Reasonflux-prm: Trajectory-aware prms for long chain-of-thought reasoning in llms.arXiv preprint arXiv:2506.18896, 2025.

StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models Reasonflux-prm: Trajectory-aware prms for long chain-of-thought reasoning in llms.arXiv preprint arXiv:2506.18896, 2025

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T23:29:17.090613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:29:17.090613Z digest=sha256:cbad05958160d96a04b918f6c4aafb144b1be010d32d8a7ad9ef348c402f4d68

Observation 73badfde-fc94-4126-a8ee-a7377e5d6975 · outbound

This paper cites JudgeLM: Fine-tuned Large Language Models are Scalable Judges.

StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models JudgeLM: Fine-tuned Large Language Models are Scalable Judges

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T23:29:17.093775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:29:17.093775Z digest=sha256:d4e5d4fe7d53aade616f8414008219f84c23fccba36b618d76e2ec6219e34f73

Observation 62a5a20b-5f9b-40b0-8555-230aa236155b · outbound

This paper cites Rm-r1: Reward modeling as reasoning.arXiv preprint arXiv:2505.02387, 2025.

StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models Rm-r1: Reward modeling as reasoning.arXiv preprint arXiv:2505.02387, 2025

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T23:29:17.096872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:29:17.096872Z digest=sha256:c8e64dbafce38ce6ed4d4d3da481557fe32205223e8afc168532a00e35cc00a8

Observation 0e9bb82c-f8a2-4e86-8f86-93fc22aa75d7 · outbound

This paper cites Gram: A generative foundation reward model for reward generalization.arXiv preprint arXiv:2506.14175, 2025.

StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models Gram: A generative foundation reward model for reward generalization.arXiv preprint arXiv:2506.14175, 2025

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T23:29:17.099452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:29:17.099452Z digest=sha256:7c96eaea7447a0d833d89928b7922568fdedaf9ae6cee5e9d33415751b4c535c

Observation fa019db4-f9cb-4ba2-9b10-625f4ed9f6b5 · outbound

This paper cites ReasonGRM: Enhancing Generative Reward Models through Large Reasoning Models.

StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models ReasonGRM: Enhancing Generative Reward Models through Large Reasoning Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T23:29:17.102230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:29:17.102230Z digest=sha256:37d09711c17735a7df922a8f3095a0fd4f39ea9374b62965745acb59f3c77cb3

Pith citing papers

No inbound Pith citation observations are available.