Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T23:29:17.102230Z
Paper Citation Record · LEDGER
As of 23 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 0 inbound Pith citation observations for arXiv:2508.05383.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T23:29:17.102230Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
42 of 42 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 80ad6a69-6d28-470a-a1e7-4609fbb6fd38 · outbound
StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models Seed1.5-VL Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66f41efe-fd1e-40ee-b23d-52eda21325e7 · outbound
StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models Learning to reason with llms, 2024
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7a78e0a-2303-4ced-8e0a-f587fe6416ab · outbound
StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models Gemini 2.5: Our most intelligent ai model, 2025
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7ee3099-9567-4146-a212-0661e622810a · outbound
StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 773c28a8-45e9-49ff-ab11-7031579c1417 · outbound
StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In European Conference on Computer Vision, pages 169–186
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd97cdbd-84ef-44a5-8e05-c975d9995c0a · outbound
StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models SOLIDGEO: Measuring Multimodal Spatial Math Reasoning in Solid Geometry
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17d1e031-80e0-44bd-a927-b968a3401fa6 · outbound
StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf2837e5-30b2-4a4f-84b8-8c7912e0264c · outbound
StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models Polymath: A Challenging Multi-modal Mathematical Reasoning Benchmark
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65d639d0-9eae-4eb5-a164-98cfa1390041 · outbound
StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models Mathscape: Evaluating mllms in multimodal math scenarios through a hierarchical benchmark
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18cd1628-2a4b-4354-868e-acbceb063b63 · outbound
StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models R-Bench: Are your Large Multimodal Model Robust to Real-world Corruptions?
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9231b1b3-8ec4-4a54-9baa-29eb6e09b9dc · outbound
StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad0ade1a-1244-40c8-b05e-4bc7fdd1d6a3 · outbound
StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models Scibench: Evaluating college-level scientific problem-solving abilities of large language models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4bbd3274-5a24-4402-afc5-6da961b84d39 · outbound
StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models Scemqa: A scientific college entrance level multimodal question answering benchmark
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29b6a9d1-578d-4c3e-a2fa-cfa6d2e16e32 · outbound
StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models VNHSGE: VietNamese High School Graduation Examination Dataset for Large Language Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 242b9482-5532-4ca9-b0a8-f34ff3e88156 · outbound
StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models Chemvlm: Exploring the power of multimodal large language models in chemistry area
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c71bf73-9339-4b45-8e0b-b473c37bc94d · outbound
StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models ChemQuests: A Curated Chemistry Question-Answer Database Extracted from ChemRxiv papers
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d9b7562-70e3-4b43-8bd4-31a4e10348c3 · outbound
StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models Evaluating the symbol binding ability of large language models for multiple-choice questions in vietnamese general education
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48e8fdc2-78e4-493b-8bf5-b8d8d4134587 · outbound
StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc2651ec-3676-420e-bd47-10df60088fe7 · outbound
StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models Reason-rft: Reinforcement fine-tuning for visual reasoning.arXiv preprint arXiv:2503.20752, 2025
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc0e6dd3-3e5a-460d-9af1-a39a376e33f0 · outbound
StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e7e4c2b-afef-4f04-8e41-48b0bdee1539 · outbound
StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models Visual-RFT: Visual Reinforcement Fine-Tuning
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca86a00a-d62e-4365-a261-53a173b04924 · outbound
StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models Perception-R1: Pioneering Perception Policy with Reinforcement Learning
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e138b8bc-4358-41a5-ad59-8ba1c35ef888 · outbound
StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models R1-Zero's "Aha Moment" in Visual Reasoning on a 2B Non-SFT Model
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f123ce55-4bbd-4388-82a6-b03a540296aa · outbound
StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d551d9dd-76e3-45e6-81fb-cf68ba2e104f · outbound
StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56d5df1a-0304-408c-9ae2-c3c583c00b7b · outbound
StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models Demystifying Long Chain-of-Thought Reasoning in LLMs
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd7b64ca-ab64-45e6-ad91-8132ecccf183 · outbound
StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models Open r1: A fully open reproduction of deepseek-r1, january 2025
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a00dfc3d-725a-4f49-ba37-4f81fabc7e24 · outbound
StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models Rethinking RL Scaling for Vision Language Models: A Transparent, From-Scratch Framework and Comprehensive Evaluation Scheme
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdd19e39-4745-4c4a-990f-a61fd25109b0 · outbound
StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da953aed-ea43-48d5-a0f7-abd7a5884468 · outbound
StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0cda74db-115e-4711-927d-1919d9f1a0fb · outbound
StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8969b9a4-7393-42e5-95c1-4d046f3c9338 · outbound
StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models RLPR: Extrapolating RLVR to General Domains without Verifiers
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b1af818-1681-4458-bc65-be417c1fb148 · outbound
StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models Grpo-lead: A difficulty-aware reinforcement learning approach for concise mathematical reasoning in language models.arXiv preprint arXiv:2504.09696, 2025
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28beb4d5-350f-44e3-a2c8-bcf574f38a30 · outbound
StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f30538b-8588-4cba-add5-7cc3256747de · outbound
StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 718a23d0-7989-4f4c-9022-a75a6ae2d9d8 · outbound
StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models R-PRM: Reasoning-Driven Process Reward Modeling
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 875fc884-4447-4b56-a5d6-a8046ecc9574 · outbound
StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models MM-PRM: Enhancing Multimodal Mathematical Reasoning with Scalable Step-Level Supervision
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47e5b366-9243-453c-a533-9142e13422e2 · outbound
StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models Reasonflux-prm: Trajectory-aware prms for long chain-of-thought reasoning in llms.arXiv preprint arXiv:2506.18896, 2025
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73badfde-fc94-4126-a8ee-a7377e5d6975 · outbound
StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models JudgeLM: Fine-tuned Large Language Models are Scalable Judges
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62a5a20b-5f9b-40b0-8555-230aa236155b · outbound
StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models Rm-r1: Reward modeling as reasoning.arXiv preprint arXiv:2505.02387, 2025
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e9bb82c-f8a2-4e86-8f86-93fc22aa75d7 · outbound
StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models Gram: A generative foundation reward model for reward generalization.arXiv preprint arXiv:2506.14175, 2025
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa019db4-f9cb-4ba2-9b10-625f4ed9f6b5 · outbound
StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models ReasonGRM: Enhancing Generative Reward Models through Large Reasoning Models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.