Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T20:29:03.988127Z
Paper Citation Record · LEDGER
As of 15 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 0 inbound Pith citation observations for arXiv:2608.09271.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T20:29:03.988127Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
48 of 48 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 48b53655-8338-455c-9072-cc43d51e5d70 · outbound
SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Maximum a posteriori policy optimisation
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation b55a5c17-6a6d-44f5-821f-3c6cfdbe5ed4 · outbound
SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation On-policy distillation of language models: Learning from self-generated mistakes
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 525ca485-55f0-49a7-b7e6-8fad8d4166b5 · outbound
SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Escaping the Verifier: Learning to Reason via Demonstrations
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c081b2c-b1cc-43ae-826a-154f1b865aa2 · outbound
SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e21b2de8-8e72-4244-a467-cbf86dfbdcbb · outbound
SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Training Verifiers to Solve Math Word Problems
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e92ab6a6-f12e-4c5c-bf29-15f1bc327298 · outbound
SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation What is the objective of reasoning with reinforcement learning?arXiv preprint arXiv:2510.13651, 2025
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f13f1dd0-43c6-4c6d-9695-30ca959afb0a · outbound
SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Imagenet: A large-scale hierarchical image database
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84f6e10f-4685-4e7f-8cc3-fb2d32f8e0f2 · outbound
SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Openvlthinker: Complex vision-language reasoning via iterative sft-rl cycles
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 20687639-1a63-447a-a10a-f63fbced52b2 · outbound
SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Cold-startreinforcementlearningwithsoftmaxpolicygradient.AdvancesinNeuralInformation Processing Systems, 30, 2017
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 26e738b6-6dd2-424a-8a0f-45140ae1f74a · outbound
SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7beda436-9439-4ca8-a003-b6981708edcb · outbound
SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation OpenThoughts: Data Recipes for Reasoning Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a17ffd8-7427-4f11-9cf7-79adc4121cf3 · outbound
SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Learning to reason for long-form story generation
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 48f9ec6e-9262-4274-afaa-6fb7501c6178 · outbound
SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation RLP: Reinforcement as a pretraining objective
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 05494d81-0622-43c3-b6f8-6e8f374f2ce4 · outbound
SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Deep residual learning for image recognition
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6cfee7c8-6f4c-48d2-9d30-2b7b50364774 · outbound
SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Shah, Qihua Dong, Zilin Xiao, Jaywon Koo, and Vicente Ordonez
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 35b152db-d9c5-4e67-be3b-5f387a809020 · outbound
SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Deepmath-103k: A large-scale, challenging, decontaminated, and verifiable mathematical dataset for advancing reasoning
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 9b185181-cef0-4488-85ec-28312844286e · outbound
SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Measuring massive multitask language understanding
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 15b122f3-3271-47d1-a85e-1056e4c2a181 · outbound
SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation MeetingBank: A benchmark dataset for meeting summarization
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c36242e-055c-4d68-baa6-1c03b07921da · outbound
SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Answer-consistent chain-of-thought reinforcement learning for multi-modal large langauge models.arXiv preprint arXiv:2510.10104, 2025
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 795c7606-06d9-4595-9f60-48d13b5bff96 · outbound
SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation OpenAI o1 System Card
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa3ffc40-e62f-4a9e-ba83-fd47bd5756da · outbound
SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Understanding r1- zero-like training: A critical perspective
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation bbf87a8c-ddbe-4498-8be6-dde95ef68651 · outbound
SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Decoupled Weight Decay Regularization
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4fd2436-2fc2-4385-b832-dbd69df5c953 · outbound
SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation On-policy distillation.Thinking Machines Lab: Connectionism, 2025
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52b88f80-ef63-4ff9-8573-7d2bbb375abc · outbound
SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation General-Reasoner: Advancing LLM Reasoning Across All Domains
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8ec9b47-7531-485b-acc9-1ae53493c044 · outbound
SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Reward augmented maximum likelihood for neural structured prediction
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a1ef6423-ea6a-4ad1-a018-93043bdfcc5d · outbound
SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Iterative reasoning preference optimization
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 51f6a1b5-835f-47b5-8420-966b7dbfdea3 · outbound
SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Rethinking the Trust Region in LLM Reinforcement Learning
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55fe61b9-3264-4cfd-b253-47e751c87af8 · outbound
SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Unresolved cited work
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 179e6939-895e-4b35-80e7-2111e926cea2 · outbound
SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Direct preference optimization: Your language model is secretly a reward model
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 3fe1534d-def2-40eb-a0e7-b2d07c77d067 · outbound
SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Unresolved cited work
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 453e334e-fed7-4919-985c-4d93a7cf781a · outbound
SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Optimal completion distillation for sequence learning
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation c449c312-c2ba-4e89-8204-79e2a2a4606b · outbound
SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Proximal Policy Optimization Algorithms
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f039705-4656-4b94-8dfe-0fa5e4a2df2b · outbound
SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0f34d5e-7215-45a8-a781-3358ef13b455 · outbound
SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation HybridFlow: A Flexible and Efficient RLHF Framework
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c91d0b00-6ba4-407d-9136-bf043f5db81e · outbound
SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1e7ee22-279d-4e3f-b327-ce76786d6063 · outbound
SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Maximum likelihood reinforcement learning.arXiv preprint arXiv:2602.02710, 2026
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3f0d3d6-0693-4bf6-9d57-009fef5ea153 · outbound
SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Vl-rethinker: Incentivizing self- reflection of vision-language models with reinforcement learning
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 5b43ec2a-b9e5-4ed5-a44e-8c85cc8456b3 · outbound
SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Sota with less: Mcts-guided sample selection for data-efficient visual reasoning self-improvement
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 002ece37-9b4a-47f5-9574-3e38074cefb9 · outbound
SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Thoughts are all over the place: On the underthinking of long reasoning models
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 5eaaa617-1b79-453c-91ba-0f1d303e262f · outbound
SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Sportr: A benchmark for multimodal large language model reasoning in sports, 2026
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d40d4dbe-5ab3-4b13-98ad-cca8b8785645 · outbound
SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Proxythinker: Test-time guidance through small visual reasoners
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 2e12425e-c792-4fb2-82e7-bcae2d64769a · outbound
SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Qwen3 Technical Report
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3a193c2-0a9a-41db-833d-8db2ba8ccc8a · outbound
SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation DAPO: An open-source LLM reinforcement learning system at scale
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d210592-509d-4f11-8a92-9c657cc31ef6 · outbound
SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation STar: Bootstrapping reasoning with reasoning
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b00b7bd6-e5af-4951-b3f8-389e1fe0f17d · outbound
SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation KAT-V1: Kwai-AutoThink Technical Report
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b41180f-ec8e-476d-ad7e-6b1f00cae5c5 · outbound
SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Reinforcing general reasoning without verifiers
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation bb7d2340-b2b4-46a5-866f-c891faa92bc3 · outbound
SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Unresolved cited work
Reference 2024
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 0ab855d8-2657-4615-997b-dada58adeac7 · outbound
SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Please reason step by step, and put your final answer within \boxed{}
Reference 2026
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
No inbound Pith citation observations are available.