Pith. sign in

Paper Citation Record · LEDGER

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation

As of 15 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 0 inbound Pith citation observations for arXiv:2608.09271.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.09271 v1

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T20:29:03.988127Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

48 of 48 outbound references displayed

  • verified exact2
  • verified fuzzy18
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 48b53655-8338-455c-9072-cc43d51e5d70 · outbound

This paper cites Maximum a posteriori policy optimisation.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Maximum a posteriori policy optimisation

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:29:05.154815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:29:03.752867Z digest=sha256:c9e73e0f284b285758a46018cb933601d158a7ee99b67e6ec0bab18fd9c409d7

Observation b55a5c17-6a6d-44f5-821f-3c6cfdbe5ed4 · outbound

This paper cites On-policy distillation of language models: Learning from self-generated mistakes.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation On-policy distillation of language models: Learning from self-generated mistakes

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T20:29:03.758603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:29:03.758603Z digest=sha256:a3e97bd6448b6cdb8fb339b12ac60b666bd8be6ed0683618ef568e9332ffaff4

Observation 525ca485-55f0-49a7-b7e6-8fad8d4166b5 · outbound

This paper cites Escaping the Verifier: Learning to Reason via Demonstrations.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Escaping the Verifier: Learning to Reason via Demonstrations

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T20:29:03.762979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:29:03.762979Z digest=sha256:50f755f5c5f85c5f052031c259a35f972aac78f7d836a3d1e5c2a362e2200603

Observation 9c081b2c-b1cc-43ae-826a-154f1b865aa2 · outbound

This paper cites MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T20:29:03.767752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:29:03.767752Z digest=sha256:7ecefdfae5883b664239f15ad3ed0b5d2b144931e2652e905b0afd5a6e1100d7

Observation e21b2de8-8e72-4244-a467-cbf86dfbdcbb · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Training Verifiers to Solve Math Word Problems

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T20:29:03.773626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:29:03.773626Z digest=sha256:33e01a5d06d55704c192735a1c36889bc1725031f8342ed81c8d1bcb71149e7a

Observation e92ab6a6-f12e-4c5c-bf29-15f1bc327298 · outbound

This paper cites What is the objective of reasoning with reinforcement learning?arXiv preprint arXiv:2510.13651, 2025.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation What is the objective of reasoning with reinforcement learning?arXiv preprint arXiv:2510.13651, 2025

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T20:29:03.779243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:29:03.779243Z digest=sha256:acbb8a2c3caf9f7b023d2c1262997cf2ae339289d444fac90adaea57ae1775f6

Observation f13f1dd0-43c6-4c6d-9695-30ca959afb0a · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Imagenet: A large-scale hierarchical image database

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T20:29:03.783941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:29:03.783941Z digest=sha256:a9b23d2e5fc2fe801d66fa39f607a7cc0711cb69c57911985b4ef166b226b9c8

Observation 84f6e10f-4685-4e7f-8cc3-fb2d32f8e0f2 · outbound

This paper cites Openvlthinker: Complex vision-language reasoning via iterative sft-rl cycles.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Openvlthinker: Complex vision-language reasoning via iterative sft-rl cycles

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:29:05.111003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:29:03.790434Z digest=sha256:cfe32fcc2150b9781565afbc4e899e9e51def44615bd1b747f92635e59de63d5

Observation 20687639-1a63-447a-a10a-f63fbced52b2 · outbound

This paper cites Cold-startreinforcementlearningwithsoftmaxpolicygradient.AdvancesinNeuralInformation Processing Systems, 30, 2017.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Cold-startreinforcementlearningwithsoftmaxpolicygradient.AdvancesinNeuralInformation Processing Systems, 30, 2017

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:29:05.094454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:29:03.794665Z digest=sha256:1525b25d39424fc74091ecca5e35f3d0a3ee877aa80d728fd48e49911b8bed64

Observation 26e738b6-6dd2-424a-8a0f-45140ae1f74a · outbound

This paper cites Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T20:29:03.802470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:29:03.802470Z digest=sha256:8dfac744e4ffb915da8a096af69e2ac6e188bc612296c8766d5b2b3b20322e11

Observation 7beda436-9439-4ca8-a003-b6981708edcb · outbound

This paper cites OpenThoughts: Data Recipes for Reasoning Models.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation OpenThoughts: Data Recipes for Reasoning Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T20:29:03.807561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:29:03.807561Z digest=sha256:d2adc25d9cf3cef4409d242dc052d543d3ef361617189850845856eac1454e19

Observation 4a17ffd8-7427-4f11-9cf7-79adc4121cf3 · outbound

This paper cites Learning to reason for long-form story generation.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Learning to reason for long-form story generation

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:29:05.075028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:29:03.813234Z digest=sha256:9c51c8817d900915c23d3e2af64e6fabaca7a1352617b2c70292682f01f48bba

Observation 48f9ec6e-9262-4274-afaa-6fb7501c6178 · outbound

This paper cites RLP: Reinforcement as a pretraining objective.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation RLP: Reinforcement as a pretraining objective

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:29:05.054235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:29:03.817892Z digest=sha256:c7b90e52b2be565eccc7c7feafcaa2dcc35c11e52d94ff3727cdfd3715636c72

Observation 05494d81-0622-43c3-b6f8-6e8f374f2ce4 · outbound

This paper cites Deep residual learning for image recognition.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Deep residual learning for image recognition

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T20:29:03.823062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:29:03.823062Z digest=sha256:61a98f0db3c2f8887267c8fccc18154003db17e52890d5c175a42e2b62dea12e

Observation 6cfee7c8-6f4c-48d2-9d30-2b7b50364774 · outbound

This paper cites Shah, Qihua Dong, Zilin Xiao, Jaywon Koo, and Vicente Ordonez.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Shah, Qihua Dong, Zilin Xiao, Jaywon Koo, and Vicente Ordonez

Reference 15

Resolution
verified exact
raw_fallback, observed 2026-08-11T20:29:04.545909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:29:03.827802Z digest=sha256:995444bc2986ca8aa3c986c7d841b027249c2d88fac0644c7bd42609ebfd0b6b

Observation 35b152db-d9c5-4e67-be3b-5f387a809020 · outbound

This paper cites Deepmath-103k: A large-scale, challenging, decontaminated, and verifiable mathematical dataset for advancing reasoning.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Deepmath-103k: A large-scale, challenging, decontaminated, and verifiable mathematical dataset for advancing reasoning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:29:05.022024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:29:03.832205Z digest=sha256:45ad539faf32f75e4b13da4a0e8c018362a364946fdfd4fd604bf3fef75466e3

Observation 9b185181-cef0-4488-85ec-28312844286e · outbound

This paper cites Measuring massive multitask language understanding.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Measuring massive multitask language understanding

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:29:05.003190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:29:03.837328Z digest=sha256:c8607387d0fa25d901b3c68910558585acf279b877082f9a457d9649e741798f

Observation 15b122f3-3271-47d1-a85e-1056e4c2a181 · outbound

This paper cites MeetingBank: A benchmark dataset for meeting summarization.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation MeetingBank: A benchmark dataset for meeting summarization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T20:29:03.841893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:29:03.841893Z digest=sha256:4cc01812c72e9f424b6f160d8a8027addc39bdb76afc8862d34b0f726ab24627

Observation 3c36242e-055c-4d68-baa6-1c03b07921da · outbound

This paper cites Answer-consistent chain-of-thought reinforcement learning for multi-modal large langauge models.arXiv preprint arXiv:2510.10104, 2025.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Answer-consistent chain-of-thought reinforcement learning for multi-modal large langauge models.arXiv preprint arXiv:2510.10104, 2025

Reference 19

Resolution
verified exact
raw_fallback, observed 2026-08-11T20:29:04.467803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:29:03.846255Z digest=sha256:9e135b0368476ff956fccd9aa6ad06def1fbdcbde6377ecf9bfbfa80d31f3780

Observation 795c7606-06d9-4595-9f60-48d13b5bff96 · outbound

This paper cites OpenAI o1 System Card.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation OpenAI o1 System Card

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T20:29:03.849865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:29:03.849865Z digest=sha256:ae14566a504f02eb2a05e6ed02981575e9a13ba0ca6c9e2f1521e116574d2153

Observation aa3ffc40-e62f-4a9e-ba83-fd47bd5756da · outbound

This paper cites Understanding r1- zero-like training: A critical perspective.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Understanding r1- zero-like training: A critical perspective

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:29:04.984900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:29:03.854986Z digest=sha256:3525ed262140a4c279f2092da05389426ed1dde82638d4c7c17da23a665d310b

Observation bbf87a8c-ddbe-4498-8be6-dde95ef68651 · outbound

This paper cites Decoupled Weight Decay Regularization.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Decoupled Weight Decay Regularization

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T20:29:03.858538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:29:03.858538Z digest=sha256:a009d78c4be7023265dfdd0ec452cd24e77c543b14dc9367c4003cbb874a6f80

Observation e4fd2436-2fc2-4385-b832-dbd69df5c953 · outbound

This paper cites On-policy distillation.Thinking Machines Lab: Connectionism, 2025.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation On-policy distillation.Thinking Machines Lab: Connectionism, 2025

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T20:29:03.861966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:29:03.861966Z digest=sha256:5176b7625223e13efdac4b216be07f1b31c596b114842e8d9cf319cd9591667f

Observation 52b88f80-ef63-4ff9-8573-7d2bbb375abc · outbound

This paper cites General-Reasoner: Advancing LLM Reasoning Across All Domains.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation General-Reasoner: Advancing LLM Reasoning Across All Domains

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T20:29:03.866650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:29:03.866650Z digest=sha256:060fee9410a9efdb40b042b8dfb05cab453cac49432e4db9bc8c4cea13c449c5

Observation f8ec9b47-7531-485b-acc9-1ae53493c044 · outbound

This paper cites Reward augmented maximum likelihood for neural structured prediction.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Reward augmented maximum likelihood for neural structured prediction

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:29:04.967759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:29:03.872377Z digest=sha256:f3b8e841d0e078ece85d12026db58f08bc845d1096091d672100dc8d1c33b4af

Observation a1ef6423-ea6a-4ad1-a018-93043bdfcc5d · outbound

This paper cites Iterative reasoning preference optimization.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Iterative reasoning preference optimization

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:29:04.951654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:29:03.876731Z digest=sha256:d62c435b5131f5defb29deb534175bf49c7f42742c0741b0915f477257754d83

Observation 51f6a1b5-835f-47b5-8420-966b7dbfdea3 · outbound

This paper cites Rethinking the Trust Region in LLM Reinforcement Learning.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Rethinking the Trust Region in LLM Reinforcement Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T20:29:03.881817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:29:03.881817Z digest=sha256:639180e2be0b89492ef0c47609c47ae5ffb1d64eac22e09f606a222330e858f8

Observation 55fe61b9-3264-4cfd-b253-47e751c87af8 · outbound

This paper cites an unresolved cited work.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T20:29:03.889681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:29:03.889681Z digest=sha256:a187e9b5853ad25a430417d5f62eabf5b34e835f96cd3555405cb37cfe36ab4b

Observation 179e6939-895e-4b35-80e7-2111e926cea2 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Direct preference optimization: Your language model is secretly a reward model

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:29:04.911990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:29:03.900830Z digest=sha256:cd1c1369cfce43154250a92c501bb46ef675f5a76fff914fb86b5d8daba05621

Observation 3fe1534d-def2-40eb-a0e7-b2d07c77d067 · outbound

This paper cites an unresolved cited work.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-11T20:29:04.890952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:29:03.905739Z digest=sha256:ded3f0b3665376b5363a6b1aa970bc44d06a927110905e3cf43776e527bf79cb

Observation 453e334e-fed7-4919-985c-4d93a7cf781a · outbound

This paper cites Optimal completion distillation for sequence learning.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Optimal completion distillation for sequence learning

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:29:04.856233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:29:03.916418Z digest=sha256:bc66e0ecd62ee482ac665d53af515d6ad36aa8c6fa1c048bc753598f2d5f272b

Observation c449c312-c2ba-4e89-8204-79e2a2a4606b · outbound

This paper cites Proximal Policy Optimization Algorithms.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Proximal Policy Optimization Algorithms

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T20:29:03.921502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:29:03.921502Z digest=sha256:4db2f0a9933348d2f0ffb5f701e4a0793fadca6fdc272f8dcbe191905dd118fc

Observation 8f039705-4656-4b94-8dfe-0fa5e4a2df2b · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T20:29:03.926196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:29:03.926196Z digest=sha256:244d8c8a17398158d72e1c32a1a1003f018917cd1d27cde2e50ba45a5dbbf1d3

Observation a0f34d5e-7215-45a8-a781-3358ef13b455 · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation HybridFlow: A Flexible and Efficient RLHF Framework

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T20:29:03.931344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:29:03.931344Z digest=sha256:d95dbd6ecf9a665f4b6cca9634c7ed2e8402482690354a1205319350764ee49e

Observation c91d0b00-6ba4-407d-9136-bf043f5db81e · outbound

This paper cites Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T20:29:03.935373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:29:03.935373Z digest=sha256:16e42a15195a7a66d242d511d0085807c40b644685179e55024242b308e43762

Observation d1e7ee22-279d-4e3f-b327-ce76786d6063 · outbound

This paper cites Maximum likelihood reinforcement learning.arXiv preprint arXiv:2602.02710, 2026.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Maximum likelihood reinforcement learning.arXiv preprint arXiv:2602.02710, 2026

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T20:29:03.940136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:29:03.940136Z digest=sha256:086535d09c9a6d83a11e94a5572c1ece6a10829456ffb36591782003e2ea43eb

Observation c3f0d3d6-0693-4bf6-9d57-009fef5ea153 · outbound

This paper cites Vl-rethinker: Incentivizing self- reflection of vision-language models with reinforcement learning.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Vl-rethinker: Incentivizing self- reflection of vision-language models with reinforcement learning

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:29:04.844077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:29:03.945359Z digest=sha256:160c6285c2035a8e00e5d67469113d94fedc61f27a711a72640b60c6e962beb4

Observation 5b43ec2a-b9e5-4ed5-a44e-8c85cc8456b3 · outbound

This paper cites Sota with less: Mcts-guided sample selection for data-efficient visual reasoning self-improvement.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Sota with less: Mcts-guided sample selection for data-efficient visual reasoning self-improvement

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:29:04.830899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:29:03.949476Z digest=sha256:df0ccdbecf6be2dd95b29ff15259e4f2cd0b46e3f30850afbe97c3c75c10589b

Observation 002ece37-9b4a-47f5-9574-3e38074cefb9 · outbound

This paper cites Thoughts are all over the place: On the underthinking of long reasoning models.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Thoughts are all over the place: On the underthinking of long reasoning models

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:29:04.812235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:29:03.953507Z digest=sha256:fb2885cbb2967048c3d235124e9108852c3def6348509c812c64e34bc5052751

Observation 5eaaa617-1b79-453c-91ba-0f1d303e262f · outbound

This paper cites Sportr: A benchmark for multimodal large language model reasoning in sports, 2026.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Sportr: A benchmark for multimodal large language model reasoning in sports, 2026

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T20:29:03.957406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:29:03.957406Z digest=sha256:302bfe03226f6e3435a7db4836b8044ab40a89c654c9062fea59fe0e094bd893

Observation d40d4dbe-5ab3-4b13-98ad-cca8b8785645 · outbound

This paper cites Proxythinker: Test-time guidance through small visual reasoners.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Proxythinker: Test-time guidance through small visual reasoners

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:29:04.797570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:29:03.961440Z digest=sha256:6e64295a9ce246dd6efafa9b0eac217b67d065b74978f0f5dda7e210d09175fd

Observation 2e12425e-c792-4fb2-82e7-bcae2d64769a · outbound

This paper cites Qwen3 Technical Report.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Qwen3 Technical Report

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T20:29:03.965787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:29:03.965787Z digest=sha256:72ad7c94b9b3a2d4e40ceb3a7f009dedfb41cf3b4e2526608e78f7321c60509f

Observation b3a193c2-0a9a-41db-833d-8db2ba8ccc8a · outbound

This paper cites DAPO: An open-source LLM reinforcement learning system at scale.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation DAPO: An open-source LLM reinforcement learning system at scale

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T20:29:03.970117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:29:03.970117Z digest=sha256:49f2e3b21268cc5e56bc06c9ee94facd5c567ec73d3fcb4af1d95b2bb5c5d0f9

Observation 0d210592-509d-4f11-8a92-9c657cc31ef6 · outbound

This paper cites STar: Bootstrapping reasoning with reasoning.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation STar: Bootstrapping reasoning with reasoning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T20:29:03.974838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:29:03.974838Z digest=sha256:6f5a7973ae8f67012a2a820ffedd974adb9c15be58b6c0e51936020a50c45739

Observation b00b7bd6-e5af-4951-b3f8-389e1fe0f17d · outbound

This paper cites KAT-V1: Kwai-AutoThink Technical Report.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation KAT-V1: Kwai-AutoThink Technical Report

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T20:29:03.979079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:29:03.979079Z digest=sha256:4ac9c870fccdc9c457f7fa36c5d669c8a5c92bcef28af6af67ae4fa71bba0fa2

Observation 7b41180f-ec8e-476d-ad7e-6b1f00cae5c5 · outbound

This paper cites Reinforcing general reasoning without verifiers.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Reinforcing general reasoning without verifiers

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:29:04.757871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:29:03.983724Z digest=sha256:84f18fbbeecdb847e0fec362f99224dbfbd08325969fbd56873e7d28a32281fb

Observation bb7d2340-b2b4-46a5-866f-c891faa92bc3 · outbound

This paper cites an unresolved cited work.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Unresolved cited work

Reference 2024

Resolution
unresolved
raw_fallback, observed 2026-08-11T20:29:04.871224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:29:03.911135Z digest=sha256:3fbb1ef75c43f72293d14d6e9d9270e239f80e45a6f3cb6c7cf3540dec25285a

Observation 0ab855d8-2657-4615-997b-dada58adeac7 · outbound

This paper cites Please reason step by step, and put your final answer within \boxed{}.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Please reason step by step, and put your final answer within \boxed{}

Reference 2026

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:29:04.742339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:29:03.988127Z digest=sha256:2994867a2f0af7ec7d322f2567091f05c96734b94ece3700092231a1a72be9d0

Pith citing papers

No inbound Pith citation observations are available.