Pith. sign in

Paper Citation Record · LEDGER

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

As of 19 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 35 inbound Pith citation observations for arXiv:2505.12346.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.12346 v1

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:40:13.165308Z

measured 77 of 77 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 35 of 35 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:07:05.170825Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T13:49:52.427933Z

Reference resolution

42 of 42 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved34
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2ee99a11-40e6-45e4-bc87-f4d4f5fe37d7 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.011959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.011959Z digest=sha256:5bfc997360f3eea6fb4a341603ee58f6c81e6d4f9026215e1619eecc28c0e35a

Observation 93197f6b-9e95-4cab-a620-cd82fcce7263 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.016281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.016281Z digest=sha256:48422a237c09c58e249220fd3d5257a2422ea58b3226b0be3fa7cfa13dbd6d3f

Observation 806637a3-f613-4a80-90d1-8ea4896e31ef · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Understanding R1-Zero-Like Training: A Critical Perspective

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.020339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.020339Z digest=sha256:827bda4654ef1093f6eec2864c81e288234d67c12654f3c54af0256f8c1a2f1a

Observation d72c7687-aa36-4f31-9f98-c9dde95e9424 · outbound

This paper cites GPT-4 Technical Report.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization GPT-4 Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.024490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.024490Z digest=sha256:4badb3387ac3ad3fd81572af9724a8f6df9911c10e07921eb77f56ba06ad15ad

Observation 060360f2-1a43-4706-b4f0-d4dbcd2858d1 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.028343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.028343Z digest=sha256:94bb8907514a439d36dd92b4d9c9c5340d8ef3ace026bc9f63495e058ba5e96b

Observation 2502b848-428b-4b18-bfe3-feaa41f62102 · outbound

This paper cites Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.032436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.032436Z digest=sha256:2898544084255c8bcedbdaeddca657d47ed08cb4f94a96172695eab0c8fdd021

Observation f557445a-6657-4e71-8261-a5188b2cbcd7 · outbound

This paper cites Cppo: Accelerating the training of group relative policy optimization-based reasoning models.arXiv preprint arXiv:2503.22342, 2025.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Cppo: Accelerating the training of group relative policy optimization-based reasoning models.arXiv preprint arXiv:2503.22342, 2025

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.036887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.036887Z digest=sha256:bcfb60e788978bc7f0566b514729cb8853300f7eb1d6ab2aef38131287f24feb

Observation 7825dc23-4a1e-4b4f-85c6-2153f653c563 · outbound

This paper cites A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.040337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.040337Z digest=sha256:a697cec7d64490e18e506e7bf8d8ec40296269ebdfdd6b2d9dca2a73ca6769b4

Observation 788c11cf-4a69-464b-a656-4f9e647dda6b · outbound

This paper cites VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.044129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.044129Z digest=sha256:bda0923ac5ce6c0fa69ced1ceb975e1fc779e79d4a4563912a8766076b1beffc

Observation 2c595027-37fa-465e-83eb-0a25c13f14b6 · outbound

This paper cites Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.047975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.047975Z digest=sha256:327155b1cf4e8dc903883a8b6facdcf87fcef55507756df81f4f6d84e78eb053

Observation 8f197b82-6299-4f5e-ba2f-a2d960eaec36 · outbound

This paper cites Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.051993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.051993Z digest=sha256:4fdfa4564a9ef55abd6b065dae9c456fdcacd2be9334469e55c4ca4cb4a0ca65

Observation 3e992d9d-dbcf-4ad6-ae0e-c713228e4c24 · outbound

This paper cites OpenAI o1 System Card.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization OpenAI o1 System Card

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.056166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.056166Z digest=sha256:b23e6a6d52ad70ce6f1194feddb26a34641bcfd62d6123619c88654a3bfec116

Observation d74b4e0c-42e0-49e2-a291-df256a281dc7 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Gemini: A Family of Highly Capable Multimodal Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.059812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.059812Z digest=sha256:92acf07d4831961874ef7bf5e3bc7652cc5f4348c36a36677bd1a15a1645a4a6

Observation 26c306b2-9fc6-4406-96a2-3acc9fc7fdf8 · outbound

This paper cites The claude 3 model family: Opus, sonnet, haiku.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization The claude 3 model family: Opus, sonnet, haiku

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:40:13.847921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:40:13.063624Z digest=sha256:e0a7d5d78ebe9c84efc071f195da9e950a5f76a4c55e45dd5fbf6d13df8fd73f

Observation 5677145c-b958-48a5-8a0f-6cfe2a052da7 · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.067375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.067375Z digest=sha256:f943c01218c1af2b9547945e456a8b81651e40f0380813da3cac2bf103887b0f

Observation 17b33c8d-64ff-47c5-883b-a9402670cbb0 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.071250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.071250Z digest=sha256:096543425555302fcab712a69def11a622a6a0c97188cb7983d6379f6acb0968

Observation c9ac5c6f-f0bf-4be6-8bb9-114314a10633 · outbound

This paper cites Gpg: A simple and strong reinforcement learning baseline for model reasoning.arXiv preprint arXiv:2504.02546, 2025.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Gpg: A simple and strong reinforcement learning baseline for model reasoning.arXiv preprint arXiv:2504.02546, 2025

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.074768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.074768Z digest=sha256:1ed1f6d999be22fb159a4539714baeae3eb73e8a3b644764f34e028e2e01a767

Observation 4d41709b-58e5-4e78-9b78-03a7a7b7d1a2 · outbound

This paper cites Grpo-lead: A difficulty-aware reinforcement learn- ing approach for concise mathematical reasoning in language models.arXiv preprint arXiv:2504.09696, 2025.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Grpo-lead: A difficulty-aware reinforcement learn- ing approach for concise mathematical reasoning in language models.arXiv preprint arXiv:2504.09696, 2025

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.078247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.078247Z digest=sha256:bc267723c575dbe44ee56d0c1c99d077d6bb78446b8ef5e52e59f2f499341151

Observation e58237ea-71b5-4bb8-b946-dbe383eb6477 · outbound

This paper cites SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.081592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.081592Z digest=sha256:c855d99c04bd0267136ff6dbedb8f3ecc575d866d430871d7ab032896430c2c5

Observation 411aa082-0154-4bea-993d-b0c547b36a80 · outbound

This paper cites Detecting hallucinations in large language models using semantic entropy.Nature, 630(8017):625–630, 2024.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Detecting hallucinations in large language models using semantic entropy.Nature, 630(8017):625–630, 2024

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.085686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.085686Z digest=sha256:7bdae73e5ba7c73a3eff39dea3cdb3637409a871c47b613cc2048768425aea5d

Observation 9da99666-9fcb-4d92-a774-359878e57ad3 · outbound

This paper cites Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:40:13.829638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:40:13.089360Z digest=sha256:36970b0d89584e6149adcc1d420be5bbf68e98b2ba7b21af933597feb401b0ab

Observation 62ca4571-c604-40a6-98ad-08bf2d5c3b01 · outbound

This paper cites Griffiths, Yuan Cao, and Karthik R Narasimhan.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Griffiths, Yuan Cao, and Karthik R Narasimhan

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:40:13.817290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:40:13.092868Z digest=sha256:6758503fc0894dcc2cc9468c9b10d02c3d45b72f436b23f93d7d87b14f61c771

Observation 4de2cb60-2685-4eb1-94b8-529d495a03cc · outbound

This paper cites Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:40:13.805391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:40:13.096326Z digest=sha256:5a7405c8a8776c972415da9956c2624a2b2b8cefce0449ae414de5b3350ce15d

Observation fe917928-6605-40af-9bea-2bc25baec380 · outbound

This paper cites Semantic Entropy Probes: Robust and Cheap Hallucination Detection in LLMs.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Semantic Entropy Probes: Robust and Cheap Hallucination Detection in LLMs

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.099891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.099891Z digest=sha256:b9427d18833ee37c2b65d2bac2657c5d06535f8d3bf97d56366f4e262e04e39e

Observation 1cb3a471-1774-4bd8-968a-38b86e82858b · outbound

This paper cites Curriculum learning.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Curriculum learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.103957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.103957Z digest=sha256:91c4d7945a3f48a772dab93d4595102321b67281787a6ba45cba62f73cf3c4bb

Observation 656e5340-d350-417b-b530-0b29a7b9794d · outbound

This paper cites A mathematical theory of communication.The Bell system technical journal, 27(3):379–423, 1948.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization A mathematical theory of communication.The Bell system technical journal, 27(3):379–423, 1948

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.107434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.107434Z digest=sha256:513f35ebd00800d37c207881bf7d7d7fc657dfbcc25949c1d41a34a942982c97

Observation 0597bbb4-e21c-4d68-a9f7-20826ecbbe40 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Chain-of-thought prompting elicits reasoning in large language models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.111000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.111000Z digest=sha256:cdd3fcd51cc8d90d4697038cbfbb8181e937a358a2eaefa613a701ec04a82226

Observation 76b47959-eda1-4b51-91e6-1a27791f76c6 · outbound

This paper cites LIMO: Less is More for Reasoning.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization LIMO: Less is More for Reasoning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.115113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.115113Z digest=sha256:c8f595cf8df472f16f6a2b54676cffb21042665783fd208c8b2c339ca5a1f415

Observation 80079864-e5b5-42e5-8828-cfe4073c139c · outbound

This paper cites ReST- MCTS*: LLM self-training via process reward guided tree search.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization ReST- MCTS*: LLM self-training via process reward guided tree search

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:40:13.771843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:40:13.118997Z digest=sha256:b58a86323c20a04c9551c5e0ab037e5fd98da6bb2e144b3952772ba8107dca58

Observation be79bc08-9e40-47b3-818d-42df74bdf59f · outbound

This paper cites Proximal Policy Optimization Algorithms.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Proximal Policy Optimization Algorithms

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.122315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.122315Z digest=sha256:b6794a1da91ef9cb352de42db8ddcf503429f736a6409b6161199f1dc4250cb9

Observation 53cb17cb-c18c-4078-8e31-ddab9477267e · outbound

This paper cites Open r1: A fully open reproduction of deepseek-r1, January 2025.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Open r1: A fully open reproduction of deepseek-r1, January 2025

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.125943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.125943Z digest=sha256:e8e22b0ed6f667dec874187df980df097b3d513bb732e95bcc77ace97cfe8afb

Observation 3cc60ea2-ddc0-474e-9acc-a62709c82019 · outbound

This paper cites Measuring mathematical problem solving with the MATH dataset.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Measuring mathematical problem solving with the MATH dataset

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:40:13.751927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:40:13.129548Z digest=sha256:e13810ebe727644d94d19cd8629a0f097c06c3d237ebb65668af8b0d90b6b173

Observation 80a5c625-2dd1-43d7-bb23-158363ec38c9 · outbound

This paper cites Solving quantitative reasoning problems with language models.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Solving quantitative reasoning problems with language models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:40:13.740603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:40:13.132960Z digest=sha256:3dff3e635e1fbe0e9c0bf69fb019e0fd7e7e8f2af1784a365b4d58a1bf6ef133

Observation 3c608ec6-defa-464d-9b97-74487315f1df · outbound

This paper cites Olympicarena: Benchmarking multi-discipline cognitive reasoning for superintelligent AI.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Olympicarena: Benchmarking multi-discipline cognitive reasoning for superintelligent AI

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:40:13.728314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:40:13.136266Z digest=sha256:e735c77050c60ad9941ffe672107c34de5d28a59326fc198b1c18c11d562a462

Observation ae2d74c8-0d7e-42d9-83b0-244a2a88dffc · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.139709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.139709Z digest=sha256:fd50dd049b7c195c7ab21f9f4a5437c3df48f86f86e2e4deb4ddd2bf4e78d072

Observation 879da7f4-6451-4c9d-8d98-f5e0fac47948 · outbound

This paper cites SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.143452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.143452Z digest=sha256:30bff222404d72b1aef1731859f851a5407f2d1f2718006b11afd3f97a2fdc5c

Observation f2527d3e-6789-464e-b0c6-c7f67fa7a817 · outbound

This paper cites Advancing LLM Reasoning Generalists with Preference Trees.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Advancing LLM Reasoning Generalists with Preference Trees

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.147339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.147339Z digest=sha256:151d2b21a4487c3c899679fb767f624d00b2a104eea9ffae98edf0d71f3e790f

Observation e18e2a4c-53a7-4fbb-91f4-f5ad7d9d0b9e · outbound

This paper cites Qwen2.5 Technical Report.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Qwen2.5 Technical Report

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.151066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.151066Z digest=sha256:764ede2cc8a0d677aabc902c388dbc8ca3e4555770dde71b2969f248007b1ac4

Observation 00c90c4c-6483-41a1-bd50-b4ab1ce0c251 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Evaluating Large Language Models Trained on Code

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.154555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.154555Z digest=sha256:5e934856e177f47cd47c74fa2638e454dd83e191c26144d269c53b190dd7471e

Observation ee56f8a5-890a-45f2-9dae-01ef05def261 · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.157996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.157996Z digest=sha256:e310766a79e70226d332e0cdb0b121964be7968c1719c9cb55cc86da60978845

Observation c22926a8-31ba-4a2b-9709-4450195aa258 · outbound

This paper cites Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.161587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.161587Z digest=sha256:da5dc7d3d8545751e40c5ad17b40b38026b7c27f63bd92ec411c427a8b95beeb

Observation 69e40fbb-3f69-438d-b86e-5541bab89966 · outbound

This paper cites Making Monolingual Sentence Embeddings Multilingual using Knowledge Distillation.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Making Monolingual Sentence Embeddings Multilingual using Knowledge Distillation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.165308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.165308Z digest=sha256:a8d5102ae44a3aa7f4659196eff777bc3136e5d602ccb30af5a3e8efb5993d3c

Pith citing papers

Observation 7b4aba80-d73a-4225-914e-cd2b8ce53e20 · inbound

Reinforcing Video Reasoning with Focused Thinking cites this paper.

Reinforcing Video Reasoning with Focused Thinking SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:12.190500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:12.190500Z digest=sha256:c2003adfc5903295b70dc97bc672f9a02367c3091363168a5d84078c2eb6536d

Observation 0b9f6000-6e9b-48f9-87c5-38a7c22a23d0 · inbound

Socratic RL: A Novel Framework for Efficient Knowledge Acquisition through Iterative Reflection and Viewpoint Distillation cites this paper.

Socratic RL: A Novel Framework for Efficient Knowledge Acquisition through Iterative Reflection and Viewpoint Distillation SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T20:07:05.170825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:07:05.170825Z digest=sha256:07e499c2db4f84b30713f7e1bdff3692c355cc06e37aa417e49377f8ba8539b7

Observation 51f79782-9ec8-4f33-83a3-77ede35621bd · inbound

Adaptive Termination for Multi-round Parallel Reasoning: An Universal Semantic Entropy-Guided Framework cites this paper.

Adaptive Termination for Multi-round Parallel Reasoning: An Universal Semantic Entropy-Guided Framework SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T18:59:40.639865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:59:40.639865Z digest=sha256:c3cd59aa14fd3f3042399f2051190c7d6fc96fd2b51fad4a13b091764efa3776

Observation 69b01b73-13c2-4459-890e-334ee896733b · inbound

EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity cites this paper.

EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T12:26:11.122616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:26:11.122616Z digest=sha256:c6063a8374749574749509960eaf06e84030cfa713a6baf130cf760c35f2c872

Observation 892fbaba-8018-4613-907c-bbe0cdaaa6cc · inbound

Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting cites this paper.

Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 1988

Resolution
unresolved
no resolver link, observed 2026-08-05T23:10:24.452911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:10:24.452911Z digest=sha256:33d2e70f26f0fafe25be96a518c61e770c213f97e7e36d26d00c574e72364d87

Observation 1d7da6ce-8da8-4a9a-8688-18fe4b8d55e8 · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 64

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T00:02:25.084591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:d399be38f217d85d19eba70ea21597c369122c77767127ea13f9565e57f63af8

Observation 2e325a43-c496-45a7-97ee-128721ef66f1 · inbound

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents cites this paper.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:56.258016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:56.258016Z digest=sha256:66f3fbc85eee3479ffc19e9e8e0eb4cf616ecb30fd10a5500165225c688869b5

Observation fbe99f8e-e2ae-4851-85b9-d275d025b811 · inbound

AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning cites this paper.

AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:20.055373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:20.055373Z digest=sha256:34db04f48b1f66351720006ca386ed9ac5c3b33fa1cfba2a55ab1fbea44ac37a

Observation 9d457360-c1b1-41e6-8da8-77d54a0f8373 · inbound

Unlocking Exploration in RLVR: Uncertainty-aware Advantage Shaping for Deeper Reasoning cites this paper.

Unlocking Exploration in RLVR: Uncertainty-aware Advantage Shaping for Deeper Reasoning SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-18T07:21:04.407632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-18T07:20:01.505216Z digest=sha256:3d055666e72508477ca4a40173e5eb8b646a4902a03a329a7dae9d6dc6310400

Observation f273ed6b-ec28-433b-b6e8-b982d2353ee8 · inbound

Mimic Human Cognition, Master Multi-Image Reasoning: A Meta-Action Framework for Enhanced Visual Understanding cites this paper.

Mimic Human Cognition, Master Multi-Image Reasoning: A Meta-Action Framework for Enhanced Visual Understanding SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T11:13:16.768253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:13:16.768253Z digest=sha256:e1098c7f37d324492a23888561ce00cf6238739078a282e654471ba8a83060fb

Observation 4b031acc-1e60-4a7e-9592-a5d988094944 · inbound

EMO-R3: Reflective Reinforcement Learning for Emotional Reasoning in Multimodal Large Language Models cites this paper.

EMO-R3: Reflective Reinforcement Learning for Emotional Reasoning in Multimodal Large Language Models SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T20:14:03.940645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:14:03.940645Z digest=sha256:6cde7916c28ebf9fdada45ce16ab0eca6875ed522685c0b42bac59a40efec6a6

Observation da19f184-186d-4aed-b54c-74ac41b3cf70 · inbound

Self-Distilled RLVR cites this paper.

Self-Distilled RLVR SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-13T19:48:11.563241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-13T19:43:46.623267Z digest=sha256:c4744267b24b933ef4e543f2ebeb91a5219b7599a7d0b4a8afb195f7d9f19cf3

Observation 97aa497d-3246-4818-8110-ce4dfb077a56 · inbound

LASER: A Data-Centric Method for Low-Cost and Efficient SQL Rewriting based on SQL-GRPO cites this paper.

LASER: A Data-Centric Method for Low-Cost and Efficient SQL Rewriting based on SQL-GRPO SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:16:02.244538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T17:15:04.225059Z digest=sha256:09340268f248b0dfb96a3d62a5bc24fb96aae817ee938c4b7003c7f196da26a8

Observation caa9ebf4-db47-47db-b183-b28db1c89bbc · inbound

HEALing Entropy Collapse: Enhancing Exploration in Few-Shot RLVR via Hybrid-Domain Entropy Dynamics Alignment cites this paper.

HEALing Entropy Collapse: Enhancing Exploration in Few-Shot RLVR via Hybrid-Domain Entropy Dynamics Alignment SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:25:55.199930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-10T05:23:08.478393Z digest=sha256:5c27f335d7b0b152ecd3d3252042b5dcaa967cb7091c7ee80353bd024e9f8589

Observation 40b4c59d-42a6-46d5-a2f3-85242df039df · inbound

Walking Through Uncertainty: An Empirical Study of Uncertainty Estimation for Audio-Aware Large Language Models cites this paper.

Walking Through Uncertainty: An Empirical Study of Uncertainty Estimation for Audio-Aware Large Language Models SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 79

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T00:41:26.300159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-07T14:29:18.348031Z digest=sha256:801d327674adcf374dd1e7416296dcdf908534e410d60ddec5effd6b840f458f

Observation 131bf09f-eb00-4823-b093-0ab1106f3131 · inbound

T$^2$PO: Uncertainty-Guided Exploration Control for Stable Multi-Turn Agentic Reinforcement Learning cites this paper.

T$^2$PO: Uncertainty-Guided Exploration Control for Stable Multi-Turn Agentic Reinforcement Learning SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-09T05:45:22.128262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T19:36:00.359351Z digest=sha256:bf74223c1256abdc34483733c4ffe743023bec61e9583d82f0425b272f0f5491

Observation c9ae2cc9-c05b-4afd-b031-5b719c166158 · inbound

Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning cites this paper.

Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:15:48.864199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-10T19:15:27.406778Z digest=sha256:d0a168774066b4f3a6d9a67277cbc408192612c6e45ad9095f0c58af55ee9991

Observation 7ddcc6ed-9c0a-4a53-b513-ecf095e389f3 · inbound

Self-Induced Outcome Potential: Turn-Level Credit Assignment for Agents without Verifiers cites this paper.

Self-Induced Outcome Potential: Turn-Level Credit Assignment for Agents without Verifiers SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T17:46:09.550349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T17:11:31.826233Z digest=sha256:f79fa22bd87abecfa5df4df89e87d0b8db9b29679ee820ff0d167fca73f510a4

Observation 6c96d400-c4e2-4a8f-8d3c-f44cd106e081 · inbound

Rethinking Entropy Minimization in Test-Time Adaptation for Autoregressive Models cites this paper.

Rethinking Entropy Minimization in Test-Time Adaptation for Autoregressive Models SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:41:24.629634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-12T00:48:47.213681Z digest=sha256:c47b41de6d5a53a5d0d2738d63e5498dcfc67827aa7000f5838473dfa1bf5e9c

Observation 7f265df0-44b6-4a12-a82c-653cfcab9e64 · inbound

DUET: Optimize Token-Budget Allocation for Reinforcement Learning with Verifiable Rewards cites this paper.

DUET: Optimize Token-Budget Allocation for Reinforcement Learning with Verifiable Rewards SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:46:29.447553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-12T01:57:11.065744Z digest=sha256:d9254dff424268c5d32912c0bd30172e33d13587510ea0f0be48006f13bf8451

Observation da4eb23b-efe3-4d94-bfdc-eed8427c032c · inbound

Holder Policy Optimisation cites this paper.

Holder Policy Optimisation SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:12:22.790348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-13T06:08:28.855671Z digest=sha256:a65d76d36300403a96e7a7509772d26856872fd5e70eab9b01f00d782e3dcf42

Observation 94c188f2-641f-4ea3-8252-6faeb9c094f9 · inbound

Holder Policy Optimisation cites this paper.

Holder Policy Optimisation SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-22T10:01:23.115200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-22T10:00:58.600743Z digest=sha256:a752f120faee004f091c62d77e5e799fe6f98d14b05c58cc4396e120a029d24b

Observation 0a337026-3afb-47eb-956a-f88f1dcfd852 · inbound

What and When to Distill: Selective Hindsight Distillation for Multi-Turn Agents cites this paper.

What and When to Distill: Selective Hindsight Distillation for Multi-Turn Agents SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:43:05.721502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T05:41:23.712146Z digest=sha256:58c0a1fc2b24254c2af1502b5a691a4324013b964682fa5a3d99e9047e1a8cfd

Observation 82fbf2f1-0959-46e2-a70b-f94cd95608b1 · inbound

Why Semantic Entropy Fails: Geometry-Aware and Calibrated Uncertainty for Policy Optimization cites this paper.

Why Semantic Entropy Fails: Geometry-Aware and Calibrated Uncertainty for Policy Optimization SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-22T08:54:45.757498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-22T08:53:02.051453Z digest=sha256:53be7dcfa7556ffb1d8db230abf3ed739286e61b97020239a3a3f28cab9a6688

Observation 45242034-0215-4725-8b8e-0e4ae85a8d30 · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 95

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:56:13.407524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:316d8cd2479cdcd59ad49cff348638c3cf7f76bbe0a96c2604b219c8a16d69af

Observation 7416aa03-8007-4fb4-a17f-6c2dc8e53bbf · inbound

FGRPO: Federated GRPO with Adaptive Aggregation on Non-IID Data cites this paper.

FGRPO: Federated GRPO with Adaptive Aggregation on Non-IID Data SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-02T01:36:25.608391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T11:44:53.211503Z digest=sha256:05cc5e94bb6fc6fca8cf605afc84c02d15ede91884044173b868900fd7b49089

Observation 2383a679-d435-4953-9f81-f7ece23e5234 · inbound

Self-Distilled Policy Gradient cites this paper.

Self-Distilled Policy Gradient SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-07-02T01:56:27.906762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-28T11:24:11.878439Z digest=sha256:566ea57fe96a34cf6c4ff485e19e4a45a26d289f7aa8a92802a91be769578d16

Observation d12d0f00-9328-4da1-be46-364c0af13427 · inbound

Reinforcement Learning from Rich Feedback with Distributional DAgger cites this paper.

Reinforcement Learning from Rich Feedback with Distributional DAgger SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-02T07:46:45.915273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T06:44:56.667364Z digest=sha256:231c825543b0308df8eb096274f72e5e81edf96c8bde49d946f4a89c72a5ec99

Observation d8439ce1-be73-4c88-a3ba-b2bbd9d8e9e5 · inbound

Back on Track: Aligning Rewards and States for Reasoning in Diffusion Large Language Models cites this paper.

Back on Track: Aligning Rewards and States for Reasoning in Diffusion Large Language Models SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:57:26.585403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-27T18:31:21.493677Z digest=sha256:345241e39c8cdecb6b13aaa74baf478ac2b97dee2821b75633d485236d3e569b

Observation 40290d98-3b46-434d-8d96-eb9c56da9001 · inbound

Architecture-Aware Reinforcement Learning Makes Sliding-Window Attention Competitive in Math Reasoning cites this paper.

Architecture-Aware Reinforcement Learning Makes Sliding-Window Attention Competitive in Math Reasoning SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-03T09:47:59.902366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-27T10:18:54.163862Z digest=sha256:96e5eee79450329e7c4d0f4b0b8218f475e2846f7e47c2b93250e84a4cacdd26

Observation eaa9e059-daff-422f-8897-f28c490542eb · inbound

Empowering Economic Simulation Through Situation-Aware Llm-Driven Generative System cites this paper.

Empowering Economic Simulation Through Situation-Aware Llm-Driven Generative System SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T23:49:02.837064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T21:43:01.949350Z digest=sha256:000c7300a6c25c857998f47299c3a67511a3b55df53afa2cf9e951bea1707f6c

Observation 14275d2a-83f5-485b-b2f0-ea9c0e1c32d6 · inbound

What are Key Factors for Updates in RL for LLM Reasoning? cites this paper.

What are Key Factors for Updates in RL for LLM Reasoning? SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:09:43.560962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-26T10:24:53.245739Z digest=sha256:8f7b3422367b137547fdc7f588c56ad85381074104a642f5e538d39e14bd9cd8

Observation 8829c247-63e7-43ce-b82f-14c140911edd · inbound

GEOALIGN: Geometric Rollout Curation for Robust LLM Reinforcement Learning cites this paper.

GEOALIGN: Geometric Rollout Curation for Robust LLM Reinforcement Learning SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:49:52.429607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T04:49:28.598430Z digest=sha256:78ceefd939469e24258bd08a2ab2bd0dd42d2edf1e8112fe133492d23c2b4f13

Observation 00ded8ff-7d33-420b-a0fd-003270913020 · inbound

Which Tokens Matter? Adaptive Token Selection for RLVR with the Relative Surprisal Index cites this paper.

Which Tokens Matter? Adaptive Token Selection for RLVR with the Relative Surprisal Index SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-01T10:25:42.387378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-01T05:27:22.527220Z digest=sha256:0ee8f57a37da774fab8083df23d15dc257c321093b611bc87dc3def73db6acc4

Observation 7c8263bc-0673-4a8f-83b1-d2cf75e97c99 · inbound

Reinforcement Learning with Metacognitive Feedback Elicits Faithful Uncertainty Expression in LLMs cites this paper.

Reinforcement Learning with Metacognitive Feedback Elicits Faithful Uncertainty Expression in LLMs SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-01T10:35:42.043252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-01T05:22:38.232552Z digest=sha256:a455bde7335114cc90384b9e4c97730c16daf07c97d4c61afbe5031af2745395