Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:15:35.267794Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 30 inbound Pith citation observations for arXiv:2506.03106.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:15:35.267794Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T13:47:19.009023Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-04T00:29:15.119817Z
32 of 32 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation a579f800-b35c-4935-90c5-4a76dec37ed3 · outbound
Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback needle in a haystack
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8f534efd-4ea8-42f6-8d0b-09cda7837070 · outbound
Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback While the worst-case complexity remains dimT E(H, ℓ, ϵ)≈O(|S|L), the critique acts as a pruning signal
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b7ad4a6d-2136-4acf-a79a-1b4869b9c1c7 · outbound
Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback Unresolved cited work
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 90b9b898-9e89-4f9c-ac0c-dbff6c9078c9 · outbound
Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback Wang, Y ., Yue, X., and Chen, W
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad29d770-e3e6-4bc6-877a-0c21c185e548 · outbound
Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback The correct maximum value, as derived from a proper analysis, should be 10 3
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ef39f899-e63f-45a1-b692-4364d4fae2ea · outbound
Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback Qwen3 Technical Report
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39b23d80-f7f2-4675-ad68-235586a58802 · outbound
Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback Zhang, X., Peng, B., Li, K., Zhou, J., and Meng, H
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d54d17d-9fa2-43c1-bf9a-a3de9d99d77a · outbound
Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback correct” and “incorrect
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f77a2b8-0532-498f-aa9c-3d95e421f3f1 · outbound
Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback As illustrated in Figure 8, this function is bounded between (0,1) , where x represents the token probability of the policy
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7ffef02b-96c2-48ce-a18e-b35fb89a5553 · outbound
Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback In this setting, the feedback function is simply the reward itself, fη(a) =r(a)
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f008ba4f-474c-4083-8873-e2fee8bd2e6d · outbound
Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback The probability of finding the unique optimal solution a∗ is equivalent to sampling the correct element from a set of sizedwithout replacement
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1b7aa275-33cb-445b-9173-d95966e02ccb · outbound
Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback Since γ≪1 , the gradient magnitude is significantly dampened
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9f3ca082-9d4f-44c2-8694-c307e8481ab4 · outbound
Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback weaker refinement,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 119a2374-8a19-4e5b-94a6-3c2f71405470 · outbound
Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback Unresolved cited work
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c9c10e5b-01ee-4e8c-9b10-4407e713a172 · outbound
Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback Unresolved cited work
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7b28c74e-0063-4c28-baa5-b1582b90f90b · outbound
Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback Conclusion:
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 459c337b-f614-44a2-b336-4b4a4742ad76 · outbound
Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback Calculate the cosine of the angle of the axial section of the cone at the vertex which is also the apex of the cone
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b3b5fbd5-10bb-4732-bb84-ca5c18c21d16 · outbound
Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback Wait, let me check that again
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 28757a41-8f8a-4084-a7b3-6f5e700fdf57 · outbound
Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback Unresolved cited work
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 13b69af0-af44-488c-b3ad-25abf3a3b693 · outbound
Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback Therefore, the exact value is − 9 100, and the approximate decimal is −0.09
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f59290c6-e1c1-4789-ab3b-44eb96d8ac3c · outbound
Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback Alternatively, if I think about angles: A is arcsin(0.4), which is in the first quadrant, B is arcsin(0.5) which is π/6, also first quadrant
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 93a02610-6565-4c74-b972-b2944de36dba · outbound
Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback Alternatively, using complex numbers or other methods? Maybe not necessary
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f43a9e87-3503-4e60-8d7d-88999f07724e · outbound
Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback The trigonometric substitution should be used more carefully, ensuring that the constraint is satisfied throughout
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 38bb0bc3-ed9d-43a2-a587-2e083088dddc · outbound
Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback The identities used do not lead to a valid simplification of the expression
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1857b7b4-4f3e-4372-939d-e47bc3240fc5 · outbound
Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback The derivative should be taken with respect to the correct variables, and the critical points should be found accurately
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 47df313b-3674-4c95-a2a8-6db297b6c724 · outbound
Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback The values chosen fora,b, andcdo not satisfy the constraintabc+a+c=b
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fe9f3e00-79b7-48ae-8afb-028fe77f3064 · outbound
Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback Therefore, the value of the original expression is −9 100
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9a8caf7d-b4f7-4f1e-96d6-2abf3d3b9665 · outbound
Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback Wait,
Reference 150
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 25931b4a-6bc6-4e11-96a3-e127664f2d9a · outbound
Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback Training language models to follow instructions with human feedback
Reference 155
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ffcab6a-b67c-4d0d-9558-9ec2d5b83103 · outbound
Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d350a58a-8bf7-4c42-8595-7e3dbfa237f4 · outbound
Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback Understanding R1-Zero-Like Training: A Critical Perspective
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45364875-fb18-435c-8de7-197aa7aacc8a · outbound
Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback Learning to Reason under Off-Policy Guidance
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef08bf37-6769-454f-b0b7-bcb326090e57 · inbound
Video-R1: Reinforcing Video Reasoning in MLLMs Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 890f6533-5d3c-4937-8e75-fd63a407581d · inbound
MoL-RL: Distilling Multi-Step Environmental Feedback into LLMs for Feedback-Independent Reasoning Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5df3f36-6a6e-4d02-a269-d7f95777e969 · inbound
CLPO: Curriculum Learning meets Policy Optimization for LLM Reasoning Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a04af6ba-5ecc-4578-9d72-008d996753aa · inbound
XRPO: Pushing the limits of GRPO with Targeted Exploration and Exploitation Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27f1b512-88a9-4cbb-929a-98e9068fa0ce · inbound
TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83044b5c-6cf1-4702-a64e-c12687fbd697 · inbound
OneThinker: All-in-one Reasoning Model for Image and Video Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9af02442-d983-4e7b-b991-45aa307c7da0 · inbound
AdaTooler-V: Adaptive Tool-Use for Images and Videos Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3baccb9b-70e2-4b70-be63-a0c3d0af0ffd · inbound
AdvJudge-Zero: Binary Decision Flips in LLM-as-a-Judge via Adversarial Control Tokens Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9fcaacfb-fa23-4a17-aee4-351c5a09750b · inbound
CPMobius: Iterative Coach-Player Reasoning for Data-Free Reinforcement Learning Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00596522-fd42-4123-951a-941771bc4d7c · inbound
Gen-Searcher: Reinforcing Agentic Search for Image Generation Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 61bce089-f9ff-404c-b03d-b31eef3ffdeb · inbound
Gen-Searcher: Reinforcing Agentic Search for Image Generation Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 266e7451-b303-4c0b-a695-4a57d85bcf95 · inbound
Large Language Model Post-Training: A Unified View of Off-Policy and On-Policy Learning Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback
Reference 106
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 11c880c2-cb3d-4c40-8b8a-947b0597696c · inbound
Hidden States Know Where Reasoning Diverges: Credit Assignment via Span-Level Wasserstein Distance Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0be93c8b-7545-4a9b-938e-05314ec92922 · inbound
Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 48110f5b-b2ff-4350-abed-0302f3f83bdf · inbound
Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 49fb552d-fa19-4b23-9a97-b8ac7eeb583e · inbound
Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c92ccc7a-cd02-4fe1-bf07-cb4b694efa8a · inbound
The Cancellation Hypothesis in Critic-Free RL: From Outcome Rewards to Token Credits Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 88a1d0f1-93d7-4601-875a-9bc7a2fa9e35 · inbound
Multi-Rollout On-Policy Distillation via Peer Successes and Failures Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1079fdf3-9892-4d3b-833e-04649fba1d4d · inbound
ICRL: Learning to Internalize Self-Critique with Reinforcement Learning Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 655f3118-46d8-4a0a-bbeb-1c7b7de284f2 · inbound
PAIR: Prefix-Aware Internal Reward Model for Multi-Turn Agent Optimization Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 257ced7d-efc5-48ce-92a7-13ba8301d3d2 · inbound
PAIR: Prefix-Aware Internal Reward Model for Multi-Turn Agent Optimization Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 903fc1e0-9e56-49c6-8d49-e5453ff4554e · inbound
ReCrit: Transition-Aware Reinforcement Learning for Scientific Critic Reasoning Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0b8c0dfe-6fa1-4a06-81dd-689eda059d4e · inbound
STRIDE: Learnable Stepwise Language Feedback for LLM Reasoning Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 83749c53-f9af-400e-8a4e-5f7b7a0b53fd · inbound
Reinforcing Human Behavior Simulation via Verbal Feedback Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2c37c286-2716-4642-a16e-14e9f9a86d98 · inbound
RL with Learnable Textual Feedback: A Bilevel Approach Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 48d2b711-066b-41b2-9cc9-1acf83b95f91 · inbound
Credit Assignment with Resets in Language Model Reasoning Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 660eefdb-b031-4778-b04d-c84b18bac9e8 · inbound
Trust Region On-Policy Distillation Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback
Reference 189
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5fe53e6d-5bfa-441c-8a6d-0c93514215e6 · inbound
REVES: REvision and VErification--Augmented Training for Test-Time Scaling Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c34e7e5f-a8a7-40f4-b444-3fab060e95ae · inbound
Agent Reinforcement Learning via Pivotal-Aware Self-Feedback Retry Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 483e6977-b08e-4d13-b184-6573aeae10e3 · inbound
LLM-as-a-Coach: Experiential Learning for Non-Verifiable Tasks Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.