Pith. sign in

Paper Citation Record · LEDGER

Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback

As of 7 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 30 inbound Pith citation observations for arXiv:2506.03106.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.03106 v7

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:15:35.267794Z

measured 62 of 62 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 30 of 30 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T13:47:19.009023Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T00:29:15.119817Z

Reference resolution

32 of 32 outbound references displayed

  • verified exact0
  • verified fuzzy19
  • unresolved11
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a579f800-b35c-4935-90c5-4a76dec37ed3 · outbound

This paper cites needle in a haystack.

Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback needle in a haystack

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:15:35.779353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:15:35.192749Z digest=sha256:53a3551cf685dfa6bcab55d090751bcbb433fe2c905566c73ba144a4548c179e

Observation 8f534efd-4ea8-42f6-8d0b-09cda7837070 · outbound

This paper cites While the worst-case complexity remains dimT E(H, ℓ, ϵ)≈O(|S|L), the critique acts as a pruning signal.

Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback While the worst-case complexity remains dimT E(H, ℓ, ϵ)≈O(|S|L), the critique acts as a pruning signal

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:15:35.769972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:15:35.195999Z digest=sha256:9cbe73a671909328edb0ac76c6fb3df7a23dbcc9e11ede3d7d396fcb35966f63

Observation b7ad4a6d-2136-4acf-a79a-1b4869b9c1c7 · outbound

This paper cites an unresolved cited work.

Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:15:35.690194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:15:35.222317Z digest=sha256:eadd17b260757691a4d59e61f0b3bb0fc237215048663c140a31a987bc0f9675

Observation 90b9b898-9e89-4f9c-ac0c-dbff6c9078c9 · outbound

This paper cites Wang, Y ., Yue, X., and Chen, W.

Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback Wang, Y ., Yue, X., and Chen, W

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:15:35.170943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:15:35.170943Z digest=sha256:9152103d8a7f620bac1b8edf2bc11659cbc975f9d0f4e4be31979ea4fb25ff08

Observation ad29d770-e3e6-4bc6-877a-0c21c185e548 · outbound

This paper cites The correct maximum value, as derived from a proper analysis, should be 10 3.

Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback The correct maximum value, as derived from a proper analysis, should be 10 3

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:15:35.556971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:15:35.267794Z digest=sha256:8b4bee9956dfa34b8acaf8fd659426f1608fa9e4735ab7cbb469c60a1e944727

Observation ef39f899-e63f-45a1-b692-4364d4fae2ea · outbound

This paper cites Qwen3 Technical Report.

Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback Qwen3 Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:15:35.178558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:15:35.178558Z digest=sha256:2f93dbc05feba2524cc7e3e01f08184268d6791513feaf8c4f54a3cfdba1e12a

Observation 39b23d80-f7f2-4675-ad68-235586a58802 · outbound

This paper cites Zhang, X., Peng, B., Li, K., Zhou, J., and Meng, H.

Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback Zhang, X., Peng, B., Li, K., Zhou, J., and Meng, H

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:15:35.181921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:15:35.181921Z digest=sha256:a29fa40c8df3be4a2e85b23fbfe264acf359b99e0798a63823d8e6d1a5812a04

Observation 0d54d17d-9fa2-43c1-bf9a-a3de9d99d77a · outbound

This paper cites correct” and “incorrect.

Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback correct” and “incorrect

Reference 8

Resolution
malformed identifier
no resolver link, observed 2026-08-07T11:15:35.185214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:15:35.185214Z digest=sha256:32af0e3e05c35b4bdb62e5c55197deccd63fd280dac772c8a8cb0cd59f858dec

Observation 5f77a2b8-0532-498f-aa9c-3d95e421f3f1 · outbound

This paper cites As illustrated in Figure 8, this function is bounded between (0,1) , where x represents the token probability of the policy.

Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback As illustrated in Figure 8, this function is bounded between (0,1) , where x represents the token probability of the policy

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:15:35.789361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:15:35.189419Z digest=sha256:5b3e567178844f3fee38c60a14bcdc6f90b87ea5303b9b91ad7a89acd535fa77

Observation 7ffef02b-96c2-48ce-a18e-b35fb89a5553 · outbound

This paper cites In this setting, the feedback function is simply the reward itself, fη(a) =r(a).

Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback In this setting, the feedback function is simply the reward itself, fη(a) =r(a)

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:15:35.760582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:15:35.199170Z digest=sha256:e6f9702173d795ab401eed014bf589177f7f72ecf94eb109845ced09529b542f

Observation f008ba4f-474c-4083-8873-e2fee8bd2e6d · outbound

This paper cites The probability of finding the unique optimal solution a∗ is equivalent to sampling the correct element from a set of sizedwithout replacement.

Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback The probability of finding the unique optimal solution a∗ is equivalent to sampling the correct element from a set of sizedwithout replacement

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:15:35.750597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:15:35.202422Z digest=sha256:9a695d68c5a17e61b6cb1ac6e6e0ae0cee58461c9f737105305b6c9325d76d23

Observation 1b7aa275-33cb-445b-9173-d95966e02ccb · outbound

This paper cites Since γ≪1 , the gradient magnitude is significantly dampened.

Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback Since γ≪1 , the gradient magnitude is significantly dampened

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:15:35.739780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:15:35.205872Z digest=sha256:3fea4f13cd44fa4c4413d10d66fb435cd4e5d893d36ed18f750a8d4ca150523a

Observation 9f3ca082-9d4f-44c2-8694-c307e8481ab4 · outbound

This paper cites weaker refinement,.

Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback weaker refinement,

Reference 15

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T11:15:35.729061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:15:35.209183Z digest=sha256:360c1d0a8471a70b16339c09488c5d0b415416fe7a806d1b13b7665b311c74cf

Observation 119a2374-8a19-4e5b-94a6-3c2f71405470 · outbound

This paper cites an unresolved cited work.

Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:15:35.708975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:15:35.215979Z digest=sha256:b7e9fd9fa360f520c084bf5ceeb85ade5fc88f41a6bdc71b525030866c27af08

Observation c9c10e5b-01ee-4e8c-9b10-4407e713a172 · outbound

This paper cites an unresolved cited work.

Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:15:35.699270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:15:35.219150Z digest=sha256:406f18ab68b1a81155f82dc9e41b10d0068075e1ef54f6038bc5bdde92509fd8

Observation 7b28c74e-0063-4c28-baa5-b1582b90f90b · outbound

This paper cites Conclusion:.

Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback Conclusion:

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:15:35.680966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:15:35.225726Z digest=sha256:82579427e86f83e5e9b709eccb7a1299db4e6d41f962a0937f16ab96d8e27677

Observation 459c337b-f614-44a2-b336-4b4a4742ad76 · outbound

This paper cites Calculate the cosine of the angle of the axial section of the cone at the vertex which is also the apex of the cone.

Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback Calculate the cosine of the angle of the axial section of the cone at the vertex which is also the apex of the cone

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:15:35.671451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:15:35.229585Z digest=sha256:ba4ead5b32dd80dc12db9d2aa95602df6f749f16e10707bde2cd74a2ca4db589

Observation b3b5fbd5-10bb-4732-bb84-ca5c18c21d16 · outbound

This paper cites Wait, let me check that again.

Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback Wait, let me check that again

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:15:35.661156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:15:35.233555Z digest=sha256:cd592c1a105f1bf889894a3ba296b4c684186b6909ed21e90bac76077e320c24

Observation 28757a41-8f8a-4084-a7b3-6f5e700fdf57 · outbound

This paper cites an unresolved cited work.

Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:15:35.651769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:15:35.237179Z digest=sha256:73e8b2bc7f7507a3cc1b2788a2dd9d5d2e740a6bdd00692929830c953ab149c0

Observation 13b69af0-af44-488c-b3ad-25abf3a3b693 · outbound

This paper cites Therefore, the exact value is − 9 100, and the approximate decimal is −0.09.

Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback Therefore, the exact value is − 9 100, and the approximate decimal is −0.09

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:15:35.632214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:15:35.244109Z digest=sha256:85131b8d7e4fc23cac737fd7bd7ba62f7ec4157e7d0d7ae2a34c121a158e1641

Observation f59290c6-e1c1-4789-ab3b-44eb96d8ac3c · outbound

This paper cites Alternatively, if I think about angles: A is arcsin(0.4), which is in the first quadrant, B is arcsin(0.5) which is π/6, also first quadrant.

Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback Alternatively, if I think about angles: A is arcsin(0.4), which is in the first quadrant, B is arcsin(0.5) which is π/6, also first quadrant

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:15:35.621418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:15:35.247160Z digest=sha256:4751eb239212422350fbe40a02a3c50b8891bce86953c79fbd740355e9ef0e86

Observation 93a02610-6565-4c74-b972-b2944de36dba · outbound

This paper cites Alternatively, using complex numbers or other methods? Maybe not necessary.

Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback Alternatively, using complex numbers or other methods? Maybe not necessary

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:15:35.610767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:15:35.250430Z digest=sha256:660af48f556346c22a6b2011c3f4848f2673851cdaaec838c6ddfc98ae0d720c

Observation f43a9e87-3503-4e60-8d7d-88999f07724e · outbound

This paper cites The trigonometric substitution should be used more carefully, ensuring that the constraint is satisfied throughout.

Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback The trigonometric substitution should be used more carefully, ensuring that the constraint is satisfied throughout

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:15:35.600273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:15:35.254352Z digest=sha256:b45e777896ce766903b7ca3f0ed6088fe6a14415b813e3c3a4c89c460e1ef907

Observation 38bb0bc3-ed9d-43a2-a587-2e083088dddc · outbound

This paper cites The identities used do not lead to a valid simplification of the expression.

Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback The identities used do not lead to a valid simplification of the expression

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:15:35.589374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:15:35.257744Z digest=sha256:c281a301a61c3a3f1784a8ec9503e41399ee0aafe77d9fbcd0f643bab508906c

Observation 1857b7b4-4f3e-4372-939d-e47bc3240fc5 · outbound

This paper cites The derivative should be taken with respect to the correct variables, and the critical points should be found accurately.

Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback The derivative should be taken with respect to the correct variables, and the critical points should be found accurately

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:15:35.579589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:15:35.261000Z digest=sha256:49aa6bbcc09d44f299f613d9f9c232ca30d698d1dd1e0a6dde9e1ed9e6d5cc72

Observation 47df313b-3674-4c95-a2a8-6db297b6c724 · outbound

This paper cites The values chosen fora,b, andcdo not satisfy the constraintabc+a+c=b.

Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback The values chosen fora,b, andcdo not satisfy the constraintabc+a+c=b

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:15:35.568539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:15:35.264265Z digest=sha256:e50662760d645795c313d881326723c75fcc28f4416099a31f83d6ca61adb01e

Observation fe9f3e00-79b7-48ae-8afb-028fe77f3064 · outbound

This paper cites Therefore, the value of the original expression is −9 100.

Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback Therefore, the value of the original expression is −9 100

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:15:35.641688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:15:35.240692Z digest=sha256:5f1e1b5cd73fbfe2c96ce0aa9234a9ae24d29af3ee1c42678b9de1e219a9c308

Observation 9a8caf7d-b4f7-4f1e-96d6-2abf3d3b9665 · outbound

This paper cites Wait,.

Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback Wait,

Reference 150

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:15:35.718737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:15:35.212533Z digest=sha256:2d4e5891111afbde537a1703b5071bae107a71e92a8230bb1c5fc33758d18146

Observation 25931b4a-6bc6-4e11-96a3-e127664f2d9a · outbound

This paper cites Training language models to follow instructions with human feedback.

Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback Training language models to follow instructions with human feedback

Reference 155

Resolution
unresolved
no resolver link, observed 2026-08-07T11:15:35.163327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:15:35.163327Z digest=sha256:d5027929308f16c98ea59dfa8a7908d8b1f567a31cfd80a3c25db941e47b7321

Observation 8ffcab6a-b67c-4d0d-9558-9ec2d5b83103 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-07T11:15:35.167088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:15:35.167088Z digest=sha256:3bd3c895527909f50e65622cfa9c3a31f65b064a5de608b0e9af4841fe5cfd00

Observation d350a58a-8bf7-4c42-8595-7e3dbfa237f4 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback Understanding R1-Zero-Like Training: A Critical Perspective

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T11:15:35.158495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:15:35.158495Z digest=sha256:787e6f0c4aeaa05f48fb3feb0c9323b0d36a4f5aa4e8cd7f515c72d796ff442f

Observation 45364875-fb18-435c-8de7-197aa7aacc8a · outbound

This paper cites Learning to Reason under Off-Policy Guidance.

Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback Learning to Reason under Off-Policy Guidance

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T11:15:35.175237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:15:35.175237Z digest=sha256:f87149351012ff932e7e8ea62c3767d591a0c7f13a52871bd8b9d536ef5235cd

Pith citing papers

Observation ef08bf37-6769-454f-b0b7-bcb326090e57 · inbound

Video-R1: Reinforcing Video Reasoning in MLLMs cites this paper.

Video-R1: Reinforcing Video Reasoning in MLLMs Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-06-09T02:06:03.715193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T09:43:00.208065Z digest=sha256:e6425e91970ff9a0402283646706e216acda5debbed2654373b4b71361d89758

Observation 890f6533-5d3c-4937-8e75-fd63a407581d · inbound

MoL-RL: Distilling Multi-Step Environmental Feedback into LLMs for Feedback-Independent Reasoning cites this paper.

MoL-RL: Distilling Multi-Step Environmental Feedback into LLMs for Feedback-Independent Reasoning Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T13:47:19.009023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T13:47:19.009023Z digest=sha256:6515809d72d615e599d6603685e985d1398baf030ee1383e83fdeea65c58a595

Observation d5df3f36-6a6e-4d02-a269-d7f95777e969 · inbound

CLPO: Curriculum Learning meets Policy Optimization for LLM Reasoning cites this paper.

CLPO: Curriculum Learning meets Policy Optimization for LLM Reasoning Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T13:53:36.062617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:53:36.062617Z digest=sha256:e780edc3c92e19ad59fece83ed9cd75fad5539db827788a20b03f006a2a12837

Observation a04af6ba-5ecc-4578-9d72-008d996753aa · inbound

XRPO: Pushing the limits of GRPO with Targeted Exploration and Exploitation cites this paper.

XRPO: Pushing the limits of GRPO with Targeted Exploration and Exploitation Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T11:11:12.915233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:11:12.915233Z digest=sha256:5ddad9819823d4afd3c4336263eff1e137b3ec6f11cefd8d6732ef8f913b6712

Observation 27f1b512-88a9-4cbb-929a-98e9068fa0ce · inbound

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance cites this paper.

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T10:53:10.877236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:53:10.877236Z digest=sha256:f083c7363299a46e40d7e353dce2eab82f883a8524950327a11182eaca43c90f

Observation 83044b5c-6cf1-4702-a64e-c12687fbd697 · inbound

OneThinker: All-in-one Reasoning Model for Image and Video cites this paper.

OneThinker: All-in-one Reasoning Model for Image and Video Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-06-09T02:06:03.715193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T02:09:39.820651Z digest=sha256:403e2ffca26ade456cde195861986b4fdf51e2652984554d00f8225bb277f845

Observation 9af02442-d983-4e7b-b991-45aa307c7da0 · inbound

AdaTooler-V: Adaptive Tool-Use for Images and Videos cites this paper.

AdaTooler-V: Adaptive Tool-Use for Images and Videos Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback

Reference 80

Resolution
verified exact
arxiv_id, observed 2026-06-09T02:06:03.715193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T21:23:33.598026Z digest=sha256:3d5c7d477f0eefe13b605fd5bdd0b5700dd0d7c8ff2b2daeabe4e5f47d82296f

Observation 3baccb9b-70e2-4b70-be63-a0c3d0af0ffd · inbound

AdvJudge-Zero: Binary Decision Flips in LLM-as-a-Judge via Adversarial Control Tokens cites this paper.

AdvJudge-Zero: Binary Decision Flips in LLM-as-a-Judge via Adversarial Control Tokens Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T15:21:55.984824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T15:21:55.984824Z digest=sha256:c9b19b857e5e1b18c9de4f196d43cc800ab753bd6354d93a5d074a181feff7cd

Observation 9fcaacfb-fa23-4a17-aee4-351c5a09750b · inbound

CPMobius: Iterative Coach-Player Reasoning for Data-Free Reinforcement Learning cites this paper.

CPMobius: Iterative Coach-Player Reasoning for Data-Free Reinforcement Learning Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-03T05:14:22.180784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T05:14:22.180784Z digest=sha256:08899d2a3bed83ca1472a3b3dfa39787bc88894dffbaf0031031435e716d603f

Observation 00596522-fd42-4123-951a-941771bc4d7c · inbound

Gen-Searcher: Reinforcing Agentic Search for Image Generation cites this paper.

Gen-Searcher: Reinforcing Agentic Search for Image Generation Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-06-09T02:06:03.715193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T21:18:10.087258Z digest=sha256:e64bc726f9131db3d84d1618de64bb63c07e0887cd4caa588503e553da7b56da

Observation 61bce089-f9ff-404c-b03d-b31eef3ffdeb · inbound

Gen-Searcher: Reinforcing Agentic Search for Image Generation cites this paper.

Gen-Searcher: Reinforcing Agentic Search for Image Generation Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-06-09T02:06:03.715193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-25T06:31:13.232880Z digest=sha256:69009ccb25ea77ac7d160a6b5ca630253e1adadba99ff5766c8a330a0db96b63

Observation 266e7451-b303-4c0b-a695-4a57d85bcf95 · inbound

Large Language Model Post-Training: A Unified View of Off-Policy and On-Policy Learning cites this paper.

Large Language Model Post-Training: A Unified View of Off-Policy and On-Policy Learning Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback

Reference 106

Resolution
verified exact
arxiv_id, observed 2026-06-09T02:06:03.715193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:28:58.515666Z digest=sha256:389052f4f4dcb48bb8d241489f0cf9a7c365d7de5368ed6cd615c57d62b74ad7

Observation 11c880c2-cb3d-4c40-8b8a-947b0597696c · inbound

Hidden States Know Where Reasoning Diverges: Credit Assignment via Span-Level Wasserstein Distance cites this paper.

Hidden States Know Where Reasoning Diverges: Credit Assignment via Span-Level Wasserstein Distance Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-06-09T02:06:03.715193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T08:05:36.206946Z digest=sha256:af9f5dc7e57b36a99ee22f1912318d1e8bcb4f476e92f1e8e32674808c2d0f5f

Observation 0be93c8b-7545-4a9b-938e-05314ec92922 · inbound

Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning cites this paper.

Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback

Reference 92

Resolution
metadata mismatch
arxiv_id, observed 2026-06-09T02:06:03.715193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T10:23:52.522238Z digest=sha256:fd118eda2a08c707139139827277c1fc7bfe3ebd5a0f59315b803518b2f6975d

Observation 48110f5b-b2ff-4350-abed-0302f3f83bdf · inbound

Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning cites this paper.

Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback

Reference 92

Resolution
metadata mismatch
arxiv_id, observed 2026-06-09T02:06:03.715193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-11T02:00:00.663355Z digest=sha256:5d4d0e1e9868c503bfba349c8af4287f80481563201f8993d0516cf186f00308

Observation 49fb552d-fa19-4b23-9a97-b8ac7eeb583e · inbound

Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning cites this paper.

Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback

Reference 92

Resolution
metadata mismatch
arxiv_id, observed 2026-06-09T02:06:03.715193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-13T07:17:13.708752Z digest=sha256:5cbdff471850f7af4149c7208c09eb531c6cbcba809fc4f2572733ad36b9e8e5

Observation c92ccc7a-cd02-4fe1-bf07-cb4b694efa8a · inbound

The Cancellation Hypothesis in Critic-Free RL: From Outcome Rewards to Token Credits cites this paper.

The Cancellation Hypothesis in Critic-Free RL: From Outcome Rewards to Token Credits Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-06-09T02:06:03.715193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:24:03.186413Z digest=sha256:5910ba70dd380200a15f27b05db568e41e18bdeb2fc11624c682360d8fcde92b

Observation 88a1d0f1-93d7-4601-875a-9bc7a2fa9e35 · inbound

Multi-Rollout On-Policy Distillation via Peer Successes and Failures cites this paper.

Multi-Rollout On-Policy Distillation via Peer Successes and Failures Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-06-09T02:06:03.715193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-14T21:29:36.803832Z digest=sha256:97d7e1c446ba37c0196a9c10e0fd0adee79cd9d4fed12f8c4245fe29b42b7472

Observation 1079fdf3-9892-4d3b-833e-04649fba1d4d · inbound

ICRL: Learning to Internalize Self-Critique with Reinforcement Learning cites this paper.

ICRL: Learning to Internalize Self-Critique with Reinforcement Learning Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-06-09T02:06:03.715193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-19T17:58:05.817581Z digest=sha256:9e983950fa4b20a514244408cc2133f3492834490594bc1d4125dc6c87fb34d1

Observation 655f3118-46d8-4a0a-bbeb-1c7b7de284f2 · inbound

PAIR: Prefix-Aware Internal Reward Model for Multi-Turn Agent Optimization cites this paper.

PAIR: Prefix-Aware Internal Reward Model for Multi-Turn Agent Optimization Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-06-09T02:06:03.715193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T10:41:25.205368Z digest=sha256:85045e6018344b2e60c8742067742ddeca1adfcde1d38d2bc7d926ae0d6c4117

Observation 257ced7d-efc5-48ce-92a7-13ba8301d3d2 · inbound

PAIR: Prefix-Aware Internal Reward Model for Multi-Turn Agent Optimization cites this paper.

PAIR: Prefix-Aware Internal Reward Model for Multi-Turn Agent Optimization Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T02:23:32.244687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:23:32.244687Z digest=sha256:fcc45f974e21124ed2ecfb100bf30dfd529dde03815acf131d74c9998c6a5d47

Observation 903fc1e0-9e56-49c6-8d49-e5453ff4554e · inbound

ReCrit: Transition-Aware Reinforcement Learning for Scientific Critic Reasoning cites this paper.

ReCrit: Transition-Aware Reinforcement Learning for Scientific Critic Reasoning Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-06-09T02:06:03.715193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T22:51:56.666980Z digest=sha256:a090996b19da6ba1977372e3d495d2293ac47d1fadf70aafab955d4007c5b78b

Observation 0b8c0dfe-6fa1-4a06-81dd-689eda059d4e · inbound

STRIDE: Learnable Stepwise Language Feedback for LLM Reasoning cites this paper.

STRIDE: Learnable Stepwise Language Feedback for LLM Reasoning Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-06-09T02:06:03.715193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T20:47:16.236629Z digest=sha256:79bed78f25492934c0e1e2d69c3c77288a9498a43f23b9f3831633f989b6b996

Observation 83749c53-f9af-400e-8a4e-5f7b7a0b53fd · inbound

Reinforcing Human Behavior Simulation via Verbal Feedback cites this paper.

Reinforcing Human Behavior Simulation via Verbal Feedback Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback

Reference 79

Resolution
metadata mismatch
arxiv_id, observed 2026-06-09T02:06:03.715193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-21T07:21:48.649289Z digest=sha256:19d8b8242805289f004516479b730c883b7daf1a0d29215a84e8853d5d2230ae

Observation 2c37c286-2716-4642-a16e-14e9f9a86d98 · inbound

RL with Learnable Textual Feedback: A Bilevel Approach cites this paper.

RL with Learnable Textual Feedback: A Bilevel Approach Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-06-30T14:44:45.361285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T14:36:31.047856Z digest=sha256:2e0943ae26e1f5db6e8ae8012c538bef769884e7b5f21f92104ba026343a02db

Observation 48d2b711-066b-41b2-9cc9-1acf83b95f91 · inbound

Credit Assignment with Resets in Language Model Reasoning cites this paper.

Credit Assignment with Resets in Language Model Reasoning Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-06-29T21:53:59.255224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T21:50:18.827822Z digest=sha256:b15d79bfd158008cf062d41e522bf1c40a6790a966499196c4bb4af358b59195

Observation 660eefdb-b031-4778-b04d-c84b18bac9e8 · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback

Reference 189

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T20:56:13.705760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:0dc57877c8992c64fd0a4583488f67a61bdc2d16dc952d2b4f01a9703b691524

Observation 5fe53e6d-5bfa-441c-8a6d-0c93514215e6 · inbound

REVES: REvision and VErification--Augmented Training for Test-Time Scaling cites this paper.

REVES: REvision and VErification--Augmented Training for Test-Time Scaling Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback

Reference 76

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:29:15.122126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-26T21:14:15.337979Z digest=sha256:660682e9c957a76ab0b4ebd74f98fd6beb4970f3ee1856496f76a73dad88681f

Observation c34e7e5f-a8a7-40f4-b444-3fab060e95ae · inbound

Agent Reinforcement Learning via Pivotal-Aware Self-Feedback Retry cites this paper.

Agent Reinforcement Learning via Pivotal-Aware Self-Feedback Retry Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-12T00:33:06.488657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:33:06.488657Z digest=sha256:09a2dbad27e8b0bbafc1acff7f3e4791d9707299aff534a991aa75cdf3eacd11

Observation 483e6977-b08e-4d13-b184-6573aeae10e3 · inbound

LLM-as-a-Coach: Experiential Learning for Non-Verifiable Tasks cites this paper.

LLM-as-a-Coach: Experiential Learning for Non-Verifiable Tasks Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T16:07:59.623378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T16:07:59.623378Z digest=sha256:fdd57399cf07474b76a842d1b730831778272f97550d319ddd8eb8a6122313d7