Pith. sign in

Paper Citation Record · LEDGER

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle

As of 20 August 2026, this Paper Citation Record lists 100 of 263 outbound references and 11 inbound Pith citation observations for arXiv:2509.16679.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.16679 v1

Coverage vector

measured 100 of 263 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T16:07:33.995648Z

measured 111 of 111 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T20:14:04.060992Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

100 of 263 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved100
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 75da94fe-2a7c-4b6d-8726-f40e5b02ced5 · outbound

This paper cites NeMo RL: A Scalable and Efficient Post-Training Library.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle NeMo RL: A Scalable and Efficient Post-Training Library

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:27.677675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:27.677675Z digest=sha256:133aa81ed42fa2fb7599df232cfbee411e995fc578b5c974aa567cfaed16926c

Observation 7b0aae59-024d-4ad4-bded-1c6c3f5b4e68 · outbound

This paper cites L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:27.711633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:27.711633Z digest=sha256:c915beb3a42501576918370804cdd1d613c462c9c9516488059f5dce704f7082

Observation dc88b0d3-dfa3-46bd-8d37-5731a868c0b8 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:27.751920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:27.751920Z digest=sha256:a8e1ba9f1a616d777c37c151ac3b87b04f114611374f14a8c1fb55e94d6aeaa3

Observation ae598da4-298a-45ac-841d-6fc736736556 · outbound

This paper cites A General Language Assistant as a Laboratory for Alignment.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle A General Language Assistant as a Laboratory for Alignment

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:27.793366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:27.793366Z digest=sha256:b8ad966bfa9166ea359b056c10f8ad283820fcaccab4bfcc074fbb34762da7c7

Observation 59a51287-89e9-49b0-8d36-c8fa5e843dfa · outbound

This paper cites an unresolved cited work.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:27.831930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:27.831930Z digest=sha256:8d81ccc86a5d97aed567d570351f27236fc97328396c3149eedeb3c0a68947fd

Observation 4b3e214e-4027-44e3-ad3e-539f0a8410da · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:27.872089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:27.872089Z digest=sha256:165761e9b04c45289b3239b2b4d812819c4d56c2868d03e37cf5f5c300ee59bc

Observation c3bdfe05-bfcc-4e36-90e5-fe91b5390176 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Constitutional AI: Harmlessness from AI Feedback

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:27.917251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:27.917251Z digest=sha256:a27cbf3a9f1a7b51fca5ffb609fbce40b94a579e000123883809042e4b04cc59

Observation 75c084d6-33c5-44d8-81c1-3e1257b7ef9d · outbound

This paper cites Thinking Machines: A Survey of LLM based Reasoning Strategies.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Thinking Machines: A Survey of LLM based Reasoning Strategies

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:27.952573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:27.952573Z digest=sha256:e2fa3abcf9708557729011a152033250c5c8db59f09e20b8590b33f15c8b96fd

Observation 1f7f2e04-4a7c-4d6a-8a75-32e8e1cf9336 · outbound

This paper cites Do, Yan Xu, and Pascale Fung.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Do, Yan Xu, and Pascale Fung

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:27.987505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:27.987505Z digest=sha256:5a0ed32c28711de80b80f847769484d5b718f49e1f870ee7460b74b72615eb4b

Observation cd246f83-02c2-4b8d-b500-eaf06e80367f · outbound

This paper cites $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:28.031833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:28.031833Z digest=sha256:a9e655204eceb0af2512dbcc1ba664856af331b108439b66f284ef2eefab075c

Observation d82380ff-a5ea-4668-81f8-f646c6492e78 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:28.068731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:28.068731Z digest=sha256:7b6af4b13eabafca9c90b386048018a150c66e519d6365956a5f2fcb71f6179a

Observation b1916877-41dd-489e-98f8-b63f121d0c27 · outbound

This paper cites Reasoning Language Models: A Blueprint.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Reasoning Language Models: A Blueprint

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:28.107428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:28.107428Z digest=sha256:60db3fbaacc11e69adb592b0c2643e38d13ea13e418a31c3c56483eb5f9d474e

Observation de9561ea-e3fb-4436-a269-e5d4cf91808e · outbound

This paper cites an unresolved cited work.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:28.145291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:28.145291Z digest=sha256:268bcdf947ee925c769529e4e39759a5e1ec1bc149ff1e333a4d290812a849a7

Observation a8d8175f-0578-4fcc-a3dd-97a8dab4cdfe · outbound

This paper cites On the Opportunities and Risks of Foundation Models.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle On the Opportunities and Risks of Foundation Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:28.187255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:28.187255Z digest=sha256:d04594a568509d5e68396217a83598c2ea9b0e0907e5417b11f606db34ce04b4

Observation 04f9b39a-d183-48f0-97f9-8ea6c5e4c965 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:28.221722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:28.221722Z digest=sha256:cf7b657cece2afffc98e08ad8d1813e7b1f7583175ae30ff6bc627bd83b73af5

Observation 07a5c395-e85b-47e6-8944-e6f28aab8d59 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:28.259721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:28.259721Z digest=sha256:122c3fdaab1c502d45c6560e9fe3aadab2c2a1e1014fec8dbe9d4641f67de58a

Observation 6f11dde3-5b1c-4a67-a18c-3dace7c545f6 · outbound

This paper cites MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:28.295561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:28.295561Z digest=sha256:1df37b3fb32e20e05d399ec56a0828b8849ce40f0be57b949005d1dbe73fa3d1

Observation 64a5b613-97f8-40e5-bd7e-3f5e9678c71d · outbound

This paper cites an unresolved cited work.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:28.332269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:28.332269Z digest=sha256:53ab96413eab7de7f7d72a3e05a5a10c2b46445e34de52c9397f33f5295858cc

Observation fe009cb9-2b50-4598-83d3-acbabef4f6b3 · outbound

This paper cites SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:28.361696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:28.361696Z digest=sha256:0afa7760e3d350f39f33a9becfcf1c2342f70e5161d6ac365e05305abadf018b

Observation d8b56924-bdac-44a4-8057-504c782c8133 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:28.398802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:28.398802Z digest=sha256:15cea73119b9739bfdd084088fb1cc3aed04476d37464cc5259d4747c3c3be68

Observation 5722edcc-a5bd-4e16-9781-3eceeb8c3177 · outbound

This paper cites Detecting and Evaluating Medical Hallucinations in Large Vision Language Models.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Detecting and Evaluating Medical Hallucinations in Large Vision Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:28.414469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:28.414469Z digest=sha256:14c981bf8637c093e5e74a9fad383a4440d0f272257f14c54bf770d3ab2e5555

Observation a7273204-5657-4801-966e-990054462abe · outbound

This paper cites an unresolved cited work.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:28.452193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:28.452193Z digest=sha256:3eb2b3762db444020638e7481aa3a3bb49aa2cbb6e50d43f212f9d4bf8c68c87

Observation 37c08c5d-0f4e-4c78-a89e-f2ee7125aaf4 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:28.487822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:28.487822Z digest=sha256:9baa1896ab472894689eb3aff17059fc999cb593c1b11e9348ed6747cc232ce3

Observation 204d6904-3f0c-4a47-b150-6d88ee4dda7f · outbound

This paper cites Compile Scene Graphs with Reinforcement Learning.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Compile Scene Graphs with Reinforcement Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:28.527593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:28.527593Z digest=sha256:0ffc7b221cc89a7ceb551853096eab8341a68abe546253a3877bbaf348112987

Observation 0336c218-8eeb-441a-85a1-478bcc30418e · outbound

This paper cites Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:28.561884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:28.561884Z digest=sha256:fff3ad5d44e006c5eec0f3b30c40a146bcf72227498ad723696a37a8db13e2f0

Observation 13e3e9ae-e1a9-492a-8e82-60b49c83205e · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Training Verifiers to Solve Math Word Problems

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:28.596194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:28.596194Z digest=sha256:1183b14cda4c20e493160bacc847c09fe2bccdf145f12f5f381d034c6bbc485c

Observation 8a572717-c055-49e3-927a-3930807cca81 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:28.635398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:28.635398Z digest=sha256:cf665439cf75eeb0ecf1a0d33a9a64ed1b6e086be5c46c14fa2ba1cdd4f90f69

Observation 6506b4ce-8fa9-4950-8f8a-c64a4fefe82d · outbound

This paper cites The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:28.674483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:28.674483Z digest=sha256:e31a587d2e00101944e3a4d756d0d420e7fd155ce14cf77533e0d982d95c81f2

Observation 994cb6c9-e3dc-4725-8a93-657b246f0fd7 · outbound

This paper cites Boosting the Generalization and Reasoning of Vision Language Models with Curriculum Reinforcement Learning.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Boosting the Generalization and Reasoning of Vision Language Models with Curriculum Reinforcement Learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:28.711477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:28.711477Z digest=sha256:aa09e0b6d78bf327c82a0c1a6f2586ec6e33c4607c0a4fcbafc33f217e58265d

Observation c39a617e-8a38-456c-a1b9-3b25bf595665 · outbound

This paper cites Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:28.753936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:28.753936Z digest=sha256:c0f52e05f82099f622e5fff9aa22a6836faf685d4cc9ac6674b73584e7a6f8c3

Observation 481bd3f4-b9cc-419d-be04-f5ae99ef2ace · outbound

This paper cites Reinforcement Pre-Training.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Reinforcement Pre-Training

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:28.792504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:28.792504Z digest=sha256:c08b33cbbc4f21189af2550577868b6ccac65931fcc309e297aa0375f28ff0f9

Observation d8be9316-60c5-4dd4-8320-403e4905d74c · outbound

This paper cites an unresolved cited work.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:28.829289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:28.829289Z digest=sha256:dfcbb9109be1c60260bc8d703e30bf6255dc0783e5e801c9b945cc575a66988a

Observation 303355bc-dd3c-43de-b77b-526cd55a09f5 · outbound

This paper cites SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:28.872411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:28.872411Z digest=sha256:6108cdaa9f347b521118a23f064c734f7fb807c06346e627ab15c6a83c89575a

Observation cf3cca11-4862-4e85-a765-c753ea3879df · outbound

This paper cites an unresolved cited work.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:28.907544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:28.907544Z digest=sha256:1e09cb0f4619710ac8384cd1173a08fd356b76405580e21320da924ae808b146

Observation 1db69603-4fa2-49dd-86c8-5e16e100152c · outbound

This paper cites an unresolved cited work.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:28.916051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:28.916051Z digest=sha256:a1e1a817d441a723e56d68a57f94eb9c36f8e6526c0d9d69d926eea0eba5a570

Observation b052fbbe-70e8-48c3-a101-990c49522374 · outbound

This paper cites Thinkless: LLM Learns When to Think.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Thinkless: LLM Learns When to Think

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:28.955737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:28.955737Z digest=sha256:6e4230400eeded0ee572c2cf7ef975b88332dccf74d4d3aed77a009f9adc567f

Observation 10e14da2-97a7-4a3f-b6da-1696f3588219 · outbound

This paper cites ReTool: Reinforcement Learning for Strategic Tool Use in LLMs.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle ReTool: Reinforcement Learning for Strategic Tool Use in LLMs

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:28.991482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:28.991482Z digest=sha256:562ed0d395d1ccb7a49303014a302a88ceb9ad1a877b232ee53994e9764bb894

Observation c001a136-ba18-4657-9af6-862cfa25bf17 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:29.027428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:29.027428Z digest=sha256:4f783635bd397ea0c3337cb2070652c40cc9c9be632b9022b693c6faa8ab7b3d

Observation b3b4b132-0618-47eb-a124-a9f0de000942 · outbound

This paper cites Group-in-Group Policy Optimization for LLM Agent Training.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Group-in-Group Policy Optimization for LLM Agent Training

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:29.094583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:29.094583Z digest=sha256:eaba515d16262655c7f74634316570798724704de03bbd505cd14e203ec27a90

Observation 1d89261b-1454-43a4-a0b5-9ea6d18d6b1d · outbound

This paper cites an unresolved cited work.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:29.131748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:29.131748Z digest=sha256:084e6528091488db2d1184665e28d3f7c6a66da73245c137038f2d96f6b9d85d

Observation b3d2f39a-8b98-457d-9cc5-1bf124dc3f08 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:29.163549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:29.163549Z digest=sha256:98fe866278762a78406f10bc545ad381064c697e2e7ad3d9946d1d0652b1e1c4

Observation d8d6583d-a411-4896-ac62-6ef44ab2e356 · outbound

This paper cites AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:29.199481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:29.199481Z digest=sha256:38efe4ad5e54aaee1b9a301014b4b15566fa74297d08017201ca0b8bdb17def1

Observation 0a302257-1715-428a-ad23-c15f21868d2b · outbound

This paper cites an unresolved cited work.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:29.231088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:29.231088Z digest=sha256:f68f3798afeeaabe865d3790bfa886131affaf831c059a0008a52f25564be8e7

Observation 74598496-fe99-4539-bf71-53871d05cb86 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Unresolved cited work

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:29.276202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:29.276202Z digest=sha256:38d65aa39f12a60bbbb62b29d0e5ceea156ef4e3ff68df7366d3955d3b495830

Observation 3770ca00-2666-4ab3-8baf-d1275f82cd24 · outbound

This paper cites Visual Pre-Training on Unlabeled Images using Reinforcement Learning.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Visual Pre-Training on Unlabeled Images using Reinforcement Learning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:29.332714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:29.332714Z digest=sha256:c4b38d8b74b545bebe8c7f158524fb4deb19f960f312b592724f5c6fcac66172

Observation 1eb84988-09be-4dc8-9337-a1b723f9e5a5 · outbound

This paper cites Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:29.383625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:29.383625Z digest=sha256:fa3c1b4faadf5f91e2f3f6efee327b2980cda17247c7ef0c27b707aae31a26b6

Observation 2d72453d-d5e8-45fc-a2a1-35a9ac617340 · outbound

This paper cites CAD-Coder: Text-to-CAD Generation with Chain-of-Thought and Geometric Reward.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle CAD-Coder: Text-to-CAD Generation with Chain-of-Thought and Geometric Reward

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:29.460017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:29.460017Z digest=sha256:c6b3e684ac06d8c86c50f3b10e5b296ad9b163229cc8f1a2b016a27fb0154b80

Observation 319648f6-4b28-432e-a38d-9b768458a376 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:29.511684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:29.511684Z digest=sha256:460be9136ca354c91d0d38638c07c075e52047652bcad3be1f1bde046deddbcd

Observation c3900aa1-9feb-460d-9ceb-9b19a2dc5c8d · outbound

This paper cites Reward Reasoning Model.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Reward Reasoning Model

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:29.589806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:29.589806Z digest=sha256:5d3c04ad663142805efcb89e3df218786c3916b2b5a4a8e4d76dc4684b38858d

Observation e9aff45a-0e39-4acf-9076-ba84fa7fa73c · outbound

This paper cites Synthetic Data RL: Task Definition Is All You Need.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Synthetic Data RL: Task Definition Is All You Need

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:29.655928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:29.655928Z digest=sha256:4ccb9b2502b9abb4e020e9241d602bb1290a0ea21f8b78bbd9c399264a702e8f

Observation 3bb229a9-44a0-47da-acb9-7bc50e4de767 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Unresolved cited work

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:29.717653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:29.717653Z digest=sha256:1dc505696918aaf43afb03d084b235ca0714b8b0df33e25361eed8aa5d54aad9

Observation bd906e2e-ba42-45cf-b81b-8cf72f147eca · outbound

This paper cites an unresolved cited work.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Unresolved cited work

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:29.775830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:29.775830Z digest=sha256:2ece1b5004fdcbfb4bea700dab7238f2f4f6d21e849ed0094ddfcebdf196b3f3

Observation 9442083a-9dd8-4aa2-9f9f-c10435c2f283 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Unresolved cited work

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:29.891730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:29.891730Z digest=sha256:5362035872d1d5b1439aaafbcab9297d030c93b9691c779f58b38a0e9eed04f1

Observation d2e83bd1-7106-43b0-9502-5f17c29e3246 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Unresolved cited work

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:30.022962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:30.022962Z digest=sha256:a5603908b76804a494bb9ea6a6318bab89f195199af66f5afc9cba4136304f60

Observation cff38d40-0384-4dfd-b50f-4b7a1be48da5 · outbound

This paper cites Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:30.092931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:30.092931Z digest=sha256:f3de731555e6f6c808b32208557f78ea4aafc5faf8dfc17e8dc936a86d6393d6

Observation 192b405b-a9bc-482a-98b8-9088abfc0528 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Unresolved cited work

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:30.153042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:30.153042Z digest=sha256:8bd18fc175d3505c0d55a6eeee633e83f027cdcf026223c734676af7ab36156e

Observation 3e89d2e0-c61a-481a-acad-d83d5fdb580e · outbound

This paper cites an unresolved cited work.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Unresolved cited work

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:30.213135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:30.213135Z digest=sha256:9ef86fdf6e6feb60c12deb19e30e7f637a86f214611adadbef28cb2d1a808a7c

Observation fe72ef61-5397-4f90-b97a-143a3c5125c4 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Unresolved cited work

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:30.275148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:30.275148Z digest=sha256:1331e7beee4134ad223b340abbf15aa87bd3febeb1b535352e294496e6c14679

Observation 73fda0c8-0b55-468b-9458-326729616f46 · outbound

This paper cites GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:30.345279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:30.345279Z digest=sha256:04a1864ed965a0787c54dcd2d2407662239428e406352f37067bf6415291540e

Observation 51a23c9c-8d3c-461a-beaa-16414fcc4226 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Unresolved cited work

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:30.515852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:30.515852Z digest=sha256:12a99143ae0bd077a55f2919ed8ab7ac6c8c55f79503281c731401d42304a471

Observation 31792538-d656-4233-b016-9db69d3c341a · outbound

This paper cites REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:30.575799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:30.575799Z digest=sha256:49a5ddc607cc7071196e3ef0f72c7b73932dedeeea953b8e1043a72b10b48110

Observation 00ae4d2d-d592-4a4c-ab4d-5ea56b6be92b · outbound

This paper cites OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:30.691620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:30.691620Z digest=sha256:92b0eda64b9c0a1c89408e7ca967d93f71e3f901ea3b2f9f705c821b744cdb8e

Observation 8ee7e32b-36b7-4bf1-9db7-2c0e8e53d450 · outbound

This paper cites Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:30.755317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:30.755317Z digest=sha256:bdaa61fcbbd754e457e13b48e28d9eef1ccfe49ff01f2cfb19b632d6a2b6bb21

Observation 8fb8a73f-287e-4ea6-a7d3-064619ab26af · outbound

This paper cites SLOT: Sample-specific Language Model Optimization at Test-time.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle SLOT: Sample-specific Language Model Optimization at Test-time

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:30.821966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:30.821966Z digest=sha256:99e299ec3bec2cab478498d6c4964fc4d3fdbb486352a9909a709cc5a0a17a26

Observation fe90b436-1748-42c1-a9ec-edf95282059b · outbound

This paper cites an unresolved cited work.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Unresolved cited work

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:30.883862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:30.883862Z digest=sha256:6238d302d42b901ffd54b645af7914f65dfee6dec7652d2bfe61c1776018fdb6

Observation 3966b4c6-aee4-4b0f-a038-e38f848f33ca · outbound

This paper cites Boosting MLLM Reasoning with Text-Debiased Hint-GRPO.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Boosting MLLM Reasoning with Text-Debiased Hint-GRPO

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:30.943890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:30.943890Z digest=sha256:efab8c300ba1f4b1acfbac86d4caf6d7a1a2d278e712849f7aa71266c9e4aff2

Observation 460d4a8a-d80b-43fe-bcf5-c16251f52e35 · outbound

This paper cites 3D-R1: Enhancing Reasoning in 3D VLMs for Unified Scene Understanding.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle 3D-R1: Enhancing Reasoning in 3D VLMs for Unified Scene Understanding

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:31.073651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:31.073651Z digest=sha256:ba7cdc6c2c465ee8c0b110860932b6a033f091d5c2e0a4c48597b66eddabd37c

Observation 423caa36-317e-440e-bc6b-5f65775b5463 · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:31.183592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:31.183592Z digest=sha256:385c779c27dec1245508a37cd1165487d6bd43f9936f027060ce9dc3ef266f06

Observation 5285291a-8157-4889-a271-cdd9944ca92b · outbound

This paper cites GPT-4o System Card.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle GPT-4o System Card

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:31.253966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:31.253966Z digest=sha256:bae2290843c780bc043eab67995d0462a65f3f5b00e317b90cb82bbe4df1dade

Observation dc91860a-5043-4c37-a387-c2d39244a6e2 · outbound

This paper cites OpenAI o1 System Card.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle OpenAI o1 System Card

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:31.317615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:31.317615Z digest=sha256:1d7fc9826dbd68b9d346c11966e8981f7c6d535bc0995a9fe73a9b386eaee3d0

Observation 04086387-639e-4a0b-b1bf-1820ffbbd2a5 · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:31.387671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:31.387671Z digest=sha256:e2519f9d90737e067cb1a39b194ce4412bcdea10cb1bb1e0a479362e3a8b7d51

Observation 6f62552a-42ae-461a-a31e-025c52a5da8e · outbound

This paper cites A Survey on Progress in LLM Alignment from the Perspective of Reward Design.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle A Survey on Progress in LLM Alignment from the Perspective of Reward Design

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:31.458650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:31.458650Z digest=sha256:2a277d5b47c026d9ca0238dcfde0c9c35af88af09d4fbb572b65259034290128

Observation 2497ffc9-72aa-4900-8f5f-06a21319f463 · outbound

This paper cites T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:31.520711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:31.520711Z digest=sha256:132f7803dc763222e7f171abc052838367f97526b39d868b283d5dd00fb75994

Observation 11a487ad-e7a2-4edf-88fa-750f4e1bae50 · outbound

This paper cites Think Only When You Need with Large Hybrid-Reasoning Models.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Think Only When You Need with Large Hybrid-Reasoning Models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:31.525481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:31.525481Z digest=sha256:464135becb1db96376892a4c8fa171d9f214c61e0e472186052ff016c2d0083c

Observation c595a2f4-a866-47f8-9ec6-70a2fc92d731 · outbound

This paper cites A Survey on Human Preference Learning for Large Language Models.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle A Survey on Human Preference Learning for Large Language Models

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:31.560030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:31.560030Z digest=sha256:e5226b4dee5928ece9573f6b0209ed72eddf5140fcb21817984cec4cebfcc424

Observation ff833ea8-a32c-4dd3-9eb5-cbd452a1d7f4 · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:31.622495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:31.622495Z digest=sha256:1a2696766bda9cd7aea52b9e02cb24ab246f882eae1e845274a43688ea79f6c0

Observation 8d877f3d-03f5-4ab8-b772-c70b75a319ca · outbound

This paper cites an unresolved cited work.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Unresolved cited work

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:31.682882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:31.682882Z digest=sha256:b85da53db63a781627f17af4deab799e92afca629cbb27cb575059f4f7f110d8

Observation 02979552-fac8-4515-b377-b34adf2a43c8 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Unresolved cited work

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:31.731791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:31.731791Z digest=sha256:7326c71b6980ec93cde70cb1c94de38326d1bf7295f86392b89170d7119c241a

Observation 5ff89a9f-4a92-4457-9983-5683a446d207 · outbound

This paper cites BIG-Bench Extra Hard.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle BIG-Bench Extra Hard

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:31.880939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:31.880939Z digest=sha256:cad69325bb9c87902457e2becd7e27c394a8125cf630bb1d5e38ec9096fa2b11

Observation 4fca4bac-f0e6-4cbd-8166-6baed38a03e1 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Unresolved cited work

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:32.003360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:32.003360Z digest=sha256:26e2c2fa9536138c3471be4e556ba37cd3d17b918c9e44fb232e3e6926938b0c

Observation 06574205-1244-4c8b-b693-d2455cf9930e · outbound

This paper cites Alignment of Language Agents.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Alignment of Language Agents

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:32.253183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:32.253183Z digest=sha256:a62e5f24317fd44006a95c1b391c6a521f7160989a78d1226f2add9b70d42f6e

Observation a69d7b46-fc1f-49b9-8fc5-fe3c25a26548 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Unresolved cited work

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:32.342978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:32.342978Z digest=sha256:9d07d863e09066bad8b950088b917bcf4db55fd7a1d6c2035532ac4c5e5acfac

Observation 7c8251bc-4df4-40ef-98e5-96a26da71e88 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Unresolved cited work

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:32.117928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:32.117928Z digest=sha256:f7a8498cb426b7976e62ba446811fd0583c955c0df5cde7438be96f4cab04ab6

Observation c7cbecd9-fc24-49bb-99c5-b6f5dd58c1ed · outbound

This paper cites an unresolved cited work.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Unresolved cited work

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:32.558509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:32.558509Z digest=sha256:7e2e926c431b058b3b2dd00f5fed2ac7af9f4c90e081b80c7d5991fc21c0ee76

Observation d1255068-badb-4245-80ee-4c78eba503a5 · outbound

This paper cites LLM Post-Training: A Deep Dive into Reasoning Large Language Models.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle LLM Post-Training: A Deep Dive into Reasoning Large Language Models

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:32.684144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:32.684144Z digest=sha256:0837cf3e451c0beb7559e6a34a19abe8343451191ba12e00c7795942d18e572d

Observation 648ae64a-db90-4202-b7ff-5b8212cf3953 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Unresolved cited work

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:32.450828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:32.450828Z digest=sha256:bce0332727075d482e4a1b404223f7fad9627dd8560d3d1fc4a80ba93234fd5d

Observation 1d7a65e5-8847-4f1d-9a0e-78372b7d687d · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:32.864814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:32.864814Z digest=sha256:1cadc93d1e36bd5e9ddd37cc42cd1b22a476ffa63e5e75659ee4992f578cec74

Observation 84e29578-dfb1-4258-aeb6-5a4d37540d91 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Unresolved cited work

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:32.939767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:32.939767Z digest=sha256:1e1221076a375b89d93e171a2d0c7390840c2c60971df125ac6cc58776fc6dc6

Observation 08dbc974-e252-4e5d-ad6b-9a2077e914ce · outbound

This paper cites an unresolved cited work.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Unresolved cited work

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:32.796991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:32.796991Z digest=sha256:ce2859ed020083d45180b1d01429f3193186594b9bf12ea9b44dbf0a2c6abb51

Observation cfe83f2a-4ba7-47b6-aaac-e62ec57aaf9d · outbound

This paper cites an unresolved cited work.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Unresolved cited work

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:33.051469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:33.051469Z digest=sha256:6a01bcc69f5686645cc39908929bb302a51e41e15d3c38194c0857ae8514a1e3

Observation 6a0e690e-1d98-4c45-b51b-c1a304d95dfb · outbound

This paper cites an unresolved cited work.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Unresolved cited work

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:33.128629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:33.128629Z digest=sha256:3af9dbd485d743baee59dd6f78f31b15c353a264fd384a5dbc7cc37c0283989a

Observation b8cfb525-bb79-4e86-b03a-f8abf63d5b85 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Unresolved cited work

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:32.978372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:32.978372Z digest=sha256:b90538eb2a582123ceac7b93e38d4b101d3c7c3060216095388b226f850e9bca

Observation d8a6185a-fc02-4d71-8a46-355ab7680a42 · outbound

This paper cites From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:33.334341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:33.334341Z digest=sha256:bce5f8980aa30b9062c4dd356e44678745dd62e1130f22e82ce721e4afad844a

Observation ac6d0e7b-4034-494d-b8cf-bf0226db0283 · outbound

This paper cites ToRL: Scaling Tool-Integrated RL.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle ToRL: Scaling Tool-Integrated RL

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:33.433285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:33.433285Z digest=sha256:d2190a6d78783804362da4ae488b1d02cff591066f9ffd2a20884af1dbb9968b

Observation 0657594b-20a4-42e4-84e5-f1d2760997d2 · outbound

This paper cites Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:33.243783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:33.243783Z digest=sha256:e5ce43fea12a80abce658ee6ef95edc2998e80413e87e39636b38a246aff1bf6

Observation 0cb9acb6-6d08-4dee-9e45-d2c4442bcc98 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Unresolved cited work

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:33.621256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:33.621256Z digest=sha256:772591b0c8e11ea84c6c49c958c489a94fba38f79d6b64598ca8ddb5a13a84d3

Observation fb5a63db-4682-4902-abd3-5b5346672270 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Unresolved cited work

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:33.701387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:33.701387Z digest=sha256:f8595c2bc5d205cb94b4b11da76634b9dca4ee1840da296bd2738773e6d1fbc3

Observation 6f1922a0-0415-4eae-b74f-083630c02ee7 · outbound

This paper cites Generalist Reward Models: Found Inside Large Language Models.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Generalist Reward Models: Found Inside Large Language Models

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:33.529325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:33.529325Z digest=sha256:ce31ea080fd55d9c18a36dc2c823391c31d106dfee4a5a236129e5a41d3d3192

Observation 6776643c-97a0-4199-9eb2-0fd15a3547b6 · outbound

This paper cites SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:33.895807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:33.895807Z digest=sha256:c3f1c1d3d8c65a8347c398540ba8359ef5ab0bd588c778fa6cc1d91f5d1bbbc5

Observation 2a7ed2cc-0db1-417c-8b68-aae1f1a6ccf7 · outbound

This paper cites Improved Visual-Spatial Reasoning via R1-Zero-Like Training.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Improved Visual-Spatial Reasoning via R1-Zero-Like Training

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:33.995648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:33.995648Z digest=sha256:ffd0443e32f1bcbc5dbe3c72c30a4985320bd8504d022517f46904f25f690e05

Pith citing papers

Observation 39461bab-9f44-4b56-9a4f-484a762757cb · inbound

Aligning Perception, Reasoning, Modeling and Interaction: A Survey on Physical AI cites this paper.

Aligning Perception, Reasoning, Modeling and Interaction: A Survey on Physical AI Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-18T10:01:13.986090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T09:56:36.716680Z digest=sha256:dd01cb3e6fe014b273c0d4a05dbad4b388966c3a0e6f5ed1281b75985a0163d1

Observation 0e13ce9d-1ac6-4f46-ba5e-ece760393791 · inbound

EMO-R3: Reflective Reinforcement Learning for Emotional Reasoning in Multimodal Large Language Models cites this paper.

EMO-R3: Reflective Reinforcement Learning for Emotional Reasoning in Multimodal Large Language Models Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T20:14:04.060992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:14:04.060992Z digest=sha256:89e47a303a157a590f388e559308c8941478807691e4c17e12e0c41e03c7a624

Observation 8f5b7743-6b98-41d1-b985-12754fba7734 · inbound

StaRPO: Stability-Augmented Reinforcement Policy Optimization cites this paper.

StaRPO: Stability-Augmented Reinforcement Policy Optimization Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:20:59.998819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T18:11:49.056805Z digest=sha256:d38d91a9e39d84fcb1af0fc3d5798f9a40df100b126f32c37a68bf355f06991e

Observation abc577b2-6a79-417c-8b8f-d5ddc2dc6734 · inbound

Reinforcement Learning for LLM-based Multi-Agent Systems through Orchestration Traces cites this paper.

Reinforcement Learning for LLM-based Multi-Agent Systems through Orchestration Traces Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:15:37.954804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T18:44:27.685266Z digest=sha256:5cae6e1c8781714c84411014017d5e5ab0cada599562c67cac0b9912587d2783

Observation 4a6c27ff-3c2d-4266-a81b-35f03d40a2b4 · inbound

Tailoring Teaching to Aptitude: Direction-Adaptive Self-Distillation for LLM Reasoning cites this paper.

Tailoring Teaching to Aptitude: Direction-Adaptive Self-Distillation for LLM Reasoning Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-22T08:11:17.717331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T08:06:30.862911Z digest=sha256:a73c1d4f6c2ebe6f23f08a5979cdf9b1244b54bbc82a816b595dc4cb34bb28de

Observation 5ef49bc2-7813-4ba6-8c46-9dddd3f78362 · inbound

TRACER: Turn-level Regret Matching with Inner Reinforcement Credit for Cooperative Multi-LLM Reasoning cites this paper.

TRACER: Turn-level Regret Matching with Inner Reinforcement Credit for Cooperative Multi-LLM Reasoning Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-06-29T12:23:23.797662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-29T12:20:16.121824Z digest=sha256:88b219e9cd743ba7d2c9dacadf4afec52b51411c422e1f0659ac68fde362e5d7

Observation bb33b206-c858-4eb3-88ed-e13ab3645cda · inbound

Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs cites this paper.

Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-03T21:08:58.849123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T00:48:56.892634Z digest=sha256:d6f66852212abb6565eaa67fa006a83d87cffdc6e18fa1844ea3ed788df49f96

Observation 541498bc-8b4c-49a6-8fac-e794c3e3ff6b · inbound

ThinkDeception: A Progressive Reinforcement Learning Framework for Interpretable Multimodal Deception Detection cites this paper.

ThinkDeception: A Progressive Reinforcement Learning Framework for Interpretable Multimodal Deception Detection Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:49:18.220604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-26T20:58:30.855338Z digest=sha256:297476960afcdfaeaa475ce98ab6d366c79bf575af17d65dde46cfe9eab3c71a

Observation bda8178b-d2b3-48f7-adc5-2f88c545bbff · inbound

Beyond Entropy: Learning from Token-Level Distributional Deviations for LLM Reasoning cites this paper.

Beyond Entropy: Learning from Token-Level Distributional Deviations for LLM Reasoning Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-04T03:49:30.086889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-26T17:37:43.856000Z digest=sha256:7cbc51d956eef7a506a47434097cd60630ff22445455bbeb93f97e192f2c1ca2

Observation 3a2061e6-be90-4961-b661-0e466d79821e · inbound

Switch-Reasoner: Learn When to Think in Multitask Mixtures via Reinforcement Learning cites this paper.

Switch-Reasoner: Learn When to Think in Multitask Mixtures via Reinforcement Learning Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-10T05:16:48.097060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T05:07:40.106099Z digest=sha256:1e598d91395301fe755d3343fd28f6b9f8de225b2e6d7ff49782de452c7b2baa

Observation 725f57e9-41a6-4f37-b26a-26ef05d0775f · inbound

Spectral Origins of the Self-Correction Blind Spot in Autoregressive Generation cites this paper.

Spectral Origins of the Self-Correction Blind Spot in Autoregressive Generation Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-14T15:30:23.485228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T15:30:23.485228Z digest=sha256:865869c18f1ed04c9a043c60fb47bc33e7f227722dab1bbbcb59987f5b8965c6