Pith. sign in

Paper Citation Record · LEDGER

Effective Reinforcement Learning for Reasoning in Language Models

As of 13 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 2 inbound Pith citation observations for arXiv:2505.17218.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.17218 v1

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:56:18.716371Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T12:06:50.901959Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T08:04:29.012591Z

Reference resolution

41 of 41 outbound references displayed

  • verified exact2
  • verified fuzzy0
  • unresolved39
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f2fa5f58-4f1f-4daf-b021-5e534733bb6f · outbound

This paper cites online" 'onlinestring :=.

Effective Reinforcement Learning for Reasoning in Language Models online" 'onlinestring :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:14.897391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:56:14.897391Z digest=sha256:c51b1564e5d985fec2f27a5c410862ebc53bb82298dfa8d77a25a336558542d4

Observation 97841ca2-85be-437f-b2a4-1f9298a55adc · outbound

This paper cites write newline.

Effective Reinforcement Learning for Reasoning in Language Models write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:14.982183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:56:14.982183Z digest=sha256:8f8c06e163b8fa951bccd301d5b62db61c10c4b6e8b23aea1b6df8364a32ec9b

Observation 8e0692bc-14ce-48d8-b019-878a3bfab9f4 · outbound

This paper cites Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs.

Effective Reinforcement Learning for Reasoning in Language Models Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:15.099420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:56:15.099420Z digest=sha256:3dd024f732b0f1b4cb570da5df065b24b041855e518056fa2c4df7a83d0a83f8

Observation 301d928e-21e4-4ad1-be3a-a639961f7515 · outbound

This paper cites Program Synthesis with Large Language Models.

Effective Reinforcement Learning for Reasoning in Language Models Program Synthesis with Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:15.184375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:56:15.184375Z digest=sha256:f101fff5aafbcb4d1a227be17ae7753fc1f1418fc9f21ee5ca51a424ffe57fcc

Observation 81179bea-7776-4a74-bbd2-e377439404e0 · outbound

This paper cites an unresolved cited work.

Effective Reinforcement Learning for Reasoning in Language Models Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:15.274117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:56:15.274117Z digest=sha256:00a21fc83889d023c61fc0eaee1fdccffb83882e73636273a0e1b2cc938feccc

Observation 3a52da57-1d21-4abd-889f-50538f316495 · outbound

This paper cites FireAct: Toward Language Agent Fine-tuning.

Effective Reinforcement Learning for Reasoning in Language Models FireAct: Toward Language Agent Fine-tuning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:15.384955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:56:15.384955Z digest=sha256:4a618298fd24c39144fa7fa5af18d4aff0b060fa90ddd24c487aef1d43f82709

Observation f7537b4e-1f81-47a5-bbc7-9a03136413ba · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Effective Reinforcement Learning for Reasoning in Language Models Evaluating Large Language Models Trained on Code

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:15.576588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:56:15.576588Z digest=sha256:220282b5a314da9c2c25e5ab508ac0d7e32c51487ad92218ad97c1827353086a

Observation 39563835-0abf-4359-9613-2f588b7937b3 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Effective Reinforcement Learning for Reasoning in Language Models Training Verifiers to Solve Math Word Problems

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:15.755879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:56:15.755879Z digest=sha256:87f009c34457375b7e809912edc0906f5f9634be32314f7d6a8b4afa05ba4ded

Observation f966a498-ec37-4cf1-9e01-3eccea43e577 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Effective Reinforcement Learning for Reasoning in Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:15.955208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:56:15.955208Z digest=sha256:4502d3a9ba4e5466efebc5aac8e4dcc09b5fdebec89dccfdd0307d5850603c9e

Observation 5c1df620-03e1-4f3e-8d84-d1bd9aef3e45 · outbound

This paper cites DeepSeek-V3 Technical Report.

Effective Reinforcement Learning for Reasoning in Language Models DeepSeek-V3 Technical Report

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:16.075367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:56:16.075367Z digest=sha256:5e0318a2f7d6986f45f0bd59e2f1282029457f1ff3b9972ec57d85689e98fb2a

Observation c0f4e25b-0026-4598-9290-4d4737d2310e · outbound

This paper cites FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI.

Effective Reinforcement Learning for Reasoning in Language Models FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:16.257566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:56:16.257566Z digest=sha256:351377aa3dba419eb08b58288ebbb8444f5b0367a463ec0b5e978b4f3beff752

Observation 050c12db-bcab-42ea-804c-9a2547a3080e · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Effective Reinforcement Learning for Reasoning in Language Models Measuring Mathematical Problem Solving With the MATH Dataset

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:16.401585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:56:16.401585Z digest=sha256:18a3c2dcb43b7a4ef5e0f7cc427457f4d4570b578e76b22cb09b2aa16a6b7b3c

Observation 6eccfac8-72d1-48ae-a614-79ee08b6f4d0 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Effective Reinforcement Learning for Reasoning in Language Models LoRA: Low-Rank Adaptation of Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:16.492420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:56:16.492420Z digest=sha256:519cb65e548b8da9cf8abb31533d0ff20f4433c1ed4048784a00abc25d56db49

Observation 36ce54ae-9fc4-43f3-b569-e7d0796fb5b8 · outbound

This paper cites Way Off-Policy Batch Deep Reinforcement Learning of Implicit Human Preferences in Dialog.

Effective Reinforcement Learning for Reasoning in Language Models Way Off-Policy Batch Deep Reinforcement Learning of Implicit Human Preferences in Dialog

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:16.588380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:56:16.588380Z digest=sha256:80b4d184681500b2c6ca1af68eb83ab6faa39cc5230ddbf340a92ed081ddb26a

Observation 63b5de8a-0e45-4078-b03c-0c1fcbee1a0e · outbound

This paper cites VinePPO: Refining Credit Assignment in RL Training of LLMs.

Effective Reinforcement Learning for Reasoning in Language Models VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:16.694006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:56:16.694006Z digest=sha256:5132e6546cc426000789e4ced55821d34e495d6a8975f435320fad17e10579ec

Observation 7b12b335-90f4-4257-9d3f-91771ac5fa9d · outbound

This paper cites Efficient Memory Management for Large Language Model Serving with PagedAttention.

Effective Reinforcement Learning for Reasoning in Language Models Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:16.814005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:56:16.814005Z digest=sha256:ffd7d894a83c76ccf7692fc248d254840c3140df109184fdc729e24b259881a0

Observation 733bbca4-b1e2-4ed9-8c6a-204e78ee2a8f · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

Effective Reinforcement Learning for Reasoning in Language Models Gonzalez, Hao Zhang, and Ion Stoica

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:16.939493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:56:16.939493Z digest=sha256:772769662e7daa4f20ef6e2cdac8fb2a69dbe1fdefe420dd006f460cfb84f5be

Observation e517c5bc-b108-4c77-9ac9-9608e45ca82c · outbound

This paper cites Reducing Variance in Temporal-Difference Value Estimation via Ensemble of Deep Networks.

Effective Reinforcement Learning for Reasoning in Language Models Reducing Variance in Temporal-Difference Value Estimation via Ensemble of Deep Networks

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:56:19.192058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T14:56:17.028265Z digest=sha256:7f177e28bca7cde6400299a91e9a28724f4238ba9a89daa874690e5f95540238

Observation 270e01a2-50f5-47df-a1bd-76d15e791718 · outbound

This paper cites Let's Verify Step by Step.

Effective Reinforcement Learning for Reasoning in Language Models Let's Verify Step by Step

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:17.083444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:56:17.083444Z digest=sha256:d6da6e4590e1fb4f68f47b47307355f1608b72da19c0a601d9b68c353285b2b9

Observation 042a3e37-ed0d-4921-8777-21f179c4531d · outbound

This paper cites an unresolved cited work.

Effective Reinforcement Learning for Reasoning in Language Models Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:17.162071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:56:17.162071Z digest=sha256:14b2958390d03a5f64a10cccb3c8ec8f588d372cfbfbb1851e911775804e7ac2

Observation 4e601eb6-26b0-4b54-bddb-8822e369e2dd · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

Effective Reinforcement Learning for Reasoning in Language Models Understanding R1-Zero-Like Training: A Critical Perspective

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:17.239533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:56:17.239533Z digest=sha256:1ced880921bba6d292bb54d4f8ded28e1ac1fe7e32a48a4792c51458e02271b0

Observation 814140a9-8ba0-4f56-bf05-c33cc0f80a66 · outbound

This paper cites s1: Simple test-time scaling.

Effective Reinforcement Learning for Reasoning in Language Models s1: Simple test-time scaling

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:17.321965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:56:17.321965Z digest=sha256:7f62a8c735430f309220d469c968c9179fb4315438554086c2fd9e78f0fc5893

Observation 43182d4f-8861-4cee-83b8-1500f8c0036a · outbound

This paper cites Self-Imitation Learning.

Effective Reinforcement Learning for Reasoning in Language Models Self-Imitation Learning

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:56:19.033963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T14:56:17.392954Z digest=sha256:71c8a7e1f7d7d1cea6d3a30e28a719a3295d0958ef288c906db5249831dce861

Observation 3f68f129-c728-4c2e-a650-2d86c599d465 · outbound

This paper cites an unresolved cited work.

Effective Reinforcement Learning for Reasoning in Language Models Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:17.461424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:56:17.461424Z digest=sha256:785a198edfa65bb394247ffb4e0d36d0e72f250ab7f1f5e27830e825bc702506

Observation 545294a5-5926-460a-b977-543dba9b152a · outbound

This paper cites CodeNet: A Large-Scale AI for Code Dataset for Learning a Diversity of Coding Tasks.

Effective Reinforcement Learning for Reasoning in Language Models CodeNet: A Large-Scale AI for Code Dataset for Learning a Diversity of Coding Tasks

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:17.548731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:56:17.548731Z digest=sha256:ac568e1c5ff207bab4a7b08b7b2ce0e41325cd2ed822ee27ee4c4b993c2664fe

Observation 780eb4ef-b32b-40c1-95a6-d78ffba76438 · outbound

This paper cites Qwen2.5 Technical Report.

Effective Reinforcement Learning for Reasoning in Language Models Qwen2.5 Technical Report

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:17.628210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:56:17.628210Z digest=sha256:c15781c9092f3b9e13f683f12b563914c2f20637e1b8c8e689c534053d99e887

Observation 8e56e107-4a44-4340-ba18-a94071a5042e · outbound

This paper cites ZeRO: Memory Optimizations Toward Training Trillion Parameter Models.

Effective Reinforcement Learning for Reasoning in Language Models ZeRO: Memory Optimizations Toward Training Trillion Parameter Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:17.692371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:56:17.692371Z digest=sha256:9d516aea524a7f062f3b218c70bd5f99ed7c45d5909246ea6018bcf4903d069b

Observation 639f44f9-78aa-41e0-aa57-6ad4818c4c78 · outbound

This paper cites Trust Region Policy Optimization.

Effective Reinforcement Learning for Reasoning in Language Models Trust Region Policy Optimization

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:17.748538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:56:17.748538Z digest=sha256:507d8b4197af65ef29d26f1ed42fdc70410ad55e47cfa8418acd5b0b664546a8

Observation 2dfa1d65-2157-41e4-9722-60757a500681 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Effective Reinforcement Learning for Reasoning in Language Models Proximal Policy Optimization Algorithms

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:17.830451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:56:17.830451Z digest=sha256:5c8af88a00c1735d86d3ad0219e7001555b2fa48d7f49b1f707a51491d8ee328

Observation ac161ae9-b028-48b7-a580-d042a309d5dd · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Effective Reinforcement Learning for Reasoning in Language Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:17.886290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:56:17.886290Z digest=sha256:ab01a673f36cc77869fead2f2959fb8354e38904a3386f1290c295beb687e83c

Observation 4a691dc3-28f2-4255-be38-cd450d3a1558 · outbound

This paper cites Reflexion: Language Agents with Verbal Reinforcement Learning.

Effective Reinforcement Learning for Reasoning in Language Models Reflexion: Language Agents with Verbal Reinforcement Learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:17.978553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:56:17.978553Z digest=sha256:f727d157095f11b3bd392849cf1b4afd9e1fc536aa8e9398e311baf077a9421b

Observation 7cbbb8d0-7fd2-4f2a-b8aa-45c296fd114c · outbound

This paper cites an unresolved cited work.

Effective Reinforcement Learning for Reasoning in Language Models Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:18.023760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:56:18.023760Z digest=sha256:1a713cdb57d427e6733a2a8257db053b181cf19275a500ebc5c3373ff4f54065

Observation 5a2659db-8d4f-42df-ae76-e81e7824c3df · outbound

This paper cites Sutton and Andrew G.

Effective Reinforcement Learning for Reasoning in Language Models Sutton and Andrew G

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:18.097191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:56:18.097191Z digest=sha256:9d305a5df03859c264900ad948b0d2aaea1436318d117295e002718a7cfc77ec

Observation 6da98b5e-23a4-40ed-b66b-21c67f6ce3bd · outbound

This paper cites an unresolved cited work.

Effective Reinforcement Learning for Reasoning in Language Models Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:19.435126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T14:56:18.186529Z digest=sha256:4165e4956acbdbe3d72d819a907f0178b1b78939fac9b8599418bd68ca4e5cb2

Observation 9fbd850b-24ad-4320-b377-21e8ef4a88a9 · outbound

This paper cites Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations.

Effective Reinforcement Learning for Reasoning in Language Models Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:18.266807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:56:18.266807Z digest=sha256:101ed08c5b161bc565a1e6b6105f994e8c0306272707f2d3ce13c14f8f4e1182

Observation c16523a4-1438-4de7-833b-e12e8e102578 · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

Effective Reinforcement Learning for Reasoning in Language Models Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:18.335444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:56:18.335444Z digest=sha256:5fffb3234f20028874ac68e29cbb953f35c68cfbc0a6a10bdec5ae1ab9c98fef

Observation fe47ffdc-66f5-4879-8e9c-6fde23034d61 · outbound

This paper cites Base Models Beat Aligned Models at Randomness and Creativity.

Effective Reinforcement Learning for Reasoning in Language Models Base Models Beat Aligned Models at Randomness and Creativity

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:18.386633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:56:18.386633Z digest=sha256:ffb65e5d8863662ef1805d7b6acd0af4753aec9f2949f44c2ce293aaa29ba04a

Observation fd47ed01-d16f-429a-a201-48e1db98f6d8 · outbound

This paper cites Tree of Thoughts: Deliberate Problem Solving with Large Language Models.

Effective Reinforcement Learning for Reasoning in Language Models Tree of Thoughts: Deliberate Problem Solving with Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:18.475326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:56:18.475326Z digest=sha256:45ccadd8963b39c140d84a15f080598a7ece3bdd621fa90d90379bd0db9a5d64

Observation 26c10c2e-0fc2-438c-8532-43fa80c9e67c · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

Effective Reinforcement Learning for Reasoning in Language Models ReAct: Synergizing Reasoning and Acting in Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:18.537200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:56:18.537200Z digest=sha256:a749f96fba2460dfc8c556d80ca4d95cc109e7a33f28acb9e1e96724386128ca

Observation 4eca98ce-8d50-4895-9a5f-eb59a07aa06b · outbound

This paper cites AgentTuning: Enabling Generalized Agent Abilities for LLMs.

Effective Reinforcement Learning for Reasoning in Language Models AgentTuning: Enabling Generalized Agent Abilities for LLMs

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:18.612262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:56:18.612262Z digest=sha256:c92f48d2173b6d1c7c2c605680ca04914ffc2c598ef179eb749abd901d022ad9

Observation 25824708-5506-4a7e-bee0-b07e9bc130f3 · outbound

This paper cites SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild.

Effective Reinforcement Learning for Reasoning in Language Models SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:18.716371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:56:18.716371Z digest=sha256:e06168746a7658a49cfe4c3cdc1523514d905de55c260bea69366441e68e645e

Pith citing papers

Observation 994ae51c-3fec-4000-86d4-ddd81495384d · inbound

PS-PPO: Prefix-Sampling PPO for Critic-Free RLHF cites this paper.

PS-PPO: Prefix-Sampling PPO for Critic-Free RLHF Effective Reinforcement Learning for Reasoning in Language Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-06-30T08:04:29.013956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-30T07:36:39.835616Z digest=sha256:629ae639df54638929ca2600a3ba18448c29575c24f85643a95c01affe76a432

Observation c7e51b63-acdb-44de-a9fd-2741712cc433 · inbound

SLPO: Scaling Latent Reasoning via a Surrogate Policy cites this paper.

SLPO: Scaling Latent Reasoning via a Surrogate Policy Effective Reinforcement Learning for Reasoning in Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T12:06:50.901959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:06:50.901959Z digest=sha256:09b41a6f11d37f285d3b11baef3773911de7a5914ac38d0650c9a7e32461430b