Pith. sign in

Paper Citation Record · LEDGER

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning

As of 7 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 2 inbound Pith citation observations for arXiv:2507.14295.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.14295 v2

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T16:14:09.490536Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T00:13:26.854490Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-28T23:22:47.413783Z

Reference resolution

57 of 57 outbound references displayed

  • verified exact2
  • verified fuzzy1
  • unresolved54
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 17eae788-e7b5-4906-b24d-02c7ea37ba91 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:02.177595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:02.177595Z digest=sha256:4ed5e97f1ac2e692998c6c871d594b523aa203ec8d739d77a1c1d40c91636d27

Observation ebcb6810-5c35-407c-b9ff-cab89baa01bc · outbound

This paper cites GPT-4 Technical Report.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:02.291984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:02.291984Z digest=sha256:842069962cee4ff8ed2d49798ca9242417d8707559eb17f236ae4a993d3927df

Observation 3df69451-78be-4620-ae58-f933f98096bf · outbound

This paper cites Qwen2.5 Technical Report.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning Qwen2.5 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:02.416492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:02.416492Z digest=sha256:932e2f9a1ce846f4c397929e2073dfeda8565b88361bb4a874c6c4984dddbfc4

Observation 4160e1db-202c-4c32-ac9c-371e39e8193d · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning Gemini: A Family of Highly Capable Multimodal Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:02.648526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:02.648526Z digest=sha256:4a77cb237de257034a2e1e403fdbfe863b8aba4b4c1b7141fcbe9dc782fcebe2

Observation ce9e97a8-c8be-469b-b310-bf62adf99a20 · outbound

This paper cites Proximal Policy Optimization Algorithms.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning Proximal Policy Optimization Algorithms

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:03.913893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:03.913893Z digest=sha256:fc31af2808fd391e890f3862fdc1cb6a68c6592725c26b5a22851ee086b3294d

Observation 6de5f99d-b149-450f-a171-fb3dcc223bb0 · outbound

This paper cites ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:04.049831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:04.049831Z digest=sha256:2cf65d033900195a7ed3eacf1bea94644f1d7aa2b677bdf3b33a8c096b9cab90

Observation 69b166d7-93b9-4d6c-8956-1bad9a1ff0c0 · outbound

This paper cites RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:04.197657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:04.197657Z digest=sha256:9533e929b7ea0799e4a83c2b4a3f887815f098e3b9f7f868723c2a644bf674a6

Observation 7a687a27-e2c6-419d-8403-c693e0e42247 · outbound

This paper cites OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:04.352421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:04.352421Z digest=sha256:60d18649789b7430e607700d9a4617d94ff58b591121a0358e6dfa20d430cb15

Observation 47cd9019-307f-493d-8c78-e3e39d7fd96c · outbound

This paper cites Training Software Engineering Agents and Verifiers with SWE-Gym.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning Training Software Engineering Agents and Verifiers with SWE-Gym

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:04.482017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:04.482017Z digest=sha256:02bb6802ecaf01cab47ecbcccbf8e5ef6fc0c326cebfecf518a0b9bca9f69efc

Observation 1cdbd948-6853-4fe3-8dc6-4761fff620e4 · outbound

This paper cites WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:04.685622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:04.685622Z digest=sha256:63c6bffa92e39dd191a611c5e073a1ccd99a102d79bdb2a724d418a864d79986

Observation 871f8b7a-38b5-4a3b-ba7f-27d4ff4df2bb · outbound

This paper cites ALFWorld: Aligning Text and Embodied Environments for Interactive Learning.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning ALFWorld: Aligning Text and Embodied Environments for Interactive Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:04.795371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:04.795371Z digest=sha256:f7d671b6f830fa635de707b0aca0d7c54fe95ac9cb4326a64bb76cc7831de0b0

Observation 7e0cb956-51fa-47cc-be48-1410d9573aba · outbound

This paper cites MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:04.918820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:04.918820Z digest=sha256:3ecf6b4ddbb17ffa5da3bf236b2c602518e0fabb25204b563831d5fb1c486927

Observation ce9125b9-71b3-4eb9-9426-ede0b7403805 · outbound

This paper cites Simworld: A world simulator for scaling photorealistic multi-agent interactions, 2025.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning Simworld: A world simulator for scaling photorealistic multi-agent interactions, 2025

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:14:10.660770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T16:14:05.133546Z digest=sha256:2630f086fb571797d747152b4e861044e856e145c7ae0a0b22f13f289a421903

Observation 3ddd5e56-30c2-468b-955b-142c29dd4d9b · outbound

This paper cites Gonzalez, and Ion Stoica.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning Gonzalez, and Ion Stoica

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:05.285192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:05.285192Z digest=sha256:da02eb7bf2ecce795613a03cf8b742582367aad5fa3a32c2a9de8674e241b743

Observation ebae5172-5e07-4063-b061-45322dd9e06b · outbound

This paper cites Training language models to follow instructions with human feedback.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning Training language models to follow instructions with human feedback

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:05.398814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:05.398814Z digest=sha256:0c6500cf8c551ea05e31c0972be86cc8b7fef01b54085fb2a9ed5a14cad4f433

Observation 5ae237ca-edbb-438a-b45f-14e8dbdf840c · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:05.544503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:05.544503Z digest=sha256:7057f04052c83b744f6b762038477883d6ea48030dc76af20345a735d67c1d16

Observation 2cac3e1e-7103-4240-bed3-e2854c459812 · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:05.665511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:05.665511Z digest=sha256:62dd67dffff7f8f24c29fdc385d02dd5bdbd9d1444b1af87a4fc725d4c6ffe80

Observation 5ae4401e-7a2f-4c3c-9d3a-980c350a22ea · outbound

This paper cites ReTool: Reinforcement Learning for Strategic Tool Use in LLMs.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning ReTool: Reinforcement Learning for Strategic Tool Use in LLMs

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:05.766444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:05.766444Z digest=sha256:128f1fa0cf30f27e3449867a723c1968f4491e072331384d7305c5689d3ddbc4

Observation 8574136e-8b3e-4922-bc44-b2c01fd2e3a9 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:05.889070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:05.889070Z digest=sha256:040e543f4aee784fa9d51c17c7d4014e81fa779e15e7c4d2a4f941f8324df6fe

Observation 298c3469-3378-4fea-bdf6-944685f4ca02 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning Understanding R1-Zero-Like Training: A Critical Perspective

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:06.014532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:06.014532Z digest=sha256:6febe427f41157e9f644c2ee9aa2e5ad6fdc85bcd6a5e853068142fd15c168ec

Observation a1fbccc2-743f-43f9-b951-47742968c637 · outbound

This paper cites Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:06.124607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:06.124607Z digest=sha256:65cfe1b5363e0bb707a94d3101223f721174ba0bbee9ce6c6387b311d10df5d4

Observation 181581e7-9293-4cf1-95a6-cbf1d8f68053 · outbound

This paper cites HEART-felt Narratives: Tracing Empathy and Narrative Style in Personal Stories with LLMs.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning HEART-felt Narratives: Tracing Empathy and Narrative Style in Personal Stories with LLMs

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:06.230944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:06.230944Z digest=sha256:71bbb31c88dc751ebc199e8ba7a9273319b2fb007f9388d4496956fe7124a95f

Observation 73dc1681-ec6e-4b8f-9a9b-84e118096ee2 · outbound

This paper cites TheoremQA: A Theorem-driven Question Answering dataset.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning TheoremQA: A Theorem-driven Question Answering dataset

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:06.336764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:06.336764Z digest=sha256:18916fb6c8830223f1df48c55024259f434ad2dd7c31b6da4154c4812f0d5697

Observation daabcfd0-3f11-408b-af82-a64257058c1c · outbound

This paper cites HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:06.442279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:06.442279Z digest=sha256:f5e52c9a35f7258b5289f505a270afc13072f2ae408ddcf0cdd21f11927328f3

Observation e41177a8-a2d0-47bd-937b-617b59d41f6e · outbound

This paper cites Reasoning over Public and Private Data in Retrieval-Based Systems.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning Reasoning over Public and Private Data in Retrieval-Based Systems

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:06.552532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:06.552532Z digest=sha256:f12940ac050aaff1d92c3a748e6f7d229c03d71ceefebeade6758edd54a97738

Observation 49182346-dc44-4c8e-b8e1-319a2d0820e8 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning Measuring Massive Multitask Language Understanding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:06.625049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:06.625049Z digest=sha256:14f6325863eea456bf3ac64af663674313ece583a584b1e71e5ec3fefaba5de1

Observation afae9fdf-48b2-4f3c-856e-656045a2e324 · outbound

This paper cites MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:06.701604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:06.701604Z digest=sha256:fcb1aa991c2a85be836aaab5ff3f8e0b615479a3c6e0471d6386dbd22aa0ced6

Observation 951c7202-42bb-417f-9270-96f43c8cebe4 · outbound

This paper cites Graph of Thoughts: Solving Elaborate Problems with Large Language Models.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning Graph of Thoughts: Solving Elaborate Problems with Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:06.824346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:06.824346Z digest=sha256:d8233cc6d9899b9f18798e4dae3837b8875386fd6e518d6ed40d2c6efe9e8400

Observation bc84f671-1e7f-4305-8939-c34d514f0e3a · outbound

This paper cites Reflexion: Language Agents with Verbal Reinforcement Learning.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning Reflexion: Language Agents with Verbal Reinforcement Learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:06.901577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:06.901577Z digest=sha256:29ded6cea0a168c84ab585bb365c1461217300eca995355dfba61c1fe5094243

Observation 5aeb5db2-1b6c-4c40-84b4-26c61b1f9262 · outbound

This paper cites Accessing GPT-4 level Mathematical Olympiad Solutions via Monte Carlo Tree Self-refine with LLaMa-3 8B.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning Accessing GPT-4 level Mathematical Olympiad Solutions via Monte Carlo Tree Self-refine with LLaMa-3 8B

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:06.999304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:06.999304Z digest=sha256:093877609bf1c6875ed30e5a3ecf57555cf982e5b98f9b898e8bb8ccd49cc686

Observation 7d33ccf3-a93e-4e24-98cf-a71390969d14 · outbound

This paper cites Self-Refine: Iterative Refinement with Self-Feedback.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning Self-Refine: Iterative Refinement with Self-Feedback

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:07.142246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:07.142246Z digest=sha256:a4013c7f4e7a7fdda4f58680a031ae7d3994a0134b0bbb03ddd9f17cdb398b66

Observation ccc5a4b1-fbe3-434b-bd90-85dc181f763e · outbound

This paper cites CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:07.252826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:07.252826Z digest=sha256:32f45fbd463f392d0860e358efaef3eee3a9a518e0444043528295a9fd1f0dd7

Observation 3c18fef3-5c6c-42f0-bad3-34cd8047f5e7 · outbound

This paper cites Large Language Models Prompting With Episodic Memory.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning Large Language Models Prompting With Episodic Memory

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:14:10.255823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T16:14:07.377051Z digest=sha256:bc6d1d45baf6ebd0a1e09347bf65a9b8f062b889d0af09953e85dae52f8e6818

Observation 3cb608d4-b8b8-4cbe-9b45-b488bf86d5c5 · outbound

This paper cites Larimar: Large Language Models with Episodic Memory Control.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning Larimar: Large Language Models with Episodic Memory Control

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:07.482480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:07.482480Z digest=sha256:1056b06951e41e797349213527f99ea33733b314cbef8f59216eed1a99aab2fc

Observation 4344a43d-7ea8-4d60-8166-252265d65b68 · outbound

This paper cites Deep reinforcement learning from human preferences.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning Deep reinforcement learning from human preferences

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:07.581685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:07.581685Z digest=sha256:7961be613f7817dbb15ec32a04c35eca1a3538eec40a9c7fd56e0a5a28e962a8

Observation 2e6002a7-da98-419a-a470-c755754598ea · outbound

This paper cites RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:07.594492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:07.594492Z digest=sha256:296fd99ca21062e8409a63e7ca09c17683301b33bf1c9226685c6b48ea4a7d79

Observation 48114e8b-2682-435f-b466-1d18e8140ded · outbound

This paper cites On scalable oversight with weak LLMs judging strong LLMs.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning On scalable oversight with weak LLMs judging strong LLMs

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:07.651045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:07.651045Z digest=sha256:1abd89966b08f8451e708a63771dea64496bdc77958acf4aca9f3ea4cf1bb689

Observation 194dec1d-eefe-415e-b1a5-da239a12b62b · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:07.746870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:07.746870Z digest=sha256:ae7046988213e8b167f9b53f164d2e1ecf2a35bd68474ba60c28356481297fdf

Observation 4ac7c2ab-2fef-4fce-a820-1cde53e3f2d8 · outbound

This paper cites Parameter Efficient Reinforcement Learning from Human Feedback.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning Parameter Efficient Reinforcement Learning from Human Feedback

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:07.820412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:07.820412Z digest=sha256:c9f036dad349cb4971d0fd2a40a9279588d035e3fb3a7b96ed281ee73628ad06

Observation 10ca8d94-9bbf-4789-a5a4-e5ad3429da5e · outbound

This paper cites Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:07.896851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:07.896851Z digest=sha256:3df89ecbeb9928f9f59dd515e3daac74f36cda223a5e0063b797a14b0e6bcaa6

Observation 4fc4a1fe-2bfc-4f79-8670-f9e59b038f75 · outbound

This paper cites UNO Arena for Evaluating Sequential Decision-Making Capability of Large Language Models.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning UNO Arena for Evaluating Sequential Decision-Making Capability of Large Language Models

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:14:09.915290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T16:14:08.019297Z digest=sha256:e0d71da5166b26ad16f2cd32e450968ef80993162557dc272a834031d0b54a03

Observation a65de3cd-6f1a-45ac-a9de-7ab66f3fa671 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning Training Verifiers to Solve Math Word Problems

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:08.106325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:08.106325Z digest=sha256:75deb66eaf5e1c3fb812f499f5dc7119949e0ede891424e0591356f0097f9a43

Observation 767d4daa-4474-44d3-9b51-91a5bf50c929 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning Measuring Mathematical Problem Solving With the MATH Dataset

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:08.189265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:08.189265Z digest=sha256:9ffab190372cd3b79c6c61bb882bf343e1b2ce0f6a413b0797ce9adb4bb27596

Observation 88d476c0-e02e-47c5-a815-cd4a96ff86cf · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:08.303551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:08.303551Z digest=sha256:4cea1d6af773b6b7141921e0afb597f8d737328d0e8f37a4fb89ba486349de33

Observation 6b3fe19b-6338-4233-8b83-17bb197f16bb · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:08.374932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:08.374932Z digest=sha256:4622bb9b12d708c6ff217b42af02a0502f8d5248d17c74679630585bc6c5c277

Observation 7b333d75-1361-4f5f-93e1-2b7aa7042db9 · outbound

This paper cites Tree of Thoughts: Deliberate Problem Solving with Large Language Models.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning Tree of Thoughts: Deliberate Problem Solving with Large Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:08.466387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:08.466387Z digest=sha256:758835a26290a5f8d1ed5f7cae796aa842c99cccedb793f234cc7452a55b5b42

Observation 86811f4b-5b09-410d-a948-d896f4aa6c56 · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning ReAct: Synergizing Reasoning and Acting in Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:08.565734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:08.565734Z digest=sha256:0b9896a8babb0a78bfab7c974df3f3beb73d3645976dd783b344363bee3ae80e

Observation f9866ada-72fb-404f-a504-70cb42a61217 · outbound

This paper cites PAL: Program-aided Language Models.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning PAL: Program-aided Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:08.667441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:08.667441Z digest=sha256:def7e4de448740f1db9f52ad33deec68626b516ecae27f05bc7383471557c82c

Observation 5e7c3e03-f7f3-46b8-bced-357d7694b112 · outbound

This paper cites Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:08.740558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:08.740558Z digest=sha256:3960e9f308022ac9c706c73aa109fe7f9b33248866c6c12295385a2585d36ffa

Observation 15b7b6c7-8e84-4df7-a904-398435224e19 · outbound

This paper cites Toolformer: Language Models Can Teach Themselves to Use Tools.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning Toolformer: Language Models Can Teach Themselves to Use Tools

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:08.865042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:08.865042Z digest=sha256:1cd964e17e407ebe1d167d964bf7d1f8dca46dd0a71c12483a9a95f88b14b59a

Observation 4413c3a9-3758-45c7-8799-5a54cd77fcc4 · outbound

This paper cites Let's Verify Step by Step.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning Let's Verify Step by Step

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:08.959676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:08.959676Z digest=sha256:c9637f6b3665a233bd598218ea2e093995f619b4bb4adc426c577a404d5bfc7f

Observation 144f1987-5e3c-4b69-8ac1-a239733e5a99 · outbound

This paper cites AutoPSV: Automated Process-Supervised Verifier.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning AutoPSV: Automated Process-Supervised Verifier

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:09.056162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:09.056162Z digest=sha256:7a2b91205c4beaf36ab7fe22a08301f2ada14e63ec814c4bc3cd78d9065bea04

Observation d16cba06-c5ee-465b-89a1-82f9c7880e9c · outbound

This paper cites Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:09.175411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:09.175411Z digest=sha256:2a95c9cff35001e960f01b8f766dcb7e37cf7a724a67a14da4213e93f369813c

Observation 972a76c5-c47e-4911-9642-3e5d1b12bd12 · outbound

This paper cites Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:09.245253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:09.245253Z digest=sha256:82abd7ca3d8852861e23fb9dc92a623b9e434b9285e6c7175342458e49bc9aa3

Observation acf55dc0-563b-4636-8500-950e53af27a0 · outbound

This paper cites Ursa: Understanding and verifying chain-of-thought reasoning in large language models, 2025.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning Ursa: Understanding and verifying chain-of-thought reasoning in large language models, 2025

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:09.305497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:09.305497Z digest=sha256:3dee649704dccc17840c273d973cc94d6de477c9ca1fe5b7b9add709e6e28f3a

Observation 35ee2956-b30f-49e9-a549-c2ea74da1a31 · outbound

This paper cites Large Language Models Cannot Self-Correct Reasoning Yet.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning Large Language Models Cannot Self-Correct Reasoning Yet

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:09.393008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:09.393008Z digest=sha256:ad6f2283760937cdd690d263640122a8591ac665863b337c7a2498554f9ef1f3

Observation 1dfd7eaa-d8bf-485e-9537-a01cb4c3251c · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:09.490536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:09.490536Z digest=sha256:1afe943446d219a0be01a3e49ef57172c1302deb2cf4c9e7262afdf1e597fd58

Pith citing papers

Observation f75ca2ed-f3e2-4825-8c23-754567efa4a3 · inbound

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization cites this paper.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:26.854490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:26.854490Z digest=sha256:b710537f79d04f22373784e5429742132eb5f4a3eb8a7a0b169e4af4a9825033

Observation 223af9c4-7897-4da4-9fb1-83288a4d7796 · inbound

DRIFT: Decoupled Rollouts and Importance-Weighted Fine-Tuning for Efficient Multi-Turn Optimization cites this paper.

DRIFT: Decoupled Rollouts and Importance-Weighted Fine-Tuning for Efficient Multi-Turn Optimization A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-06-28T23:22:47.415781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T23:16:49.358792Z digest=sha256:babd638bf086c1796b3b44f130e130e9d65db9f40ab388d5bc18eca74255c848