Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T11:57:49.117837Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 84 of 84 outbound references and 14 inbound Pith citation observations for arXiv:2502.02508.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T11:57:49.117837Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-08T14:40:36.038438Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T05:57:41.429120Z
84 of 84 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 1229900a-f24d-4041-8375-33a12896dd33 · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search ( s + 2) ( 2.4− t 60 ) = 9 Let’s solve these equations step by step
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ad843070-ac30-4804-a4dd-5e40acc01e59 · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Self- consistency improves chain of thought reasoning in lan- guage models,
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22b49c8f-5b87-473f-b689-220da3b35a17 · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search O1 Replication Journey -- Part 2: Surpassing O1-preview through Simple Distillation, Big Progress or Bitter Lesson?
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4c20d2e-04df-4ed4-ad19-030622c80607 · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f07620ed-7ca1-4ac4-9df6-786b49367384 · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Mutual Reasoning Makes Smaller LLMs Stronger Problem-Solvers
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2ac63d0-277d-4136-855b-3533293a6cef · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c692f436-95e2-4d79-8f9f-7afaea91f86f · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c1d70402-b030-47bc-b71a-db2d3020a1d0 · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a4e8493c-1a30-45a9-b0bd-7bf9cab9f7c4 · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7f7acff6-6204-4b09-a876-a4dc6805cb13 · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d6d1d30-2b9b-442a-8479-e85dc2cdf8fd · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search FIMO: A Challenge Formal Dataset for Automated Theorem Proving
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf00df94-b20b-4b68-8a09-efa45ca0b981 · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search LLaMA-Berry: Pairwise Optimization for O1-like Olympiad-Level Mathematical Reasoning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17b21789-6b6d-49df-8266-65ab88b727f1 · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search The final answer is: 7 Figure 6: Math Domain Example
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4951b2ab-7486-4d87-abd6-a331eddc6b91 · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search - For p = 2: The exponent in 17! is 15
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 90b6fbf8-c42e-458b-be63-406341098dd8 · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e10216c5-ea86-4ad2-832d-5319d6692a62 · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f2b2bf08-6002-4209-9eba-b358a6efedc4 · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e92b71bf-6983-4a18-854a-7f2ad207b046 · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation abb21582-f2a0-4546-bc1c-a6fcfe9e8b39 · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 777bdcce-ada4-49fc-97fb-1b9de9789362 · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a8672318-e70f-4ff6-8aea-480f3952d479 · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 87438d9e-6158-4009-9c43-d64daa440b31 · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ce5e6b8e-8bcf-492d-850e-299148b484f4 · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search - The liger is a physiotherapist
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 02121093-8594-4fab-969a-83c14b37d159 · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search - The dimensions of the box are 52.3 x 43.6 x 36.1 inches
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9523bf24-98d7-47d6-bb58-8f9ecaa43a15 · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search - The seal hides the cards that she has from the bee but does not build a power plant near the green fields of the husky
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation bf68543f-2169-4be7-8dd4-c5c382ab41d0 · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Therefore, the final answer is: True
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 44c8e436-d893-49ea-870e-02b0eb7a74ef · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3571d180-50a5-4a2b-9753-e55d5e86d7e4 · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ea7231ba-b934-4bd9-aba9-6f1de521ebc8 · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Based on the facts above, answer the following question
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 99591b1a-de19-4422-93e4-dbfbe821be49 · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d4f71e77-8491-4029-ad36-21663d827ae3 · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9c233a02-91d4-4681-a439-8cff21c8f94a · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Christopher Reeve’s spinal cord injury was severe, and he required specialized medical equipment and ongoing treatment
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation fd1a1377-5d53-4f12-9b72-8680a4a240c4 · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search The molar mass of Mg is 24.31 g/mol
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0dcf4edd-ccde-4e4f-907c-e2be8e9c3cae · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 69830b52-fb84-451f-88ba-3d86a1614d41 · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Since the reaction occurs in a beaker and the volume change is significant, we need to consider the external pressure and the change in volume
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2c19af94-8fab-42ae-8de9-effd2345dfb4 · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Here, text = ’ertubwi’ , sep = ’p’ , and maxsplit = 5
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ba60185a-c98c-40c9-a26a-1ba0eee35b54 · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ec314509-d6d9-4609-9331-5d69742a99f6 · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b6ca5d4f-7b25-4fdd-856f-60c266f740d1 · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 349d8217-d6a6-45f3-870e-cd1e5e910ed8 · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Let’s consider the correct approach:
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f816ac15-1d9a-4c7d-a835-3929024227bd · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Given the function’s behavior and the input, the correct approach is to split the string into two equal parts and reverse the first part
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 86edcf3a-a8eb-42be-8f1e-bd830876c99a · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8ecbffa9-686f-4070-b627-d7e3800d4444 · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 05b90b56-5fab-4fb8-aca9-12731434e318 · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Therefore, the final answer is: uertpbwi
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1ca2df08-b4d2-4b80-aee0-0bf7aad63402 · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c5f63040-891e-405e-b8bd-d9d959bff08f · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation abb1883f-114f-4317-9126-d9c4c5187dea · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3e88025a-aafd-41ca-a416-84d63f7d2690 · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ceeaa65e-5ae9-4037-8c03-29b00904140d · outbound
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 825bd4f2-9bbf-4372-8d44-9d73c49354da · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f374d70c-f487-4960-b004-b9010dd17816 · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 440fd80b-05ef-4887-94d2-c5c0950a5070 · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search $x$\", (high, 0), E); label(\
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a98f65f2-ce8b-4bd4-9cd3-4a3eb0c65d48 · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6acb6765-8bdd-49c4-8d36-86741e19e657 · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8cc3729f-f38e-4e13-af50-12adacdcab2c · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation fdf46ab8-92e1-40e4-87dd-0dd07027af26 · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3bb8c525-c608-422b-bad7-b8df7e631eb8 · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search The prompt templates for these situations are detailed in Appendix D.1.1
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 73504001-38ec-4431-a69c-b15472a3f8fd · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f9036570-0158-46fc-a653-4793b71f957f · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2be8cf18-cc91-498b-8c66-97132679f4f4 · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a31b254e-eaf5-4416-b473-50d75e761889 · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 21746f7e-f1ff-437a-9a0f-a5188a9b3e90 · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Therefore, the final answer is: \(\boxed{answer}\)
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 74f75386-9f21-4dbb-a897-5a9759c8299b · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Verify: [brief explanation of why you are correct with one sentence]
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b8140fec-c184-44ac-ba4f-ba4d59b635a4 · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8e954bae-307e-4996-a927-807d6d650c62 · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search ground truth solution
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5145fcb8-f4d6-446b-b393-24040ba1a788 · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Your task is to carefully review your own solution to a math problem, and adhere to the following guidelines:
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 001184a8-ab1d-465f-bf9a-98620a7691d4 · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search In Step <id>: [brief explanation of the mistake with one sentence]
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 175599de-f265-4d54-94bb-a9a071bf276b · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Alternatively: [your suggested step with one sentence]
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5ec0956c-14ff-4e15-9f1b-8914bc491e99 · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 12b31eee-316f-4c46-a17b-ac1890d98f6e · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search ground truth solution
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 87e3eeb2-83dd-4dff-89fa-3a4ca5684bce · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search You are collaborating with a partner to solve math problems
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d0f29297-4655-444c-9f67-15c26dcb8c94 · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Your partner’s partial solution
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 68dcfcd3-ac7a-4ee1-bcfc-b54d0dd5352a · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Your partner’s partial solution
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 49e2bf81-44b6-4ce0-853b-db0cf2e0b213 · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Alternatively: [your suggested step with one sentence]
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1eb586c8-faf3-4076-939e-fca111bc0d06 · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b8929a89-83dd-49ab-9288-3c0f08af1672 · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search ground truth solution
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 343bcb62-c7e8-47c8-acc6-b67a5f3a413d · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search DO NOT refer to any mistake in your partner’s partial solution
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f3f5fad2-c131-4f0f-be50-f619e458dcd5 · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search three two five
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2bd64f1d-4468-408b-a609-8a5c9002de74 · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Proximal Policy Optimization Algorithms
Reference 668
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48295069-fb9a-405b-bdbc-fd320cc7f24f · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Efficient reductions for imitation learning,
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25f66ac8-423b-486c-9a92-82531b4a572c · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search The CLRS-Text Algorithmic Reasoning Language Benchmark
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b043971-f640-4552-aee1-adc2149f58f7 · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Im- itation learning: A survey of learning methods,
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74fa24fa-07e2-48f0-99c3-4e3cb1325270 · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search OpenMathInstruct-2: Accelerating AI for Math with Massive Open-Source Instruction Data
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18646795-6beb-43db-b12b-d781a9d6c42c · outbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge
Reference 3634
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3dce8e4f-0f9d-4d11-a579-5c0aabe028d2 · inbound
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9a615c3-b6df-4fb9-b809-2dc43c1778f4 · inbound
From System 1 to System 2: A Survey of Reasoning Large Language Models Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search
Reference 247
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 44ac4dce-7654-4d1a-88b3-5134cc8e77f0 · inbound
Divide-Fuse-Conquer: Eliciting "Aha Moments" in Multi-Scenario Games Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4ecac6f-8fb3-4663-a6a9-f731e82367d8 · inbound
Satori-SWE: Evolutionary Test-Time Scaling for Sample-Efficient Software Engineering Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9fce8929-7c3b-4dd7-9c7f-21a23129442c · inbound
A Survey on Large Language Models for Mathematical Reasoning Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f604a5e2-42d6-4142-893b-39539c96e298 · inbound
AdapThink: Adaptive Thinking Preferences for Reasoning Language Model Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3877bd5-b1fc-4b05-8349-1711b53a8446 · inbound
Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2465e9b8-0ca7-4897-9838-026eacf3cd3d · inbound
Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c5bfa9d-980d-4cbb-980e-0309ac158e4e · inbound
Confidence-Aware Alignment Makes Reasoning LLMs More Reliable Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 95abce7f-b9cf-4393-a279-22f842fa3615 · inbound
STRIDE: Learnable Stepwise Language Feedback for LLM Reasoning Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation fb1c88c3-4325-4065-858e-d40580c05100 · inbound
REFLECTOR: Internalizing Step-wise Reflection against Indirect Jailbreak Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4dae21f0-5cb3-4497-ac38-039c04402437 · inbound
Efficient Agentic Reasoning Through Self-Regulated Simulative Planning Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d65c1290-9258-4ad1-accd-ea2af3017a66 · inbound
On the Generalization Gap in Self-Evolving Language Model Reasoning Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 545b8966-bd2a-4b23-b2e0-185011057ac0 · inbound
The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search
Reference 208
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.