Pith. sign in

Paper Citation Record · LEDGER

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search

As of 10 August 2026, this Paper Citation Record lists 84 of 84 outbound references and 14 inbound Pith citation observations for arXiv:2502.02508.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.02508 v3

Coverage vector

measured 84 of 84 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T11:57:49.117837Z

measured 98 of 98 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T14:40:36.038438Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T05:57:41.429120Z

Reference resolution

84 of 84 outbound references displayed

  • verified exact0
  • verified fuzzy31
  • unresolved53
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1229900a-f24d-4041-8375-33a12896dd33 · outbound

This paper cites ( s + 2) ( 2.4− t 60 ) = 9 Let’s solve these equations step by step.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search ( s + 2) ( 2.4− t 60 ) = 9 Let’s solve these equations step by step

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:57:50.209342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:57:48.824996Z digest=sha256:b1ba3b80567f295c8404e5904c7cec6e10951e67349a4b930ea800522c7e3c0a

Observation ad843070-ac30-4804-a4dd-5e40acc01e59 · outbound

This paper cites Self- consistency improves chain of thought reasoning in lan- guage models,.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Self- consistency improves chain of thought reasoning in lan- guage models,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-09T11:57:48.754123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:57:48.754123Z digest=sha256:77803d48261da9e545401e1587d3e8e7ec2379014df51f7ab9dcc2cb22dc7bed

Observation 22b49c8f-5b87-473f-b689-220da3b35a17 · outbound

This paper cites O1 Replication Journey -- Part 2: Surpassing O1-preview through Simple Distillation, Big Progress or Bitter Lesson?.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search O1 Replication Journey -- Part 2: Surpassing O1-preview through Simple Distillation, Big Progress or Bitter Lesson?

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-09T11:57:48.758762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:57:48.758762Z digest=sha256:6f3b8a01d384882bf7b1c159b6763a16bc737c2a47b938b7e0f4e8bcc6837e41

Observation e4c20d2e-04df-4ed4-ad19-030622c80607 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-09T11:57:48.763506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:57:48.763506Z digest=sha256:c8aa0847edca99ac2c04413008da5d5a31deac09af8847bd1bb9b844b7e8d047

Observation f07620ed-7ca1-4ac4-9df6-786b49367384 · outbound

This paper cites Mutual Reasoning Makes Smaller LLMs Stronger Problem-Solvers.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Mutual Reasoning Makes Smaller LLMs Stronger Problem-Solvers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-09T11:57:48.773155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:57:48.773155Z digest=sha256:d31d3f0ecc24575548f696629e859c759a7509cc6515712c932b9c2204457b54

Observation d2ac63d0-277d-4136-855b-3533293a6cef · outbound

This paper cites Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T11:57:48.778636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:57:48.778636Z digest=sha256:c1cc46c42f8e7f3bbc34af1b5efaa551c706f7e7de849c328ec486525625e4d7

Observation c692f436-95e2-4d79-8f9f-7afaea91f86f · outbound

This paper cites an unresolved cited work.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:50.066840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:57:48.873047Z digest=sha256:9b5529b550828e9fd395f8e1946e5ac0bee367e0f67a1db8fccbf9b347e2ca6c

Observation c1d70402-b030-47bc-b71a-db2d3020a1d0 · outbound

This paper cites an unresolved cited work.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:50.053463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:57:48.877128Z digest=sha256:c24096d61c214ef2df12d05721830322f3b2773a8f227eedc882b88b780c382e

Observation a4e8493c-1a30-45a9-b0bd-7bf9cab9f7c4 · outbound

This paper cites an unresolved cited work.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:50.040040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:57:48.881086Z digest=sha256:565e615600a43f4fe643bb9734da5f79ca566445573351ca907bac848d205cfc

Observation 7f7acff6-6204-4b09-a876-a4dc6805cb13 · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T11:57:48.797404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:57:48.797404Z digest=sha256:b50d0d6dd7c52f20ae709c9c3ed6f7d345b6ec330e3d422fb89f58b7b653f15f

Observation 9d6d1d30-2b9b-442a-8479-e85dc2cdf8fd · outbound

This paper cites FIMO: A Challenge Formal Dataset for Automated Theorem Proving.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search FIMO: A Challenge Formal Dataset for Automated Theorem Proving

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T11:57:48.802192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:57:48.802192Z digest=sha256:c961cb184fc431221f0bbef089ba25ded0134aa0fae270ba5b24739211f6f780

Observation cf00df94-b20b-4b68-8a09-efa45ca0b981 · outbound

This paper cites LLaMA-Berry: Pairwise Optimization for O1-like Olympiad-Level Mathematical Reasoning.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search LLaMA-Berry: Pairwise Optimization for O1-like Olympiad-Level Mathematical Reasoning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T11:57:48.815737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:57:48.815737Z digest=sha256:c6a4b39cfa8a921de44b5f157569e7676f25246e31064b4bdb9932f0f722edda

Observation 17b21789-6b6d-49df-8266-65ab88b727f1 · outbound

This paper cites The final answer is: 7 Figure 6: Math Domain Example.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search The final answer is: 7 Figure 6: Math Domain Example

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T11:57:48.820421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:57:48.820421Z digest=sha256:d34b7a0b976609d90063f0836e8dfeee3651bb503233fae885f4467d2c0e6bd0

Observation 4951b2ab-7486-4d87-abd6-a331eddc6b91 · outbound

This paper cites - For p = 2: The exponent in 17! is 15.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search - For p = 2: The exponent in 17! is 15

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:57:50.170634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:57:48.838795Z digest=sha256:6b06f05e208a8271fe1eb8353a1dc4f9f40d1caebbc04231cb354eb1dd2a40c7

Observation 90b6fbf8-c42e-458b-be63-406341098dd8 · outbound

This paper cites an unresolved cited work.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:50.196434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:57:48.829328Z digest=sha256:d6db1ec609ba619722c5a3787c48c5597ea04b3310926984d6eed61b070dd4bd

Observation e10216c5-ea86-4ad2-832d-5319d6692a62 · outbound

This paper cites an unresolved cited work.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:50.157546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:57:48.843189Z digest=sha256:248e48170c12c91b146b4b4402ec9990c2b74c8554b58502d284d1b50684d1ef

Observation f2b2bf08-6002-4209-9eba-b358a6efedc4 · outbound

This paper cites an unresolved cited work.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:50.144768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:57:48.847508Z digest=sha256:f5540c4bcbfc63a9bd85aa119b952e739eb1378c4bfff655b4de962c01ff6c42

Observation e92b71bf-6983-4a18-854a-7f2ad207b046 · outbound

This paper cites an unresolved cited work.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:50.131974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:57:48.851794Z digest=sha256:3a197c110eaf27523f5ec0c2164c2036ccd2de8568ae90c872725cdf6851d67c

Observation abb21582-f2a0-4546-bc1c-a6fcfe9e8b39 · outbound

This paper cites an unresolved cited work.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:50.119153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:57:48.856150Z digest=sha256:d296e337ff65b3f94c151f6c3d109e539f201ef88b65be492f2c4d7884c8723a

Observation 777bdcce-ada4-49fc-97fb-1b9de9789362 · outbound

This paper cites an unresolved cited work.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:50.106481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:57:48.860446Z digest=sha256:cd8dfca071c981213a282c5cc3c91e6efd8342d93b36460460c9edd42ac4ef1d

Observation a8672318-e70f-4ff6-8aea-480f3952d479 · outbound

This paper cites an unresolved cited work.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:50.093349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:57:48.864677Z digest=sha256:af85d593d389a37d7c9d4e9fd46b095adecf686cd7e584bf0d3fa91db8639525

Observation 87438d9e-6158-4009-9c43-d64daa440b31 · outbound

This paper cites an unresolved cited work.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:50.080011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:57:48.868939Z digest=sha256:ea8a5f44c81844d6da0df2f278aa0f63706ef2115850c39c89e89cd922705415

Observation ce5e6b8e-8bcf-492d-850e-299148b484f4 · outbound

This paper cites - The liger is a physiotherapist.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search - The liger is a physiotherapist

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:57:50.027233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:57:48.885162Z digest=sha256:cbf6639a2ba417a06e03f228fa5401dc5444da053b522cc62984f494ff19c371

Observation 02121093-8594-4fab-969a-83c14b37d159 · outbound

This paper cites - The dimensions of the box are 52.3 x 43.6 x 36.1 inches.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search - The dimensions of the box are 52.3 x 43.6 x 36.1 inches

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:57:50.013742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:57:48.889206Z digest=sha256:64e406a20de54dbf5a233547ae6dbe73ea57b750cc208cc062da88c273ddc2c2

Observation 9523bf24-98d7-47d6-bb58-8f9ecaa43a15 · outbound

This paper cites - The seal hides the cards that she has from the bee but does not build a power plant near the green fields of the husky.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search - The seal hides the cards that she has from the bee but does not build a power plant near the green fields of the husky

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:57:50.000605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:57:48.893289Z digest=sha256:7da2194b3185b03c57e7011983a09e7870114250bc5898c8e1dd030a3435116a

Observation bf68543f-2169-4be7-8dd4-c5c382ab41d0 · outbound

This paper cites Therefore, the final answer is: True.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Therefore, the final answer is: True

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:57:49.987513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:57:48.897435Z digest=sha256:bb6aeb1ec92c841ddb63969b460311c36b2fe792e23a2edfc28554443655ee83

Observation 44c8e436-d893-49ea-870e-02b0eb7a74ef · outbound

This paper cites an unresolved cited work.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:49.974480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:57:48.901882Z digest=sha256:e40c36f4c6a9d2d729ffc25d6898832f8a0a547f050384747493fa3b15e8d16b

Observation 3571d180-50a5-4a2b-9753-e55d5e86d7e4 · outbound

This paper cites an unresolved cited work.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:49.961712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:57:48.906295Z digest=sha256:711dbce46edee4bf404423bc94294c00c659aa802f7e0013548f903fab0c499a

Observation ea7231ba-b934-4bd9-aba9-6f1de521ebc8 · outbound

This paper cites Based on the facts above, answer the following question.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Based on the facts above, answer the following question

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:57:49.948650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:57:48.910561Z digest=sha256:f439ace02690263c7569e58cd2f2499b37997bb72449b1a82b6cd238a864483f

Observation 99591b1a-de19-4422-93e4-dbfbe821be49 · outbound

This paper cites an unresolved cited work.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:49.935352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:57:48.914960Z digest=sha256:4e5318a92da3ba751e062449d4e191846c9cc3fa5bde75f8dcf5251cf26f1d87

Observation d4f71e77-8491-4029-ad36-21663d827ae3 · outbound

This paper cites an unresolved cited work.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:49.922571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:57:48.919175Z digest=sha256:9a8f83ab2ce71294ce245aa37857ad5809c3f7bae6bc7eb98c908950c4bcaaa8

Observation 9c233a02-91d4-4681-a439-8cff21c8f94a · outbound

This paper cites Christopher Reeve’s spinal cord injury was severe, and he required specialized medical equipment and ongoing treatment.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Christopher Reeve’s spinal cord injury was severe, and he required specialized medical equipment and ongoing treatment

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:57:49.909191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:57:48.923486Z digest=sha256:9b95ea5712e7f95946cc77eadd51cce07fad12afd024484a95353343f6eb11e4

Observation fd1a1377-5d53-4f12-9b72-8680a4a240c4 · outbound

This paper cites The molar mass of Mg is 24.31 g/mol.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search The molar mass of Mg is 24.31 g/mol

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:57:49.895865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:57:48.927807Z digest=sha256:09021f835973684f3b868ec49f311abe9fbfa5b8a1a152d20f8d33b98eb9f044

Observation 0dcf4edd-ccde-4e4f-907c-e2be8e9c3cae · outbound

This paper cites an unresolved cited work.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:49.882898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:57:48.931959Z digest=sha256:5be5b85b0b8c1338038261c9ba29a3deb0441d0f2dd0e4d7b53249c6246f2bdd

Observation 69830b52-fb84-451f-88ba-3d86a1614d41 · outbound

This paper cites Since the reaction occurs in a beaker and the volume change is significant, we need to consider the external pressure and the change in volume.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Since the reaction occurs in a beaker and the volume change is significant, we need to consider the external pressure and the change in volume

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:57:49.869824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:57:48.936376Z digest=sha256:483ea109f8f6a5cf140e87ef9d63295af9c9056e1327f0fab0bdddb1a67715e5

Observation 2c19af94-8fab-42ae-8de9-effd2345dfb4 · outbound

This paper cites Here, text = ’ertubwi’ , sep = ’p’ , and maxsplit = 5.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Here, text = ’ertubwi’ , sep = ’p’ , and maxsplit = 5

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:57:49.856599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:57:48.940664Z digest=sha256:5f5eaaae2a20bb8c4827b0472dfbd7efb5d1963e2ec960673a3b16acc4789d41

Observation ba60185a-c98c-40c9-a26a-1ba0eee35b54 · outbound

This paper cites an unresolved cited work.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:49.843271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:57:48.945072Z digest=sha256:25d27a717f3c7936aba8991bcee5237c5485e71d4cd46d19f79dfb0c3a0c4b73

Observation ec314509-d6d9-4609-9331-5d69742a99f6 · outbound

This paper cites an unresolved cited work.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:49.830571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:57:48.949270Z digest=sha256:b27e211565690098ec81060373f77b6d3761077ad9daac46439a6ff0143d1e11

Observation b6ca5d4f-7b25-4fdd-856f-60c266f740d1 · outbound

This paper cites an unresolved cited work.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:49.817832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:57:48.953461Z digest=sha256:9300935e4326badd9fbd21a25df9e3abb4006febb0d8b299f179d70535a0579e

Observation 349d8217-d6a6-45f3-870e-cd1e5e910ed8 · outbound

This paper cites Let’s consider the correct approach:.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Let’s consider the correct approach:

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:57:49.805212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:57:48.957469Z digest=sha256:75688ff5b619d9a0e3436ecda7ca632f23acd655c4b9abfceedcfe265f6222af

Observation f816ac15-1d9a-4c7d-a835-3929024227bd · outbound

This paper cites Given the function’s behavior and the input, the correct approach is to split the string into two equal parts and reverse the first part.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Given the function’s behavior and the input, the correct approach is to split the string into two equal parts and reverse the first part

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:57:49.792615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:57:48.961799Z digest=sha256:d486c563fa49ab491c86437ea6fff1ee860e54d55a984ac29339b7c54d8b8744

Observation 86edcf3a-a8eb-42be-8f1e-bd830876c99a · outbound

This paper cites an unresolved cited work.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:49.779702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:57:48.966290Z digest=sha256:6c2a81d939778fcc16e6f20212c41e86fc05e79a058939d87c8c7c06c9e87ec5

Observation 8ecbffa9-686f-4070-b627-d7e3800d4444 · outbound

This paper cites an unresolved cited work.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:49.766859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:57:48.970794Z digest=sha256:dcdde57fc247e122857d8b0da7a245cea7261f0973dd6b4ede3bcdddd07b14b2

Observation 05b90b56-5fab-4fb8-aca9-12731434e318 · outbound

This paper cites Therefore, the final answer is: uertpbwi.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Therefore, the final answer is: uertpbwi

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:57:49.754569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:57:48.974976Z digest=sha256:6106235d89babe2a87ed515b1bb2e26bf9484f8b931733f51777f875cbfc5ab2

Observation 1ca2df08-b4d2-4b80-aee0-0bf7aad63402 · outbound

This paper cites an unresolved cited work.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:49.741549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:57:48.979367Z digest=sha256:c56cd3530723ac15333cdf453f0a29e662fc721f45d0579f03428d98cfdd0f17

Observation c5f63040-891e-405e-b8bd-d9d959bff08f · outbound

This paper cites an unresolved cited work.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:49.729279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:57:48.983809Z digest=sha256:2b5db5d2ba4ea8fcb7e89997987bf591778aa0f0509d89f5554294b2e64922e2

Observation abb1883f-114f-4317-9126-d9c4c5187dea · outbound

This paper cites an unresolved cited work.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:49.716694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:57:48.988466Z digest=sha256:45f38242d4507dad7c42a87d19ba31572aff9db076a041430e2297f54c4d91f5

Observation 3e88025a-aafd-41ca-a416-84d63f7d2690 · outbound

This paper cites an unresolved cited work.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:49.703827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:57:48.992777Z digest=sha256:0f8d651402bf76cb315cf123ef145f589ea4a09b60c391ff802a50b2fb94e5f2

Observation ceeaa65e-5ae9-4037-8c03-29b00904140d · outbound

This paper cites superhuman.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search superhuman

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:57:49.690490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:57:48.996928Z digest=sha256:7b74181fbc682f95fb81bab0a09505ac321ffa6945b37dcde20dde81ed2e539d

Observation 825bd4f2-9bbf-4372-8d44-9d73c49354da · outbound

This paper cites an unresolved cited work.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:49.677070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:57:49.001527Z digest=sha256:fbfd43978e72abbb3012b88dcdd1b2f30ffcfc097e88952f22744fa7057961be

Observation f374d70c-f487-4960-b004-b9010dd17816 · outbound

This paper cites an unresolved cited work.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:49.663891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:57:49.006055Z digest=sha256:7fb75ffce2779ebc53da10f0cf10248045250b06dd943e8c7c1c3239923e67e1

Observation 440fd80b-05ef-4887-94d2-c5c0950a5070 · outbound

This paper cites $x$\", (high, 0), E); label(\.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search $x$\", (high, 0), E); label(\

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:57:50.183460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:57:48.833684Z digest=sha256:e7cb7a75baf697d45202f3a349035fad111083d41cc6dfd3e0b229d670b3a01b

Observation a98f65f2-ce8b-4bd4-9cd3-4a3eb0c65d48 · outbound

This paper cites an unresolved cited work.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:49.650716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:57:49.010468Z digest=sha256:7d2e581c87294d35933d572d9399080878621949649344259878a8a6e1ee0bc1

Observation 6acb6765-8bdd-49c4-8d36-86741e19e657 · outbound

This paper cites an unresolved cited work.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:49.637406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:57:49.014684Z digest=sha256:0efb0ce3e06ab0a9a71fb7fc6783ec4886f049175c580d977bce48ad4016c594

Observation 8cc3729f-f38e-4e13-af50-12adacdcab2c · outbound

This paper cites an unresolved cited work.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:49.624675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:57:49.019125Z digest=sha256:c11faf40bee159cae7fdecdf8e9ee08a86ad706cc0db913084170d607e919e2e

Observation fdf46ab8-92e1-40e4-87dd-0dd07027af26 · outbound

This paper cites an unresolved cited work.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:49.612216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:57:49.023421Z digest=sha256:d1c5b6816e156f2ab34989c4d2e33681d8729e3e01e4d0b630e706ee8c11cb37

Observation 3bb8c525-c608-422b-bad7-b8df7e631eb8 · outbound

This paper cites The prompt templates for these situations are detailed in Appendix D.1.1.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search The prompt templates for these situations are detailed in Appendix D.1.1

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:57:49.599320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:57:49.027700Z digest=sha256:ea7fd57b4b1337ca516ecdb9a21f1cae14d3d7561b202dedfd7492b1e98a9852

Observation 73504001-38ec-4431-a69c-b15472a3f8fd · outbound

This paper cites an unresolved cited work.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:49.586401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:57:49.031908Z digest=sha256:df5905dbee64efc07f1ed31df05e774c960afa400234858d82025e54b43409f0

Observation f9036570-0158-46fc-a653-4793b71f957f · outbound

This paper cites an unresolved cited work.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:49.573534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:57:49.036034Z digest=sha256:19834b27ce45dce8c408624d4f696249ac7ed132168feb0a85ba91018395ad79

Observation 2be8cf18-cc91-498b-8c66-97132679f4f4 · outbound

This paper cites an unresolved cited work.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:49.560445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:57:49.040666Z digest=sha256:8351e200a7fa90c8ffd3d0c120b795293cb3ae79e59dd66636acb85cce88f531

Observation a31b254e-eaf5-4416-b473-50d75e761889 · outbound

This paper cites an unresolved cited work.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work

Reference 69

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:49.547163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:57:49.044888Z digest=sha256:78d16b46c284632103b3d8551f9842b050b4e5977f672687637a3be590b1bcef

Observation 21746f7e-f1ff-437a-9a0f-a5188a9b3e90 · outbound

This paper cites Therefore, the final answer is: \(\boxed{answer}\).

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Therefore, the final answer is: \(\boxed{answer}\)

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:57:49.533673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:57:49.048976Z digest=sha256:62b6f69f07293eea91f40dc17e03299fb7ce4216fa1a5f0d5a65e9151aef929a

Observation 74f75386-9f21-4dbb-a897-5a9759c8299b · outbound

This paper cites Verify: [brief explanation of why you are correct with one sentence].

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Verify: [brief explanation of why you are correct with one sentence]

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:57:49.520413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:57:49.053555Z digest=sha256:3f45efa7e0002423f79dd3e5b763df688687eb57beca73776e5592d6ea771d0c

Observation b8140fec-c184-44ac-ba4f-ba4d59b635a4 · outbound

This paper cites an unresolved cited work.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work

Reference 72

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:49.506941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:57:49.057793Z digest=sha256:aa44f7cf784d74cca4712a2476779154f454070184c1fd00605f8fe835eb95c5

Observation 8e954bae-307e-4996-a927-807d6d650c62 · outbound

This paper cites ground truth solution.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search ground truth solution

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:57:49.492947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:57:49.061965Z digest=sha256:f7134ae78fce9b6cb65ef9cbd1a3090d3ef2e50066046ef5981e2ae4f7bc93fa

Observation 5145fcb8-f4d6-446b-b393-24040ba1a788 · outbound

This paper cites Your task is to carefully review your own solution to a math problem, and adhere to the following guidelines:.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Your task is to carefully review your own solution to a math problem, and adhere to the following guidelines:

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:57:49.479751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:57:49.066144Z digest=sha256:22748cbf60f3f1dd7683b6545757be16e7a49aa10b0dbc60eace0dd53c38e3a8

Observation 001184a8-ab1d-465f-bf9a-98620a7691d4 · outbound

This paper cites In Step <id>: [brief explanation of the mistake with one sentence].

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search In Step <id>: [brief explanation of the mistake with one sentence]

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:57:49.466269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:57:49.070386Z digest=sha256:286ea1369d6e04c9db468eb4bf04856a5af6dbaa4c703e540140bc0675437127

Observation 175599de-f265-4d54-94bb-a9a071bf276b · outbound

This paper cites Alternatively: [your suggested step with one sentence].

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Alternatively: [your suggested step with one sentence]

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:57:49.452554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:57:49.074909Z digest=sha256:58da6bed551ec5f25c987c18b47b97f05d697ed29e00896e13894cc1745e357c

Observation 5ec0956c-14ff-4e15-9f1b-8914bc491e99 · outbound

This paper cites an unresolved cited work.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work

Reference 77

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:49.438562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:57:49.078980Z digest=sha256:56fab1795a44856e05920ffa844a1f81de89b2c36d0926fa109d8f5afe93529b

Observation 12b31eee-316f-4c46-a17b-ac1890d98f6e · outbound

This paper cites ground truth solution.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search ground truth solution

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:57:49.425042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:57:49.083128Z digest=sha256:997faf412e758034334d05d6491b24a74e8588a9d7c02dfb6177f268ffe9a502

Observation 87e3eeb2-83dd-4dff-89fa-3a4ca5684bce · outbound

This paper cites You are collaborating with a partner to solve math problems.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search You are collaborating with a partner to solve math problems

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:57:49.411650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:57:49.087244Z digest=sha256:0d82bb3ed5d4d7ea5abc17f3e7cfe5ac88d8d065caaf641fedfdc27bf3e41a5a

Observation d0f29297-4655-444c-9f67-15c26dcb8c94 · outbound

This paper cites Your partner’s partial solution.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Your partner’s partial solution

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:57:49.398107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:57:49.092476Z digest=sha256:d7dbc10a5302ba014867c41dc6f16cafc8499b4c77ecf14b390dcb006a571141

Observation 68dcfcd3-ac7a-4ee1-bcfc-b54d0dd5352a · outbound

This paper cites Your partner’s partial solution.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Your partner’s partial solution

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:57:49.383765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:57:49.096709Z digest=sha256:4ab11385e2d099ce9e700dba050f2440148631b2d95f6fef394da25c36fb1c76

Observation 49e2bf81-44b6-4ce0-853b-db0cf2e0b213 · outbound

This paper cites Alternatively: [your suggested step with one sentence].

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Alternatively: [your suggested step with one sentence]

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:57:49.368878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:57:49.100902Z digest=sha256:6fc054fd081be4eb685722271842c913cf14b259c93bb8364be725e7a93499eb

Observation 1eb586c8-faf3-4076-939e-fca111bc0d06 · outbound

This paper cites an unresolved cited work.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work

Reference 83

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:49.354962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:57:49.105098Z digest=sha256:7b73548970d8a30f7f8809957bd3415133bba99802e73a3fd1839a2cb20d0e85

Observation b8929a89-83dd-49ab-9288-3c0f08af1672 · outbound

This paper cites ground truth solution.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search ground truth solution

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:57:49.340855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:57:49.109751Z digest=sha256:daacd97fabdb1526d3a63e987eca31908cd86ffc07f876a55ab7d62af6aaad19

Observation 343bcb62-c7e8-47c8-acc6-b67a5f3a413d · outbound

This paper cites DO NOT refer to any mistake in your partner’s partial solution.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search DO NOT refer to any mistake in your partner’s partial solution

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:57:49.327010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:57:49.113775Z digest=sha256:0ce5c69923614c9f3c3aa4a56708f912ce5065cdf0e5b71f5d9fc7e73beea20c

Observation f3f5fad2-c131-4f0f-be50-f619e458dcd5 · outbound

This paper cites three two five.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search three two five

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:57:49.311495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:57:49.117837Z digest=sha256:087ea467f3fd202256a27783924b10557f3a743bf6226a40cbc4111f1e779f25

Observation 2bd64f1d-4468-408b-a609-8a5c9002de74 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Proximal Policy Optimization Algorithms

Reference 668

Resolution
unresolved
no resolver link, observed 2026-08-09T11:57:48.792502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:57:48.792502Z digest=sha256:b2fee1d48100ae9d84fcc871928394361d44a49d7402d12372ba120016f8d106

Observation 48295069-fb9a-405b-bdbc-fd320cc7f24f · outbound

This paper cites Efficient reductions for imitation learning,.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Efficient reductions for imitation learning,

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-09T11:57:48.788137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:57:48.788137Z digest=sha256:6017edd7e99fd61ed0db61fb445e4fb87a6f3cfe4246e1622a542193d164f90d

Observation 25f66ac8-423b-486c-9a92-82531b4a572c · outbound

This paper cites The CLRS-Text Algorithmic Reasoning Language Benchmark.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search The CLRS-Text Algorithmic Reasoning Language Benchmark

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-09T11:57:48.811186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:57:48.811186Z digest=sha256:811a8506c9d7c2059c9b1f54ebd907d80a6c16f89f8e8623dc0bce45ca8f7236

Observation 9b043971-f640-4552-aee1-adc2149f58f7 · outbound

This paper cites Im- itation learning: A survey of learning methods,.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Im- itation learning: A survey of learning methods,

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-09T11:57:48.783650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:57:48.783650Z digest=sha256:033d01db9de17bc7b5cef28e9c0da38a1f90af7ff44ca3bca5c034f4ff6486a5

Observation 74fa24fa-07e2-48f0-99c3-4e3cb1325270 · outbound

This paper cites OpenMathInstruct-2: Accelerating AI for Math with Massive Open-Source Instruction Data.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search OpenMathInstruct-2: Accelerating AI for Math with Massive Open-Source Instruction Data

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-09T11:57:48.748398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:57:48.748398Z digest=sha256:d58244f50e38581222f51b37b90bfd6d4788d1c2206b6ff440afdb789c86fda5

Observation 18646795-6beb-43db-b12b-d781a9d6c42c · outbound

This paper cites CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge

Reference 3634

Resolution
unresolved
no resolver link, observed 2026-08-09T11:57:48.806698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:57:48.806698Z digest=sha256:b3f1d7cd6f742b84137a5dba8c1b165550912a6036655dcd088d064afc36292e

Pith citing papers

Observation 3dce8e4f-0f9d-4d11-a579-5c0aabe028d2 · inbound

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling cites this paper.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-08T14:40:36.038438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:40:36.038438Z digest=sha256:dcc11d719f51cd356b10a6a9e9bc97069ee4c386e9e25475c1aa64077cfea671

Observation f9a615c3-b6df-4fb9-b809-2dc43c1778f4 · inbound

From System 1 to System 2: A Survey of Reasoning Large Language Models cites this paper.

From System 1 to System 2: A Survey of Reasoning Large Language Models Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search

Reference 247

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:36:24.423192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T01:36:23.845366Z digest=sha256:bcc64283c2841d50329554ebc25c217e052ebb208f8f2c93c175bc88b7dda92d

Observation 44ac4dce-7654-4d1a-88b3-5134cc8e77f0 · inbound

Divide-Fuse-Conquer: Eliciting "Aha Moments" in Multi-Scenario Games cites this paper.

Divide-Fuse-Conquer: Eliciting "Aha Moments" in Multi-Scenario Games Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:04.324930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:04.324930Z digest=sha256:aaa9a0d2e9d8be62047604d33517b296e7476021c77338652c45d7950907f4b6

Observation e4ecac6f-8fb3-4663-a6a9-f731e82367d8 · inbound

Satori-SWE: Evolutionary Test-Time Scaling for Sample-Efficient Software Engineering cites this paper.

Satori-SWE: Evolutionary Test-Time Scaling for Sample-Efficient Software Engineering Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:47:22.078456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:47:22.078456Z digest=sha256:1789d26b7a455faeb660ed8c34afff5ecaeb38247543a0769c3e9bf4d6a4c943

Observation 9fce8929-7c3b-4dd7-9c7f-21a23129442c · inbound

A Survey on Large Language Models for Mathematical Reasoning cites this paper.

A Survey on Large Language Models for Mathematical Reasoning Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T05:14:47.425928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:14:47.425928Z digest=sha256:868c111b4f9b6c871e5dcd4bc1fd6c3db95af69418cdeadcb68230835d4fd8f4

Observation f604a5e2-42d6-4142-893b-39539c96e298 · inbound

AdapThink: Adaptive Thinking Preferences for Reasoning Language Model cites this paper.

AdapThink: Adaptive Thinking Preferences for Reasoning Language Model Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T23:28:39.674751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:28:39.674751Z digest=sha256:a998172ab610bd7643b5eaea6f76241d004b13e2e348c99e1750ad0feebffcbb

Observation d3877bd5-b1fc-4b05-8349-1711b53a8446 · inbound

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL cites this paper.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:06.395257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:06.395257Z digest=sha256:44fef1b5018a85757f50177d5089ebc332014549cf08cab5862d35041304bf62

Observation 2465e9b8-0ca7-4897-9838-026eacf3cd3d · inbound

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization cites this paper.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:28.211988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:28.211988Z digest=sha256:cacb444fcfb6ce01c5375c8c4914f4edf4353bd0424a011001d4d63a9701ca87

Observation 6c5bfa9d-980d-4cbb-980e-0309ac158e4e · inbound

Confidence-Aware Alignment Makes Reasoning LLMs More Reliable cites this paper.

Confidence-Aware Alignment Makes Reasoning LLMs More Reliable Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:55:53.932850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T02:10:40.020460Z digest=sha256:d50c65a96a6b96d1e59016cc2be221e808f467c10439515331b1b13bcd403e28

Observation 95abce7f-b9cf-4393-a279-22f842fa3615 · inbound

STRIDE: Learnable Stepwise Language Feedback for LLM Reasoning cites this paper.

STRIDE: Learnable Stepwise Language Feedback for LLM Reasoning Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-20T20:49:00.696829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T20:47:16.236629Z digest=sha256:60137d35e018e48985452d3848d738b3ae1a8c009cf157bb5036b8f4b331657f

Observation fb1c88c3-4325-4065-858e-d40580c05100 · inbound

REFLECTOR: Internalizing Step-wise Reflection against Indirect Jailbreak cites this paper.

REFLECTOR: Internalizing Step-wise Reflection against Indirect Jailbreak Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T06:19:41.990511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-21T06:16:01.040236Z digest=sha256:340736b108b452341a88767c0b22750bfc35bf90e8a776acd663230065763751

Observation 4dae21f0-5cb3-4497-ac38-039c04402437 · inbound

Efficient Agentic Reasoning Through Self-Regulated Simulative Planning cites this paper.

Efficient Agentic Reasoning Through Self-Regulated Simulative Planning Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-05-22T06:34:41.015581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T06:33:36.846345Z digest=sha256:7de1885658ed48e4555f3822ccc3d8cac3c4b1f181328aaea1ae0095571082be

Observation d65c1290-9258-4ad1-accd-ea2af3017a66 · inbound

On the Generalization Gap in Self-Evolving Language Model Reasoning cites this paper.

On the Generalization Gap in Self-Evolving Language Model Reasoning Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-06-28T17:22:24.890131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T17:18:37.671369Z digest=sha256:65e90384801367db66846439ba4f86b4b351b7f3c96635843581fa63fad66203

Observation 545b8966-bd2a-4b23-b2e0-185011057ac0 · inbound

The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes cites this paper.

The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search

Reference 208

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:57:41.430647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-27T12:59:51.091008Z digest=sha256:ebbb932136219a87f95c4f0eaecd5808e43f138a90a85b6dc9df31739bdfcbed