Pith. sign in

Paper Citation Record · LEDGER

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning

As of 21 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 2 inbound Pith citation observations for arXiv:2511.02130.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2511.02130 v2

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T00:18:57.117932Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T23:14:56.115834Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T07:55:58.659031Z

Reference resolution

61 of 61 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved61
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f5f1fd78-4430-4e0e-ac5e-84d4fb6dc945 · outbound

This paper cites Ai agents as universal task solvers.arXiv preprint arXiv:2510.12066, 2025.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Ai agents as universal task solvers.arXiv preprint arXiv:2510.12066, 2025

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:49.533694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:49.533694Z digest=sha256:c604b90dd398aee7bede93cdbb6e15c96fbb9c8a37083432500bab99d468304c

Observation 1db28879-e705-457d-83c9-00e7e1ecf568 · outbound

This paper cites Weitzman.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Weitzman

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:49.599504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:49.599504Z digest=sha256:82c35cb0275d5e8a06b0f7959364883e7a0afa97d78a5ee961b90a6feca185a9

Observation bd9cd0bd-75a9-48d6-97a4-71af8ceeefd8 · outbound

This paper cites Adaptive inference-time compute: Llms can predict if they can do better, even mid-generation, 2024.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Adaptive inference-time compute: Llms can predict if they can do better, even mid-generation, 2024

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:49.699031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:49.699031Z digest=sha256:fa502415c30a30ce167c3a9e0769cb1765b4cf75aac7f4c1011c21418a616da4

Observation ccf52c75-4cac-443f-a8b2-64846505a4cd · outbound

This paper cites Learning how hard to think: Input-adaptive allocation of lm computation, 2024.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Learning how hard to think: Input-adaptive allocation of lm computation, 2024

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:49.775735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:49.775735Z digest=sha256:cee938613f2c1a00bf6ffb33dbcb833329642f619a57a7635c681238502f7061

Observation 53c0d873-76fc-450a-8992-79d35b225291 · outbound

This paper cites Reasoning models know when they’re right: Probing hidden states for self-verification, 2025.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Reasoning models know when they’re right: Probing hidden states for self-verification, 2025

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:49.852359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:49.852359Z digest=sha256:de7d1b29bb5e6e56318baf167f48ce03fb9e66e80230a612e762bae2c74a086f

Observation 53cf6160-9e1d-4a0c-9523-245d0286e0ba · outbound

This paper cites Reasoning models better express their confidence.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Reasoning models better express their confidence

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:49.927869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:49.927869Z digest=sha256:2db39602b1508d335a93f7d0f7b1d7c70d4eb010e5356f6fce67174211a15798

Observation 359544d3-9c10-4bf9-a124-4a4cd27c01d5 · outbound

This paper cites Are the hidden states hiding something? testing the limits of factuality-encoding capabilities in llms, 2025.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Are the hidden states hiding something? testing the limits of factuality-encoding capabilities in llms, 2025

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:50.016662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:50.016662Z digest=sha256:3648f60b3925d7d2bca66b5f37ecc599bb44f1f6f184fbfef59b3ff087d98b22

Observation 34f2f98f-f6b3-428b-b689-b550b46683a8 · outbound

This paper cites When do llms admit their mistakes? understanding the role of model belief in retraction, 2025.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning When do llms admit their mistakes? understanding the role of model belief in retraction, 2025

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:50.155089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:50.155089Z digest=sha256:4c5c5b26e884b6a8312a9a1458545df457f71272accc80fd01c9f7bf47dae95c

Observation 5a724d05-908e-4bee-8398-45c60272b5ef · outbound

This paper cites Latts: Locally adaptive test-time scaling.arXiv preprint arXiv:2509.20368, 2025.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Latts: Locally adaptive test-time scaling.arXiv preprint arXiv:2509.20368, 2025

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:50.231644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:50.231644Z digest=sha256:a24c14898ff25e14db800b255e9ecb9de5d48e85f0163c775883dddb48520b07

Observation 72151895-ca52-432b-9685-02026352fb2f · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:50.369476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:50.369476Z digest=sha256:299762f9c9cc434e4cf959d5c5eb12fef504eff0fac559a5df5756dcd77eb43c

Observation 1e54afe1-de7c-44d9-906e-b06c25480fb8 · outbound

This paper cites Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:50.448154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:50.448154Z digest=sha256:239df92e2b7e9d4d9c6f7db3b3771151d0b24f4481a52c03cfeedd4a4883d88e

Observation aa48d2f6-f530-4deb-8886-e05eb718ca7d · outbound

This paper cites Scaling LLM test-time com- pute optimally can be more effective than scaling parameters for reasoning.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Scaling LLM test-time com- pute optimally can be more effective than scaling parameters for reasoning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:50.548821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:50.548821Z digest=sha256:d5e58e5215c6a356b4f41e423692525b723598dd3bf28309dd59f65b834b30ad

Observation f8c97bc6-fce8-421c-b125-72039548985e · outbound

This paper cites Adaptive test-time reasoning via reward-guided dual-phase search, 2025.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Adaptive test-time reasoning via reward-guided dual-phase search, 2025

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:50.669264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:50.669264Z digest=sha256:7749fd9baf2ac37ead6e4ea9662d9929911484c0f54e62605adf053c025a1ef2

Observation ae526caf-3e4d-4de6-b8b5-eb699ef6a94b · outbound

This paper cites Large language model guided tree-of-thought, 2023.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Large language model guided tree-of-thought, 2023

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:50.797975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:50.797975Z digest=sha256:01273638e1530a52a385b6ed9e6d28ff79222435248703bb92b776aa31a6c1cf

Observation 5d41821a-9f60-4e60-b9fe-846a52318582 · outbound

This paper cites Demystifying chains, trees, and graphs of thoughts.IEEE Transactions on Pattern Analysis and Machine Intelligence, page 1–20, 2025.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Demystifying chains, trees, and graphs of thoughts.IEEE Transactions on Pattern Analysis and Machine Intelligence, page 1–20, 2025

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:50.950726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:50.950726Z digest=sha256:a1f9c899c9f0b6d413af1a1156a8717d350fd9c2be54d8e6f964ba572927dd6f

Observation 1af5f92b-81d8-444e-b9a8-c22d24af9fd0 · outbound

This paper cites Fractured chain-of-thought reasoning, 2025.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Fractured chain-of-thought reasoning, 2025

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:51.109409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:51.109409Z digest=sha256:f4ea41b71f8d216d4a9367f6194ab98d07428b4fc880aaacd96d516439c24352

Observation cbb725a0-d7fd-455e-bbb9-b7d8ba2d640d · outbound

This paper cites Don’t get lost in the trees: Streamlining llm reasoning by overcoming tree search exploration pitfalls, 2025.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Don’t get lost in the trees: Streamlining llm reasoning by overcoming tree search exploration pitfalls, 2025

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:51.200061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:51.200061Z digest=sha256:fa4a1bb66d3c4244a708f73b0aed2fafc940cdf4251c777e3decdfe330913d68

Observation 7edc85e4-7976-4528-b532-6a187fd755e5 · outbound

This paper cites Bartoldson, Bhavya Kailkhura, Guillaume Lajoie, Glen Berseth, Nikolay Malkin, and Moksh Jain.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Bartoldson, Bhavya Kailkhura, Guillaume Lajoie, Glen Berseth, Nikolay Malkin, and Moksh Jain

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:51.332541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:51.332541Z digest=sha256:b1ee1da78a1c200f5597883bac2e8d5be3ec16b80f8f0416e76f20a0d08bdec4

Observation 1c9fc4d1-f4ae-4db1-8b35-db2c43c619c8 · outbound

This paper cites Self-consistency improves chain of thought reasoning in language models, 2023.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Self-consistency improves chain of thought reasoning in language models, 2023

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:51.452615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:51.452615Z digest=sha256:6fce68eb8b9eac9bc031df31775375f3ef61478b0a32637a99eb63170ab83236

Observation 5ee55051-6514-46e4-ae26-65709cad83f3 · outbound

This paper cites Escape sky-high cost: Early-stopping self-consistency for multi-step reasoning, 2024.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Escape sky-high cost: Early-stopping self-consistency for multi-step reasoning, 2024

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:51.600118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:51.600118Z digest=sha256:676f0156fa4a30c87d27d9e5404081ddb80103f208cc8f2d775c6e5e34368c00

Observation 9b90644e-ca83-4203-95db-ef85405e3824 · outbound

This paper cites Answer convergence as a signal for early stopping in reasoning, 2025.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Answer convergence as a signal for early stopping in reasoning, 2025

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:51.746637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:51.746637Z digest=sha256:2464b724c34efa83427ebd509f1ce5e9b242a7dad2c41e818dbfa115230d2e6d

Observation 6e3f106b-c4fb-46cc-bf93-d812c01ccb67 · outbound

This paper cites Confidence improves self-consistency in llms.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Confidence improves self-consistency in llms

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:51.816973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:51.816973Z digest=sha256:1f41f986f0bdc4f2a481894a72132dcdfc6521490b31d2dffb60e7011ffb6af9

Observation 4730e179-e366-476e-9708-c5e5dc94f702 · outbound

This paper cites Best-of-∞ – asymptotic performance of test-time compute, 2025.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Best-of-∞ – asymptotic performance of test-time compute, 2025

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:51.977267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:51.977267Z digest=sha256:c1be11d8dd9825469fe8b9e040bf790a0217f77216eabccb72185496a9ac96ce

Observation 14182385-0dc7-4179-946d-681866ef5ecf · outbound

This paper cites Reasoning at the right length: Adaptive budget forcing for efficient and accu- rate LLM inference.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Reasoning at the right length: Adaptive budget forcing for efficient and accu- rate LLM inference

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:52.103823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:52.103823Z digest=sha256:8d508388391dd382102ae1ab86d4b99d5162a954f7519991fefd41ad0cca094f

Observation 50983a61-6450-40ed-9b1b-860b734e1e3b · outbound

This paper cites Stop when enough: Adaptive early-stopping for chain-of-thought reasoning, 2025.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Stop when enough: Adaptive early-stopping for chain-of-thought reasoning, 2025

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:52.246411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:52.246411Z digest=sha256:8f431342be1efeb56d6e7d736771e9e3c7c46b648e20a4f9fbfff631449400a6

Observation 786d86cd-dad1-49bf-ac6f-aa3878570218 · outbound

This paper cites Dynamic early exit in reasoning models, 2025.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Dynamic early exit in reasoning models, 2025

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:52.376737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:52.376737Z digest=sha256:e674a7f0d7dbadae0a92e4d67ce56ab6fc4ace67a2a350ee3ff246f6ea19230c

Observation e90d3606-b78c-489a-81de-17a3eeb9a798 · outbound

This paper cites Forest-of-Thought: Scaling Test-Time Compute for Enhancing LLM Reasoning.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Forest-of-Thought: Scaling Test-Time Compute for Enhancing LLM Reasoning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:52.530409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:52.530409Z digest=sha256:9ca689e0c5fac0950690b9513261176507c722f701960bcef66f820ac18de816

Observation 8fac33f9-86a4-40b1-9d76-1a8c282ee0ec · outbound

This paper cites Tran, Yi Tay, and Donald Metzler.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Tran, Yi Tay, and Donald Metzler

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:52.631413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:52.631413Z digest=sha256:afa40ff5eb100bf7d05047765074d72bb3a2b3f8ec57477061e5187558176731

Observation a5ba5b92-8e66-4071-9a31-398b90ffd545 · outbound

This paper cites Learning when to plan: Efficiently allocating test-time compute for llm agents, 2025.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Learning when to plan: Efficiently allocating test-time compute for llm agents, 2025

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:52.737449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:52.737449Z digest=sha256:bcd022530625b0723870d770edb829480ddf8f580f06d383b31e6e36a9c29794

Observation 29133d63-4fb1-4801-b915-ed33260db601 · outbound

This paper cites Can past experience accelerate llm reasoning?, 2025.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Can past experience accelerate llm reasoning?, 2025

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:53.001575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:53.001575Z digest=sha256:241d09baf7684f30797febdf8ef277732a6c3fde9925ab53cb5f15576b196ae5

Observation 42fde987-d5d5-4d4e-82bc-766c6e28def9 · outbound

This paper cites React: Synergizing reasoning and acting in language models, 2023.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning React: Synergizing reasoning and acting in language models, 2023

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:53.107261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:53.107261Z digest=sha256:4ef151dedfc28cd4843041adf9ca0ed18f6df5a0e950a19d4438fa937d9f552c

Observation 502e8052-b4a4-4805-ac2a-00ddbdb5d318 · outbound

This paper cites Least-to-most prompting enables complex reasoning in large language models, 2023.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Least-to-most prompting enables complex reasoning in large language models, 2023

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:53.250671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:53.250671Z digest=sha256:43e0045a0331f76818e06522c72c3377fe0170ec36cd80a87ec39f9b9a72b43f

Observation d98bd29e-8f15-4604-956c-bcb3abe0727e · outbound

This paper cites Plan-and-solve prompting: Improving zero-shot chain-of-thought reasoning by large language models, 2023.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Plan-and-solve prompting: Improving zero-shot chain-of-thought reasoning by large language models, 2023

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:53.423313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:53.423313Z digest=sha256:50698bf1bc88869de1df97183be8f07a807758d6c811f1f8a6722fdeeb433c76

Observation 8de295ef-7984-4c33-bc11-e214f6843e6d · outbound

This paper cites Universal model routing for efficient llm inference, 2025.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Universal model routing for efficient llm inference, 2025

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:53.463678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:53.463678Z digest=sha256:84e882d99d1c5d67aa944fb5d8fdffded4e6889a1601d7460848fcf8ebcfc2b1

Observation 52014968-b680-4f54-a462-cd365e95e27d · outbound

This paper cites Chen, Trevor Chow, Ishan S.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Chen, Trevor Chow, Ishan S

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:53.570113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:53.570113Z digest=sha256:1577f11712a31df4b0458d019a10159eb0e72973f460ffb40c4e10ae71aa0a72

Observation c1b60305-9969-42f8-904b-927d5a51531a · outbound

This paper cites an unresolved cited work.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:53.732069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:53.732069Z digest=sha256:dccaf627335d9fa04745876a9e57211c754044765eeb445b35dc470116aca966

Observation ca0799a4-3655-4411-b409-7ab04b4e427b · outbound

This paper cites Masrouter: Learning to route llms for multi-agent systems, 2025.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Masrouter: Learning to route llms for multi-agent systems, 2025

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:53.845979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:53.845979Z digest=sha256:7b502223f77917eb09b6a1830c9a891d007c35b8cd583d4827d0e001dc4dcde9

Observation 49389cfa-613c-4591-9906-3d8ea5d3bb7b · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:53.941103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:53.941103Z digest=sha256:377b5cc24be8d8359d1a7ce9205412bc81d9d8885d0f47feb98ca54b7e3f4b5b

Observation 461006de-6502-44d5-9090-9cc057fb52ca · outbound

This paper cites s1: Simple test-time scaling.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning s1: Simple test-time scaling

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:54.094174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:54.094174Z digest=sha256:444835ade674ffbdeed6b8fb8021a53087e1ef5017be2da2701186892b4618aa

Observation 5aeae402-b595-4ac6-9205-079dbeecfc93 · outbound

This paper cites L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:54.255096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:54.255096Z digest=sha256:cb63cb9d41f9e3d0e4604a7e1ba03fb18234a5ebf5329903fb13738eeea54ba3

Observation 46964432-08f6-40e2-a2e4-ce7ad8318d7b · outbound

This paper cites Principles of metareasoning.Artificial Intelligence, 49(1):361– 395, 1991.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Principles of metareasoning.Artificial Intelligence, 49(1):361– 395, 1991

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:54.377086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:54.377086Z digest=sha256:da358ef931d12a6015f0e7429519654007ad525c655ec2008267bb0ed363467f

Observation 7ce641b3-6d8b-4131-809f-8e601a04b7b2 · outbound

This paper cites The pandora’s box problem with sequential inspections, 2025.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning The pandora’s box problem with sequential inspections, 2025

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:54.497296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:54.497296Z digest=sha256:e27ed3bad0831009dfdedacde7a7666f4ba1bea4eff16091d8fd32a10fa3bbbd

Observation b88aa7cf-08a9-41a3-a704-77c681a8e40b · outbound

This paper cites Deep think with confidence, 2025.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Deep think with confidence, 2025

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:54.620933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:54.620933Z digest=sha256:b4e7f76048305ffa3650b4d1ce791bcbd20e4e4f8b944cd0ae2e5004ac9e7155

Observation a166893a-e424-4399-b681-c8af805d54fc · outbound

This paper cites Ai agents as universal task solvers, 2025.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Ai agents as universal task solvers, 2025

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:54.816578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:54.816578Z digest=sha256:f3fd5a0a2cb0a9e6930d7aaaefd9ffe37940eb125a019bcb0575059bb56bf872

Observation 88938956-2678-4966-9976-67728ff3f64c · outbound

This paper cites e1: Learning adaptive control of reasoning effort.arXiv preprint arXiv:2510.27042, 2025.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning e1: Learning adaptive control of reasoning effort.arXiv preprint arXiv:2510.27042, 2025

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:54.922794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:54.922794Z digest=sha256:d5de7a032719a2ff51b573a5c8551d2e6969994b9393e91d2273fa9757c0f319

Observation 239240b7-4603-44f3-822b-e4624ea5bcb9 · outbound

This paper cites The Gittins Index: A Design Principle for Decision-Making Under Uncertainty.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning The Gittins Index: A Design Principle for Decision-Making Under Uncertainty

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:55.114608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:55.114608Z digest=sha256:67bdebd68a4458a613bf294c2cd9c4538d93e28fba732451059af17ee2942bbd

Observation c6ed65eb-277e-4737-803b-9ef798223146 · outbound

This paper cites Universal sequential search problems.Problems of information transmission, 9(3):265–266, 1973.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Universal sequential search problems.Problems of information transmission, 9(3):265–266, 1973

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:55.255548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:55.255548Z digest=sha256:234f44f62473859a78522b373a4eabb4e4d93ab040d9ab98313925c7dfefcf2e

Observation b562c7ab-0e52-477a-a83e-14e5b193da90 · outbound

This paper cites Department of Energy, 1978.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Department of Energy, 1978

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:55.395824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:55.395824Z digest=sha256:5663ac1b07b67903545e40d2bbd1017756a276991866dde9610c092c19ea9d88

Observation a548a6c5-dc39-4541-aa40-c141511a0b3e · outbound

This paper cites Cost-aware bayesian optimization via the pandora’s box gittins index.Advances in Neural Information Processing Systems, 37:115523–115562, 2024.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Cost-aware bayesian optimization via the pandora’s box gittins index.Advances in Neural Information Processing Systems, 37:115523–115562, 2024

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:55.472999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:55.472999Z digest=sha256:3d7e46b76b0e55bdefc45c2c7055ed10a75dd440c9572531ceef48562cdace14

Observation 18455bb8-e886-4175-b0d9-53b30b46fa39 · outbound

This paper cites Qwen3 Technical Report.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Qwen3 Technical Report

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:55.559048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:55.559048Z digest=sha256:7e916525e8b01b6882e5fc70030a5837b610ab1e9a6f487215676d836585a10c

Observation b01a01f2-1f88-415d-b189-f12636591218 · outbound

This paper cites Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:55.692856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:55.692856Z digest=sha256:2fbab404f10dc98eb6a70fbd9badf126b3144c9cd02772a14531145f0b06e3c9

Observation 70cf59c5-d901-4c73-8691-fbc6a7a6970c · outbound

This paper cites 2024 amc 12b — problems and solutions.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning 2024 amc 12b — problems and solutions

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:55.809995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:55.809995Z digest=sha256:21d2e5899708069e6785c616ff56abb91540ecdb3b4b2f18ffd0c48e4374023a

Observation 4de28b56-f962-4df9-a4b5-7642a9198805 · outbound

This paper cites 2024 amc 12a — problems and solutions.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning 2024 amc 12a — problems and solutions

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:55.923446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:55.923446Z digest=sha256:d182ecc4152830de4119d89e5c282f68a2287dcd190e4e8295f5b3cad5e5c97e

Observation 2219d362-364b-43fe-bfb3-b9b214c1ab7d · outbound

This paper cites 2024 amc 10b — problems and solutions.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning 2024 amc 10b — problems and solutions

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:56.049118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:56.049118Z digest=sha256:a8c0ea8c854395736ac46bc9af263647351a1c9cc3d34308af8ef7d5fa40041f

Observation 4292be6a-bbac-4e26-b594-7ec140183df4 · outbound

This paper cites 2024 amc 10a — problems and solutions.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning 2024 amc 10a — problems and solutions

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:56.171172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:56.171172Z digest=sha256:0a6f9cae610ae12251aef8509213e42b89f1d8ffbde092ddb65a12b41e0a8da3

Observation 43dff967-fed9-4e4c-b231-887db9ce705a · outbound

This paper cites Solving quantitative reasoning problems with language models.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Solving quantitative reasoning problems with language models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:56.300191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:56.300191Z digest=sha256:944c831cfe32e8442869b24b97f11cd06905bbb818f5dbc29803d91ce25abdce

Observation e53748f6-761d-4301-a84b-647029e547c8 · outbound

This paper cites Let’s verify step by step, 2023.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Let’s verify step by step, 2023

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:56.479317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:56.479317Z digest=sha256:ffa67f1e30e338fff7b7a1b6cdb89daaddbf3e7a75db9bdcc075026926daf67e

Observation 30257503-5cd5-49c0-9a65-d9a58e30e44c · outbound

This paper cites 2024 aime i — problems and solutions.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning 2024 aime i — problems and solutions

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:56.651338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:56.651338Z digest=sha256:e0e7e0b032bfe99d072869c839801946882e90b1d2331fcf0855ad60cf4a88e7

Observation 8e18e2b8-bc6b-42ce-b133-dc20326ce6ff · outbound

This paper cites 2024 aime ii — problems and solutions.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning 2024 aime ii — problems and solutions

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:56.825467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:56.825467Z digest=sha256:b0e9ce88c7d66c981321297dab4a5a601cb0b1f3fbdfef55f0a8efa226f57c5e

Observation 80d71ef5-bd30-4ff2-ad38-578f17acc464 · outbound

This paper cites 2025 aime i — problems and solutions.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning 2025 aime i — problems and solutions

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:56.973166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:56.973166Z digest=sha256:31cc7e8a7c35c543dbf55b560fd8da1495419c0b21b9c60c7a8cbfabc4935579

Observation fcb2e03b-2cad-49f6-a32f-1fe44b5feecc · outbound

This paper cites 2025 aime ii — problems and solutions.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning 2025 aime ii — problems and solutions

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:57.117932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:57.117932Z digest=sha256:b0de962e472712659f686dff0fcfe04333d85b9d0b23991781ccf940beedcb8f

Pith citing papers

Observation f0dec1ea-6bae-4e41-9e5c-fc8c872d320b · inbound

ExecTune: Effective Steering of Black-Box LLMs with Guide Models cites this paper.

ExecTune: Effective Steering of Black-Box LLMs with Guide Models Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-07-27T01:19:51.926662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-10T16:55:34.091812Z digest=sha256:8d9083f67677ab264ca47b1aa57a246b7b3777170fd4db38f583a70bca45c9eb

Observation bb75c710-b8c2-4b17-a3fb-8b8755367b24 · inbound

Cache-Aware Prompt Compression:A Two-Tier Cost Model for LLM API Caching cites this paper.

Cache-Aware Prompt Compression:A Two-Tier Cost Model for LLM API Caching Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T23:14:56.115834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:14:56.115834Z digest=sha256:99dd5009501bcdbf437fc09a5da1c00d8831a4d62907c7a421b6e186514131e9