Pith. sign in

Paper Citation Record · LEDGER

When Does In-Context Search Help? A Sampling-Complexity Theory of Reflection-Driven Reasoning

As of 21 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 0 inbound Pith citation observations for arXiv:2607.06720.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.06720 v1

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-10T22:46:24.057572Z

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

38 of 38 outbound references displayed

  • verified exact7
  • verified fuzzy6
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch23

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0ddcbfe1-5b30-4caf-aaaa-d2e6461a3fc4 · outbound

This paper cites Scheherazade: Evaluating Chain-of-Thought Math Reasoning in LLMs with Chain-of-Problems.

When Does In-Context Search Help? A Sampling-Complexity Theory of Reflection-Driven Reasoning Scheherazade: Evaluating Chain-of-Thought Math Reasoning in LLMs with Chain-of-Problems

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T22:47:36.800236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-10T22:46:24.057572Z digest=sha256:7aac7279d22aef19acffbf404ea8d6b5f28e6b4270940963fb43d16c8969f946

Observation c74fb849-5df2-442c-bc08-abea72af0a35 · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?.

When Does In-Context Search Help? A Sampling-Complexity Theory of Reflection-Driven Reasoning Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T22:47:36.908710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-10T22:46:24.057572Z digest=sha256:71e809dff07e7c6b3e26f21f07205f894e3b6c450866b30b816bdef67bcce2f8

Observation 1538791e-a5a4-47e0-9fee-7081a57d4038 · outbound

This paper cites Reasoning with Sampling: Your Base Model is Smarter Than You Think.

When Does In-Context Search Help? A Sampling-Complexity Theory of Reflection-Driven Reasoning Reasoning with Sampling: Your Base Model is Smarter Than You Think

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T22:47:36.926198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-10T22:46:24.057572Z digest=sha256:df8276be233650070c8670819a53af1ae9243a12687fa563e68ea507d7c58047

Observation 1b0eb939-1108-44a2-bdef-b26479d5cea1 · outbound

This paper cites Information gain-based policy optimization: A simple and effective approach for multi-turn llm agents.

When Does In-Context Search Help? A Sampling-Complexity Theory of Reflection-Driven Reasoning Information gain-based policy optimization: A simple and effective approach for multi-turn llm agents

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-07-10T22:47:36.783304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-10T22:46:24.057572Z digest=sha256:96aa47230daad567c6d3b76a45aed6a7c27cb1336de9a86ecd7ee629fa93db4b

Observation 5b8c275b-e331-46c3-86ef-3512950af717 · outbound

This paper cites Demystifying reinforcement learning in agentic reasoning.

When Does In-Context Search Help? A Sampling-Complexity Theory of Reflection-Driven Reasoning Demystifying reinforcement learning in agentic reasoning

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-07-10T22:47:37.015172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-10T22:46:24.057572Z digest=sha256:39810c4199006039fb695d400964580c6cbbad334cf55ee48db5773958325cf6

Observation 32c50b34-3160-4960-80d1-44575255273c · outbound

This paper cites Verifying Meta-Awareness via Predictive Rewards in Reasoning Models.

When Does In-Context Search Help? A Sampling-Complexity Theory of Reflection-Driven Reasoning Verifying Meta-Awareness via Predictive Rewards in Reasoning Models

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T22:47:37.015452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-10T22:46:24.057572Z digest=sha256:8ecedf55158c5f2991496d1ad0aff7741a934da6c43a2c105066f50b7cf7362e

Observation 72b37363-3d26-4015-824f-cb4fe426bdf5 · outbound

This paper cites Attention as a Compass: Efficient Exploration for Process-Supervised RL in Reasoning Models.

When Does In-Context Search Help? A Sampling-Complexity Theory of Reflection-Driven Reasoning Attention as a Compass: Efficient Exploration for Process-Supervised RL in Reasoning Models

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-07-10T22:47:37.115519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-10T22:46:24.057572Z digest=sha256:e2268d116a5f8664fcef1fd5c8351a6c908fdb8f21e76abb3f9c88b4ed9d786d

Observation dbbfc7b0-236f-4916-9693-d74153e63ae4 · outbound

This paper cites LLM-guided Hierarchical Search for End-to-end Reasoning Intensive Retrieval.

When Does In-Context Search Help? A Sampling-Complexity Theory of Reflection-Driven Reasoning LLM-guided Hierarchical Search for End-to-end Reasoning Intensive Retrieval

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T22:47:36.798485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-10T22:46:24.057572Z digest=sha256:f28bf59224e52d5718af92cc79d0f178b16f8747cbd1b7b6d849adfb2f92b73f

Observation c631b2ca-cb05-41de-8bb4-4df31e7b8f85 · outbound

This paper cites On Provable Length and Compositional Generalization.

When Does In-Context Search Help? A Sampling-Complexity Theory of Reflection-Driven Reasoning On Provable Length and Compositional Generalization

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-10T22:47:36.689542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-10T22:46:24.057572Z digest=sha256:ec8c78766aadd3a2555060f34dc00f3c435049abda204c4dd8899e9a989e197c

Observation 666564c5-8afa-4cf7-9ad9-e5551f5b1fda · outbound

This paper cites Sub-Task Decomposition Enables Learning in Sequence to Sequence Tasks.

When Does In-Context Search Help? A Sampling-Complexity Theory of Reflection-Driven Reasoning Sub-Task Decomposition Enables Learning in Sequence to Sequence Tasks

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T22:47:36.943776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-10T22:46:24.057572Z digest=sha256:5d466070c39af081a10a89c1236715b89fd1bc71c8c564192425853e4c22c100

Observation 047bdbaf-cc9e-4a62-ae54-500580972956 · outbound

This paper cites The Expressive Power of Transformers with Chain of Thought.

When Does In-Context Search Help? A Sampling-Complexity Theory of Reflection-Driven Reasoning The Expressive Power of Transformers with Chain of Thought

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T22:47:37.053420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-10T22:46:24.057572Z digest=sha256:97426e1a36c0545a1754a576058e9bf0c998675fa57c82b6ec232b9b5d7db427

Observation 3d6c804b-ca24-4ec5-b231-b2d8c0e290a6 · outbound

This paper cites Auto-Regressive Next-Token Predictors are Universal Learners.

When Does In-Context Search Help? A Sampling-Complexity Theory of Reflection-Driven Reasoning Auto-Regressive Next-Token Predictors are Universal Learners

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-10T22:47:36.921943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-10T22:46:24.057572Z digest=sha256:0205bb301346a4e1cb77eac3fdb524bf095e4ab22c1659b4cc18ee1deafe644a

Observation d6714c33-2067-40cf-a0cb-d857c98367b2 · outbound

This paper cites On Limitations of the Transformer Architecture.

When Does In-Context Search Help? A Sampling-Complexity Theory of Reflection-Driven Reasoning On Limitations of the Transformer Architecture

Reference 13

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T22:47:36.634843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-10T22:46:24.057572Z digest=sha256:3bb882bb37a933801dd1e9e8188e4c2d5b1d0c9ec853e84538e73fbbd86f1503

Observation b15747ea-c3e3-4f09-b2d6-c31e310eea0e · outbound

This paper cites ICLR 2024 Workshop on Mathematical and Empirical Understanding of Foundation Models , year=.

When Does In-Context Search Help? A Sampling-Complexity Theory of Reflection-Driven Reasoning ICLR 2024 Workshop on Mathematical and Empirical Understanding of Foundation Models , year=

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T22:47:37.632953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-10T22:46:24.057572Z digest=sha256:6f0a14070b64292ddb89016a54d39fcb5f63717e1cfd0627314a089334cfc0c5

Observation 5474dfe3-6953-48dc-ae87-7d5f9231a3de · outbound

This paper cites arXiv preprint arXiv:2503.17523 , year =.

When Does In-Context Search Help? A Sampling-Complexity Theory of Reflection-Driven Reasoning arXiv preprint arXiv:2503.17523 , year =

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-10T22:47:36.840385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-10T22:46:24.057572Z digest=sha256:431c6a2772bdbdd92c05191db6b8947a1369473f2803b95a4ebbcdf483bd12ed

Observation 075f5c22-8e17-4dbe-90ae-4d3499790df0 · outbound

This paper cites arXiv preprint arXiv:2509.10739 , year=.

When Does In-Context Search Help? A Sampling-Complexity Theory of Reflection-Driven Reasoning arXiv preprint arXiv:2509.10739 , year=

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-10T22:47:36.747392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-10T22:46:24.057572Z digest=sha256:9040e036a3851dc46d9702c2f91d135c86ad0929014580cc677186e5d31534c9

Observation 12edb697-4695-413a-96a7-9379223637b6 · outbound

This paper cites Beyond markovian: Reflective exploration via bayes-adaptive rl for llm reasoning.

When Does In-Context Search Help? A Sampling-Complexity Theory of Reflection-Driven Reasoning Beyond markovian: Reflective exploration via bayes-adaptive rl for llm reasoning

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-07-10T22:47:36.854345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-10T22:46:24.057572Z digest=sha256:40a8b4cd8e7a7a4087af2eef7b5b133de6a5555e9c9cba049e422c7cf842a7cf

Observation fc7db696-43b6-463a-8437-49580cce97e8 · outbound

This paper cites arXiv preprint arXiv:2509.25810 , year=.

When Does In-Context Search Help? A Sampling-Complexity Theory of Reflection-Driven Reasoning arXiv preprint arXiv:2509.25810 , year=

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-10T22:47:36.819392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-10T22:46:24.057572Z digest=sha256:d57ee64ef0744a95a08f9147b2d0eeb839cdcaa3893c76d3f0e960c245422750

Observation fcf30ff1-767a-4c01-8a94-e4a4f45bb65f · outbound

This paper cites From Reasoning to Super-Intelligence: A Search-Theoretic Perspective.

When Does In-Context Search Help? A Sampling-Complexity Theory of Reflection-Driven Reasoning From Reasoning to Super-Intelligence: A Search-Theoretic Perspective

Reference 19

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T22:47:36.961264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-10T22:46:24.057572Z digest=sha256:676a4d7271a9965407e7e3112cb3924d1612928219079ccd1baf5ecadb549ce7

Observation aa66c9f2-0212-4de1-9d4c-7a6497078044 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

When Does In-Context Search Help? A Sampling-Complexity Theory of Reflection-Driven Reasoning Evaluating Large Language Models Trained on Code

Reference 20

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T22:47:36.891002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-10T22:46:24.057572Z digest=sha256:49382b9de3d919b6f12b507a4338af0b2308875fb402a62d5b54e896c68003e9

Observation 56d04362-da81-4788-aaec-ec92f030ef70 · outbound

This paper cites Science , volume=.

When Does In-Context Search Help? A Sampling-Complexity Theory of Reflection-Driven Reasoning Science , volume=

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T22:47:37.648981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-10T22:46:24.057572Z digest=sha256:901994e443e1ab6a793db218b8ae8ed1bf16bcf7b73fadf94df07515835984dc

Observation 1d8be571-b92c-4a62-9b7d-ece1011e40a4 · outbound

This paper cites an unresolved cited work.

When Does In-Context Search Help? A Sampling-Complexity Theory of Reflection-Driven Reasoning Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-07-10T22:47:37.697477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-10T22:46:24.057572Z digest=sha256:42cae7c54f325d7bf5ccabaaec783487a7ad7986ea4c4c948d372e2e3824e55d

Observation cc7f2098-a735-4c47-aeb1-d47aa7771757 · outbound

This paper cites Code Generation with AlphaCodium: From Prompt Engineering to Flow Engineering.

When Does In-Context Search Help? A Sampling-Complexity Theory of Reflection-Driven Reasoning Code Generation with AlphaCodium: From Prompt Engineering to Flow Engineering

Reference 23

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T22:47:37.033476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-10T22:46:24.057572Z digest=sha256:f64cdaf620b8318bff797efc19c52ac76bd42cb9e37fbdde7bace1d077c7a01e

Observation 79c40c40-e39a-4813-993e-40c01865b8e1 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

When Does In-Context Search Help? A Sampling-Complexity Theory of Reflection-Driven Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 24

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T22:47:37.091834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-10T22:46:24.057572Z digest=sha256:a875ca100edc7e2dc4fa5f7acc73b099764f681721ea15a943da4014f936e8c6

Observation 715cae4c-d91c-46ef-a8f0-4ab76e33cbec · outbound

This paper cites 2024 , publisher=.

When Does In-Context Search Help? A Sampling-Complexity Theory of Reflection-Driven Reasoning 2024 , publisher=

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T22:47:37.611645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-10T22:46:24.057572Z digest=sha256:4cc763a8d949e88c3bc4fb0e1288a30b61413e10f2ade1dd2b9b3256faa4a622

Observation 154e4261-7417-4538-9619-8b308517adc7 · outbound

This paper cites Advances in neural information processing systems , volume=.

When Does In-Context Search Help? A Sampling-Complexity Theory of Reflection-Driven Reasoning Advances in neural information processing systems , volume=

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T22:47:37.681147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-10T22:46:24.057572Z digest=sha256:0e183c00b8a59ef8d532ba875d12694504ce657d4097ea06f8339706706c9a01

Observation e273d32a-704a-4d20-9c8d-5731ed8669bb · outbound

This paper cites OpenAI o1 System Card.

When Does In-Context Search Help? A Sampling-Complexity Theory of Reflection-Driven Reasoning OpenAI o1 System Card

Reference 27

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T22:47:37.153870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-10T22:46:24.057572Z digest=sha256:243ea57cf9f3528ce59bef247680d40a64da14b367d0b8d63309cabc685013f5

Observation 520111fb-09ab-4729-ab92-edf28a61e26a · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

When Does In-Context Search Help? A Sampling-Complexity Theory of Reflection-Driven Reasoning Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 28

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T22:47:36.997958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-10T22:46:24.057572Z digest=sha256:9605c9b24100e9e99fb19f7aea902015492dbfdcf4d57e45475153750f0a2e16

Observation b112f2d3-a095-4b9a-86e3-209e239b050c · outbound

This paper cites Demystifying Long Chain-of-Thought Reasoning in LLMs.

When Does In-Context Search Help? A Sampling-Complexity Theory of Reflection-Driven Reasoning Demystifying Long Chain-of-Thought Reasoning in LLMs

Reference 29

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T22:47:37.135482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-10T22:46:24.057572Z digest=sha256:0a7cf08030cb599cd7e351e19ae2852ddd9ed1dfbc0be78789728252cbb4ca74

Observation 25514781-7d34-4b67-81b1-270ff1d85a37 · outbound

This paper cites Reflexion: Language Agents with Verbal Reinforcement Learning.

When Does In-Context Search Help? A Sampling-Complexity Theory of Reflection-Driven Reasoning Reflexion: Language Agents with Verbal Reinforcement Learning

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-07-10T22:47:36.708014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-10T22:46:24.057572Z digest=sha256:ddeee1ee8143bdd21101f911dfc7518bc22cf1548ace439ad604a43969165c4b

Observation 00c0a389-7f18-4d62-a670-35a928300823 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

When Does In-Context Search Help? A Sampling-Complexity Theory of Reflection-Driven Reasoning Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 31

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T22:47:37.071726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-10T22:46:24.057572Z digest=sha256:6c3e71b024185387d566afecb07284d0c08ba1738938dce6761ebc0e6de63c54

Observation f50e1aa2-1f74-4ff7-b340-f7e988bc6cc5 · outbound

This paper cites Advances in neural information processing systems , volume=.

When Does In-Context Search Help? A Sampling-Complexity Theory of Reflection-Driven Reasoning Advances in neural information processing systems , volume=

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T22:47:37.715384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-10T22:46:24.057572Z digest=sha256:051efc30e717285aa0f7723760bc856d3876aeb0a23012da49f4253949298ed9

Observation 0707d643-2523-4897-883c-f703c7f958ae · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

When Does In-Context Search Help? A Sampling-Complexity Theory of Reflection-Driven Reasoning Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-07-10T22:47:36.979600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-10T22:46:24.057572Z digest=sha256:132f16b5e41f6e1ce5f971cee9f59b117cfa9705174cf4444f026f6939eb7493

Observation 92397945-18f4-4439-8c67-99b0c1dbb7f5 · outbound

This paper cites Rethinking the Unsolvable: When In-Context Search Meets Test-Time Scaling.

When Does In-Context Search Help? A Sampling-Complexity Theory of Reflection-Driven Reasoning Rethinking the Unsolvable: When In-Context Search Meets Test-Time Scaling

Reference 34

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T22:47:36.979178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-10T22:46:24.057572Z digest=sha256:a98a4636a418ef9387862b935f0b2c55d8015e82080e817df644e80861939846

Observation 76da442d-4f17-4027-a65d-13edee3e41db · outbound

This paper cites Qwen3 Technical Report.

When Does In-Context Search Help? A Sampling-Complexity Theory of Reflection-Driven Reasoning Qwen3 Technical Report

Reference 35

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T22:47:37.051667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-10T22:46:24.057572Z digest=sha256:3ff2b76406efddf9fa9b1ac0630d273b6dce6a0c768e914080e820724ce4ef1a

Observation d08c656a-e749-459f-bf30-08c6c569346e · outbound

This paper cites an unresolved cited work.

When Does In-Context Search Help? A Sampling-Complexity Theory of Reflection-Driven Reasoning Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-07-10T22:47:37.664062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-10T22:46:24.057572Z digest=sha256:2e2ce0448bda663bdd768de49f164de4bbc1ec1aa77c39e9a006a0fcf75852ea

Observation 858cd4ae-daa6-4bae-83d1-ee2133bf73a1 · outbound

This paper cites Towards System 2 Reasoning in LLMs: Learning How to Think With Meta Chain-of-Thought.

When Does In-Context Search Help? A Sampling-Complexity Theory of Reflection-Driven Reasoning Towards System 2 Reasoning in LLMs: Learning How to Think With Meta Chain-of-Thought

Reference 37

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T22:47:36.883086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-10T22:46:24.057572Z digest=sha256:c8a2ae98ef29d71863105c260cd403a4d2f218eaeb6a1c00bd17160edc747b80

Observation 7a794ba1-4f77-43bd-8acf-95e2d277c0ab · outbound

This paper cites International Conference on Machine Learning , pages=.

When Does In-Context Search Help? A Sampling-Complexity Theory of Reflection-Driven Reasoning International Conference on Machine Learning , pages=

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T22:47:37.733191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-10T22:46:24.057572Z digest=sha256:93fc7598165818e09edaf2c547a2d280c543dcb379ce91522517bc5b4fee1669

Pith citing papers

No inbound Pith citation observations are available.