Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T22:13:08.130979Z
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 5 inbound Pith citation observations for arXiv:2505.07787.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T22:13:08.130979Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T18:59:41.908963Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-10T08:53:03.650861Z
48 of 48 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 734f7bea-eeab-45ad-abd8-006d4402c301 · outbound
Learning from Peers in Reasoning Models Learning to reason with llms, 2024
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 6bf98da7-f546-4aa5-9ca7-886c4289e752 · outbound
Learning from Peers in Reasoning Models Openai o1 system card, 2024
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation ad3bf63e-a3db-46db-9eaa-619348128eb5 · outbound
Learning from Peers in Reasoning Models Openai o3 mini, 2025
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation dc6284a3-b2d5-4ecb-964f-b3e11ac25fd5 · outbound
Learning from Peers in Reasoning Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55a2075c-9408-440e-a9b7-c37b4432f74f · outbound
Learning from Peers in Reasoning Models Qwq-32b: Embracing the power of reinforcement learning, March 2025
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea969c82-05f6-49ee-b5a9-89dc96260bd9 · outbound
Learning from Peers in Reasoning Models A Survey on Test-Time Scaling in Large Language Models: What, How, Where, and How Well?
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cfcec97f-116b-433a-a896-1a9e4e90c1ca · outbound
Learning from Peers in Reasoning Models Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09db910c-f15c-49d6-ba1e-66a7977444c0 · outbound
Learning from Peers in Reasoning Models Revisiting the Test-Time Scaling of o1-like Models: Do they Truly Possess Test-Time Scaling Capabilities?
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8787dd33-e9be-4d8b-ac31-7778051f82ba · outbound
Learning from Peers in Reasoning Models Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d52a79d9-5ea9-4ef2-b9ad-d9f9f52aafa0 · outbound
Learning from Peers in Reasoning Models Peer instruction enhanced student performance on qualitative problem-solving questions
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation d962cc0d-933a-413f-9a97-8a4c4ac2ce11 · outbound
Learning from Peers in Reasoning Models Implementation of the peer-led team-learning instructional model as a stopgap measure improves student achievement for students opting out of laboratory
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 668fc0d9-66f7-4f22-9851-baab6bf8dbee · outbound
Learning from Peers in Reasoning Models Clean evidence on peer effects
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 4b1d8937-a5e7-4a33-a8b8-dc05abfedb25 · outbound
Learning from Peers in Reasoning Models American Invitational Mathematics Examination - AIME 2024, February 2024
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation e6a39753-6514-4de5-9f66-4d5b16d0de1a · outbound
Learning from Peers in Reasoning Models American Invitational Mathematics Examination - AIME 2025, February 2025
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 324d651b-d730-4cee-ba33-3d0cb07721dc · outbound
Learning from Peers in Reasoning Models Smith, Kevin Buzzard, Timothy Gowers, Peter J
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 1102c8d7-144f-4d9d-9ae9-f5aa9067b6c0 · outbound
Learning from Peers in Reasoning Models Unresolved cited work
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed5b9d85-e626-42da-9891-434067d8f8f7 · outbound
Learning from Peers in Reasoning Models Binary codes capable of correcting deletions, insertions, and reversals
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14fef261-e6d3-4dea-a634-fde1821a4a1c · outbound
Learning from Peers in Reasoning Models Gonzalez, Hao Zhang, and Ion Stoica
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9eaa6296-24cd-4872-b0a1-d0f9e59a6567 · outbound
Learning from Peers in Reasoning Models Self-Consistency Improves Chain of Thought Reasoning in Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ae1e96c-8db6-4c78-af79-b139d745ed3b · outbound
Learning from Peers in Reasoning Models SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6061d5e9-d45b-4285-b794-51ba9b7f10bb · outbound
Learning from Peers in Reasoning Models START: Self-taught Reasoner with Tools
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d42b39d2-6305-404b-bafe-abd1e7e22b27 · outbound
Learning from Peers in Reasoning Models An Empirical Study on Eliciting and Improving R1-like Reasoning Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7d1478e-7632-4161-ad6e-09f8f8287111 · outbound
Learning from Peers in Reasoning Models Mixture-of-Agents Enhances Large Language Model Capabilities
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97c2489e-81db-49c0-9ee6-36d761fd84a5 · outbound
Learning from Peers in Reasoning Models Hello gpt-4o
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation d68fd940-c6af-4903-be10-13b66fe522e7 · outbound
Learning from Peers in Reasoning Models Deepseek-r1 thoughtology: Let’s< think> about llm reasoning
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45dc9fe9-d3a1-4268-a9c6-d49d17559e2a · outbound
Learning from Peers in Reasoning Models Scaling of Search and Learning: A Roadmap to Reproduce o1 from Reinforcement Learning Perspective
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9764a76c-7263-479a-8c0f-ab31e1acd5b5 · outbound
Learning from Peers in Reasoning Models Making language models better reasoners with step-aware verifier
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation d39af146-99d7-47b3-ad95-18d0435a955d · outbound
Learning from Peers in Reasoning Models Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ab01dbb-9d85-4df6-8754-b45ab425cf99 · outbound
Learning from Peers in Reasoning Models The Good, The Bad, and The Greedy: Evaluation of LLMs Should Not Ignore Non-Determinism
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56dbcd6f-9a2e-401a-8862-95a69533dda7 · outbound
Learning from Peers in Reasoning Models When is the consistent prediction likely to be a correct prediction?
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d09c0608-debb-4b3a-a87c-b55311d75a6a · outbound
Learning from Peers in Reasoning Models Scaling laws for reward model overoptimization
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98317615-8b66-4c69-8820-7c5f07c491cf · outbound
Learning from Peers in Reasoning Models Training Verifiers to Solve Math Word Problems
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98a9b889-2cc8-4112-8684-8984e0af0143 · outbound
Learning from Peers in Reasoning Models Fast Best-of-N Decoding via Speculative Rejection
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a651ec4-6066-4a89-aa06-d865f167084e · outbound
Learning from Peers in Reasoning Models BoNBoN Alignment for Large Language Models and the Sweetness of Best-of-n Sampling
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc3b0bdc-a4ad-4032-95cf-58a4c1bb206b · outbound
Learning from Peers in Reasoning Models Variational Best-of-N Alignment
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a494a453-106e-4256-927c-1307cbe87ec0 · outbound
Learning from Peers in Reasoning Models BOND: Aligning LLMs with Best-of-N Distillation
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 176eed5c-228b-4cd1-899e-d44f7633ce10 · outbound
Learning from Peers in Reasoning Models Improv- ing factuality and reasoning in language models through multiagent debate
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9dd506f2-cd2e-4aaf-b279-c7f4132f68e4 · outbound
Learning from Peers in Reasoning Models ReConcile: Round-Table Conference Improves Reasoning via Consensus among Diverse LLMs
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 924fb163-8900-4ee8-a037-f88f59c158df · outbound
Learning from Peers in Reasoning Models Corex: Pushing the Boundaries of Complex Reasoning through Multi-Model Collaboration
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a1ef6dc-b383-42ac-b0f4-7c6b6473adf2 · outbound
Learning from Peers in Reasoning Models CoMM: Collaborative Multi-Agent, Multi-Reasoning-Path Prompting for Complex Problem Solving
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b6d9edd-0f8f-4f5f-819b-a990fabe5336 · outbound
Learning from Peers in Reasoning Models Malt: Improving reasoning with multi-agent llm training
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d172b947-ab7d-426d-802c-3b912c557430 · outbound
Learning from Peers in Reasoning Models SMoA: Improving Multi-agent Large Language Models with Sparse Mixture-of-Agents
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea781724-dee2-47fd-ae7e-c2aedc302a9b · outbound
Learning from Peers in Reasoning Models The First Few Tokens Are All You Need: An Efficient and Effective Unsupervised Prefix Fine-Tuning Method for Reasoning Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4694d4ce-ee0d-40cf-9e13-7d1c26b971ef · outbound
Learning from Peers in Reasoning Models Unresolved cited work
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation daa131d4-2068-4c73-abd5-d20f3c261cfa · outbound
Learning from Peers in Reasoning Models To summarize, my recent findings are: I attempted to expressp ands in terms ofq and r using the conditionpr +qs = 0, leading top =kq and s =−kr... Peer 4:
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation c17113c2-b587-4e50-9fa8-84ea57b0297e · outbound
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 3a498150-0106-41f2-ac55-37b4d3e79e79 · outbound
Reference 1190
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation efd968be-ce2e-42cf-8aef-04a1e45fe9d6 · outbound
Learning from Peers in Reasoning Models Aha” moments Figure 18: We illustrate the average number of tokens and “Aha
Reference 2025
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation e66e0fbc-e017-4a28-bfda-b3c83cc3bc04 · inbound
Adaptive Termination for Multi-round Parallel Reasoning: An Universal Semantic Entropy-Guided Framework Learning from Peers in Reasoning Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63245719-83c1-421f-965f-4e7e4baeb7c0 · inbound
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Learning from Peers in Reasoning Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7824f657-254f-4a13-b8c3-96efb2d142a9 · inbound
Sticker-TTS: Learn to Utilize Historical Experience with a Sticker-driven Test-Time Scaling Framework Learning from Peers in Reasoning Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c27b8e1e-b6de-48bb-b66c-bb51be424721 · inbound
EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance Learning from Peers in Reasoning Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70809acb-c836-4335-8fd2-839cd2eee6eb · inbound
Cut Your Losses! Learning to Prune Paths Early for Efficient Parallel Reasoning Learning from Peers in Reasoning Models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.