Pith. sign in

Paper Citation Record · LEDGER

Learning from Peers in Reasoning Models

As of 22 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 5 inbound Pith citation observations for arXiv:2505.07787.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.07787 v1

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T22:13:08.130979Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:59:41.908963Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-10T08:53:03.650861Z

Reference resolution

48 of 48 outbound references displayed

  • verified exact0
  • verified fuzzy15
  • unresolved32
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 734f7bea-eeab-45ad-abd8-006d4402c301 · outbound

This paper cites Learning to reason with llms, 2024.

Learning from Peers in Reasoning Models Learning to reason with llms, 2024

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:13:09.223355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:13:07.855064Z digest=sha256:4b95c56f7f633a681619b86a4e395175b087633e47b51c90e8066c6d0578b5f0

Observation 6bf98da7-f546-4aa5-9ca7-886c4289e752 · outbound

This paper cites Openai o1 system card, 2024.

Learning from Peers in Reasoning Models Openai o1 system card, 2024

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:13:09.208189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:13:07.861273Z digest=sha256:04e390023165b93693b6a7aec58f742ed9792a95ab6850927177a64885379962

Observation ad3bf63e-a3db-46db-9eaa-619348128eb5 · outbound

This paper cites Openai o3 mini, 2025.

Learning from Peers in Reasoning Models Openai o3 mini, 2025

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:13:09.186264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:13:07.866487Z digest=sha256:c67152d44dfdd334d9e997668760fffdad7bd9513da83d808cb0695bbb64e7a1

Observation dc6284a3-b2d5-4ecb-964f-b3e11ac25fd5 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Learning from Peers in Reasoning Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T22:13:07.871741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:13:07.871741Z digest=sha256:2635fb073f747c2837c9c5f9842c9b1eabf943640ed13f5c48ba2274139bdfc7

Observation 55a2075c-9408-440e-a9b7-c37b4432f74f · outbound

This paper cites Qwq-32b: Embracing the power of reinforcement learning, March 2025.

Learning from Peers in Reasoning Models Qwq-32b: Embracing the power of reinforcement learning, March 2025

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T22:13:07.877747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:13:07.877747Z digest=sha256:02632f8425da544f0d2ac6db1da271b126165de540c70ce607a030624717d8f5

Observation ea969c82-05f6-49ee-b5a9-89dc96260bd9 · outbound

This paper cites A Survey on Test-Time Scaling in Large Language Models: What, How, Where, and How Well?.

Learning from Peers in Reasoning Models A Survey on Test-Time Scaling in Large Language Models: What, How, Where, and How Well?

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T22:13:07.889708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:13:07.889708Z digest=sha256:1722480191ebb74a714dbe9861de471ccc7ff393db8bb196693ea01f5d259f70

Observation cfcec97f-116b-433a-a896-1a9e4e90c1ca · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Learning from Peers in Reasoning Models Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T22:13:07.895199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:13:07.895199Z digest=sha256:aca5b97b7c558ff6f2ac2759fbb8c88bb59f87fc2e0ad97a226ac26240e5ce5f

Observation 09db910c-f15c-49d6-ba1e-66a7977444c0 · outbound

This paper cites Revisiting the Test-Time Scaling of o1-like Models: Do they Truly Possess Test-Time Scaling Capabilities?.

Learning from Peers in Reasoning Models Revisiting the Test-Time Scaling of o1-like Models: Do they Truly Possess Test-Time Scaling Capabilities?

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T22:13:07.899674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:13:07.899674Z digest=sha256:7ce01fd88a54859b614a90554edcf2cc5e0fb4124433acf6298a0f42d9e3db7d

Observation 8787dd33-e9be-4d8b-ac31-7778051f82ba · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Learning from Peers in Reasoning Models Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T22:13:07.904679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:13:07.904679Z digest=sha256:26e57c0399836846ff86e13a9f2a21b1687c0dd3be843f7ef0b8a372e360f339

Observation d52a79d9-5ea9-4ef2-b9ad-d9f9f52aafa0 · outbound

This paper cites Peer instruction enhanced student performance on qualitative problem-solving questions.

Learning from Peers in Reasoning Models Peer instruction enhanced student performance on qualitative problem-solving questions

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:13:09.151022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:13:07.909850Z digest=sha256:ff0b4185c8af06c2cf06b3aaa145c0e9eac4f70d2e6d6ba432b52fca08d464a5

Observation d962cc0d-933a-413f-9a97-8a4c4ac2ce11 · outbound

This paper cites Implementation of the peer-led team-learning instructional model as a stopgap measure improves student achievement for students opting out of laboratory.

Learning from Peers in Reasoning Models Implementation of the peer-led team-learning instructional model as a stopgap measure improves student achievement for students opting out of laboratory

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:13:09.134160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:13:07.915099Z digest=sha256:9bd1401e65c9afd8c67135e35848292a50f8bbb55468e647a3a878c29c23ec93

Observation 668fc0d9-66f7-4f22-9851-baab6bf8dbee · outbound

This paper cites Clean evidence on peer effects.

Learning from Peers in Reasoning Models Clean evidence on peer effects

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:13:09.111363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:13:07.920055Z digest=sha256:0555c0725f1dfe899c0fde0276812352b6e7f03267dbdacec99595e8252934e6

Observation 4b1d8937-a5e7-4a33-a8b8-dc05abfedb25 · outbound

This paper cites American Invitational Mathematics Examination - AIME 2024, February 2024.

Learning from Peers in Reasoning Models American Invitational Mathematics Examination - AIME 2024, February 2024

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:13:09.095477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:13:07.924952Z digest=sha256:2aafda75e06332b18ddb42449347d9375801d7e3cce6d0397bc670f57d1c417e

Observation e6a39753-6514-4de5-9f66-4d5b16d0de1a · outbound

This paper cites American Invitational Mathematics Examination - AIME 2025, February 2025.

Learning from Peers in Reasoning Models American Invitational Mathematics Examination - AIME 2025, February 2025

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:13:09.072866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:13:07.929936Z digest=sha256:c44a8a0bcfe34c2768e9cacba47d17cc7356a5931473d27802646b573adc35c8

Observation 324d651b-d730-4cee-ba33-3d0cb07721dc · outbound

This paper cites Smith, Kevin Buzzard, Timothy Gowers, Peter J.

Learning from Peers in Reasoning Models Smith, Kevin Buzzard, Timothy Gowers, Peter J

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:13:09.052024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:13:07.934051Z digest=sha256:534f63ba5c8eb7439eba6be65f15791d31d3370b4ed4387deed79abb8382da50

Observation 1102c8d7-144f-4d9d-9ae9-f5aa9067b6c0 · outbound

This paper cites an unresolved cited work.

Learning from Peers in Reasoning Models Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T22:13:07.939382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:13:07.939382Z digest=sha256:79bc29b457597c35bac0205a7966a63a91596334bfd048bb025a9b10ebf7cbe2

Observation ed5b9d85-e626-42da-9891-434067d8f8f7 · outbound

This paper cites Binary codes capable of correcting deletions, insertions, and reversals.

Learning from Peers in Reasoning Models Binary codes capable of correcting deletions, insertions, and reversals

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T22:13:07.947670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:13:07.947670Z digest=sha256:0cafc22d3385623f86c516c011563ab90316824ca6c768cba85fc70a2b18fc58

Observation 14fef261-e6d3-4dea-a634-fde1821a4a1c · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

Learning from Peers in Reasoning Models Gonzalez, Hao Zhang, and Ion Stoica

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T22:13:07.963086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:13:07.963086Z digest=sha256:1aae11ce27a8a15500e4926e9a827fb25a26a88a22b55cc69323f978baf2b20a

Observation 9eaa6296-24cd-4872-b0a1-d0f9e59a6567 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Learning from Peers in Reasoning Models Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T22:13:07.967962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:13:07.967962Z digest=sha256:70a6cae9bb2f90a2247e98ebbb0a76bd0d8ab7456f11eadda29ce17361f1041d

Observation 5ae1e96c-8db6-4c78-af79-b139d745ed3b · outbound

This paper cites SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild.

Learning from Peers in Reasoning Models SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T22:13:07.972963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:13:07.972963Z digest=sha256:445cea72003f11cbf5b7b576d58830cac09c3e048713aa9bc951aa59cb3da22a

Observation 6061d5e9-d45b-4285-b794-51ba9b7f10bb · outbound

This paper cites START: Self-taught Reasoner with Tools.

Learning from Peers in Reasoning Models START: Self-taught Reasoner with Tools

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T22:13:07.977913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:13:07.977913Z digest=sha256:cd850ea84ea26fbd167ad61abfc0d72a46fe3286a03e5266ee03e514af098c8a

Observation d42b39d2-6305-404b-bafe-abd1e7e22b27 · outbound

This paper cites An Empirical Study on Eliciting and Improving R1-like Reasoning Models.

Learning from Peers in Reasoning Models An Empirical Study on Eliciting and Improving R1-like Reasoning Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T22:13:07.983704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:13:07.983704Z digest=sha256:5dd0b6ccf3f7d51c0748743d578c3b469ab44c37716520143e3ca74105cf6f09

Observation f7d1478e-7632-4161-ad6e-09f8f8287111 · outbound

This paper cites Mixture-of-Agents Enhances Large Language Model Capabilities.

Learning from Peers in Reasoning Models Mixture-of-Agents Enhances Large Language Model Capabilities

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T22:13:07.989152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:13:07.989152Z digest=sha256:f9ec3d427b6a821aff9bd074b8880bea24a2302d36cfc111265937fc6313e5d7

Observation 97c2489e-81db-49c0-9ee6-36d761fd84a5 · outbound

This paper cites Hello gpt-4o.

Learning from Peers in Reasoning Models Hello gpt-4o

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:13:09.008256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:13:07.993708Z digest=sha256:fdd6eff403b830242db1050a2e1d987d3fab73b86bf44e83f840e86adba464a4

Observation d68fd940-c6af-4903-be10-13b66fe522e7 · outbound

This paper cites Deepseek-r1 thoughtology: Let’s< think> about llm reasoning.

Learning from Peers in Reasoning Models Deepseek-r1 thoughtology: Let’s< think> about llm reasoning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T22:13:07.997993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:13:07.997993Z digest=sha256:68d65b44a0cdb87abd5a374ceb4752b9e8f9f04839f0e4a2e49d8d9359f77d7b

Observation 45dc9fe9-d3a1-4268-a9c6-d49d17559e2a · outbound

This paper cites Scaling of Search and Learning: A Roadmap to Reproduce o1 from Reinforcement Learning Perspective.

Learning from Peers in Reasoning Models Scaling of Search and Learning: A Roadmap to Reproduce o1 from Reinforcement Learning Perspective

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T22:13:08.003507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:13:08.003507Z digest=sha256:6a92354aac084a26a97df3945f45f9de143d8dde3a420989fcae0fb8d4d3cf77

Observation 9764a76c-7263-479a-8c0f-ab31e1acd5b5 · outbound

This paper cites Making language models better reasoners with step-aware verifier.

Learning from Peers in Reasoning Models Making language models better reasoners with step-aware verifier

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:13:08.993840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:13:08.008065Z digest=sha256:da3cb5cb8573a834f46149e84c008de0befeb846d72c0987e35817550b32837e

Observation d39af146-99d7-47b3-ad95-18d0435a955d · outbound

This paper cites Large Language Monkeys: Scaling Inference Compute with Repeated Sampling.

Learning from Peers in Reasoning Models Large Language Monkeys: Scaling Inference Compute with Repeated Sampling

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T22:13:08.012995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:13:08.012995Z digest=sha256:f4b4f6a03d2badaa4e0fae0d231c0e16476d142e92420345e808a49344ea8cda

Observation 9ab01dbb-9d85-4df6-8754-b45ab425cf99 · outbound

This paper cites The Good, The Bad, and The Greedy: Evaluation of LLMs Should Not Ignore Non-Determinism.

Learning from Peers in Reasoning Models The Good, The Bad, and The Greedy: Evaluation of LLMs Should Not Ignore Non-Determinism

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T22:13:08.017364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:13:08.017364Z digest=sha256:ca927ce76b1853ae14880da98a79c019009f907419b6921a05d79323ff4366e6

Observation 56dbcd6f-9a2e-401a-8862-95a69533dda7 · outbound

This paper cites When is the consistent prediction likely to be a correct prediction?.

Learning from Peers in Reasoning Models When is the consistent prediction likely to be a correct prediction?

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T22:13:08.023383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:13:08.023383Z digest=sha256:ac99c22a798167bfce18eb9aa79cb67f91a8a323dac2635f138ac57c3e3c0125

Observation d09c0608-debb-4b3a-a87c-b55311d75a6a · outbound

This paper cites Scaling laws for reward model overoptimization.

Learning from Peers in Reasoning Models Scaling laws for reward model overoptimization

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T22:13:08.028296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:13:08.028296Z digest=sha256:0150187e0194e40c387430210ebd2cf0e78da63c8ea149ea846335603650afea

Observation 98317615-8b66-4c69-8820-7c5f07c491cf · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Learning from Peers in Reasoning Models Training Verifiers to Solve Math Word Problems

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T22:13:08.033813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:13:08.033813Z digest=sha256:d333da93587abecd022de2fa4fd7b5d92f1005c7cb652aad76dbfca0c7d43dd1

Observation 98a9b889-2cc8-4112-8684-8984e0af0143 · outbound

This paper cites Fast Best-of-N Decoding via Speculative Rejection.

Learning from Peers in Reasoning Models Fast Best-of-N Decoding via Speculative Rejection

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T22:13:08.039933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:13:08.039933Z digest=sha256:e1f4624ef852bb728d2597f98cb35a7dc592397a7d5687e98771ca04ca9c87bf

Observation 8a651ec4-6066-4a89-aa06-d865f167084e · outbound

This paper cites BoNBoN Alignment for Large Language Models and the Sweetness of Best-of-n Sampling.

Learning from Peers in Reasoning Models BoNBoN Alignment for Large Language Models and the Sweetness of Best-of-n Sampling

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T22:13:08.046407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:13:08.046407Z digest=sha256:3ca80184e90b5854cebef632ad59855fb0059707b285f4454d415805836d9511

Observation dc3b0bdc-a4ad-4032-95cf-58a4c1bb206b · outbound

This paper cites Variational Best-of-N Alignment.

Learning from Peers in Reasoning Models Variational Best-of-N Alignment

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T22:13:08.052702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:13:08.052702Z digest=sha256:8646819dc1a73c846e155908483a552e53eaf49914b5cb225a14b5879ec8f5bc

Observation a494a453-106e-4256-927c-1307cbe87ec0 · outbound

This paper cites BOND: Aligning LLMs with Best-of-N Distillation.

Learning from Peers in Reasoning Models BOND: Aligning LLMs with Best-of-N Distillation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T22:13:08.058931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:13:08.058931Z digest=sha256:bba6cc764489b87d9b2e73d1675a10328d34c9404986efa10e8907680c177d52

Observation 176eed5c-228b-4cd1-899e-d44f7633ce10 · outbound

This paper cites Improv- ing factuality and reasoning in language models through multiagent debate.

Learning from Peers in Reasoning Models Improv- ing factuality and reasoning in language models through multiagent debate

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T22:13:08.065575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:13:08.065575Z digest=sha256:5d82f3c219235e2ffb226c5267d7ff2ebd91a95ac15de23bd0652e3322091f6d

Observation 9dd506f2-cd2e-4aaf-b279-c7f4132f68e4 · outbound

This paper cites ReConcile: Round-Table Conference Improves Reasoning via Consensus among Diverse LLMs.

Learning from Peers in Reasoning Models ReConcile: Round-Table Conference Improves Reasoning via Consensus among Diverse LLMs

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T22:13:08.070742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:13:08.070742Z digest=sha256:7108a998ff4679025f3634f18c7dca37f58ccd885c97c0a8b57ac10ce50a0ba5

Observation 924fb163-8900-4ee8-a037-f88f59c158df · outbound

This paper cites Corex: Pushing the Boundaries of Complex Reasoning through Multi-Model Collaboration.

Learning from Peers in Reasoning Models Corex: Pushing the Boundaries of Complex Reasoning through Multi-Model Collaboration

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T22:13:08.076659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:13:08.076659Z digest=sha256:66915dd2924ac19dfd52a70c2b920dfc4911cd0851c8a24d65987af7fc251e85

Observation 8a1ef6dc-b383-42ac-b0f4-7c6b6473adf2 · outbound

This paper cites CoMM: Collaborative Multi-Agent, Multi-Reasoning-Path Prompting for Complex Problem Solving.

Learning from Peers in Reasoning Models CoMM: Collaborative Multi-Agent, Multi-Reasoning-Path Prompting for Complex Problem Solving

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T22:13:08.084222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:13:08.084222Z digest=sha256:8753ec714ff3b3110bd227f628f9496c43fe1df4f44bf7642a5cf0a6fc16c239

Observation 1b6d9edd-0f8f-4f5f-819b-a990fabe5336 · outbound

This paper cites Malt: Improving reasoning with multi-agent llm training.

Learning from Peers in Reasoning Models Malt: Improving reasoning with multi-agent llm training

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T22:13:08.091269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:13:08.091269Z digest=sha256:d8b9189c6d8fb53773e3868a1f1cf4806533ff959efc3790bd6a51a44fc0b5c9

Observation d172b947-ab7d-426d-802c-3b912c557430 · outbound

This paper cites SMoA: Improving Multi-agent Large Language Models with Sparse Mixture-of-Agents.

Learning from Peers in Reasoning Models SMoA: Improving Multi-agent Large Language Models with Sparse Mixture-of-Agents

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T22:13:08.097520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:13:08.097520Z digest=sha256:15ba460b5d02b465aab09bd472e7587ea186a2d2bb1b92f351ec4912a6679e6b

Observation ea781724-dee2-47fd-ae7e-c2aedc302a9b · outbound

This paper cites The First Few Tokens Are All You Need: An Efficient and Effective Unsupervised Prefix Fine-Tuning Method for Reasoning Models.

Learning from Peers in Reasoning Models The First Few Tokens Are All You Need: An Efficient and Effective Unsupervised Prefix Fine-Tuning Method for Reasoning Models

Reference 43

Resolution
malformed identifier
no resolver link, observed 2026-08-15T22:13:08.104292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:13:08.104292Z digest=sha256:d07505a253af7bd8ae9729ae2b84b1add1d3a1a741ae3c66b2842e32f3447d58

Observation 4694d4ce-ee0d-40cf-9e13-7d1c26b971ef · outbound

This paper cites an unresolved cited work.

Learning from Peers in Reasoning Models Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:13:08.927962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:13:08.115458Z digest=sha256:ad34c8a7c7435d40f4c548f11a5718aa77b03d37b1a1f37925f8c19774ae6b19

Observation daa131d4-2068-4c73-abd5-d20f3c261cfa · outbound

This paper cites To summarize, my recent findings are: I attempted to expressp ands in terms ofq and r using the conditionpr +qs = 0, leading top =kq and s =−kr... Peer 4:.

Learning from Peers in Reasoning Models To summarize, my recent findings are: I attempted to expressp ands in terms ofq and r using the conditionpr +qs = 0, leading top =kq and s =−kr... Peer 4:

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:13:08.908867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:13:08.119827Z digest=sha256:76ff0a2702299a42b439ddf9a4c3b83c81a7b141e32749d9e981ab61186abb18

Observation c17113c2-b587-4e50-9fa8-84ea57b0297e · outbound

This paper cites Peer 3:.

Learning from Peers in Reasoning Models Peer 3:

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:13:08.881783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:13:08.125607Z digest=sha256:774d091fb532f6d9a10de27ecaa751356edd0eda0d1492adc47dec212491456a

Observation 3a498150-0106-41f2-ac55-37b4d3e79e79 · outbound

This paper cites Peer 4:.

Learning from Peers in Reasoning Models Peer 4:

Reference 1190

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:13:08.864664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:13:08.130979Z digest=sha256:07eedc05dead8bf05405914b3e9c7899a990f664ef703ad69542ffe601b638bc

Observation efd968be-ce2e-42cf-8aef-04a1e45fe9d6 · outbound

This paper cites Aha” moments Figure 18: We illustrate the average number of tokens and “Aha.

Learning from Peers in Reasoning Models Aha” moments Figure 18: We illustrate the average number of tokens and “Aha

Reference 2025

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:13:08.951365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:13:08.111240Z digest=sha256:cad0b94b4b00220f1978571a0578c87f3b7d3e5a783df0cf45f4fd3144ffe9d4

Pith citing papers

Observation e66e0fbc-e017-4a28-bfda-b3c83cc3bc04 · inbound

Adaptive Termination for Multi-round Parallel Reasoning: An Universal Semantic Entropy-Guided Framework cites this paper.

Adaptive Termination for Multi-round Parallel Reasoning: An Universal Semantic Entropy-Guided Framework Learning from Peers in Reasoning Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T18:59:41.908963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:59:41.908963Z digest=sha256:a80e0617a2291a9626d79a18859372624b5b4a413ee84e86b0f1c330ce37ad48

Observation 63245719-83c1-421f-965f-4e7e4baeb7c0 · inbound

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models cites this paper.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Learning from Peers in Reasoning Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:08.060154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:08.060154Z digest=sha256:9eeb6ab5faf8a07d22ccf86c580c8e6108c6ca550c819c7ed1f7f8055415cb2a

Observation 7824f657-254f-4a13-b8c3-96efb2d142a9 · inbound

Sticker-TTS: Learn to Utilize Historical Experience with a Sticker-driven Test-Time Scaling Framework cites this paper.

Sticker-TTS: Learn to Utilize Historical Experience with a Sticker-driven Test-Time Scaling Framework Learning from Peers in Reasoning Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T05:45:03.041044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T05:45:03.041044Z digest=sha256:bdb5807738d87e5658088477aa4e61d13221b35cf16bbcbd52a907bcea643802

Observation c27b8e1e-b6de-48bb-b66c-bb51be424721 · inbound

EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance cites this paper.

EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance Learning from Peers in Reasoning Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T14:44:14.634746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:44:14.634746Z digest=sha256:6b824379a6ba9b7f27e4e6ba68e6738ef7047ada673f85e8e5d44556f609e020

Observation 70809acb-c836-4335-8fd2-839cd2eee6eb · inbound

Cut Your Losses! Learning to Prune Paths Early for Efficient Parallel Reasoning cites this paper.

Cut Your Losses! Learning to Prune Paths Early for Efficient Parallel Reasoning Learning from Peers in Reasoning Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:53:03.652476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-10T08:52:46.918285Z digest=sha256:dc517a6dcf4b6673a1df27d245319113703380361b02a73344db77da959b843b