Pith. sign in

Paper Citation Record · LEDGER

Weak-Driven Learning: How Weak Agents make Strong Agents Stronger

As of 12 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 1 inbound Pith citation observation for arXiv:2602.08222.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2602.08222 v2

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T03:27:50.454980Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-13T22:47:05.132020Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-13T22:48:22.886654Z

Reference resolution

25 of 25 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved22
  • parse uncertain0
  • malformed identifier3
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 59fa4573-ea50-4015-ac82-d0c550a352f8 · outbound

This paper cites Pre” denotes before joint training, and “Stronger (Post).

Weak-Driven Learning: How Weak Agents make Strong Agents Stronger Pre” denotes before joint training, and “Stronger (Post)

Reference 1

Resolution
malformed identifier
no resolver link, observed 2026-08-03T03:27:50.401394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:27:50.401394Z digest=sha256:bacad17c7572991eff1637e3f89ca28989b0629b46c94f44015d846da75a1080

Observation 7d97815c-8e84-490e-901e-ad17479804e2 · outbound

This paper cites UNDIAL: Self-Distillation with Adjusted Logits for Robust Unlearning in Large Language Models.

Weak-Driven Learning: How Weak Agents make Strong Agents Stronger UNDIAL: Self-Distillation with Adjusted Logits for Robust Unlearning in Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T03:27:48.463612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:27:48.463612Z digest=sha256:e7fb64b4023c4a23f504aae562bf0bd0391cb4f8b2cb5e62a54d4db363f9d3df

Observation f6084749-5a4c-4117-8136-fda46866aa78 · outbound

This paper cites MiniLLM: On-Policy Distillation of Large Language Models.

Weak-Driven Learning: How Weak Agents make Strong Agents Stronger MiniLLM: On-Policy Distillation of Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T03:27:48.534063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:27:48.534063Z digest=sha256:184c501008f90ca53e4609329941f89bd931249537210c4f28be1e49e92d2314

Observation 5b682dd4-53a1-47d8-b5fe-33ba85abf028 · outbound

This paper cites The False Promise of Imitating Proprietary LLMs.

Weak-Driven Learning: How Weak Agents make Strong Agents Stronger The False Promise of Imitating Proprietary LLMs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T03:27:48.597536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:27:48.597536Z digest=sha256:ee973165c83508f6624564b8601186b074e575a841c64eac8a4fb8b364f059fb

Observation 3526dbb3-d001-47dc-9941-6a34fbb45d75 · outbound

This paper cites Reinforced Self-Training (ReST) for Language Modeling.

Weak-Driven Learning: How Weak Agents make Strong Agents Stronger Reinforced Self-Training (ReST) for Language Modeling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T03:27:48.664337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:27:48.664337Z digest=sha256:7ca3bae886b0faf83c0d1a0dda7fea7ee103d9c2020c975948a96cae5f367427

Observation 724c113c-9907-46df-8fb4-9ce810d66487 · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Weak-Driven Learning: How Weak Agents make Strong Agents Stronger Distilling the Knowledge in a Neural Network

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T03:27:48.724872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:27:48.724872Z digest=sha256:a58667959e28a0b821f6def5b9585fb1731aef429fc06eda2a562a075e45ac84

Observation c0e0a2ae-74d4-4925-a499-82e9d139bc74 · outbound

This paper cites NEFTune: Noisy Embeddings Improve Instruction Finetuning.

Weak-Driven Learning: How Weak Agents make Strong Agents Stronger NEFTune: Noisy Embeddings Improve Instruction Finetuning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T03:27:49.024742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:27:49.024742Z digest=sha256:3cdb35715c4a438fa3eff62c607cc2b70752a35b6b4f5d9498010f30fdf94b07

Observation 67026650-d548-407c-b6e3-251e952190cf · outbound

This paper cites On the Self-Verification Limitations of Large Language Models on Reasoning and Planning Tasks.

Weak-Driven Learning: How Weak Agents make Strong Agents Stronger On the Self-Verification Limitations of Large Language Models on Reasoning and Planning Tasks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T03:27:49.384386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:27:49.384386Z digest=sha256:03223db34e0e888ab68e4a00ab3b3165096f3624aa66ed9eed7c7d1216894993

Observation ce582407-695a-48b4-a401-ca4e7122c79e · outbound

This paper cites WizardLM: Empowering large pre-trained language models to follow complex instructions.

Weak-Driven Learning: How Weak Agents make Strong Agents Stronger WizardLM: Empowering large pre-trained language models to follow complex instructions

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T03:27:49.509414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:27:49.509414Z digest=sha256:9caf66c1f43886de412c351a5514fd5fb12d41ccd2ba71c8f07a623c3e5545bb

Observation 86d6066e-9e7d-4323-a480-4914ae995cd8 · outbound

This paper cites Qwen3 Technical Report.

Weak-Driven Learning: How Weak Agents make Strong Agents Stronger Qwen3 Technical Report

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T03:27:49.668034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:27:49.668034Z digest=sha256:c05863eef6f414483a8700e4e4f2a2680d6c4b061fcc00290cad38e85aaca413

Observation 847763be-7e44-4c1e-9a08-4e48fa62f5fd · outbound

This paper cites Zhang, L., Song, J., Gao, A., Chen, J., Bao, C., and Ma, K.

Weak-Driven Learning: How Weak Agents make Strong Agents Stronger Zhang, L., Song, J., Gao, A., Chen, J., Bao, C., and Ma, K

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T03:27:49.855213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:27:49.855213Z digest=sha256:7b82ae4896593b3f641a9b13a50f021b7c04dc8c90701ec55fc6b31b63b47d23

Observation b8ec002e-4c1d-4fad-ac9f-2e3ae9a599dc · outbound

This paper cites 1.4 Million Open-Source Distilled Reasoning Dataset to Empower Large Language Model Training.

Weak-Driven Learning: How Weak Agents make Strong Agents Stronger 1.4 Million Open-Source Distilled Reasoning Dataset to Empower Large Language Model Training

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T03:27:50.171399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:27:50.171399Z digest=sha256:a35e12609570ec0c988c97c6d836aa4b3b121473a39d24da63408a401285755d

Observation efaa566f-0fc5-4a8d-9bf9-4fe0d3b1d0e0 · outbound

This paper cites Transformer copilot: Learning from the mistake log in llm fine-tuning.arXiv preprint arXiv:2505.16270,.

Weak-Driven Learning: How Weak Agents make Strong Agents Stronger Transformer copilot: Learning from the mistake log in llm fine-tuning.arXiv preprint arXiv:2505.16270,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T03:27:50.268899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:27:50.268899Z digest=sha256:69584649fc9e690afc7eacbff56699e628d73ab6f1688686ef39b505415ffe3a

Observation 241cde0f-c00f-40d5-9061-76f59cdd8b0c · outbound

This paper cites an unresolved cited work.

Weak-Driven Learning: How Weak Agents make Strong Agents Stronger Unresolved cited work

Reference 22

Resolution
malformed identifier
no resolver link, observed 2026-08-03T03:27:50.311246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:27:50.311246Z digest=sha256:749cf63612a31814eeb98021efd940e5f00f648d3f0c5bdea61b68c69d2c2042

Observation 29d6226e-7335-4ca8-9378-126e17829fbc · outbound

This paper cites To answer the user’s question, you first think about the reasoning process and then provide the user with the answer.

Weak-Driven Learning: How Weak Agents make Strong Agents Stronger To answer the user’s question, you first think about the reasoning process and then provide the user with the answer

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T03:27:50.454980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:27:50.454980Z digest=sha256:a4ce7b199004942f899aec9f14c98b634d611fdaeb3762e985bbb9efb8873c32

Observation 8cb47913-d6bd-4bdc-8d4a-23c21298a438 · outbound

This paper cites Configuration C (α= 0.1, β= 0.9, γ= 0), which disables the regression-repair signal, achieves the highest accuracy on MATH 500 (70.2%).

Weak-Driven Learning: How Weak Agents make Strong Agents Stronger Configuration C (α= 0.1, β= 0.9, γ= 0), which disables the regression-repair signal, achieves the highest accuracy on MATH 500 (70.2%)

Reference 500

Resolution
malformed identifier
no resolver link, observed 2026-08-03T03:27:50.342387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:27:50.342387Z digest=sha256:d1db1ee8965f45fa34d6c41f0e87b4d77fc9e9cde86eb0150bf3abb968d62028

Observation f44f265f-9029-492d-95e7-20521ee0881d · outbound

This paper cites Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision.

Weak-Driven Learning: How Weak Agents make Strong Agents Stronger Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision

Reference 2009

Resolution
unresolved
no resolver link, observed 2026-08-03T03:27:48.234476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:27:48.234476Z digest=sha256:ae71f0fafeac15189f3d998e467abdf67d372e76c90087e0f2e6feff5250acaa

Observation 6ecea6dc-8ded-41bd-ba4f-6cf9a04ca6f2 · outbound

This paper cites Program Induction by Rationale Generation : Learning to Solve and Explain Algebraic Word Problems.

Weak-Driven Learning: How Weak Agents make Strong Agents Stronger Program Induction by Rationale Generation : Learning to Solve and Explain Algebraic Word Problems

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-03T03:27:49.162370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:27:49.162370Z digest=sha256:27d5fa84366712e46cac1445a7cb9d6d258dff44d6cbb6d6cf5a07c5ecd0d36b

Observation b091852e-94c3-471b-8806-70e2c8284acf · outbound

This paper cites Be Your Own Teacher: Improve the Performance of Convolutional Neural Networks via Self Distillation.

Weak-Driven Learning: How Weak Agents make Strong Agents Stronger Be Your Own Teacher: Improve the Performance of Convolutional Neural Networks via Self Distillation

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-03T03:27:50.058066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:27:50.058066Z digest=sha256:e9ac41a509dc7f5abe2ebd16112889ce47fe9e2089d86a83d55b12172d684c07

Observation 83d4c6e4-18aa-4afd-a077-c26593e135f9 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Weak-Driven Learning: How Weak Agents make Strong Agents Stronger Evaluating Large Language Models Trained on Code

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-03T03:27:48.347683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:27:48.347683Z digest=sha256:ccf0986748860dd1cb446b2d9ebf27bc8d6dfaa57e42247f63651e6dd5d0dc17

Observation 9a2721a2-edda-4515-ba47-6c8667c336fd · outbound

This paper cites Are NLP Models really able to Solve Simple Math Word Problems?.

Weak-Driven Learning: How Weak Agents make Strong Agents Stronger Are NLP Models really able to Solve Simple Math Word Problems?

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-03T03:27:49.270292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:27:49.270292Z digest=sha256:85fdd3a477db864556ac5be403b52c8d18780b7bab74f5761078ea5bb89894b7

Observation 54607f76-e570-42cd-a449-77b3b492c499 · outbound

This paper cites Program Synthesis with Large Language Models.

Weak-Driven Learning: How Weak Agents make Strong Agents Stronger Program Synthesis with Large Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T03:27:48.189325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:27:48.189325Z digest=sha256:e3fda122010a9c1ec09ed87b2908f8ad193ebbca65bc7815eec2855ca5dd4a2f

Observation e52eeadc-95b1-4472-a424-8987e24dc7cf · outbound

This paper cites Large Language Models Cannot Self-Correct Reasoning Yet.

Weak-Driven Learning: How Weak Agents make Strong Agents Stronger Large Language Models Cannot Self-Correct Reasoning Yet

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T03:27:48.790097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:27:48.790097Z digest=sha256:641341d860d1b75184cf1b40abcb6eefd7d99fe2d74dce647675025916ed69b2

Observation 1582f3b3-c8ad-4a09-a7d4-46aa0a24c73f · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Weak-Driven Learning: How Weak Agents make Strong Agents Stronger Training Verifiers to Solve Math Word Problems

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T03:27:48.418725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:27:48.418725Z digest=sha256:060ab80fed4a5356dfe61934cede86642321ac8975276ed0733fc7a6d07d4d77

Observation 243bab27-3301-439b-92cf-da892926afa5 · outbound

This paper cites Real-Time Aligned Reward Model beyond Semantics.

Weak-Driven Learning: How Weak Agents make Strong Agents Stronger Real-Time Aligned Reward Model beyond Semantics

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-03T03:27:48.890632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:27:48.890632Z digest=sha256:1fdf27bdc0a1c6de9d7bfd0785d030e78acbdbe1639880e7306cff9da677f2ee

Pith citing papers

Observation 5075ea1c-cdf2-4e45-8eb5-1cf11b959a1d · inbound

Policy Improvement Reinforcement Learning cites this paper.

Policy Improvement Reinforcement Learning Weak-Driven Learning: How Weak Agents make Strong Agents Stronger

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-06-09T03:07:11.059756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-13T22:47:05.132020Z digest=sha256:395aa614067efe2668bc65c68a75b44337966e5fa0fb3c40a4543b76fc8cdbfd