Pith. sign in

Paper Citation Record · LEDGER

Weak-Driven Learning: How Weak Agents make Strong Agents Stronger

As of 21 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 2 inbound Pith citation observations for arXiv:2602.08222.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2602.08222 v2

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T03:27:50.454980Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T14:28:01.622883Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-13T22:48:22.886654Z

Reference resolution

25 of 25 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved22
  • parse uncertain0
  • malformed identifier3
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 59fa4573-ea50-4015-ac82-d0c550a352f8 · outbound

This paper cites Pre” denotes before joint training, and “Stronger (Post).

Weak-Driven Learning: How Weak Agents make Strong Agents Stronger Pre” denotes before joint training, and “Stronger (Post)

Reference 1

Resolution
malformed identifier
no resolver link, observed 2026-08-03T03:27:50.401394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:27:50.401394Z digest=sha256:bacad17c7572991eff1637e3f89ca28989b0629b46c94f44015d846da75a1080

Observation 7d97815c-8e84-490e-901e-ad17479804e2 · outbound

This paper cites UNDIAL: Self-Distillation with Adjusted Logits for Robust Unlearning in Large Language Models.

Weak-Driven Learning: How Weak Agents make Strong Agents Stronger UNDIAL: Self-Distillation with Adjusted Logits for Robust Unlearning in Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T03:27:48.463612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:27:48.463612Z digest=sha256:c78361dd4e54ee4820706acc00a28cd68f8c4c73435fc6bc23e87e6a9f5c380d

Observation f6084749-5a4c-4117-8136-fda46866aa78 · outbound

This paper cites MiniLLM: On-Policy Distillation of Large Language Models.

Weak-Driven Learning: How Weak Agents make Strong Agents Stronger MiniLLM: On-Policy Distillation of Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T03:27:48.534063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:27:48.534063Z digest=sha256:17686c33ba9cd7b5aed974e57318e81d2e4cc0a494574f03d4cdb3c55ae32fb3

Observation 5b682dd4-53a1-47d8-b5fe-33ba85abf028 · outbound

This paper cites The False Promise of Imitating Proprietary LLMs.

Weak-Driven Learning: How Weak Agents make Strong Agents Stronger The False Promise of Imitating Proprietary LLMs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T03:27:48.597536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:27:48.597536Z digest=sha256:dee0a3c7d10086c91370eaacae49ecd031c76229ff401beee2e875d2b04665e1

Observation 3526dbb3-d001-47dc-9941-6a34fbb45d75 · outbound

This paper cites Reinforced Self-Training (ReST) for Language Modeling.

Weak-Driven Learning: How Weak Agents make Strong Agents Stronger Reinforced Self-Training (ReST) for Language Modeling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T03:27:48.664337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:27:48.664337Z digest=sha256:2773f3034d7f5809bbb34b3742a1fe4a2fc6d731dc51a417667b1b27e31a3e0a

Observation 724c113c-9907-46df-8fb4-9ce810d66487 · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Weak-Driven Learning: How Weak Agents make Strong Agents Stronger Distilling the Knowledge in a Neural Network

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T03:27:48.724872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:27:48.724872Z digest=sha256:c41d1fe83609c0549bf70cac9bd25a35eb549af50d41375217bf1677630357f7

Observation c0e0a2ae-74d4-4925-a499-82e9d139bc74 · outbound

This paper cites NEFTune: Noisy Embeddings Improve Instruction Finetuning.

Weak-Driven Learning: How Weak Agents make Strong Agents Stronger NEFTune: Noisy Embeddings Improve Instruction Finetuning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T03:27:49.024742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:27:49.024742Z digest=sha256:3690d5a9628b793f0c2644837adf048e38902b8cfa0df719e4bc6384fb31591c

Observation 67026650-d548-407c-b6e3-251e952190cf · outbound

This paper cites On the Self-Verification Limitations of Large Language Models on Reasoning and Planning Tasks.

Weak-Driven Learning: How Weak Agents make Strong Agents Stronger On the Self-Verification Limitations of Large Language Models on Reasoning and Planning Tasks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T03:27:49.384386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:27:49.384386Z digest=sha256:67b4df9ae055d9a412deeeba0c0f2e53c135b4261627b6e5f0181f235564abbd

Observation ce582407-695a-48b4-a401-ca4e7122c79e · outbound

This paper cites WizardLM: Empowering large pre-trained language models to follow complex instructions.

Weak-Driven Learning: How Weak Agents make Strong Agents Stronger WizardLM: Empowering large pre-trained language models to follow complex instructions

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T03:27:49.509414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:27:49.509414Z digest=sha256:6f2caa7fe834850f1a27e0057a1b6e3e7f253d8ba5de5c3accd2aa6eb6406882

Observation 86d6066e-9e7d-4323-a480-4914ae995cd8 · outbound

This paper cites Qwen3 Technical Report.

Weak-Driven Learning: How Weak Agents make Strong Agents Stronger Qwen3 Technical Report

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T03:27:49.668034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:27:49.668034Z digest=sha256:c05863eef6f414483a8700e4e4f2a2680d6c4b061fcc00290cad38e85aaca413

Observation 847763be-7e44-4c1e-9a08-4e48fa62f5fd · outbound

This paper cites Zhang, L., Song, J., Gao, A., Chen, J., Bao, C., and Ma, K.

Weak-Driven Learning: How Weak Agents make Strong Agents Stronger Zhang, L., Song, J., Gao, A., Chen, J., Bao, C., and Ma, K

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T03:27:49.855213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:27:49.855213Z digest=sha256:7b82ae4896593b3f641a9b13a50f021b7c04dc8c90701ec55fc6b31b63b47d23

Observation b8ec002e-4c1d-4fad-ac9f-2e3ae9a599dc · outbound

This paper cites 1.4 Million Open-Source Distilled Reasoning Dataset to Empower Large Language Model Training.

Weak-Driven Learning: How Weak Agents make Strong Agents Stronger 1.4 Million Open-Source Distilled Reasoning Dataset to Empower Large Language Model Training

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T03:27:50.171399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:27:50.171399Z digest=sha256:81ff985109a07ff33eb6873e4ab9e6a6dbf86d7f0ba29e302a9f9e9c52622003

Observation efaa566f-0fc5-4a8d-9bf9-4fe0d3b1d0e0 · outbound

This paper cites Transformer copilot: Learning from the mistake log in llm fine-tuning.arXiv preprint arXiv:2505.16270,.

Weak-Driven Learning: How Weak Agents make Strong Agents Stronger Transformer copilot: Learning from the mistake log in llm fine-tuning.arXiv preprint arXiv:2505.16270,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T03:27:50.268899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:27:50.268899Z digest=sha256:69584649fc9e690afc7eacbff56699e628d73ab6f1688686ef39b505415ffe3a

Observation 241cde0f-c00f-40d5-9061-76f59cdd8b0c · outbound

This paper cites an unresolved cited work.

Weak-Driven Learning: How Weak Agents make Strong Agents Stronger Unresolved cited work

Reference 22

Resolution
malformed identifier
no resolver link, observed 2026-08-03T03:27:50.311246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:27:50.311246Z digest=sha256:749cf63612a31814eeb98021efd940e5f00f648d3f0c5bdea61b68c69d2c2042

Observation 29d6226e-7335-4ca8-9378-126e17829fbc · outbound

This paper cites To answer the user’s question, you first think about the reasoning process and then provide the user with the answer.

Weak-Driven Learning: How Weak Agents make Strong Agents Stronger To answer the user’s question, you first think about the reasoning process and then provide the user with the answer

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T03:27:50.454980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:27:50.454980Z digest=sha256:a4ce7b199004942f899aec9f14c98b634d611fdaeb3762e985bbb9efb8873c32

Observation 8cb47913-d6bd-4bdc-8d4a-23c21298a438 · outbound

This paper cites Configuration C (α= 0.1, β= 0.9, γ= 0), which disables the regression-repair signal, achieves the highest accuracy on MATH 500 (70.2%).

Weak-Driven Learning: How Weak Agents make Strong Agents Stronger Configuration C (α= 0.1, β= 0.9, γ= 0), which disables the regression-repair signal, achieves the highest accuracy on MATH 500 (70.2%)

Reference 500

Resolution
malformed identifier
no resolver link, observed 2026-08-03T03:27:50.342387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:27:50.342387Z digest=sha256:d1db1ee8965f45fa34d6c41f0e87b4d77fc9e9cde86eb0150bf3abb968d62028

Observation f44f265f-9029-492d-95e7-20521ee0881d · outbound

This paper cites Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision.

Weak-Driven Learning: How Weak Agents make Strong Agents Stronger Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision

Reference 2009

Resolution
unresolved
no resolver link, observed 2026-08-03T03:27:48.234476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:27:48.234476Z digest=sha256:e2c66bb0b2081bb75ea655cda9c45ad5393a76d7bf1aa67a6e28e8ccd6caedb9

Observation 6ecea6dc-8ded-41bd-ba4f-6cf9a04ca6f2 · outbound

This paper cites Program Induction by Rationale Generation : Learning to Solve and Explain Algebraic Word Problems.

Weak-Driven Learning: How Weak Agents make Strong Agents Stronger Program Induction by Rationale Generation : Learning to Solve and Explain Algebraic Word Problems

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-03T03:27:49.162370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:27:49.162370Z digest=sha256:e259e3e68da0361b5e7cd794050d5594515e0c1f933891f4854e0aeb2bd340a3

Observation b091852e-94c3-471b-8806-70e2c8284acf · outbound

This paper cites Be Your Own Teacher: Improve the Performance of Convolutional Neural Networks via Self Distillation.

Weak-Driven Learning: How Weak Agents make Strong Agents Stronger Be Your Own Teacher: Improve the Performance of Convolutional Neural Networks via Self Distillation

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-03T03:27:50.058066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:27:50.058066Z digest=sha256:7858e67e2f0226f645080d041bfdfb0bb305461ce30b1201c738e714323ca8d4

Observation 83d4c6e4-18aa-4afd-a077-c26593e135f9 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Weak-Driven Learning: How Weak Agents make Strong Agents Stronger Evaluating Large Language Models Trained on Code

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-03T03:27:48.347683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:27:48.347683Z digest=sha256:e822814fcd4d124888d4340683986a2dfb0be222b4e10a8b32ac38ffda74f2ac

Observation 9a2721a2-edda-4515-ba47-6c8667c336fd · outbound

This paper cites Are NLP Models really able to Solve Simple Math Word Problems?.

Weak-Driven Learning: How Weak Agents make Strong Agents Stronger Are NLP Models really able to Solve Simple Math Word Problems?

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-03T03:27:49.270292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:27:49.270292Z digest=sha256:85fdd3a477db864556ac5be403b52c8d18780b7bab74f5761078ea5bb89894b7

Observation 54607f76-e570-42cd-a449-77b3b492c499 · outbound

This paper cites Program Synthesis with Large Language Models.

Weak-Driven Learning: How Weak Agents make Strong Agents Stronger Program Synthesis with Large Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T03:27:48.189325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:27:48.189325Z digest=sha256:a2e458f239cf5fd65c8040d656ebf16f398877b10c2fddaf927c57ef84d4d067

Observation e52eeadc-95b1-4472-a424-8987e24dc7cf · outbound

This paper cites Large Language Models Cannot Self-Correct Reasoning Yet.

Weak-Driven Learning: How Weak Agents make Strong Agents Stronger Large Language Models Cannot Self-Correct Reasoning Yet

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T03:27:48.790097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:27:48.790097Z digest=sha256:77969ab134b2f1ac40b3acc63f13d4f3dd2bd472e6a855c133fa568514cbb77a

Observation 1582f3b3-c8ad-4a09-a7d4-46aa0a24c73f · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Weak-Driven Learning: How Weak Agents make Strong Agents Stronger Training Verifiers to Solve Math Word Problems

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T03:27:48.418725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:27:48.418725Z digest=sha256:1bed8564ec86d37f71e84164a582b2c4232492c1586ea510e006db7d5e41246a

Observation 243bab27-3301-439b-92cf-da892926afa5 · outbound

This paper cites Real-Time Aligned Reward Model beyond Semantics.

Weak-Driven Learning: How Weak Agents make Strong Agents Stronger Real-Time Aligned Reward Model beyond Semantics

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-03T03:27:48.890632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:27:48.890632Z digest=sha256:1fdf27bdc0a1c6de9d7bfd0785d030e78acbdbe1639880e7306cff9da677f2ee

Pith citing papers

Observation 5075ea1c-cdf2-4e45-8eb5-1cf11b959a1d · inbound

Policy Improvement Reinforcement Learning cites this paper.

Policy Improvement Reinforcement Learning Weak-Driven Learning: How Weak Agents make Strong Agents Stronger

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-06-09T03:07:11.059756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-13T22:47:05.132020Z digest=sha256:ee9be8913a37c4af91bb196dc567343edcbe479f58b9024d4ab11b33facadcc3

Observation 88863e25-b0d0-4b71-ad14-d5c05ee34262 · inbound

WDL-OPD: Weak-Driven On-Policy Distillation via Mixture-Constrained Co-Training cites this paper.

WDL-OPD: Weak-Driven On-Policy Distillation via Mixture-Constrained Co-Training Weak-Driven Learning: How Weak Agents make Strong Agents Stronger

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T14:28:01.622883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:28:01.622883Z digest=sha256:dbe2d7b712e2607852545bdf4706a64fb22debff9346e7b4329f81edbd1a3604