Pith. sign in

Paper Citation Record · LEDGER

Curriculum Learning for Efficient Chain-of-Thought Distillation via Structure-Aware Masking and GRPO

As of 23 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 0 inbound Pith citation observations for arXiv:2602.17686.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2602.17686 v4

Coverage vector

measured 14 of 14 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T04:22:10.826939Z

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

14 of 14 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0d734869-655f-483c-ba19-07bf6b1aae8b · outbound

This paper cites Mcc-kd: Multi-cot consis- tent knowledge distillation.

Curriculum Learning for Efficient Chain-of-Thought Distillation via Structure-Aware Masking and GRPO Mcc-kd: Multi-cot consis- tent knowledge distillation

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:09.466073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:09.466073Z digest=sha256:6141b0bc825bc338aa89bc1a5a260606fb17d14632d936d4f9161633304c0751

Observation 11f4e598-4146-4909-aaa0-23a0b8ddf05f · outbound

This paper cites MiCoTA: Bridging the Learnability Gap with Intermediate CoT and Teacher Assistants.

Curriculum Learning for Efficient Chain-of-Thought Distillation via Structure-Aware Masking and GRPO MiCoTA: Bridging the Learnability Gap with Intermediate CoT and Teacher Assistants

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:09.812983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:09.812983Z digest=sha256:9699aae7b644495a3292f7126a4754804ebeea237a3f25712912de0bd0d6d3a7

Observation 8466175b-991c-4576-9db0-f5ce4d5952b6 · outbound

This paper cites Distilling step-by-step! outperforming larger language mod- els with less training data and smaller model sizes.

Curriculum Learning for Efficient Chain-of-Thought Distillation via Structure-Aware Masking and GRPO Distilling step-by-step! outperforming larger language mod- els with less training data and smaller model sizes

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:10.071077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:10.071077Z digest=sha256:246bfadbfd24e1262c81796a7cbf14a35fa85451fe6d61d2b853c7c58f0cbc0d

Observation d2faa3d1-e98a-4da1-86b6-48af6ce1137f · outbound

This paper cites Implicit Reasoning in Large Language Models: A Comprehensive Survey.

Curriculum Learning for Efficient Chain-of-Thought Distillation via Structure-Aware Masking and GRPO Implicit Reasoning in Large Language Models: A Comprehensive Survey

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:10.213699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:10.213699Z digest=sha256:79a742c138251f7f22265e35782b7f0538c73b656c92b5d900068c03ccba6168

Observation a2667e53-be8e-46f9-981c-14079b6fe28c · outbound

This paper cites SuperRL: Reinforcement Learning with Supervision to Boost Language Model Reasoning.

Curriculum Learning for Efficient Chain-of-Thought Distillation via Structure-Aware Masking and GRPO SuperRL: Reinforcement Learning with Supervision to Boost Language Model Reasoning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:10.295612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:10.295612Z digest=sha256:3135d20ea25f5ea103a09b56cc289a35f9700793b90284ee5251b3c6146770fa

Observation d4b873ae-18e6-42a4-8222-0d1f42210444 · outbound

This paper cites Qwen2.5 Technical Report.

Curriculum Learning for Efficient Chain-of-Thought Distillation via Structure-Aware Masking and GRPO Qwen2.5 Technical Report

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:10.495797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:10.495797Z digest=sha256:2132d78246da35065ad6d7dcd0053de9e8aef88487a389c3bf9e4d767b88e70a

Observation 9cd351bf-69f6-4c5a-86d7-067d9115c77d · outbound

This paper cites Masked-and-reordered self-supervision for reinforcement learn- ing from verifiable rewards.arXiv preprint arXiv:2511.17473,.

Curriculum Learning for Efficient Chain-of-Thought Distillation via Structure-Aware Masking and GRPO Masked-and-reordered self-supervision for reinforcement learn- ing from verifiable rewards.arXiv preprint arXiv:2511.17473,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:10.744393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:10.744393Z digest=sha256:336dee758612bc76119fd9653f430dba5f4320d2ac7f07f3e1578d36c65daa76

Observation b77da9d4-cbbc-40b2-831f-1360c70b88e5 · outbound

This paper cites Tokenskip: Controllable chain-of- thought compression in llms.

Curriculum Learning for Efficient Chain-of-Thought Distillation via Structure-Aware Masking and GRPO Tokenskip: Controllable chain-of- thought compression in llms

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:10.826939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:10.826939Z digest=sha256:db1f8cc1b95176091f651874e6d43d5e49ad562892c1d9971344917f833f26c8

Observation 7d74edbc-3bb0-44fd-93cd-4bd10ff76d40 · outbound

This paper cites Mixed Distillation Helps Smaller Language Model Better Reasoning.

Curriculum Learning for Efficient Chain-of-Thought Distillation via Structure-Aware Masking and GRPO Mixed Distillation Helps Smaller Language Model Better Reasoning

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:10.152312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:10.152312Z digest=sha256:f1f4c87f097e37eafbc5f7ed279c793c74d71668520dd52056a1abb2a4ab0376

Observation 95ee30c0-11ab-4f82-a1d1-3615c9943127 · outbound

This paper cites Implicit Chain of Thought Reasoning via Knowledge Distillation.

Curriculum Learning for Efficient Chain-of-Thought Distillation via Structure-Aware Masking and GRPO Implicit Chain of Thought Reasoning via Knowledge Distillation

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:09.720279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:09.720279Z digest=sha256:c7a8b6b2837c132d54b49ee529948cf6bdf6c70f43c44db9dfa3f49b9b675f60

Observation ebd6edff-21bb-44aa-ac5c-f3407636046c · outbound

This paper cites an unresolved cited work.

Curriculum Learning for Efficient Chain-of-Thought Distillation via Structure-Aware Masking and GRPO Unresolved cited work

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:10.378043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:10.378043Z digest=sha256:d5ef8fa28c7d59441150273bd9715d8f632e35734b17936894fd02f33856ccf4

Observation 479937a8-3623-4c43-9bf5-592299ad9f5a · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Curriculum Learning for Efficient Chain-of-Thought Distillation via Structure-Aware Masking and GRPO Training Verifiers to Solve Math Word Problems

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:09.581991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:09.581991Z digest=sha256:dad54506e1a38d40b0a0a699b7411c807a707f5a6febc33707c1d4a995162ca2

Observation 8333095e-d9a6-48d0-be76-cc7115c8e034 · outbound

This paper cites Codi: Compressing chain-of- thought into continuous space via self-distillation.

Curriculum Learning for Efficient Chain-of-Thought Distillation via Structure-Aware Masking and GRPO Codi: Compressing chain-of- thought into continuous space via self-distillation

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:10.636869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:10.636869Z digest=sha256:89150c44e17c0bcc08c47f56470aa359b8090759b6bddee51fe9d6853566b9b1

Observation e8be5630-6f60-48a6-8ba5-83962c57fced · outbound

This paper cites The Llama 3 Herd of Models.

Curriculum Learning for Efficient Chain-of-Thought Distillation via Structure-Aware Masking and GRPO The Llama 3 Herd of Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:09.986203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:09.986203Z digest=sha256:dadaac2bdf097e7d02a42ff90ca035920549efea941a52f80a35d90da99645ad

Pith citing papers

No inbound Pith citation observations are available.