Pith. sign in

Paper Citation Record · LEDGER

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation

As of 7 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 0 inbound Pith citation observations for arXiv:2507.00054.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.00054 v1

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:45:04.656219Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

61 of 61 outbound references displayed

  • verified exact4
  • verified fuzzy32
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f9f7112a-3f5c-4430-89ce-ec48996e0545 · outbound

This paper cites , Aneja, J.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Aneja, J

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:11.959401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:44:59.832083Z digest=sha256:3ca506469bd0980483f79d289dc591fa0e0ac3333fb22ecffec6799828fd6de9

Observation 5334a726-edd3-4b45-a7fd-d1a4b28a4136 · outbound

This paper cites , Vieillard, N.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Vieillard, N

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:11.756821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:44:59.887421Z digest=sha256:15e82326b043ec537d7ac4c098fa58441dd60d7631438affab15a3f315ed17d9

Observation 7976eb97-c850-406e-8b1a-15efc9efb035 · outbound

This paper cites , Vieillard, N.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Vieillard, N

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:11.433341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:44:59.969850Z digest=sha256:fd3209964caa910c4a5c2ea7da9d1b9ef43378da695d6391a25466fb672cdc7f

Observation 963a99e5-8b87-4347-bc6c-33ec5f7679fa · outbound

This paper cites Towards Understanding Distilled Reasoning Models: A Representational Approach.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation Towards Understanding Distilled Reasoning Models: A Representational Approach

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:00.068336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:00.068336Z digest=sha256:1a2b94a5938edb5d1965f505a167c7f40f17f712dbac110fbe2a8f38eca46812

Observation 9813e50b-31b0-44d5-a629-a293fc9c17f5 · outbound

This paper cites , Sastre, I.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Sastre, I

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:11.259853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:45:00.163772Z digest=sha256:58f66dd80d076743b2dbae082f748a09423f2425ceeba879c7916b54436793b0

Observation ee832064-5677-498b-93ca-96c24d0a7770 · outbound

This paper cites Smaller, Weaker, Yet Better: Training LLM Reasoners via Compute-Optimal Sampling.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation Smaller, Weaker, Yet Better: Training LLM Reasoners via Compute-Optimal Sampling

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:00.253407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:00.253407Z digest=sha256:10934d0949996e8611b9593414ae251ce89898634f5285afe04d2e0471e22150

Observation 7ca5a211-4a5a-4bd4-8f10-23e74975e1db · outbound

This paper cites , Peebles, B.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Peebles, B

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:11.016657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:45:00.323553Z digest=sha256:acb64aa9fff5b1a6982c8cf55fb89a19a8957a77b4a95a30ff523353c193aa23

Observation 8585932b-95fd-4e64-8168-4765ac6904cd · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation Training Verifiers to Solve Math Word Problems

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:00.383119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:00.383119Z digest=sha256:0f963ad8822e26d1cbc96422e1e73c610b084615043ce841d67e82ccbaeb90ee

Observation ae03c8ce-0b48-4fad-9d78-2c9f0baa78ed · outbound

This paper cites \ Ngo, C.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation \ Ngo, C

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:10.872019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:45:00.432702Z digest=sha256:492f4df04c4977464a0b86872a34bfd669ad994dff55e9df6820f683ecd3ba63

Observation b0065cb9-5d07-4763-8d75-6ddc406cc9e4 · outbound

This paper cites , See, A.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , See, A

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:10.714160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:45:00.499454Z digest=sha256:290fa6f3e9d90dc24734465f64ae2f16193df602ba11c5d0156276db64f7c2e8

Observation fda81fd6-744d-48b2-a6d0-d940a7757342 · outbound

This paper cites , Feng, B.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Feng, B

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:10.515033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:45:00.562770Z digest=sha256:9ee88148f16e9b01820fa05d1211f549d30d5898685edbe05b0f008efae3ad91

Observation 88c34840-15c4-4128-afbb-d71d4410be5d · outbound

This paper cites an unresolved cited work.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:45:10.354570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:45:00.644914Z digest=sha256:28a2cb6c727bcaec38f42490496647480a78cd44c5469ef22a26adf1a770ea73

Observation 8a236ce7-78d9-45d5-9e9a-3c7e6eb7b7cd · outbound

This paper cites an unresolved cited work.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:45:10.118810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:45:00.717936Z digest=sha256:bf57b0058340f466f5c293f24a8abaf46b7f183322af4b7528679ac8b96d203a

Observation 8af77e78-1dab-42c3-bf88-2b07b2c7012b · outbound

This paper cites Advantage-Guided Distillation for Preference Alignment in Small Language Models.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation Advantage-Guided Distillation for Preference Alignment in Small Language Models

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:45:05.664522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:45:00.774715Z digest=sha256:2e47403858dfddb9c5cb59ba43a926db0dd1ff7b49582ac6f492bc3898672f2a

Observation ba453a7e-76aa-4268-970c-cc687a2a4b18 · outbound

This paper cites MiniLLM: On-Policy Distillation of Large Language Models.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation MiniLLM: On-Policy Distillation of Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:00.840611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:00.840611Z digest=sha256:67f282d663404dbfa92c065978e3176eaacf93f9943b09594532fb809795402e

Observation ec3ea647-c2f7-4d1b-bdb7-e7e74b0b621e · outbound

This paper cites , Zhang, L L.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Zhang, L L

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:09.899083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:45:00.916470Z digest=sha256:9d14dd04d64d674d0578b51e410cb435c85d0139c153c9c712487ba2912474e5

Observation 3a330c4d-5b4f-4e1d-91db-ba0f3d6215db · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:00.969331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:00.969331Z digest=sha256:70f194743530221567c97104c897a53b7fc6b6c1833f58dd3d26129e22934fbb

Observation 2cc95be0-a373-4d5d-89c2-100f15211594 · outbound

This paper cites , Vinyals, O.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Vinyals, O

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:09.681927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:45:01.034051Z digest=sha256:86886b99011b1334cf098264b5f1fbeda9f6dcb97f1944570f9a656043a32ceb

Observation 2988ec95-7570-4805-86fe-20b740f166c6 · outbound

This paper cites , Borgeaud, S.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Borgeaud, S

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:09.484507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:45:01.104215Z digest=sha256:854bc70ee5a51c5f3a6db8e42c7f1debc06c62e3b06ce619516d287891ec7949

Observation 38769f70-6b26-4119-a44c-9bc87c2dcc74 · outbound

This paper cites , Li, C L.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Li, C L

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:09.258428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:45:01.208161Z digest=sha256:d0ad5525980dc8313d73ad321e3f0a6df0b9a23af1d1cacf6c125de8cb9ff7c8

Observation b2223c7a-1689-484d-a335-981e059a7d3c · outbound

This paper cites , Yin, Y.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Yin, Y

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:09.049575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:45:01.282141Z digest=sha256:d9d804bed2682127db944b48cc5478f26223c94d65ad0bd5ea91a8807c36dfb6

Observation c63aaa27-aeda-4fdf-b195-07a0ea7ca4b8 · outbound

This paper cites MindStar: Enhancing Math Reasoning in Pre-trained LLMs at Inference Time.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation MindStar: Enhancing Math Reasoning in Pre-trained LLMs at Inference Time

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:01.362923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:01.362923Z digest=sha256:b1d6887c35bbff9d35ebe76b7090e03f63155ffada84355b561feb9c77580ebe

Observation ba9ae507-d70d-49b0-9c49-f1052c19a884 · outbound

This paper cites \ Rush, A M.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation \ Rush, A M

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:08.841836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:45:01.446884Z digest=sha256:e5e95cf9d4447fb08300aef03d41885fc2be4dfb2f30a53294db04af8fbcdf2b

Observation d76b30f7-97b5-49f2-8da9-86c67609c337 · outbound

This paper cites DistiLLM-2: A Contrastive Approach Boosts the Distillation of LLMs.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation DistiLLM-2: A Contrastive Approach Boosts the Distillation of LLMs

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:01.511764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:01.511764Z digest=sha256:a71bf4afe461e678d304dd68ebe1e1865333c41c32f8fca52035bd7b60790a52

Observation bf1372cb-2c2d-4815-b264-30df58d36e6d · outbound

This paper cites DistiLLM: Towards Streamlined Distillation for Large Language Models.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation DistiLLM: Towards Streamlined Distillation for Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:01.602702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:01.602702Z digest=sha256:4a0ecc3c561178599ff89d929c5998f88affcc22046821820c2bcee862278ea8

Observation bbfeb791-a91d-4c43-9cde-01ea480d017f · outbound

This paper cites , Fang, L.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Fang, L

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:08.621314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:45:01.662995Z digest=sha256:84c916befb7af23fd010c1bc823e907579641b3a5104529ac1ad16028e02efa1

Observation 1f40c54b-8e47-4293-9125-d763d5183b0d · outbound

This paper cites Winning Big with Small Models: Knowledge Distillation vs. Self-Training for Reducing Hallucination in Product QA Agents.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation Winning Big with Small Models: Knowledge Distillation vs. Self-Training for Reducing Hallucination in Product QA Agents

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:01.734181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:01.734181Z digest=sha256:f1f20016d8ed08d8350a027b923b99ed16f05029444122cd28083c57fa7dbd22

Observation 1f1635b6-be41-4bcb-aa4b-a089978da199 · outbound

This paper cites GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:01.803122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:01.803122Z digest=sha256:218c0aea0a4e47e8f95ee2e1ffce93fd5c278fb97225e13ff7288702e677185a

Observation a41021f4-9434-410f-86af-7fde9090a21d · outbound

This paper cites Evolving Knowledge Distillation with Large Language Models and Active Learning.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation Evolving Knowledge Distillation with Large Language Models and Active Learning

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:45:05.399923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:45:01.899416Z digest=sha256:bdf6529bcd3c48d5cf2868991405d5b2e848b359c881acbbb16a54774b095630

Observation c9427546-55e8-4e8e-81c2-4f74cca695e4 · outbound

This paper cites , Chen, C.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Chen, C

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:08.535787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:45:01.987456Z digest=sha256:eba917af02486ae964d43b1b6e54e33969278df89de25a8743d7e0d31c1de972

Observation 663344f9-7e30-47e0-89bc-9ac38933d316 · outbound

This paper cites , Yang, Z.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Yang, Z

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:08.365808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:45:02.050079Z digest=sha256:2131ae610a56145f4d467caf6e83bb35208ec3199c077c7a1bae3bcdc06ded4b

Observation 335802d2-6c7b-402c-8a4f-1c2f5861e9a2 · outbound

This paper cites , Dani, J.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Dani, J

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:08.178769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:45:02.176949Z digest=sha256:76225f09b7542ba9d193d9a9169a38082a2238c488daaa28b6ade6c72a2b8757

Observation f6df0b60-2933-495f-86bb-c081bc56fa71 · outbound

This paper cites an unresolved cited work.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:45:08.012500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:45:02.296854Z digest=sha256:e930d863d20ec17110dc115257deeadfba586e5442acce50620afa52fd01f4ae

Observation 71d9bb01-9f8b-4ccc-a41a-b8716ccb4285 · outbound

This paper cites , Lerer, A.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Lerer, A

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:07.780048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:45:02.381998Z digest=sha256:a6bb3d5b301a41f0128f1827ae149e5f2ccbb4f74ccdbba2eaeefea6c5690e23

Observation 710f5486-52fa-4b6d-8121-0e832a6133de · outbound

This paper cites Progressive distillation induces an implicit curriculum.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation Progressive distillation induces an implicit curriculum

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:02.463708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:02.463708Z digest=sha256:a82506a9e5fddc2420f8a687b606e3d90348e7e19a3ed3a51ef67b7c87c5dc13

Observation 0d86e769-ece0-4b0e-8736-d57c1f90975d · outbound

This paper cites , Kim, D.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Kim, D

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:07.585952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:45:02.552158Z digest=sha256:70d8c780f34170861412d5a5d8d1b30d9b24fe0728fcd52c68fb8f8df45a9225

Observation 17d3ac87-dd4d-41a7-bd6d-017ad7752bd1 · outbound

This paper cites \ Xie, S.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation \ Xie, S

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:07.347692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:45:02.621187Z digest=sha256:a63ebfe8413395b4eb0fb5ce4684230274db20633554d30e6297262914af8b03

Observation dedb6313-b2d1-43cd-89cb-46c0ddf3e22b · outbound

This paper cites CourseGPT-zh: an Educational Large Language Model Based on Knowledge Distillation Incorporating Prompt Optimization.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation CourseGPT-zh: an Educational Large Language Model Based on Knowledge Distillation Incorporating Prompt Optimization

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:02.697368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:02.697368Z digest=sha256:e9c19bc7b5e497e1ede0aa37b9d37bb7c0d0847870c6a212e165b9172b04e62d

Observation ab12e167-1817-48ba-a32c-9941bc621f32 · outbound

This paper cites DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:02.785656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:02.785656Z digest=sha256:95cd473b9840286e0c19df8bfb9f84d6096d22594451d6064fbf973dfee62de4

Observation 870979c2-4e34-441d-8bad-e60b4bcbc6c5 · outbound

This paper cites Spurious Rewards: Rethinking Training Signals in RLVR.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation Spurious Rewards: Rethinking Training Signals in RLVR

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:02.881915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:02.881915Z digest=sha256:a92196c8ce02768d8874449f0cf929e4dc88bf1af4166c01f4c97e354cf5ba19

Observation f2c60319-8fc3-427f-901d-8f0079bc5860 · outbound

This paper cites , Wang, P.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Wang, P

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:07.207353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:45:02.953539Z digest=sha256:17889e4560045b7283c3577140abba34e69c2e8d8522cebc1ada0b7f498f822d

Observation 98e0f490-a8be-44cd-9845-3e4ee7a4db71 · outbound

This paper cites The Curse of Recursion: Training on Generated Data Makes Models Forget.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation The Curse of Recursion: Training on Generated Data Makes Models Forget

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:03.029566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:03.029566Z digest=sha256:b31e3bf39ce61c3063274aa0579b5d52bb3a73cb95c6ca51f13b74a1c80dbe9c

Observation 13b7323d-6a46-4c54-9d5e-0eae71e6578a · outbound

This paper cites Gemma 3 Technical Report.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation Gemma 3 Technical Report

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:03.122039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:03.122039Z digest=sha256:beb1cc818f97fa1eb44c85806990010d2fa34b800e22aae9e55310dfe9adeb31

Observation 48454789-9132-4ff1-9fe3-3a143de98d2c · outbound

This paper cites , Riviere, M.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Riviere, M

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:07.097328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:45:03.183890Z digest=sha256:2bf18ac70f5bcde1f5505353c4bf56114acc74aafcb90f12065ebaf814a3f834

Observation e8406170-4e51-474e-9ba0-a26b4c29b939 · outbound

This paper cites , Han, Y.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Han, Y

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:06.972873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:45:03.261446Z digest=sha256:a160846284cb035a213846c4e159e19317d5e6e97ae6ed2babad3481483fe4c8

Observation 8a34b152-2464-440c-a5d0-9a7c4e36895b · outbound

This paper cites On Teacher Hacking in Language Model Distillation.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation On Teacher Hacking in Language Model Distillation

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:45:05.113351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:45:03.331296Z digest=sha256:a48a445bd9fd8665b8b3d9b424a7193964892be196418f16157d8f0d4af5e72f

Observation 210ff84e-986d-4660-8fd8-a7be4c99448c · outbound

This paper cites Who Taught You That? Tracing Teachers in Model Distillation.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation Who Taught You That? Tracing Teachers in Model Distillation

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:45:04.943544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:45:03.441504Z digest=sha256:2176b4ddbe77f4751243bb1c5f6cfe26342b3274a997bb649bf614af4c95ffbf

Observation 42e753af-4bd2-485d-819e-d73112561ade · outbound

This paper cites , Deng, Y.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Deng, Y

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:06.848770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:45:03.513056Z digest=sha256:0829c9bbd21ad4c28505f03b054abfd47fd887bcd795e86be0e28f26f5ab51d1

Observation 5fc159c6-a4f4-462a-8797-6b4e79b3f834 · outbound

This paper cites A Comprehensive Survey of Small Language Models in the Era of Large Language Models: Techniques, Enhancements, Applications, Collaboration with LLMs, and Trustworthiness.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation A Comprehensive Survey of Small Language Models in the Era of Large Language Models: Techniques, Enhancements, Applications, Collaboration with LLMs, and Trustworthiness

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:03.646532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:03.646532Z digest=sha256:09027f31e48f7982dc867d24aa3862a532772b7514530608e77708980953a5f7

Observation 41d95ac2-3c60-4166-a078-0a33e91b1192 · outbound

This paper cites , Zhu, J Y.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Zhu, J Y

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:06.722937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:45:03.749143Z digest=sha256:70fe562f14290d1ef6eddcfc6d20e1893b7159676adf3567722b609a9d758b8c

Observation ba61adf8-9c15-4502-b0db-7ff5881e1db4 · outbound

This paper cites an unresolved cited work.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:45:06.608963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:45:03.850600Z digest=sha256:77104324e121a94978ae8e35d8beb74347ac1bbcdd4fddcf0c090e93468127a9

Observation d3fcc0b3-7961-4740-aadb-991067663937 · outbound

This paper cites , Wang, X.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Wang, X

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:06.488819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:45:03.920673Z digest=sha256:5dd1cb57f6e73144899825914d3906dcefb27b171db0005eca589a453a0f5e96

Observation 8eafc91f-9730-40aa-b099-dbda10b62673 · outbound

This paper cites , Bai, H.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Bai, H

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:06.376947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:45:03.977383Z digest=sha256:6e1599132bf038a885eba00d0a97cc8bcb398b1d6c4c74bc1870e46f9a56a77d

Observation 6290ed46-791e-4ba6-aeac-df36eb0ac94b · outbound

This paper cites Speculative Knowledge Distillation: Bridging the Teacher-Student Gap Through Interleaved Sampling.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation Speculative Knowledge Distillation: Bridging the Teacher-Student Gap Through Interleaved Sampling

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:04.144603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:04.144603Z digest=sha256:eba31324aa4fb2c0ca86c250f095c54d4e01a2209c02595d3ec3a2b31b28e63d

Observation c9a0cc89-1f21-4f07-be33-6c3063688a6b · outbound

This paper cites Qwen2.5 Technical Report.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation Qwen2.5 Technical Report

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:04.222611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:04.222611Z digest=sha256:24f21ff2a56db7433787e472154800d79a68e805ed2d1f9e2b55666c2346834b

Observation 5f3339e7-b708-4899-86bf-8386abf8fd21 · outbound

This paper cites Demystifying Long Chain-of-Thought Reasoning in LLMs.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation Demystifying Long Chain-of-Thought Reasoning in LLMs

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:04.300085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:04.300085Z digest=sha256:5735c85e3701dd3f337e5ab4765b68a162f1a9068d0749f4533645d816f6e13d

Observation 012f0bf2-3f3d-406b-b1c2-e8c678c75d5e · outbound

This paper cites , Wang, C.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Wang, C

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:06.275447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:45:04.408916Z digest=sha256:5e8fb0fae7cb3d147bf6f1981b6d7f5577bfbb39cfd0f3ae0978afcf8558b7ab

Observation c1d9a273-3749-4648-b408-bc667bbe660f · outbound

This paper cites , Zhu, R.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Zhu, R

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:06.151692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:45:04.485680Z digest=sha256:43b437a633a46494f2ec5ff8319c54869a03f22b304c626f275ba783cd0cd0da

Observation d6cefc1b-c2e3-4129-a043-5a2f8a170d5d · outbound

This paper cites , Zhu, R.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Zhu, R

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:06.021904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:45:04.548728Z digest=sha256:1bc8b18c11982cf56e11ea79ab8bfc5e312229554a15933ab3636970ebe335fc

Observation 5b25e6b2-6caf-47bb-b638-48ef19505fc7 · outbound

This paper cites , Shen, J.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Shen, J

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:05.904741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:45:04.612738Z digest=sha256:5180f7e809439ea90f045d0ad4942f929c2507e68ae6ad438b74377a1251ea27

Observation 5dae83d3-89e7-4d02-807a-4b57195fbd38 · outbound

This paper cites Distill Not Only Data but Also Rewards: Can Smaller Language Models Surpass Larger Ones?.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation Distill Not Only Data but Also Rewards: Can Smaller Language Models Surpass Larger Ones?

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:04.656219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:04.656219Z digest=sha256:b77f9798e5bbc084d07a2f96dc5fcc31d2849d9d2d8d883fc74972c0c6ea32ec

Pith citing papers

No inbound Pith citation observations are available.