Pith. sign in

Paper Citation Record · LEDGER

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation

As of 7 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 0 inbound Pith citation observations for arXiv:2507.00054.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.00054 v1

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:45:04.656219Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

61 of 61 outbound references displayed

  • verified exact4
  • verified fuzzy32
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f9f7112a-3f5c-4430-89ce-ec48996e0545 · outbound

This paper cites , Aneja, J.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Aneja, J

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:11.959401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:44:59.832083Z digest=sha256:ac65c5dc340b69fc67d699cbc370133f955184603326c1a97bab887e31435baa

Observation 5334a726-edd3-4b45-a7fd-d1a4b28a4136 · outbound

This paper cites , Vieillard, N.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Vieillard, N

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:11.756821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:44:59.887421Z digest=sha256:ee38e3efb0a64880e3ab0717a35af4d65c0c2fc57ace8c83560529f117e62ae7

Observation 7976eb97-c850-406e-8b1a-15efc9efb035 · outbound

This paper cites , Vieillard, N.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Vieillard, N

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:11.433341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:44:59.969850Z digest=sha256:8dc672a9ecd8b29f1aeb3a362e7fd4f6a6091a50f23761ea1b770ba59492d730

Observation 963a99e5-8b87-4347-bc6c-33ec5f7679fa · outbound

This paper cites Towards Understanding Distilled Reasoning Models: A Representational Approach.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation Towards Understanding Distilled Reasoning Models: A Representational Approach

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:00.068336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:00.068336Z digest=sha256:ef31ca46f0cb3c802626e7e20ca687d742139ee9ca1e10736c012a8bc15f8f93

Observation 9813e50b-31b0-44d5-a629-a293fc9c17f5 · outbound

This paper cites , Sastre, I.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Sastre, I

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:11.259853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:45:00.163772Z digest=sha256:fa0e4ea0146eaed0e7e0cae3f4cf1c2d691252655e4e5ad52051be64988e26ff

Observation ee832064-5677-498b-93ca-96c24d0a7770 · outbound

This paper cites Smaller, Weaker, Yet Better: Training LLM Reasoners via Compute-Optimal Sampling.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation Smaller, Weaker, Yet Better: Training LLM Reasoners via Compute-Optimal Sampling

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:00.253407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:00.253407Z digest=sha256:c178d919b857e206aacad202e86ed3f629f3c970507e709663fa656a965e58e2

Observation 7ca5a211-4a5a-4bd4-8f10-23e74975e1db · outbound

This paper cites , Peebles, B.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Peebles, B

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:11.016657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:45:00.323553Z digest=sha256:9f160c6a7b0f48952b727e8fea140222ee74e00a597c29452d21dd9fbce375c2

Observation 8585932b-95fd-4e64-8168-4765ac6904cd · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation Training Verifiers to Solve Math Word Problems

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:00.383119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:00.383119Z digest=sha256:e681395b3f963807cb95d6096205d10ae55f1da36ed20dae9b08674da9c7ee51

Observation ae03c8ce-0b48-4fad-9d78-2c9f0baa78ed · outbound

This paper cites \ Ngo, C.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation \ Ngo, C

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:10.872019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:45:00.432702Z digest=sha256:89639e95004cfa305d23050367325defeaf30553c374f7dd9b9b734dc4106707

Observation b0065cb9-5d07-4763-8d75-6ddc406cc9e4 · outbound

This paper cites , See, A.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , See, A

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:10.714160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:45:00.499454Z digest=sha256:266f4da91b38074f88e7861ee7c3fdde5d824584ffa780fd74d789de5f98405f

Observation fda81fd6-744d-48b2-a6d0-d940a7757342 · outbound

This paper cites , Feng, B.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Feng, B

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:10.515033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:45:00.562770Z digest=sha256:bc12170b5831ca228fec8c7cb2abf7cc0be318129a3287c5d8bbeb309ea8e91d

Observation 88c34840-15c4-4128-afbb-d71d4410be5d · outbound

This paper cites an unresolved cited work.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:45:10.354570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:45:00.644914Z digest=sha256:28a0e3679c8c4c638a8040a6fbaf1dbc1c0371c8312b87d7b4377077457cff3c

Observation 8a236ce7-78d9-45d5-9e9a-3c7e6eb7b7cd · outbound

This paper cites an unresolved cited work.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:45:10.118810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:45:00.717936Z digest=sha256:5524310ad6c2d275c5833f7b9e3b46b4463a4be73055cef5fbd3dc81e230b754

Observation 8af77e78-1dab-42c3-bf88-2b07b2c7012b · outbound

This paper cites Advantage-Guided Distillation for Preference Alignment in Small Language Models.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation Advantage-Guided Distillation for Preference Alignment in Small Language Models

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:45:05.664522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:45:00.774715Z digest=sha256:c4451df9d2c346f7bcc2fc85e7fa47e46c777b807f970d61b1d542400b945c03

Observation ba453a7e-76aa-4268-970c-cc687a2a4b18 · outbound

This paper cites MiniLLM: On-Policy Distillation of Large Language Models.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation MiniLLM: On-Policy Distillation of Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:00.840611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:00.840611Z digest=sha256:e25278529d0db1ac6141852b9644b2e634ada6f478cc4fd4c7478e669f329c34

Observation ec3ea647-c2f7-4d1b-bdb7-e7e74b0b621e · outbound

This paper cites , Zhang, L L.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Zhang, L L

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:09.899083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:45:00.916470Z digest=sha256:4c947c48f4dcbcd65b52a3b96a8dcb1fb1a4c0cfcd4adfe571a6cfc558e783aa

Observation 3a330c4d-5b4f-4e1d-91db-ba0f3d6215db · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:00.969331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:00.969331Z digest=sha256:510ed5f901a3cdc419e671931a9938504a28cd80631eea24992c9507abf9df69

Observation 2cc95be0-a373-4d5d-89c2-100f15211594 · outbound

This paper cites , Vinyals, O.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Vinyals, O

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:09.681927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:45:01.034051Z digest=sha256:5b19f479ac934d3728eff08b7ca84cb84f27f61eef1bcf05771ca888243ad75c

Observation 2988ec95-7570-4805-86fe-20b740f166c6 · outbound

This paper cites , Borgeaud, S.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Borgeaud, S

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:09.484507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:45:01.104215Z digest=sha256:ad302899898d0ec313eaa28d6b2d10b191f43cbfd0058845facddff8c8663734

Observation 38769f70-6b26-4119-a44c-9bc87c2dcc74 · outbound

This paper cites , Li, C L.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Li, C L

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:09.258428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:45:01.208161Z digest=sha256:11f99031898e620b213cc6d5f0490e25e783bdf72a0b1b62edf8c645b036b2d0

Observation b2223c7a-1689-484d-a335-981e059a7d3c · outbound

This paper cites , Yin, Y.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Yin, Y

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:09.049575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:45:01.282141Z digest=sha256:f409322a8f72560ff4c11c84d2d140b3ed055139e371866d2c598e8bdfc6bb60

Observation c63aaa27-aeda-4fdf-b195-07a0ea7ca4b8 · outbound

This paper cites MindStar: Enhancing Math Reasoning in Pre-trained LLMs at Inference Time.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation MindStar: Enhancing Math Reasoning in Pre-trained LLMs at Inference Time

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:01.362923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:01.362923Z digest=sha256:bc4d2bad4a382f2ff882d1e28f0462bacf1ae868dbdaeba77c27764c9c2b5ae3

Observation ba9ae507-d70d-49b0-9c49-f1052c19a884 · outbound

This paper cites \ Rush, A M.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation \ Rush, A M

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:08.841836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:45:01.446884Z digest=sha256:feb0cf6fd951e4c07243051096d3292b3acb89048978088b92de10f4c20a29c9

Observation d76b30f7-97b5-49f2-8da9-86c67609c337 · outbound

This paper cites DistiLLM-2: A Contrastive Approach Boosts the Distillation of LLMs.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation DistiLLM-2: A Contrastive Approach Boosts the Distillation of LLMs

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:01.511764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:01.511764Z digest=sha256:8a952b82e56b0e9fde362176773e67173655eacc5e678f09a1a93230158a583d

Observation bf1372cb-2c2d-4815-b264-30df58d36e6d · outbound

This paper cites DistiLLM: Towards Streamlined Distillation for Large Language Models.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation DistiLLM: Towards Streamlined Distillation for Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:01.602702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:01.602702Z digest=sha256:254521da81bc04e3be73dd0c74a12eee4be40472065917233b2a56ebd41231a7

Observation bbfeb791-a91d-4c43-9cde-01ea480d017f · outbound

This paper cites , Fang, L.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Fang, L

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:08.621314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:45:01.662995Z digest=sha256:6c64d13535f740a99f06e5bfefb8abe3385e509765fc575b60a5b1049f0b5f6b

Observation 1f40c54b-8e47-4293-9125-d763d5183b0d · outbound

This paper cites Winning Big with Small Models: Knowledge Distillation vs. Self-Training for Reducing Hallucination in Product QA Agents.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation Winning Big with Small Models: Knowledge Distillation vs. Self-Training for Reducing Hallucination in Product QA Agents

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:01.734181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:01.734181Z digest=sha256:c31c9af7973a12621a1ed0e736d1ac2d4085cc61bf04c239c643a373aead15be

Observation 1f1635b6-be41-4bcb-aa4b-a089978da199 · outbound

This paper cites GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:01.803122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:01.803122Z digest=sha256:2d90b1f571e13e61666515bc8e0a1c767c6253a59ee8ad331b75c57159070364

Observation a41021f4-9434-410f-86af-7fde9090a21d · outbound

This paper cites Evolving Knowledge Distillation with Large Language Models and Active Learning.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation Evolving Knowledge Distillation with Large Language Models and Active Learning

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:45:05.399923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:45:01.899416Z digest=sha256:8dc335257d2566770c2eec15fda235e1c862ca2f66d709f625b7d71bc7db7d1a

Observation c9427546-55e8-4e8e-81c2-4f74cca695e4 · outbound

This paper cites , Chen, C.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Chen, C

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:08.535787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:45:01.987456Z digest=sha256:486a82004175d69bfb1d1f44fcbd204e87f35d87f80a7deb3467dc5eec582c98

Observation 663344f9-7e30-47e0-89bc-9ac38933d316 · outbound

This paper cites , Yang, Z.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Yang, Z

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:08.365808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:45:02.050079Z digest=sha256:065a45112368179539cb07f6a16f143ba9cf7fd72e6dc7d1529c4b228242e765

Observation 335802d2-6c7b-402c-8a4f-1c2f5861e9a2 · outbound

This paper cites , Dani, J.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Dani, J

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:08.178769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:45:02.176949Z digest=sha256:a16c033602e7a92e7fc571b6fc1c261adfdd0efacd00e5a36ea3b8800e16d837

Observation f6df0b60-2933-495f-86bb-c081bc56fa71 · outbound

This paper cites an unresolved cited work.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:45:08.012500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:45:02.296854Z digest=sha256:554d1455c81afe65a8de4dcc7833b93dba73c84fe20cccc22dcc9f0ecb83e9f4

Observation 71d9bb01-9f8b-4ccc-a41a-b8716ccb4285 · outbound

This paper cites , Lerer, A.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Lerer, A

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:07.780048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:45:02.381998Z digest=sha256:9af209ba9862b700ded312bd085c179563044679d074f2d76f2a58e92e60d92a

Observation 710f5486-52fa-4b6d-8121-0e832a6133de · outbound

This paper cites Progressive distillation induces an implicit curriculum.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation Progressive distillation induces an implicit curriculum

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:02.463708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:02.463708Z digest=sha256:f9071776c4400f3e552a279203d77739854488e60af6b26387060ca17f642f0c

Observation 0d86e769-ece0-4b0e-8736-d57c1f90975d · outbound

This paper cites , Kim, D.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Kim, D

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:07.585952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:45:02.552158Z digest=sha256:e27c1c1e0632fa1ff094bdfd6c2846db273d258fbabca28c0905436f58543f09

Observation 17d3ac87-dd4d-41a7-bd6d-017ad7752bd1 · outbound

This paper cites \ Xie, S.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation \ Xie, S

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:07.347692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:45:02.621187Z digest=sha256:7d0c14d6b140902c132fba54c76b7f32cc3e463f19c56101738e3448b074488e

Observation dedb6313-b2d1-43cd-89cb-46c0ddf3e22b · outbound

This paper cites CourseGPT-zh: an Educational Large Language Model Based on Knowledge Distillation Incorporating Prompt Optimization.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation CourseGPT-zh: an Educational Large Language Model Based on Knowledge Distillation Incorporating Prompt Optimization

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:02.697368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:02.697368Z digest=sha256:995e3ecddc53ec36830dd04a3eb46652907064a159960fbfb0f6247ff30814ac

Observation ab12e167-1817-48ba-a32c-9941bc621f32 · outbound

This paper cites DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:02.785656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:02.785656Z digest=sha256:dab488e163a701a5b6f647c83555f6efafe268087e75532a1981bfa45b9044d6

Observation 870979c2-4e34-441d-8bad-e60b4bcbc6c5 · outbound

This paper cites Spurious Rewards: Rethinking Training Signals in RLVR.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation Spurious Rewards: Rethinking Training Signals in RLVR

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:02.881915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:02.881915Z digest=sha256:bcb24f697d7fe5ee8b9570849b6266fb17874fd7ebd2f622c601ea7918d8d1d0

Observation f2c60319-8fc3-427f-901d-8f0079bc5860 · outbound

This paper cites , Wang, P.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Wang, P

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:07.207353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:45:02.953539Z digest=sha256:db4a91d466110bfdf58d87b066010b60b2aa08c3aad5ae144ed19092fb42d9f3

Observation 98e0f490-a8be-44cd-9845-3e4ee7a4db71 · outbound

This paper cites The Curse of Recursion: Training on Generated Data Makes Models Forget.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation The Curse of Recursion: Training on Generated Data Makes Models Forget

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:03.029566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:03.029566Z digest=sha256:ade3f28246cd063f617aa015827f3ad46c477f33138f072ced69bed68cec1067

Observation 13b7323d-6a46-4c54-9d5e-0eae71e6578a · outbound

This paper cites Gemma 3 Technical Report.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation Gemma 3 Technical Report

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:03.122039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:03.122039Z digest=sha256:740b6f0f8740e2b112eed770c39e5fa88f4a3a3b5bb071eba4adf031b1d7204c

Observation 48454789-9132-4ff1-9fe3-3a143de98d2c · outbound

This paper cites , Riviere, M.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Riviere, M

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:07.097328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:45:03.183890Z digest=sha256:318ee8bf3cf1daeff08ada9f9172d2972b246dc3b294e940fe2602aafff76f40

Observation e8406170-4e51-474e-9ba0-a26b4c29b939 · outbound

This paper cites , Han, Y.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Han, Y

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:06.972873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:45:03.261446Z digest=sha256:53276980112b5e343144c528ffa2afbe5cb8c5c5c7f08fee614a206e9c32d8f7

Observation 8a34b152-2464-440c-a5d0-9a7c4e36895b · outbound

This paper cites On Teacher Hacking in Language Model Distillation.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation On Teacher Hacking in Language Model Distillation

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:45:05.113351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:45:03.331296Z digest=sha256:1521354408d8554024378e183ce6a1d8ad03ddadbca382e82e7e3b6d9aeac735

Observation 210ff84e-986d-4660-8fd8-a7be4c99448c · outbound

This paper cites Who Taught You That? Tracing Teachers in Model Distillation.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation Who Taught You That? Tracing Teachers in Model Distillation

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:45:04.943544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:45:03.441504Z digest=sha256:274720fc387f8a69f621778380d0504ea79756d7a6ab3b47581b1a7b474b321a

Observation 42e753af-4bd2-485d-819e-d73112561ade · outbound

This paper cites , Deng, Y.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Deng, Y

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:06.848770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:45:03.513056Z digest=sha256:410d90ba7781685ebacc174d489b33956a7afc19b4b2a43b1e03e3b5c07fb649

Observation 5fc159c6-a4f4-462a-8797-6b4e79b3f834 · outbound

This paper cites A Comprehensive Survey of Small Language Models in the Era of Large Language Models: Techniques, Enhancements, Applications, Collaboration with LLMs, and Trustworthiness.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation A Comprehensive Survey of Small Language Models in the Era of Large Language Models: Techniques, Enhancements, Applications, Collaboration with LLMs, and Trustworthiness

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:03.646532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:03.646532Z digest=sha256:9bb5fcf953b5f71a5bea45c8e9e44f5e6c88a69c49eb281a4576f45149ea848e

Observation 41d95ac2-3c60-4166-a078-0a33e91b1192 · outbound

This paper cites , Zhu, J Y.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Zhu, J Y

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:06.722937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:45:03.749143Z digest=sha256:076a97f428551607fac8fe8a02e102bb80692fc06bd7598eb9a28f72f7664ba7

Observation ba61adf8-9c15-4502-b0db-7ff5881e1db4 · outbound

This paper cites an unresolved cited work.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:45:06.608963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:45:03.850600Z digest=sha256:456817f7654b81065c9f6baafe022c52d6a618ba43060821962ae00462cf48be

Observation d3fcc0b3-7961-4740-aadb-991067663937 · outbound

This paper cites , Wang, X.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Wang, X

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:06.488819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:45:03.920673Z digest=sha256:4956f124b1e5b07b65b0c0bac63bd1228a1a3bec9f92d00a17fd6c8382402808

Observation 8eafc91f-9730-40aa-b099-dbda10b62673 · outbound

This paper cites , Bai, H.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Bai, H

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:06.376947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:45:03.977383Z digest=sha256:ff59414e7c37004c4fa3f204c2102157fe9359484f3408dc029a6fc27ee72172

Observation 6290ed46-791e-4ba6-aeac-df36eb0ac94b · outbound

This paper cites Speculative Knowledge Distillation: Bridging the Teacher-Student Gap Through Interleaved Sampling.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation Speculative Knowledge Distillation: Bridging the Teacher-Student Gap Through Interleaved Sampling

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:04.144603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:04.144603Z digest=sha256:671b84d3df1331692362132f1da38d9401e614b9ee64c9b1e927693163bc955c

Observation c9a0cc89-1f21-4f07-be33-6c3063688a6b · outbound

This paper cites Qwen2.5 Technical Report.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation Qwen2.5 Technical Report

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:04.222611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:04.222611Z digest=sha256:efa4667cf3e13c6c137a53f720517d663202a76db21ec463c44a7fecf664c462

Observation 5f3339e7-b708-4899-86bf-8386abf8fd21 · outbound

This paper cites Demystifying Long Chain-of-Thought Reasoning in LLMs.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation Demystifying Long Chain-of-Thought Reasoning in LLMs

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:04.300085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:04.300085Z digest=sha256:c868e850f4a9a4456abdf2c3cf3265e79ee42bb359c7daaa59d46383481080a3

Observation 012f0bf2-3f3d-406b-b1c2-e8c678c75d5e · outbound

This paper cites , Wang, C.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Wang, C

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:06.275447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:45:04.408916Z digest=sha256:254e085e73f55e6c1cd8ee6150c957a44e648c670156776afbd9b63f3a9b1f42

Observation c1d9a273-3749-4648-b408-bc667bbe660f · outbound

This paper cites , Zhu, R.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Zhu, R

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:06.151692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:45:04.485680Z digest=sha256:fec212638bba4d2f7d3ae7559f0a92f238bd88f481680d6d87648d4e227056f8

Observation d6cefc1b-c2e3-4129-a043-5a2f8a170d5d · outbound

This paper cites , Zhu, R.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Zhu, R

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:06.021904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:45:04.548728Z digest=sha256:f498328b40427d3f08e5ee7b85cd94ef4d9815c4ebfc0fc5f978a3e38c88d145

Observation 5b25e6b2-6caf-47bb-b638-48ef19505fc7 · outbound

This paper cites , Shen, J.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Shen, J

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:05.904741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:45:04.612738Z digest=sha256:2bff1bd00de5eb37068ad9bd28f7a1899012da77355053037b72d69a991bb665

Observation 5dae83d3-89e7-4d02-807a-4b57195fbd38 · outbound

This paper cites Distill Not Only Data but Also Rewards: Can Smaller Language Models Surpass Larger Ones?.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation Distill Not Only Data but Also Rewards: Can Smaller Language Models Surpass Larger Ones?

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:04.656219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:04.656219Z digest=sha256:2e3ab833fda5d95d8e79968327411fe0977bf7e26b5af5bbba48e9477a1b907f

Pith citing papers

No inbound Pith citation observations are available.