Pith. sign in

Paper Citation Record · LEDGER

AdaKP: Online Adaptive Knowledge-Point Selection for Reasoning-Oriented Reinforcement Learning

As of 7 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 0 inbound Pith citation observations for arXiv:2607.24833.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.24833 v1

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T07:34:57.251159Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

48 of 48 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved48
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 81b18e99-c7e4-4435-892c-76782d9ce4ee · outbound

This paper cites 2026 , note =.

AdaKP: Online Adaptive Knowledge-Point Selection for Reasoning-Oriented Reinforcement Learning 2026 , note =

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T07:34:52.667421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:34:52.667421Z digest=sha256:d0746feceaf62cd151d8a9aacd2ac20f058a8a6d9c85dd40104205a6e8101bfd

Observation 244961c9-8f48-4f6b-8f4c-cce26612c28c · outbound

This paper cites 2025 , note =.

AdaKP: Online Adaptive Knowledge-Point Selection for Reasoning-Oriented Reinforcement Learning 2025 , note =

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T07:34:52.720900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:34:52.720900Z digest=sha256:f6043b14a230e2bf4b9d8fa09bf15b407434b659256a022cb6601b30bab3e872

Observation 04b29a83-3f63-4861-b4e5-29a1d2c9d395 · outbound

This paper cites an unresolved cited work.

AdaKP: Online Adaptive Knowledge-Point Selection for Reasoning-Oriented Reinforcement Learning Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T07:34:52.808847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:34:52.808847Z digest=sha256:1e9d7c63f27c0c34f4c676e47d6f7252782970464d631cd1fe923dd0d2cb3912

Observation ef391fe9-6136-4169-a24a-8f0556461d0b · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

AdaKP: Online Adaptive Knowledge-Point Selection for Reasoning-Oriented Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T07:34:52.905598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:34:52.905598Z digest=sha256:6dd03abe6aff62e0f8939dd8dc80b82e937f297a2e690a841d7f450514add8ca

Observation 6bacec8d-f120-4ce3-a2b7-db03af9e9a12 · outbound

This paper cites 2024 , note =.

AdaKP: Online Adaptive Knowledge-Point Selection for Reasoning-Oriented Reinforcement Learning 2024 , note =

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T07:34:53.022906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:34:53.022906Z digest=sha256:420648ea20224a828f1f4fd25d5280eb353712295211a074702433512413bc3a

Observation 87a9d1a3-6d2d-4968-946a-4afa7233e0eb · outbound

This paper cites Proximal Policy Optimization Algorithms.

AdaKP: Online Adaptive Knowledge-Point Selection for Reasoning-Oriented Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T07:34:53.089892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:34:53.089892Z digest=sha256:51b6e68ca57cff113bd4d5a148d87d5fbb33abee1dcabb07b64e5319a0816670

Observation 4cd58a9d-cc36-4cf6-98aa-2e3635c3c67f · outbound

This paper cites Advances in Neural Information Processing Systems (NeurIPS) , year =.

AdaKP: Online Adaptive Knowledge-Point Selection for Reasoning-Oriented Reinforcement Learning Advances in Neural Information Processing Systems (NeurIPS) , year =

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T07:34:53.172408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:34:53.172408Z digest=sha256:aaf5b3114e761973452e5645537e4001391327ad985cc9401e8ddfb42a556ddd

Observation 35971f69-dcd6-420b-b6df-2ec85b50d6cf · outbound

This paper cites International Conference on Learning Representations (ICLR) , year =.

AdaKP: Online Adaptive Knowledge-Point Selection for Reasoning-Oriented Reinforcement Learning International Conference on Learning Representations (ICLR) , year =

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T07:34:53.287453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:34:53.287453Z digest=sha256:26776396005c81e5d3389790606cfbd525224a0406f28cf3be1285ecefd1af1f

Observation 56b98f94-39ec-4418-b207-d4eaaeecbfb5 · outbound

This paper cites an unresolved cited work.

AdaKP: Online Adaptive Knowledge-Point Selection for Reasoning-Oriented Reinforcement Learning Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T07:34:53.352958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:34:53.352958Z digest=sha256:2d9e9988199bb9ff0b9612d0592e3a502c6497b63518898270d8b706443932ce

Observation 03fce6f1-890c-438b-91ae-5704d09df196 · outbound

This paper cites Advances in Neural Information Processing Systems (NeurIPS) , year =.

AdaKP: Online Adaptive Knowledge-Point Selection for Reasoning-Oriented Reinforcement Learning Advances in Neural Information Processing Systems (NeurIPS) , year =

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T07:34:53.462245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:34:53.462245Z digest=sha256:2f9915a5783fe0150b7fe4a35acde2e8bf9339d45ca31a648d045438a3e088bc

Observation 6a91bba5-0c5d-40b4-823a-56c3226fb3c6 · outbound

This paper cites , booktitle =.

AdaKP: Online Adaptive Knowledge-Point Selection for Reasoning-Oriented Reinforcement Learning , booktitle =

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T07:34:53.578840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:34:53.578840Z digest=sha256:054de2b6ae89d868e78e67bf8fa5a931921a8090149421d9b175811afa85f5d3

Observation 501596e3-1294-47b5-ae47-ff65e4dff05b · outbound

This paper cites , journal =.

AdaKP: Online Adaptive Knowledge-Point Selection for Reasoning-Oriented Reinforcement Learning , journal =

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T07:34:53.652970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:34:53.652970Z digest=sha256:c4cd2a205e86af6669dfc802f0cf27b83f5ed16ca8d02db318a37f6ab5a31183

Observation ebb1ed94-b1b7-42a9-8472-a027ac6eb1e8 · outbound

This paper cites Advances in Neural Information Processing Systems (NeurIPS) , year =.

AdaKP: Online Adaptive Knowledge-Point Selection for Reasoning-Oriented Reinforcement Learning Advances in Neural Information Processing Systems (NeurIPS) , year =

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T07:34:53.721806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:34:53.721806Z digest=sha256:5cd98e6023c2f88eb5a30a4d3602554bcd5d7ad2039524b9449555e18f1321ba

Observation caa7f80d-8fa4-40ef-9fb1-2d1aef61493f · outbound

This paper cites and Zhang, Hao and Stoica, Ion , booktitle =.

AdaKP: Online Adaptive Knowledge-Point Selection for Reasoning-Oriented Reinforcement Learning and Zhang, Hao and Stoica, Ion , booktitle =

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T07:34:53.842119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:34:53.842119Z digest=sha256:4fd6b73452e511a2f757d68fc3909d0f7895e8f5b8f351ccc4d334b249ac178d

Observation 1df552fc-4245-4de6-ab8b-e5cb5d41a851 · outbound

This paper cites 2025 , note =.

AdaKP: Online Adaptive Knowledge-Point Selection for Reasoning-Oriented Reinforcement Learning 2025 , note =

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T07:34:53.929243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:34:53.929243Z digest=sha256:7776091ed72472c347493205c2231cbbb5d57649a92280fb68f58a7b155fd1b5

Observation ecadc99e-6b6c-40d1-b1cd-007ea25f2f97 · outbound

This paper cites 2023 , note =.

AdaKP: Online Adaptive Knowledge-Point Selection for Reasoning-Oriented Reinforcement Learning 2023 , note =

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T07:34:54.016895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:34:54.016895Z digest=sha256:23818e41b49ddad3f894c19f1da11efc4929e25f513c8ea9c4dbf1d63bd835da

Observation 27bc778d-a9ab-4877-a035-5bc786942a90 · outbound

This paper cites 2024 , note =.

AdaKP: Online Adaptive Knowledge-Point Selection for Reasoning-Oriented Reinforcement Learning 2024 , note =

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T07:34:54.072903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:34:54.072903Z digest=sha256:f6953f4156bc93fb6f4ba3ca4afb4dddae600fbee9a96149e963f053dab27638

Observation 5b691810-3838-4793-961b-6550e5a5eb06 · outbound

This paper cites an unresolved cited work.

AdaKP: Online Adaptive Knowledge-Point Selection for Reasoning-Oriented Reinforcement Learning Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T07:34:54.177432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:34:54.177432Z digest=sha256:c5c5c25662a9c1da74ad814e35aee73d6579b30fd4092f59d7e2526fe13144ff

Observation e7a4c51a-2945-45bc-a7bd-2a6e11a2109a · outbound

This paper cites International Conference on Learning Representations (ICLR) , year =.

AdaKP: Online Adaptive Knowledge-Point Selection for Reasoning-Oriented Reinforcement Learning International Conference on Learning Representations (ICLR) , year =

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T07:34:54.255689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:34:54.255689Z digest=sha256:b2589959ef2a83316bc9c2fb551ff1ef040d50692f3dd948295e4296c0f32693

Observation 8569224e-5a1b-4503-bc37-8c8a8ad3c12c · outbound

This paper cites International Conference on Machine Learning (ICML) , year =.

AdaKP: Online Adaptive Knowledge-Point Selection for Reasoning-Oriented Reinforcement Learning International Conference on Machine Learning (ICML) , year =

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T07:34:54.337048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:34:54.337048Z digest=sha256:9f3bcd89c981bc450a2f16db64c60bc40d43561f3c16f33fdb6244e39b3ad89e

Observation a0fd8522-fee6-4c9a-b44e-29d9c7ae87e5 · outbound

This paper cites an unresolved cited work.

AdaKP: Online Adaptive Knowledge-Point Selection for Reasoning-Oriented Reinforcement Learning Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T07:34:54.455969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:34:54.455969Z digest=sha256:2713a9c00954d35cf220fea6e8810318ba3bf2f04019954ff3c21ceebfb8ecbc

Observation 7d5b5313-5282-441e-9b9e-0622dcf199e2 · outbound

This paper cites Don't Tell the Answer, Truly Guide the Reasoning During.

AdaKP: Online Adaptive Knowledge-Point Selection for Reasoning-Oriented Reinforcement Learning Don't Tell the Answer, Truly Guide the Reasoning During

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T07:34:54.537997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:34:54.537997Z digest=sha256:275b2684cf65c444bbd669126a8028ee84cc526e95d32d6eeba302964f451822

Observation 8c9d8cfe-5ecc-4540-9d30-8741ced11c91 · outbound

This paper cites an unresolved cited work.

AdaKP: Online Adaptive Knowledge-Point Selection for Reasoning-Oriented Reinforcement Learning Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T07:34:54.601137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:34:54.601137Z digest=sha256:b435cd8f6b28558b4e447f7640bdc1acb90b16a83f53ceab37d48e494bb9ee64

Observation 45429833-a9ff-4208-938f-93cba4343624 · outbound

This paper cites an unresolved cited work.

AdaKP: Online Adaptive Knowledge-Point Selection for Reasoning-Oriented Reinforcement Learning Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T07:34:54.701839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:34:54.701839Z digest=sha256:ae1b59c725bba42b6db502a4b34617d4ee65db4f4f332dc0c5ac1af71d5a3152

Observation 26c06b32-36c7-43d3-ba09-8a252be4e03a · outbound

This paper cites Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding.

AdaKP: Online Adaptive Knowledge-Point Selection for Reasoning-Oriented Reinforcement Learning Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T07:34:54.765756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:34:54.765756Z digest=sha256:7de95ab684a256901c50dbc0c0d3273f24c0045048bdd9065d5d02ac4c7c0438

Observation b6635fc8-6287-485f-9f4f-f0b72cdaf244 · outbound

This paper cites Boosting.

AdaKP: Online Adaptive Knowledge-Point Selection for Reasoning-Oriented Reinforcement Learning Boosting

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T07:34:54.942938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:34:54.942938Z digest=sha256:0406f6d27a58be461a3eb60fc98cd4c46ac29c486cc4a8ef592ce891259fd8f3

Observation 3dfea157-c72a-4a0e-aaac-d803c40e3597 · outbound

This paper cites Train at Moving Edge: Online-Verified Prompt Selection for Efficient.

AdaKP: Online Adaptive Knowledge-Point Selection for Reasoning-Oriented Reinforcement Learning Train at Moving Edge: Online-Verified Prompt Selection for Efficient

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T07:34:55.055430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:34:55.055430Z digest=sha256:cc09eed9c665149b4e5e466375515ff34b98fbe2cd0694392c40c81336eb82f5

Observation 19fca2ce-0a71-4027-b71a-d83b82273fa7 · outbound

This paper cites an unresolved cited work.

AdaKP: Online Adaptive Knowledge-Point Selection for Reasoning-Oriented Reinforcement Learning Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T07:34:55.112071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:34:55.112071Z digest=sha256:628eee988eb3ce530654612ec3d97adde69bc4ab33dcc577b9dfad9913d3c9a5

Observation 0733587a-51d5-43fe-bd4f-8fb3e44c8c28 · outbound

This paper cites an unresolved cited work.

AdaKP: Online Adaptive Knowledge-Point Selection for Reasoning-Oriented Reinforcement Learning Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T07:34:55.220532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:34:55.220532Z digest=sha256:f5c322191f4998526a2088841a7cf3c3fe0064fe5ca896f8af7e520f2a2b195a

Observation 6156148b-691e-451a-9ee1-c8ad2f6d81d7 · outbound

This paper cites an unresolved cited work.

AdaKP: Online Adaptive Knowledge-Point Selection for Reasoning-Oriented Reinforcement Learning Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T07:34:55.301544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:34:55.301544Z digest=sha256:cae25d61ce874ccd4e6df4a4894bc79ea664136a2264a5df80014d5cad77546d

Observation c30af885-3f78-4ee6-8acf-9c45f11113fd · outbound

This paper cites Measuring Mathematical Problem Solving With the.

AdaKP: Online Adaptive Knowledge-Point Selection for Reasoning-Oriented Reinforcement Learning Measuring Mathematical Problem Solving With the

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T07:34:55.388018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:34:55.388018Z digest=sha256:4f9d9c36f8dee95efd3a5964cd4bc023c3ea80cd6bf5f74e1caba5f26538486c

Observation 2172d4fc-2eaa-4f1a-bd87-a8cd7a30081e · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

AdaKP: Online Adaptive Knowledge-Point Selection for Reasoning-Oriented Reinforcement Learning Training Verifiers to Solve Math Word Problems

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T07:34:55.532748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:34:55.532748Z digest=sha256:2da8b7965b6cc7479eced4d98cc202677aa3936adb68cefee7aa85574cfb1bfe

Observation 30538229-4d0d-43fc-8c6b-b9800ed7363b · outbound

This paper cites Advances in Neural Information Processing Systems (NeurIPS) , year =.

AdaKP: Online Adaptive Knowledge-Point Selection for Reasoning-Oriented Reinforcement Learning Advances in Neural Information Processing Systems (NeurIPS) , year =

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T07:34:55.602336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:34:55.602336Z digest=sha256:51b2b55b0032f9e68a072534c9f0b18192baaebdc6bc09057f04e9b177a24b8e

Observation 02d054e1-546e-4910-b105-ac5b7fbd7c13 · outbound

This paper cites International Conference on Learning Representations (ICLR) , year =.

AdaKP: Online Adaptive Knowledge-Point Selection for Reasoning-Oriented Reinforcement Learning International Conference on Learning Representations (ICLR) , year =

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T07:34:55.683543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:34:55.683543Z digest=sha256:10702b0cf619a46061f26eb51479bb135c5c7934829cc23bc18d86d3e06cb840

Observation 036d2c91-1866-401b-bfe9-3eb6cbb782ba · outbound

This paper cites Advances in Neural Information Processing Systems (NeurIPS) , year =.

AdaKP: Online Adaptive Knowledge-Point Selection for Reasoning-Oriented Reinforcement Learning Advances in Neural Information Processing Systems (NeurIPS) , year =

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T07:34:55.763796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:34:55.763796Z digest=sha256:e6552a8674a11153c3b2320e6c25d4d11183cce7ed321fd3bf1044c0e638de42

Observation d372dc61-e76c-4364-9a82-5502f9bf78b9 · outbound

This paper cites Advances in Neural Information Processing Systems (NeurIPS) , year =.

AdaKP: Online Adaptive Knowledge-Point Selection for Reasoning-Oriented Reinforcement Learning Advances in Neural Information Processing Systems (NeurIPS) , year =

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T07:34:55.870328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:34:55.870328Z digest=sha256:13277d19689017b19171449038d410d9027c837656bfd90ea99d76574c11a27d

Observation 72e1f0de-0a34-46c2-a54b-8f58af08b9a5 · outbound

This paper cites International Conference on Learning Representations (ICLR) , year =.

AdaKP: Online Adaptive Knowledge-Point Selection for Reasoning-Oriented Reinforcement Learning International Conference on Learning Representations (ICLR) , year =

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T07:34:55.960507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:34:55.960507Z digest=sha256:495705b00d7e6e31a9e6a5bd61f472ccb965f2ef27c460d7f415a4a1e866267d

Observation 0b6b5067-8c38-4037-b5e7-595390f58244 · outbound

This paper cites Advances in Neural Information Processing Systems (NeurIPS) , year =.

AdaKP: Online Adaptive Knowledge-Point Selection for Reasoning-Oriented Reinforcement Learning Advances in Neural Information Processing Systems (NeurIPS) , year =

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T07:34:56.046520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:34:56.046520Z digest=sha256:648aeec278104d04db6991b61e647c93703a96715cc93f0ad06bc47a8f515a75

Observation c956bc34-fda9-4db0-b9a6-724bbd67aa86 · outbound

This paper cites Back to Basics: Revisiting.

AdaKP: Online Adaptive Knowledge-Point Selection for Reasoning-Oriented Reinforcement Learning Back to Basics: Revisiting

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T07:34:56.122271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:34:56.122271Z digest=sha256:8db66fcd2f515f6dd59aaaf0f2f1b63b9c8840e6717b84b5f0c796618a509170

Observation 53302985-149c-4556-896b-0afd4b18c722 · outbound

This paper cites 2024 , note =.

AdaKP: Online Adaptive Knowledge-Point Selection for Reasoning-Oriented Reinforcement Learning 2024 , note =

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T07:34:56.231813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:34:56.231813Z digest=sha256:b58d10a7d0dc80ea8a98f56803a60a0f343852e52f31f2be078b3ad6d014555b

Observation 897bbd8e-ef9a-4d1b-a8bc-958357b59bb2 · outbound

This paper cites Advances in Neural Information Processing Systems (NeurIPS) , year =.

AdaKP: Online Adaptive Knowledge-Point Selection for Reasoning-Oriented Reinforcement Learning Advances in Neural Information Processing Systems (NeurIPS) , year =

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T07:34:56.376887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:34:56.376887Z digest=sha256:8de30878634af163ae124c00383017893f9bbf3dada7cf7aa8dddaa4315d33e7

Observation 1535bde9-33fa-4139-9187-8d966b499f80 · outbound

This paper cites Reinforced Self-Training (.

AdaKP: Online Adaptive Knowledge-Point Selection for Reasoning-Oriented Reinforcement Learning Reinforced Self-Training (

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T07:34:56.525374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:34:56.525374Z digest=sha256:24fe2b1f37c5787677713c29ac466bcfb6555283926288e9d7252d23f5c83676

Observation 2a60d578-c1c2-4792-8300-0387304d60ce · outbound

This paper cites Transactions on Machine Learning Research (TMLR) , year =.

AdaKP: Online Adaptive Knowledge-Point Selection for Reasoning-Oriented Reinforcement Learning Transactions on Machine Learning Research (TMLR) , year =

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T07:34:56.646520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:34:56.646520Z digest=sha256:1f1f574160e2a1e051fba766c1c58941bf8718e0250383f24a1246d2507a6ed1

Observation 46f0c937-c1c1-42b7-b42e-7b9169cd431c · outbound

This paper cites Teaching Large Language Models to Reason with Reinforcement Learning.

AdaKP: Online Adaptive Knowledge-Point Selection for Reasoning-Oriented Reinforcement Learning Teaching Large Language Models to Reason with Reinforcement Learning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-01T07:34:56.790463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:34:56.790463Z digest=sha256:d7446500be9d3570e3cbd4033ed8fa2f2ea376ed64fcfb60e741f0bb81624c3c

Observation 389b8b2e-45e0-4354-bbc4-4edca3600dab · outbound

This paper cites 2016 , note =.

AdaKP: Online Adaptive Knowledge-Point Selection for Reasoning-Oriented Reinforcement Learning 2016 , note =

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T07:34:56.971940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:34:56.971940Z digest=sha256:741341c2d922c38aaa3af7b413a148264b65eb2b65c6f473c994be4e1dcce40d

Observation b79782fc-1867-4cda-847b-521e8cb4e170 · outbound

This paper cites International Conference on Learning Representations (ICLR) , year =.

AdaKP: Online Adaptive Knowledge-Point Selection for Reasoning-Oriented Reinforcement Learning International Conference on Learning Representations (ICLR) , year =

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-01T07:34:57.087681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:34:57.087681Z digest=sha256:0c0488239e131af44965bc2637f9b979f859cd6af2d97d7ec42363d64ab42bf9

Observation 294c8b0e-9dfd-4cb3-8a34-8b5fa2165ff2 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

AdaKP: Online Adaptive Knowledge-Point Selection for Reasoning-Oriented Reinforcement Learning Evaluating Large Language Models Trained on Code

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-01T07:34:57.162319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:34:57.162319Z digest=sha256:fa9d55112a8a36741e2959e3c590ec47e350944b60173be8b169f87c95fd4435

Observation f533e7ed-2f38-4c6f-8b49-9056c522fab1 · outbound

This paper cites an unresolved cited work.

AdaKP: Online Adaptive Knowledge-Point Selection for Reasoning-Oriented Reinforcement Learning Unresolved cited work

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-01T07:34:57.251159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:34:57.251159Z digest=sha256:6db35c049021229fda86ecafd55a2d135d678a8987439f85d3cd1f1cfab1dda7

Pith citing papers

No inbound Pith citation observations are available.