Pith. sign in

Paper Citation Record · LEDGER

Reasoning Fine-Tuning Induces Persistent Latent Policy States

As of 16 August 2026, this Paper Citation Record lists 81 of 81 outbound references and 0 inbound Pith citation observations for arXiv:2607.18532.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.18532 v1

Coverage vector

measured 81 of 81 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T15:16:21.881919Z

measured 81 of 81 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

81 of 81 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved80
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 33a8c4d3-bf7e-42a6-b333-a605efb0fb5e · outbound

This paper cites 2023 , eprint=.

Reasoning Fine-Tuning Induces Persistent Latent Policy States 2023 , eprint=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T15:16:13.575026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:16:13.575026Z digest=sha256:d9e57b7ccc094b920700880a4a052b19cea60fdff784d9caa7b4a477aaca4243

Observation 0237123c-6c4a-4e5f-adee-38562c24232e · outbound

This paper cites ICML 2025 Workshop on Reliable and Responsible Foundation Models , year=.

Reasoning Fine-Tuning Induces Persistent Latent Policy States ICML 2025 Workshop on Reliable and Responsible Foundation Models , year=

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T15:16:13.663724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:16:13.663724Z digest=sha256:2b98c1b32d987257873d2eb12971d50f2fbe7827452bdcad7d94a41985565cea

Observation cba94a23-644f-4f6c-9b5f-3a7de9259228 · outbound

This paper cites Proceedings of the 41st International Conference on Machine Learning , articleno =.

Reasoning Fine-Tuning Induces Persistent Latent Policy States Proceedings of the 41st International Conference on Machine Learning , articleno =

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T15:16:13.791427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:16:13.791427Z digest=sha256:50e2d4c968e6035d6307266c57b9b615bbc2f3574cb603bae53134159a54ce92

Observation 825cca83-bfc2-474e-8259-d4b1cd4b9007 · outbound

This paper cites 2026 , eprint=.

Reasoning Fine-Tuning Induces Persistent Latent Policy States 2026 , eprint=

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T15:16:14.005522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:16:14.005522Z digest=sha256:f2908dad77bb88c961767dabecb8f6446bc98009038904b5602b297d279d6751

Observation 9cae750f-2587-4ca1-88f3-e75583e151c8 · outbound

This paper cites The Thirty-ninth Annual Conference on Neural Information Processing Systems , year=.

Reasoning Fine-Tuning Induces Persistent Latent Policy States The Thirty-ninth Annual Conference on Neural Information Processing Systems , year=

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T15:16:14.134693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:16:14.134693Z digest=sha256:bb63522184ba7ab6ecb65e7382981910da9b058c9e43d58c1b9eb140c13ae31b

Observation 6728a5c1-0b1f-4253-97c4-046d8771266b · outbound

This paper cites 2025 , eprint=.

Reasoning Fine-Tuning Induces Persistent Latent Policy States 2025 , eprint=

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T15:16:14.242377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:16:14.242377Z digest=sha256:9718935db732baf014c315eda0cd3570fadbb23a5f6e8ae746e5cc54f9b123e9

Observation 8650d4b9-d380-4d2a-89a0-f46024c1e7d1 · outbound

This paper cites On the Limits of.

Reasoning Fine-Tuning Induces Persistent Latent Policy States On the Limits of

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T15:16:14.327297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:16:14.327297Z digest=sha256:3ea295d3ed064dd8c8eb3c2d9776938c86fe7ea370f1764b3a71f9c56a1b8821

Observation 467d64c5-f2cd-4572-a170-e50cd354927b · outbound

This paper cites 2025 , eprint=.

Reasoning Fine-Tuning Induces Persistent Latent Policy States 2025 , eprint=

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T15:16:14.406132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:16:14.406132Z digest=sha256:53fb43d2246aa9025f9a03bce15beb98061fffd3ef768405901db695b59b55b3

Observation 042f060e-c17d-44f4-939f-c7e173964bee · outbound

This paper cites 2025 , eprint=.

Reasoning Fine-Tuning Induces Persistent Latent Policy States 2025 , eprint=

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T15:16:14.471044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:16:14.471044Z digest=sha256:7a7b5a182bc600a7f9f36f8eaad6a54351c0e3f7b791f67c2e76b263a85695bc

Observation b27331d8-64f4-47b5-b859-f09b6d073b63 · outbound

This paper cites 2025 , eprint=.

Reasoning Fine-Tuning Induces Persistent Latent Policy States 2025 , eprint=

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T15:16:14.598063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:16:14.598063Z digest=sha256:879fbeebfbadb4d0be05293c72063ee02d1ed25d8f7e8b2ef3d6b5bfa2eaea84

Observation 59874cf3-8d33-4391-937f-60ac25e2127c · outbound

This paper cites 2025 , eprint=.

Reasoning Fine-Tuning Induces Persistent Latent Policy States 2025 , eprint=

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T15:16:14.677547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:16:14.677547Z digest=sha256:4ff902f982bba50139b94d3a93738bfd16af7ff0fd190edb5b2a09ad6bfdebae

Observation ab57c9b8-672b-4e80-8a1e-80b946677828 · outbound

This paper cites 2025 , eprint=.

Reasoning Fine-Tuning Induces Persistent Latent Policy States 2025 , eprint=

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T15:16:14.747523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:16:14.747523Z digest=sha256:231b5000fa1b888dd8566acace9b767fcc86f77d7bd1138325c42d1204574f32

Observation 3ea5a647-9542-4900-a049-97d160b676a9 · outbound

This paper cites an unresolved cited work.

Reasoning Fine-Tuning Induces Persistent Latent Policy States Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T15:16:14.815561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:16:14.815561Z digest=sha256:ae1d8634d22fb70804a227ca61c5046851e4edf1b54ce5f4b65b660738f8f111

Observation 65c73470-88b7-444d-8548-0bc050766d3f · outbound

This paper cites and Spragins, John D.

Reasoning Fine-Tuning Induces Persistent Latent Policy States and Spragins, John D

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T15:16:15.019700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:16:15.019700Z digest=sha256:220c104e6d45685e177ec8672e469318376609ca9a304963e4295bd90c80b0c1

Observation 481ad52a-d756-478c-8517-25952180d99c · outbound

This paper cites On the Direction of.

Reasoning Fine-Tuning Induces Persistent Latent Policy States On the Direction of

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T15:16:15.094412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:16:15.094412Z digest=sha256:f70e39cecb360d157ce5ced9c53924f0901e91849e8e3abfc3b97cfba340933b

Observation 57f14358-b1f0-4f17-9a2d-a3fce4a535b9 · outbound

This paper cites 2026 , eprint=.

Reasoning Fine-Tuning Induces Persistent Latent Policy States 2026 , eprint=

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T15:16:15.262132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:16:15.262132Z digest=sha256:8f7e44a36e9501d08af3388e662f8247a708c2bc7b26dd7df4a012a53d510b4c

Observation a4ddcdd1-beb6-48fe-8f41-19a0e287696b · outbound

This paper cites 2025 , eprint=.

Reasoning Fine-Tuning Induces Persistent Latent Policy States 2025 , eprint=

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T15:16:15.356343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:16:15.356343Z digest=sha256:185ca7b43b6862158381a6920333e6c479446a1443080b96e7d411a3da779107

Observation 1ce70e93-0578-4921-87ec-31d8686c4f5e · outbound

This paper cites and Fu, K.

Reasoning Fine-Tuning Induces Persistent Latent Policy States and Fu, K

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T15:16:15.450612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:16:15.450612Z digest=sha256:4e249ddf0bef7bfa0c78fcda11fe80399db0557f8eb56500856143cb5713b13c

Observation ee69454c-69e7-4093-9127-01f9b8345d97 · outbound

This paper cites Neural computation , volume=.

Reasoning Fine-Tuning Induces Persistent Latent Policy States Neural computation , volume=

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T15:16:15.533323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:16:15.533323Z digest=sha256:a7f1effc0fd61ee70b79736778689a3f8a8aaef7092596b45035115db352c266

Observation 4395fd2e-2538-4e89-838d-91a570e3e027 · outbound

This paper cites 2017 , editor =.

Reasoning Fine-Tuning Induces Persistent Latent Policy States 2017 , editor =

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T15:16:15.610158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:16:15.610158Z digest=sha256:95df7fd6849636e381c25af16661eaa77e9def8dfb17f54a6407d7201d380d14

Observation cd2e82d8-aea0-439d-ab6b-df7d719fcfff · outbound

This paper cites 2025 , eprint=.

Reasoning Fine-Tuning Induces Persistent Latent Policy States 2025 , eprint=

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T15:16:15.678014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:16:15.678014Z digest=sha256:794db91c3db5a2dac2b9ba7b10c07595eada27f14a6b2479dd96b0c3ec9d2bc9

Observation 54220faa-9214-465a-8618-85660cb82660 · outbound

This paper cites The Fourteenth International Conference on Learning Representations , year=.

Reasoning Fine-Tuning Induces Persistent Latent Policy States The Fourteenth International Conference on Learning Representations , year=

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T15:16:15.801597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:16:15.801597Z digest=sha256:996fb123985654f155a338216e5716f28d684d81d7d2885d0a40a5e0fad6b422

Observation 2b926bd8-6e17-461e-bdab-6e899cd33315 · outbound

This paper cites Mechanistic Interpretability Workshop at NeurIPS 2025 , year=.

Reasoning Fine-Tuning Induces Persistent Latent Policy States Mechanistic Interpretability Workshop at NeurIPS 2025 , year=

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T15:16:15.980313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:16:15.980313Z digest=sha256:99e163679fe317567ab7918be009d04223e30ac447746505456d4e4e8bf4367c

Observation 630c407c-a66f-4480-b977-196daf945a96 · outbound

This paper cites Towards Understanding Fine-Tuning Mechanisms of.

Reasoning Fine-Tuning Induces Persistent Latent Policy States Towards Understanding Fine-Tuning Mechanisms of

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T15:16:16.039869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:16:16.039869Z digest=sha256:d708184f5dce8850b354d8a6f4bceb71ce7a2eb8c0e6eda3267d1f88b13a0c72

Observation 9e5e1ac0-cb93-4dab-8106-9df9f2059402 · outbound

This paper cites ICLR 2025 Workshop on Building Trust in Language Models and Applications , year=.

Reasoning Fine-Tuning Induces Persistent Latent Policy States ICLR 2025 Workshop on Building Trust in Language Models and Applications , year=

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T15:16:16.144526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:16:16.144526Z digest=sha256:026767746dec945a54ec160ee4dee8ef6dc89584cfa3bbb4c16069c74eef4ec7

Observation d296e457-83f4-4039-a39f-d774b220f5b7 · outbound

This paper cites Demystifying Reasoning Dynamics with Mutual Information: Thinking Tokens are Information Peaks in.

Reasoning Fine-Tuning Induces Persistent Latent Policy States Demystifying Reasoning Dynamics with Mutual Information: Thinking Tokens are Information Peaks in

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T15:16:16.228601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:16:16.228601Z digest=sha256:a96fb7636df72eb5d2e2fcfaef4849377543f46fc9f35e8f517e55848d406d97

Observation 50574a29-ae63-4fe4-bb1e-ffc014cf2002 · outbound

This paper cites 2025 , eprint=.

Reasoning Fine-Tuning Induces Persistent Latent Policy States 2025 , eprint=

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T15:16:16.343248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:16:16.343248Z digest=sha256:17124cff9d9a0f658753010e7e53d551435cf777bcce5adff05578825d5ed860

Observation 54b48bc6-6e46-434f-89c6-f4365975738d · outbound

This paper cites Workshop on Reasoning and Planning for Large Language Models , year=.

Reasoning Fine-Tuning Induces Persistent Latent Policy States Workshop on Reasoning and Planning for Large Language Models , year=

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T15:16:16.413360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:16:16.413360Z digest=sha256:dd06766a94ac32d9b2422ff0e33f673d6c7de239b7e136e06bcb72ef55fd7543

Observation da6c3141-549f-43aa-8ee2-91ea6bfdbea9 · outbound

This paper cites Measuring Chain of Thought Faithfulness by Unlearning Reasoning Steps.

Reasoning Fine-Tuning Induces Persistent Latent Policy States Measuring Chain of Thought Faithfulness by Unlearning Reasoning Steps

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T15:16:16.520411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:16:16.520411Z digest=sha256:57198e6cb8d622b4216f3becf98a30f75a34b4df055fcc14c5e7994f9d86c19f

Observation a6fa5c60-3bb7-4993-83f7-8ddb32e3a250 · outbound

This paper cites 2021 , eprint=.

Reasoning Fine-Tuning Induces Persistent Latent Policy States 2021 , eprint=

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T15:16:16.643332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:16:16.643332Z digest=sha256:07714bff9fa38ca6ce8a4e37fd0a137a7554a14a9b3e160546309ea97f90645a

Observation f8116245-94a6-411a-bc09-21ba5fa65f8c · outbound

This paper cites 2024 , eprint=.

Reasoning Fine-Tuning Induces Persistent Latent Policy States 2024 , eprint=

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T15:16:16.866739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:16:16.866739Z digest=sha256:3604fa62a4cc7c4097c72b7b88e58715de04020ced05ffd21bbd6671204f80fc

Observation 53c5c773-b076-460e-994b-3049f8fac1f0 · outbound

This paper cites 2024 , url =.

Reasoning Fine-Tuning Induces Persistent Latent Policy States 2024 , url =

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T15:16:16.981690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:16:16.981690Z digest=sha256:6b77c8b894daeb22b88797ef989201f8ddca8b3c7cae8f1efa74cd7e4d27f080

Observation 61f7b793-bbd8-4996-afd4-85a667aad898 · outbound

This paper cites QwQ-32B: Embracing the Power of Reinforcement Learning , url =.

Reasoning Fine-Tuning Induces Persistent Latent Policy States QwQ-32B: Embracing the Power of Reinforcement Learning , url =

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T15:16:17.452678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:16:17.452678Z digest=sha256:b7e7eb1afd1f175babe9557eb1bf1d7be2e793ee9c8ee3843a134b5b0c306400

Observation 9d0feb81-2050-4136-93cb-a57b02161405 · outbound

This paper cites 2025 , eprint=.

Reasoning Fine-Tuning Induces Persistent Latent Policy States 2025 , eprint=

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-01T15:16:17.707886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:16:17.707886Z digest=sha256:dc07e1f6736139be165e6f58d1491e20277c81df88f841ea9676e2d44d200176

Observation f5990075-718f-479f-9710-f153f3c414d7 · outbound

This paper cites 2025 , eprint=.

Reasoning Fine-Tuning Induces Persistent Latent Policy States 2025 , eprint=

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T15:16:17.816565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:16:17.816565Z digest=sha256:aab8e4c7d7c5293eb1c82d5234f79fcbaf5d64b4bebf4d1ba157207172d797c7

Observation fee474c1-646b-46b4-b2c0-a875375e7dd2 · outbound

This paper cites 2025 , eprint=.

Reasoning Fine-Tuning Induces Persistent Latent Policy States 2025 , eprint=

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-01T15:16:17.937789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:16:17.937789Z digest=sha256:8b289817cc849311423c5dccb0cd0781d55430d4a67d73bfa0eb87954d62a185

Observation 9f8cd739-d4a3-4d91-80f4-3e502332c454 · outbound

This paper cites 2024 , eprint=.

Reasoning Fine-Tuning Induces Persistent Latent Policy States 2024 , eprint=

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-01T15:16:18.051644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:16:18.051644Z digest=sha256:064bb74487f93797d83b847dfb992dfde233b4196b08719b5e15819d70830424

Observation c1a7d186-5f2b-4244-86a1-8d38051610ce · outbound

This paper cites 2024 , month = aug, journal =.

Reasoning Fine-Tuning Induces Persistent Latent Policy States 2024 , month = aug, journal =

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-01T15:16:18.171431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:16:18.171431Z digest=sha256:8c262d82195fb6d93cdc839feec1858d09b1cb814ec9d3989c8053e1f961982e

Observation 225151fe-8573-4ab5-b0f2-7e6505fbcb51 · outbound

This paper cites 2026 , url=.

Reasoning Fine-Tuning Induces Persistent Latent Policy States 2026 , url=

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-01T15:16:18.289118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:16:18.289118Z digest=sha256:34abceda0198ad824a791c4207a704c11caa90ac2b800940563704921899b4ad

Observation 428c6fe0-8b92-49d3-8936-030695d31b1f · outbound

This paper cites 2025 , url=.

Reasoning Fine-Tuning Induces Persistent Latent Policy States 2025 , url=

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-01T15:16:18.387352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:16:18.387352Z digest=sha256:9c04b96e4a577e80a0190f4b27eb703154a3897b9d9d5af1d1f1022bc2d4058e

Observation 7bb7bf73-2cc3-4f78-9efc-07ade3fe236f · outbound

This paper cites Ackerson and K.

Reasoning Fine-Tuning Induces Persistent Latent Policy States Ackerson and K

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-01T15:16:18.507344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:16:18.507344Z digest=sha256:2517c68c06796b75e02d0ec127d4ba50f313bc18691e96927680d09666fe7188

Observation e88f86d1-e470-4804-b859-fb76df477f91 · outbound

This paper cites Llama 3 model card.

Reasoning Fine-Tuning Induces Persistent Latent Policy States Llama 3 model card

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-01T15:16:18.625152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:16:18.625152Z digest=sha256:bdfe6eb54141c4f4861717878ff6d0a930e5cd5750d369ec2452e6e282ff9f38

Observation 44f25d29-4d44-4d87-b9ae-14a3052e1a72 · outbound

This paper cites Qwen Technical Report.

Reasoning Fine-Tuning Induces Persistent Latent Policy States Qwen Technical Report

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-01T15:16:18.727168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:16:18.727168Z digest=sha256:72214e9f0d87255ab8b296683c084cac817034770768310a04d16832e321b93e

Observation 3cbf0cc6-f2e8-4005-85a6-b52cc9a20e8a · outbound

This paper cites On the identifiability of switching dynamical systems.

Reasoning Fine-Tuning Induces Persistent Latent Policy States On the identifiability of switching dynamical systems

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-01T15:16:18.872972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:16:18.872972Z digest=sha256:68365175870cb2d2bffd1b11366859bc5d919d9b4536f0d957627458c652c7ee

Observation 5218849b-2c28-4c8b-b8ee-efb2367ce13e · outbound

This paper cites Bogdan, Uzay Macar, Neel Nanda, and Arthur Conmy.

Reasoning Fine-Tuning Induces Persistent Latent Policy States Bogdan, Uzay Macar, Neel Nanda, and Arthur Conmy

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-01T15:16:19.017904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:16:19.017904Z digest=sha256:c346d259505c15b4fc0b93e54c61c58e81f06832497907fc5c78f6c90b18c031

Observation 20f4bf1c-0488-4917-9586-2376a3cff0d4 · outbound

This paper cites A statistical physics of language model reasoning.

Reasoning Fine-Tuning Induces Persistent Latent Policy States A statistical physics of language model reasoning

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-01T15:16:19.135142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:16:19.135142Z digest=sha256:df239cd9889f04ee822e79c6acf4342f26ccf2ea1753fc843c49852d38ea1caf

Observation 512e6320-024c-4caa-8445-8b1cf574c0ab · outbound

This paper cites SEAL : Steerable reasoning calibration of large language models for free.

Reasoning Fine-Tuning Induces Persistent Latent Policy States SEAL : Steerable reasoning calibration of large language models for free

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-01T15:16:19.233415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:16:19.233415Z digest=sha256:30d067a83ffc4a1ca0bb1b03883fcf48631c964e0fa9fcfe84ee4e9c4f51101e

Observation ace043ea-d3d4-43c4-88c9-7942c4c9783b · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Reasoning Fine-Tuning Induces Persistent Latent Policy States Training Verifiers to Solve Math Word Problems

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-01T15:16:19.284845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:16:19.284845Z digest=sha256:2da715d2837ba7b16b9f9c2198105793e8f75fd34571a752f0092989fbbe88b9

Observation 6b08b06a-9f2c-4e64-af04-aab8727a9034 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Reasoning Fine-Tuning Induces Persistent Latent Policy States DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-01T15:16:19.344259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:16:19.344259Z digest=sha256:3e276db59be0569ca3a1cd1bfaaa3d6ec410efbf3cb708be830ffdd1f97e079a

Observation 70284a98-88aa-4380-a7e5-48c4bca322a4 · outbound

This paper cites Asymptotic properties of the maximum likelihood estimator in autoregressive models with markov regime.

Reasoning Fine-Tuning Induces Persistent Latent Policy States Asymptotic properties of the maximum likelihood estimator in autoregressive models with markov regime

Reference 60

Resolution
verified exact
doi, observed 2026-08-01T15:18:36.075659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-01T15:16:19.438599Z digest=sha256:0068b80716884dbf77d56041d42026233e5f1560611f71e63cae37f2fa22f0c4

Observation 77f71155-c3a6-4da7-8e08-3e46fc698976 · outbound

This paper cites Variational learning for switching state-space models.

Reasoning Fine-Tuning Induces Persistent Latent Policy States Variational learning for switching state-space models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-01T15:16:19.513576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:16:19.513576Z digest=sha256:5f043aebcfe6ae06a65ef7e25a8f13b0fedced5fe540e03263a6ba5e7d92e30c

Observation 5eaa8c9a-8439-4081-ae5c-a8f6bb238749 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Reasoning Fine-Tuning Induces Persistent Latent Policy States Measuring Mathematical Problem Solving With the MATH Dataset

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-01T15:16:19.587543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:16:19.587543Z digest=sha256:d30b29c640afcc2347f806003c5000078644561feee780611ada003e08b37472

Observation 3210d02a-0435-42ce-ac1e-c12e06c530c9 · outbound

This paper cites On the direction of RLVR updates for LLM reasoning: Identification and exploitation.

Reasoning Fine-Tuning Induces Persistent Latent Policy States On the direction of RLVR updates for LLM reasoning: Identification and exploitation

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-01T15:16:19.658340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:16:19.658340Z digest=sha256:78d0f3405a4e89cdc8420d7e07c7563962b47f47bbaf69a8ca1e01d88cfe1be2

Observation 097ff543-1f93-46fe-aa8e-85ee20435285 · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

Reasoning Fine-Tuning Induces Persistent Latent Policy States Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-01T15:16:19.738268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:16:19.738268Z digest=sha256:486f695ff6e5e1d62f2dd7d56d4f1bba39d4c004e56a9c7e566b3acc0a99cc29

Observation 2f6276b5-af19-497b-ba48-1973f23e5736 · outbound

This paper cites The CoT Encyclopedia: Analyzing, Predicting, and Controlling how a Reasoning Model will Think.

Reasoning Fine-Tuning Induces Persistent Latent Policy States The CoT Encyclopedia: Analyzing, Predicting, and Controlling how a Reasoning Model will Think

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-01T15:16:19.811207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:16:19.811207Z digest=sha256:898f498e392f338557196ae4879d221bd1fc980ddf9cba63e7f8928c17e2c56e

Observation f94be62e-2ada-4fe9-8fb0-fe73221fb720 · outbound

This paper cites Clue: Non-parametric verification from experience via hidden-state clustering.

Reasoning Fine-Tuning Induces Persistent Latent Policy States Clue: Non-parametric verification from experience via hidden-state clustering

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-01T15:16:19.877258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:16:19.877258Z digest=sha256:35b47de5963d65558955b65270e00e7d01d869baaa9625015ed5f4708eb6a2e5

Observation 148555f6-d956-4d39-8031-a158ce7cf927 · outbound

This paper cites Bayesian Learning and Inference in Recurrent Switching Linear Dynamical Systems.

Reasoning Fine-Tuning Induces Persistent Latent Policy States Bayesian Learning and Inference in Recurrent Switching Linear Dynamical Systems

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-01T15:16:19.948636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:16:19.948636Z digest=sha256:f183c9b6d472672e8b5a6d04ae56d35331fb609042e8a832e80881e9d9348b58

Observation 9a20706a-5886-4b1f-98de-690212d3e067 · outbound

This paper cites Learning a generative meta-model of llm activations, 2026.

Reasoning Fine-Tuning Induces Persistent Latent Policy States Learning a generative meta-model of llm activations, 2026

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-01T15:16:20.033854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:16:20.033854Z digest=sha256:9ebb72504669ae24738a8caa7387c1bfbe1d62e0c678aeb8d320f6aacbc3e3a6

Observation f5477018-f50b-47dc-8d13-76beba001d38 · outbound

This paper cites Thought Branches: Interpreting LLM Reasoning Requires Resampling.

Reasoning Fine-Tuning Induces Persistent Latent Policy States Thought Branches: Interpreting LLM Reasoning Requires Resampling

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-01T15:16:20.107692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:16:20.107692Z digest=sha256:8e889943f1a9788ad3685c6a01a8c82f03a4dc48ba82dd70ae631fd202f28287

Observation eb1d2b8b-99c5-40e1-8dfc-91c1afe523c8 · outbound

This paper cites Narrow finetuning leaves clearly readable traces in the activation differences.

Reasoning Fine-Tuning Induces Persistent Latent Policy States Narrow finetuning leaves clearly readable traces in the activation differences

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-01T15:16:20.156420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:16:20.156420Z digest=sha256:73e445714bce00548c1056c22ecb69844c9be8e23469702f530cd00134f26918

Observation a7fada9f-7740-45c3-aa0a-411da92ede04 · outbound

This paper cites Diab, Virginia Smith, and Dawn Song.

Reasoning Fine-Tuning Induces Persistent Latent Policy States Diab, Virginia Smith, and Dawn Song

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-01T15:16:20.213159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:16:20.213159Z digest=sha256:3e22552462cb9cac0b031a85a6e8c51b6e26ad3179735639464f2649ff855b63

Observation 05a6d9c7-35b3-44c4-89bc-836597cd8b16 · outbound

This paper cites La, Duy M.

Reasoning Fine-Tuning Induces Persistent Latent Policy States La, Duy M

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-01T15:16:20.267179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:16:20.267179Z digest=sha256:8f377b5097dabff0fafbbc7164e43867d6df612c45b0991887a8566f3bf2c978

Observation 1029cfcf-3514-4aa6-a2ee-6900768ac5b8 · outbound

This paper cites an unresolved cited work.

Reasoning Fine-Tuning Induces Persistent Latent Policy States Unresolved cited work

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-01T15:16:20.324657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:16:20.324657Z digest=sha256:f66763012cee957c40f173cbcc9529af3e4c3c6e2b2dd6e2ab24aa9f634f278f

Observation 105d9ba2-7d4e-4233-a7dc-a4bf246b3867 · outbound

This paper cites Demystifying reasoning dynamics with mutual information: Thinking tokens are information peaks in LLM reasoning.

Reasoning Fine-Tuning Induces Persistent Latent Policy States Demystifying reasoning dynamics with mutual information: Thinking tokens are information peaks in LLM reasoning

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-01T15:16:20.400811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:16:20.400811Z digest=sha256:dfd5aa8e68e1d2aeaa76d19f27beac9658678a21fe5020b106559d7d9ae5cf25

Observation ef28362d-1aac-4f39-9ed9-9a994a4888a0 · outbound

This paper cites Learnable latent embeddings for joint behavioural and neural analysis.

Reasoning Fine-Tuning Induces Persistent Latent Policy States Learnable latent embeddings for joint behavioural and neural analysis

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-01T15:16:20.504062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:16:20.504062Z digest=sha256:b9364b758185877146a9b19dd44eda3efea69a791d668b49eb56a02f9c5db540

Observation b5990ec5-5c43-40ac-a081-b8e8e931a8ec · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Reasoning Fine-Tuning Induces Persistent Latent Policy States DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-01T15:16:20.617290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:16:20.617290Z digest=sha256:3057f4da92ba1292abead2d1876d20a2975299a1dbcc2f5e358df7f73fe37e66

Observation abc2bec8-ab5a-4f28-be14-fd6ae8b393ff · outbound

This paper cites Understanding reasoning in thinking language models via steering vectors.

Reasoning Fine-Tuning Induces Persistent Latent Policy States Understanding reasoning in thinking language models via steering vectors

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-01T15:16:20.713420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:16:20.713420Z digest=sha256:dd3da003882178560c43d393b15a9af7ce49bf9f959757e2178eb57652196af3

Observation de19a86b-b3f8-46dc-a0c7-27e2c4d16e13 · outbound

This paper cites Base Models Know How to Reason, Thinking Models Learn When.

Reasoning Fine-Tuning Induces Persistent Latent Policy States Base Models Know How to Reason, Thinking Models Learn When

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-01T15:16:20.782109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:16:20.782109Z digest=sha256:1fae3b7314c2b81dc9aa9b7c46418b6d46a4e31a164e57e6a7b950a70fee10c5

Observation 68b545c0-b727-40d3-8cf8-685a9d3bdbde · outbound

This paper cites Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning.

Reasoning Fine-Tuning Induces Persistent Latent Policy States Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-01T15:16:20.893289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:16:20.893289Z digest=sha256:573d832979a2befd6caf5855913e57b76f1a76eb7c6769691d1b7f8781d14582

Observation 088cf89d-45ac-41fc-b511-ae95e7b496a1 · outbound

This paper cites Towards understanding fine-tuning mechanisms of LLM s via circuit analysis.

Reasoning Fine-Tuning Induces Persistent Latent Policy States Towards understanding fine-tuning mechanisms of LLM s via circuit analysis

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-01T15:16:20.964020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:16:20.964020Z digest=sha256:65f3a579e07ba1351595b4728217fc53327c01c7e885ebf54465c9acd840ed94

Observation ef769a69-4967-46e1-9560-0f82b6e00b25 · outbound

This paper cites MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark.

Reasoning Fine-Tuning Induces Persistent Latent Policy States MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-01T15:16:21.070905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:16:21.070905Z digest=sha256:ec08fb1f0b8747d7d0408f0681aa48fef01228292eadb5b3d1bfebee4a3d55fd

Observation d1b1a916-9964-45b0-91f3-c4a1b61dcfc8 · outbound

This paper cites Reasoning-Finetuning Repurposes Latent Representations in Base Models.

Reasoning Fine-Tuning Induces Persistent Latent Policy States Reasoning-Finetuning Repurposes Latent Representations in Base Models

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-01T15:16:21.171395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:16:21.171395Z digest=sha256:e889316b9984b22d6b94fc603496a3ca691f179a450a0e16058b6fad55c835fe

Observation a095c327-b401-4bce-ad08-19f547317a1d · outbound

This paper cites Rank-1 loras encode interpretable reasoning signals.

Reasoning Fine-Tuning Induces Persistent Latent Policy States Rank-1 loras encode interpretable reasoning signals

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-01T15:16:21.241977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:16:21.241977Z digest=sha256:f85712abe64c286d6118ca26f97a888bc7d1ea9e3cda8e04228a07179941fd25

Observation 757871bd-ffcf-445d-95be-d910dad14927 · outbound

This paper cites On the limits of RLVR : Support, entropy, and the illusion of reasoning.

Reasoning Fine-Tuning Induces Persistent Latent Policy States On the limits of RLVR : Support, entropy, and the illusion of reasoning

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-01T15:16:21.323304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:16:21.323304Z digest=sha256:113bb796e1ffb9bf724ed4297a071dbd1994c152af3a35a43fda8155db4964b9

Observation 0548a8f8-b974-4f50-87f4-108e828d447e · outbound

This paper cites CTRLS : Chain-of-thought reasoning via latent state transition.

Reasoning Fine-Tuning Induces Persistent Latent Policy States CTRLS : Chain-of-thought reasoning via latent state transition

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-01T15:16:21.396068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:16:21.396068Z digest=sha256:2642db40bf9b6fe419b2953931aa39b11898c024d49780db1624888332e1098e

Observation e79d1087-996e-49e7-b07e-f4815b7116d5 · outbound

This paper cites Yakowitz and John D.

Reasoning Fine-Tuning Induces Persistent Latent Policy States Yakowitz and John D

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-01T15:16:21.468535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:16:21.468535Z digest=sha256:7a9022388510d29eb4e076464544ce82ca4f1f9028a03777c44b070ff0e37975

Observation 11e2a532-030f-4bcb-959e-fd2c96be1b83 · outbound

This paper cites Qwen2 Technical Report.

Reasoning Fine-Tuning Induces Persistent Latent Policy States Qwen2 Technical Report

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-01T15:16:21.564657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:16:21.564657Z digest=sha256:ad1ab83f9d78b362185a36c16604b4c349f6308aa90c133e8db00788ca744514

Observation 8d4e677c-f82b-45cf-8808-8cf8d0623749 · outbound

This paper cites Qwen2.5 Technical Report.

Reasoning Fine-Tuning Induces Persistent Latent Policy States Qwen2.5 Technical Report

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-01T15:16:21.608057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:16:21.608057Z digest=sha256:771a473e10ccb89698b7f2fdcece50a74c30e0658605f1f2242aad4f87e89e82

Observation 14968147-d138-4ad7-9bf1-d7be20e534cb · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?.

Reasoning Fine-Tuning Induces Persistent Latent Policy States Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-01T15:16:21.697846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:16:21.697846Z digest=sha256:682753a6a68925c303eeea1bec5c1108783c302be8ed3ad6bca9bdcc7f8c307b

Observation 5e8206db-ed4d-4f84-bf78-8b65574e8cf3 · outbound

This paper cites 1.4 Million Open-Source Distilled Reasoning Dataset to Empower Large Language Model Training.

Reasoning Fine-Tuning Induces Persistent Latent Policy States 1.4 Million Open-Source Distilled Reasoning Dataset to Empower Large Language Model Training

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-01T15:16:21.796399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:16:21.796399Z digest=sha256:9d55ec505acb4f23c28d6f7ce6e302c7f0ff1783d215013af86183a7dfd90c2b

Observation f0555707-067a-4fbd-964d-8e4fd93601b9 · outbound

This paper cites The path not taken: Rlvr provably learns off the principals.

Reasoning Fine-Tuning Induces Persistent Latent Policy States The path not taken: Rlvr provably learns off the principals

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-01T15:16:21.881919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:16:21.881919Z digest=sha256:59f8b92dd774417351aa77dba244128ec6993646551effc12c8b175091579919

Pith citing papers

No inbound Pith citation observations are available.