Pith. sign in

Paper Citation Record · LEDGER

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs

As of 4 August 2026, this Paper Citation Record lists 80 of 80 outbound references and 1 inbound Pith citation observation for arXiv:2605.28388.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.28388 v1

Coverage vector

measured 80 of 80 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-29T12:11:00.402276Z

measured 81 of 81 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-03T06:30:56.289259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-30T07:19:00.012926Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-06-30T07:24:21.458766Z

Reference resolution

80 of 80 outbound references displayed

  • verified exact30
  • verified fuzzy0
  • unresolved47
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fab16154-6f0c-4db4-a2c9-003d60c9dece · outbound

This paper cites Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs.

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-06-29T12:13:26.556805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:0b9cbe6c7aa15f7bb2965ef5e714a6ad9d7a6722f9b8fc8fddc03ea7c582f80e

Observation 8f908df6-1388-49e4-9ebc-5ee16e0550c3 · outbound

This paper cites Online difficulty filtering for reasoning oriented reinforcement learning.

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs Online difficulty filtering for reasoning oriented reinforcement learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-29T12:11:00.402276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:34d05f35d037c62e55234c53f1daff2e37baa791bf749c6725d417792e4e89d8

Observation 811a0bf9-f40a-47e3-83bc-1897ca42416d · outbound

This paper cites Curriculum learning.

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs Curriculum learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-29T12:11:00.402276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:a8e03def3e5324ddeba0e3a48fcce23cc10c39bee94ecbb8c393636ac748d5b0

Observation 2503259f-b4a7-43d2-aa49-00eb34bf1e91 · outbound

This paper cites Temporal sparse autoencoders: Leveraging the sequential nature of language for interpretability.

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs Temporal sparse autoencoders: Leveraging the sequential nature of language for interpretability

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-29T12:11:00.402276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:b5874cc536ffdb4f63a073f4db1a119d6c88101d24d0b07c88d65668a4f9fd3d

Observation 97074281-0cc8-47ca-9187-5781ab93cb35 · outbound

This paper cites Hansen and Duo Peng and Yuhui Zhang and Alejandro Lozano and Min Woo Sun and Emma Lundberg and Serena Yeung-Levy , year=.

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs Hansen and Duo Peng and Yuhui Zhang and Alejandro Lozano and Min Woo Sun and Emma Lundberg and Serena Yeung-Levy , year=

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-06-29T12:13:26.572974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:0b1b2889620848af4c2f6985d695a42e0ecfdcc0b5e641e20cd15dba3dfd7b2f

Observation b3d2478d-512b-4561-9eea-9c11b551aafe · outbound

This paper cites Process Reinforcement through Implicit Rewards.

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs Process Reinforcement through Implicit Rewards

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-06-29T12:13:26.606024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:b5cdf23504a1dee96d6fc4c341d6d2e3e8b03bb14da9d307d03e0dc5594a4257

Observation 4af94b0d-0415-4649-8264-9320efb95736 · outbound

This paper cites Deep Think with Confidence.

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs Deep Think with Confidence

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-06-29T12:13:26.579988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:e0f5685c221803a1447b15c07624624d583a5ba97069f49be23ae6fd1aec1071

Observation e379eb8a-8ec5-433e-93e6-f4dd04041912 · outbound

This paper cites I have covered all the bases here: Interpreting reasoning features in large language models via sparse autoencoders.

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs I have covered all the bases here: Interpreting reasoning features in large language models via sparse autoencoders

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-29T12:11:00.402276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:6e3b1e7858482bb34eceb5c90506ff9b8f428aaf5a0ac95e55cd4f8403deebb4

Observation 10a882b1-b1ed-4733-8ee0-e1a7b1d76650 · outbound

This paper cites Scaling and evaluating sparse autoencoders.

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs Scaling and evaluating sparse autoencoders

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-29T12:11:00.402276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:9be721bf6aef2e28fe2ebcb270de6ffea07c53b20d6bd010f53341287b92aed3

Observation aa457bc8-42cb-4367-aaa6-baba872204b7 · outbound

This paper cites Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems.

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-29T12:11:00.402276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:fda326a95e4376aa88d496ea6827d75da88ba76646f4e24ba87a55e6c1739ce0

Observation 6f13bddb-134b-4e26-9029-79f28e8ba273 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs Measuring Mathematical Problem Solving With the MATH Dataset

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-06-29T12:13:26.568125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:cfeb90b351d004ef4846df1bc0a49781e410e67181369bce35d95942f2de514e

Observation d9ea390b-d186-4cc7-8ebd-be33fdad455c · outbound

This paper cites REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization.

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-06-29T12:13:26.528934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:ebdee4d0a33388714bc9f5fede021d4f539be9b62e1d5b1c49674c200269017c

Observation ea7dbd3e-77ad-400a-9ef3-94050c5b047c · outbound

This paper cites Sparse autoencoders find highly interpretable features in language models.

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs Sparse autoencoders find highly interpretable features in language models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-29T12:11:00.402276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:629394786454cfe708ae9141bc34b7c15af785de2c2f5b1fb87bca0e792b3123

Observation 0bbe7b15-9844-46b8-8d32-3be1876227e0 · outbound

This paper cites Vcrl: Variance-based curriculum reinforcement learning for large language models.

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs Vcrl: Variance-based curriculum reinforcement learning for large language models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-29T12:13:26.570523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:721da3a5bc83a72bc7e95907ba77bd8eba4b9e676c554b89a728406b78aeb47a

Observation a61e86aa-afd9-4c73-a3b9-d23b90cc6996 · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T12:13:26.597014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:da9c48f2c5b1121a68a9dca27785058170d9fa25688f8e53303fcbc3e3c6553a

Observation b20825b9-6f6c-4197-a99b-a4ef5a8c4463 · outbound

This paper cites Le, Myeongho Jeon, Kim Vu, Viet Dac Lai, and Eunho Yang.

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs Le, Myeongho Jeon, Kim Vu, Viet Dac Lai, and Eunho Yang

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-29T12:11:00.402276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:ac456fabe5c03d34bcd9753b6a9bf2722771ea9ec17a2de2b70b07c6f7c334e7

Observation 1217f1ce-52bb-4cdb-a40f-37884b290287 · outbound

This paper cites Solving quan- titative reasoning problems with language models.Advances in neural information processing systems, 35:3843–3857, 2022.

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs Solving quan- titative reasoning problems with language models.Advances in neural information processing systems, 35:3843–3857, 2022

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-29T12:11:00.402276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:e85c37deac73426686a76d3ad2534bf41c344d29c666fca56fd31e3f1c642672

Observation b828254b-15b1-4e9b-9e49-a67b1fe35a71 · outbound

This paper cites Numinamath: The largest public dataset in ai4maths with 860k pairs of competition math problems and solutions.Hugging Face repository, 2024.

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs Numinamath: The largest public dataset in ai4maths with 860k pairs of competition math problems and solutions.Hugging Face repository, 2024

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-29T12:11:00.402276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:179d121137ee3699b3a6caddff02601e625c4a59e0ae1284b7b512e7bcfe4d92

Observation 92fb137b-6598-4706-8eba-adf896fa57cc · outbound

This paper cites QuestA: Expanding reasoning capacity in LLMs via question augmentation.

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs QuestA: Expanding reasoning capacity in LLMs via question augmentation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-29T12:11:00.402276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:07c3a935d9cc98f4413babee099f8caf368223d4d2d2da83dfec64ead2c3267b

Observation 87373631-0bff-4bec-8c62-2e845eab17ec · outbound

This paper cites ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models.

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-06-29T12:13:26.588012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:9fac761f0e3e91082b385a404ef0e4b67f8a7cffdd6f8f63cfd558925b59aad7

Observation d76253e8-46e5-45e5-a2ef-fbbaa87f45a7 · outbound

This paper cites Beyond pass@ 1: Self-play with variational problem synthesis sustains rlvr.arXiv preprint arXiv:2508.14029.

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs Beyond pass@ 1: Self-play with variational problem synthesis sustains rlvr.arXiv preprint arXiv:2508.14029

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-06-29T12:13:26.546069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:797dc801e1a3af49c58aef3b97c334441a1116fc2c83d7ca9819334dbcad3c3e

Observation 7ec33387-212e-4c05-a071-246ab361e2e8 · outbound

This paper cites Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2.

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-06-29T12:13:26.551175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:4ace89462d19afd41aa427dba17a0736d78844ab5fb6475844c00e02d1c01a71

Observation 5aede3d3-09e1-43b6-b5e4-a483ae4db0e2 · outbound

This paper cites Let's Verify Step by Step.

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs Let's Verify Step by Step

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-06-29T12:13:26.603752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:7992580084d8b68819d014bd92c7a8fc56855f85ad0e0448da931f3946e823c8

Observation 5896e8dc-729b-4681-ab04-c871df99d358 · outbound

This paper cites Flow-GRPO: Training Flow Matching Models via Online RL.

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs Flow-GRPO: Training Flow Matching Models via Online RL

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-06-29T12:13:26.533838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:333b31b9a1a4f5ac4a6b7d75ad97cd68f671061a5c832a0c95db0f1c81d09f9f

Observation 0eeb5c27-d76b-405c-a8b8-b3496223dac5 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs Understanding R1-Zero-Like Training: A Critical Perspective

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-06-29T12:13:26.590325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:fdbce55c9f8b69f83c1970b95e2cb808f39723346fac7066cd3ff8e1b437d54d

Observation 9fd02390-6969-4edf-87dd-88317e7426e9 · outbound

This paper cites P^2O: Joint Policy and Prompt Optimization.

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs P^2O: Joint Policy and Prompt Optimization

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-06-29T12:13:26.563632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:341232afca464806bb851591a1b329ef8b912ebda036e7a75124fcafe4c338ad

Observation 8be2425e-b6f8-4692-9669-309b22e8d798 · outbound

This paper cites Deepscaler: Surpassing o1-preview with a 1.5 b model by scaling rl.Notion Blog, 2025.

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs Deepscaler: Surpassing o1-preview with a 1.5 b model by scaling rl.Notion Blog, 2025

Reference 27

Resolution
unresolved
no resolver link, observed 2026-06-29T12:11:00.402276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:01006997c1c791a4a37286d34325da29cea2c4e950d764c45fb702971302b19a

Observation a6a43801-b238-4adb-b480-1e8da677afaa · outbound

This paper cites Bissyande, Haoye Tian, and Bach Le.

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs Bissyande, Haoye Tian, and Bach Le

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-29T12:11:00.402276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:d2be659b4012728795947a05b34fc776cd07d08b5bcd3c69eab24c610c7bf82d

Observation fbc7b728-24cb-416e-bd52-b6822afa9937 · outbound

This paper cites Michaud, Yonatan Belinkov, David Bau, and Aaron Mueller.

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs Michaud, Yonatan Belinkov, David Bau, and Aaron Mueller

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-29T12:11:00.402276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:929305f5044930469da38771c98cf5aa950a36a7b80298adc5e30a63f718c6ba

Observation 667c6f5a-0239-46bb-963a-cb36ddb04a28 · outbound

This paper cites Sparse autoencoder.CS294A Lecture notes, 72(2011):1–19, 2011.

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs Sparse autoencoder.CS294A Lecture notes, 72(2011):1–19, 2011

Reference 30

Resolution
unresolved
no resolver link, observed 2026-06-29T12:11:00.402276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:aaec66be4eae808b563d12085d2e502a6e16c401ca86f64dde0d48b4d4cc3bef

Observation 52ba3a10-6510-4da7-8117-e8f853cc6d58 · outbound

This paper cites MedVLM-R1: Incentivizing Medical Reasoning Capability of Vision-Language Models (VLMs) via Reinforcement Learning.

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs MedVLM-R1: Incentivizing Medical Reasoning Capability of Vision-Language Models (VLMs) via Reinforcement Learning

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T12:13:26.541175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:e45e071ff8782d61b1604101b3f15ec8ab2cb1f952cfcf7f709a6c950f71a471

Observation 18a35d8f-1f77-455b-a899-cfadd56af530 · outbound

This paper cites Curricu- lum reinforcement learning from easy to hard tasks im- proves llm reasoning.arXiv preprint arXiv:2506.06632.

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs Curricu- lum reinforcement learning from easy to hard tasks im- proves llm reasoning.arXiv preprint arXiv:2506.06632

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-06-29T12:13:26.553644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:38804e4485181d6e51e56da30ac84cec7e2a3ac86888bdc8b91eb71a92ed4129

Observation a3a1bd52-5811-46ab-9fda-09792f946bde · outbound

This paper cites Automatically interpreting millions of features in large language models.

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs Automatically interpreting millions of features in large language models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-29T12:11:00.402276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:ef5483f280a8d0b7d8115c04c8df3a4eb51d50f9e560901836e3e7ea73e6af42

Observation 09290c32-bce7-428c-ae22-1d6cf10e6b9c · outbound

This paper cites Near-Future Policy Optimization.

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs Near-Future Policy Optimization

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-06-29T12:13:26.592555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:6ee0ad830f23a54c3049733a4260d4d69a09da5f485c0cd3a2b50391512a6f9a

Observation 6c78e222-cf37-476f-acbe-9109267146b0 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs Proximal Policy Optimization Algorithms

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-06-29T12:13:26.599172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:9e8e62f564de30cff9913b5ded5980ba289c26f295c6d773cdcbef6a32699093

Observation 7c21adc3-97e9-4ab3-8285-7e2ffd094ea2 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-06-29T12:13:26.582443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:c2f58b3ceaccc2017cff95d13121fda52c5523bb53f0e6cc075b03bbd93c484e

Observation 12fa6662-c969-417f-9940-e15496ced6a2 · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs HybridFlow: A Flexible and Efficient RLHF Framework

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-06-29T12:13:26.561427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:b05cc3199ab6a41c66df7c2a0831579ca9067bc1b5d750318e992f0bcf362674

Observation d3f621c5-34cc-4a4a-bbc8-a404eb19a703 · outbound

This paper cites Towards high data efficiency in reinforcement learning with verifiable reward.

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs Towards high data efficiency in reinforcement learning with verifiable reward

Reference 38

Resolution
unresolved
no resolver link, observed 2026-06-29T12:11:00.402276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:18fbf077fb8fa7db8daacffb09e59486fc134c1e17fa189973f268d0e831cb6f

Observation e13180f9-a5d3-4ef0-889f-94b394a5a387 · outbound

This paper cites Reinforcement Learning for Reasoning in Large Language Models with One Training Example.

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-06-29T12:13:26.565970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:435d64766ab36729cef89bd76df3a6b12af2261a009c9fb45e278f2fc61de1ce

Observation 772734fd-d24c-4b8f-9fe8-3444bb2043c6 · outbound

This paper cites Learn hard problems during rl with reference guided fine-tuning.

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs Learn hard problems during rl with reference guided fine-tuning

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-06-29T12:13:26.548687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:c1081f0aff19f8d92a4636de040dbe1438323d59e621e9a7fcce72b0c7be676e

Observation cce9635e-cd79-4155-bfc7-8c52124843ad · outbound

This paper cites DanceGRPO: Unleashing GRPO on Visual Generation.

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs DanceGRPO: Unleashing GRPO on Visual Generation

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-06-29T12:13:26.584739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:2ff7fe631a9dc620d59bce7a99545d4c7ae54dd8e9be9b8e7d25c06dd374c32e

Observation 6ab60fd2-f140-4963-a446-84a35bdc8f18 · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-06-29T12:13:26.601447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:4abb1b391d2e19dd393c9c7e40a55517e1340b3dc4a1850e62a1a2e8747c70c6

Observation ac7a7c67-77f7-46e8-9875-4e3aaeeb8c81 · outbound

This paper cites DCPO: Dynamic Clipping Policy Optimization.

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs DCPO: Dynamic Clipping Policy Optimization

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-06-29T12:13:26.575383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:4162e59e33cf0c8b7f14f33f3615688abf6584b86654d79e37c0547cb8d724c4

Observation 4c81fe84-6e9c-40d8-bf94-418c9af1ee44 · outbound

This paper cites Depth-Breadth Synergy in RLVR: Unlocking LLM Reasoning Gains with Adaptive Exploration.

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs Depth-Breadth Synergy in RLVR: Unlocking LLM Reasoning Gains with Adaptive Exploration

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-06-29T12:13:26.577872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:329138ee963df9037237e69897a047e32916b13a9859d66d5ff585df966c8811

Observation 7bead28b-7bf3-4309-87eb-e755ab489a84 · outbound

This paper cites Kwok, Zhenguo Li, Adrian Weller, and Weiyang Liu.

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs Kwok, Zhenguo Li, Adrian Weller, and Weiyang Liu

Reference 45

Resolution
unresolved
no resolver link, observed 2026-06-29T12:11:00.402276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:d8f3350270fcde85dac1f40eb2b46de947140c2b88d6b9173c41fb59d07d4a8d

Observation b115c98e-8723-4179-a9b5-1424a4da029b · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-06-29T12:13:26.594722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:b74f6fd57ca6e99f0c9f9ba257a2642c7d085a392bacccaa0d335667bf6b75e4

Observation f616f55a-bff7-4913-acb3-68dcb2fbf2a3 · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?.

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-06-29T12:13:26.531518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:3134a9ae88b6c6ec5ba8dd159939667d84a3c8d6ca83f5aff3527cfb745d4e16

Observation 5bf5ee62-0083-4b7f-b626-20ca878ee647 · outbound

This paper cites VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks.

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-06-29T12:13:26.559111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:cb30a22b0bf8ab7f66c6f2921e14378b6f93c36f0d00c3d0b0f0331272b189e9

Observation a420a429-b884-4e5f-a940-58999437a1ea · outbound

This paper cites Wong, and Yu Cheng.

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs Wong, and Yu Cheng

Reference 49

Resolution
unresolved
no resolver link, observed 2026-06-29T12:11:00.402276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:e4dcecdb3fe0259c1dec7a6aada75ba02cb8c825df5a023a6a0d6d70f7dc7731

Observation 76e2995d-9f59-4a95-b381-3755081e7a75 · outbound

This paper cites Scaf-GRPO: Scaffolded group relative policy optimization for enhancing LLM reasoning.

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs Scaf-GRPO: Scaffolded group relative policy optimization for enhancing LLM reasoning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-06-29T12:11:00.402276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:a95b612516829fef89b2b54224be6f97f5457dea5d38672f323131a19f8eda10

Observation 5245ac71-be60-4076-8529-1dbe5e835795 · outbound

This paper cites Absolute Zero: Reinforced Self-play Reasoning with Zero Data.

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs Absolute Zero: Reinforced Self-play Reasoning with Zero Data

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-06-29T12:13:26.536180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:00f41328c2b83c4f723e3013f5855289a60a16df36e70692009488ca394d0b28

Observation a8e5c49b-05a9-4933-ba14-f6e0deacab65 · outbound

This paper cites Group Sequence Policy Optimization.

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs Group Sequence Policy Optimization

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-06-29T12:13:26.543437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:b0bbdbae7cdc18af8aa856c9a43c1b6ee782448d66c75efb0ec5787f7d1d423c

Observation 3a9aeec3-b49f-4057-886a-18c342bd32b4 · outbound

This paper cites AbsTopK: Rethinking sparse au- toencoders for bidirectional features.

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs AbsTopK: Rethinking sparse au- toencoders for bidirectional features

Reference 53

Resolution
unresolved
no resolver link, observed 2026-06-29T12:11:00.402276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:7a79e6ddfe130c33a3ca0b5a992421b18c13e542d98937aa70066ec33c0c4113

Observation 7fa1a272-ff80-44a3-a4df-9bd6145b082c · outbound

This paper cites an unresolved cited work.

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs Unresolved cited work

Reference 54

Resolution
unresolved
no resolver link, observed 2026-06-29T12:11:00.402276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:d06e22b65696c0492ba52c70f2b9ccb869e83b7f44ed869540c8279224d5f0e3

Observation 5eee3410-40fe-4f4a-aa05-a375bb0bfc36 · outbound

This paper cites Therefore, she doesn ’t need to do any more situps on Wednesday to meet her goal.

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs Therefore, she doesn ’t need to do any more situps on Wednesday to meet her goal

Reference 55

Resolution
unresolved
no resolver link, observed 2026-06-29T12:11:00.402276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:abb9e416721ea2ec526d35d2800925a29b074e32a29a288d08aca223f57b788e

Observation 9da6bb6c-4acb-4868-aa2c-c881a6748d18 · outbound

This paper cites \boxed{{{situps_needed_wednesday}}}.

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs \boxed{{{situps_needed_wednesday}}}

Reference 56

Resolution
unresolved
no resolver link, observed 2026-06-29T12:11:00.402276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:483994b52010c6d6ff846b333ba2aff745db92c7527ea880c3429b1acfb42fe3

Observation fd234732-25bc-4b90-9e65-73c2fe106134 · outbound

This paper cites an unresolved cited work.

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs Unresolved cited work

Reference 57

Resolution
unresolved
no resolver link, observed 2026-06-29T12:11:00.402276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:9b65f72412f176d8ef42ede5352e8439d1023f6883cc7e4f2ccee8e0b7b6ac4d

Observation 98bcb9f6-b4f5-4091-a57e-a4a23cc8ace1 · outbound

This paper cites an unresolved cited work.

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs Unresolved cited work

Reference 58

Resolution
unresolved
no resolver link, observed 2026-06-29T12:11:00.402276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:5517d3e8b99b53186bbde56131f88ffd6dda517ffa113e5305fbc8868dc18f16

Observation ec722431-3a22-4b1d-a21b-56722b943835 · outbound

This paper cites an unresolved cited work.

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs Unresolved cited work

Reference 59

Resolution
unresolved
no resolver link, observed 2026-06-29T12:11:00.402276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:a9d5df8a4b5e10f58997632cc20b59a3a1a8d2ba3ce3ea10d09e780e00b09c51

Observation 214368da-ef18-40e5-821f-e77fbb55afea · outbound

This paper cites an unresolved cited work.

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs Unresolved cited work

Reference 60

Resolution
unresolved
no resolver link, observed 2026-06-29T12:11:00.402276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:82474a9ee8dac7382f5b5d333c9102fe4de8e23df53be75da67a53acb03b4275

Observation 001fe2f7-393f-4f73-b181-8689b3cdbde7 · outbound

This paper cites an unresolved cited work.

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs Unresolved cited work

Reference 61

Resolution
unresolved
no resolver link, observed 2026-06-29T12:11:00.402276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:5cd183c809eb2cfa2ad175abde1d030fe82d4c594efe499697d729ebd43a8869

Observation 3f098db9-22ef-47b5-a92f-85f64dcae12c · outbound

This paper cites \boxed{{{int(total_cost)}}}.

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs \boxed{{{int(total_cost)}}}

Reference 62

Resolution
unresolved
no resolver link, observed 2026-06-29T12:11:00.402276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:ee4eedf0f50acef328fceb25a6dc59dcb0f770969cb7a829d66bfa99157db300

Observation a74382fa-f50b-4c3d-87a8-b07631137756 · outbound

This paper cites With 3 foxes, the total number of weasels caught per week is \(3 \times 4 = 12\) weasels, and the total number of rabbits caught per week is \(3 \ times 2 = 6\) rabbits.

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs With 3 foxes, the total number of weasels caught per week is \(3 \times 4 = 12\) weasels, and the total number of rabbits caught per week is \(3 \ times 2 = 6\) rabbits

Reference 63

Resolution
unresolved
no resolver link, observed 2026-06-29T12:11:00.402276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:c56b829d4c86144ff31ebe855a885123620c5e3f9d1fca218d524bef461febce

Observation 84bee187-b2b4-4b9b-a8aa-6fc286268134 · outbound

This paper cites an unresolved cited work.

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs Unresolved cited work

Reference 64

Resolution
unresolved
no resolver link, observed 2026-06-29T12:11:00.402276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:95db4ea7c4ac13ad25172692522ad2606a700cbedff99d34dd1b163e92b1804f

Observation e585aa7f-0d4e-480f-aef3-c12c544c0b87 · outbound

This paper cites an unresolved cited work.

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs Unresolved cited work

Reference 65

Resolution
unresolved
no resolver link, observed 2026-06-29T12:11:00.402276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:57e3e8433019a8d079eb3c5f159e6026fef778c18d80ea1979e23577fadfcfbe

Observation 750ba8c1-6af6-4d95-bf92-d65b09734d8e · outbound

This paper cites an unresolved cited work.

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs Unresolved cited work

Reference 66

Resolution
unresolved
no resolver link, observed 2026-06-29T12:11:00.402276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:158cce785c9e19364f67de6b5c73c64ee33a8e92f600df804434e0187b3a7ed9

Observation 9d5f5a23-b6ee-4597-bdbc-c34ca34a57af · outbound

This paper cites \boxed{{{int(weasels_left)} {int(rabbits_left)}}}.

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs \boxed{{{int(weasels_left)} {int(rabbits_left)}}}

Reference 67

Resolution
unresolved
no resolver link, observed 2026-06-29T12:11:00.402276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:c834201991cf44ed9de8f295180c56474bd6dd35a1f399072afc3c75aefc67ce

Observation 567fa6ea-d752-44c0-938c-1d1c08ef4da4 · outbound

This paper cites an unresolved cited work.

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs Unresolved cited work

Reference 68

Resolution
unresolved
no resolver link, observed 2026-06-29T12:11:00.402276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:6de389d62e6a90156d56a99b67cac3103c363e510c855fa73d56b38a7107bbca

Observation fb72ff17-6878-4083-aaa5-5545091c8782 · outbound

This paper cites an unresolved cited work.

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs Unresolved cited work

Reference 69

Resolution
unresolved
no resolver link, observed 2026-06-29T12:11:00.402276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:bdc041424e53cd45f6298efcbc206b88cae676f43f20f211d3904bdd42ec5973

Observation 49b96d98-745f-4d89-ac84-609849418322 · outbound

This paper cites \boxed{{{int(difference)}}}.

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs \boxed{{{int(difference)}}}

Reference 70

Resolution
unresolved
no resolver link, observed 2026-06-29T12:11:00.402276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:23b222619e03d3e6d27b41421ecccfbdf5ba3907d5d55dcb1287f3cfe3b3ba09

Observation 173b2a80-0414-41ef-9c42-4e80ee70bc2d · outbound

This paper cites So, \( T = A - 20 \).

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs So, \( T = A - 20 \)

Reference 71

Resolution
unresolved
no resolver link, observed 2026-06-29T12:11:00.402276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:8b6d3b7f83f30b8597be103e532f12c89f04db3c1a274149450c36e63a5182b1

Observation f9e6f6d3-2e7f-4f1a-a1d4-2ce30f7c0d1c · outbound

This paper cites So, \( S = M + 10 \).

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs So, \( S = M + 10 \)

Reference 72

Resolution
unresolved
no resolver link, observed 2026-06-29T12:11:00.402276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:d3450a0de568a85b65d304e08fb1b826f364d785cf184abb26c13f7ef878d3e8

Observation 7c0242f4-a184-4ebd-b8c1-f1b89c43b2dd · outbound

This paper cites So, \( M = 70 \).

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs So, \( M = 70 \)

Reference 73

Resolution
unresolved
no resolver link, observed 2026-06-29T12:11:00.402276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:c493fe28eb0d07b61fedbe3c4f29f7c095802d5e7de62ed9b3f87c7745a35080

Observation 7f7fbaf8-7f0f-49d7-bce2-e4f67e73ac61 · outbound

This paper cites \boxed{{{int(total_marks)}}}.

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs \boxed{{{int(total_marks)}}}

Reference 74

Resolution
malformed identifier
arxiv_id, observed 2026-06-29T12:13:26.538591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:763f8f24a470f7a0f4565655a91d221baf37372a71c3c41226971a0605a8ca00

Observation 18939218-ca36-4710-822a-b59de9c1aafa · outbound

This paper cites Suppose that a+ (1/b) and b+ (1/a) are the roots of the equationx 2 −px+q= 0.

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs Suppose that a+ (1/b) and b+ (1/a) are the roots of the equationx 2 −px+q= 0

Reference 75

Resolution
unresolved
no resolver link, observed 2026-06-29T12:11:00.402276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:607fef9fe2b9136e3b3f5d5a6a45a23d24f33edd1dbd912335df4c7dd98d9d0c

Observation d761ae3c-a730-4c3e-9b46-25ab600e2630 · outbound

This paper cites Suppose that a+ (z/b) and b+ (1/a) are the roots of the equationx 2 −px+q= 0.

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs Suppose that a+ (z/b) and b+ (1/a) are the roots of the equationx 2 −px+q= 0

Reference 76

Resolution
unresolved
no resolver link, observed 2026-06-29T12:11:00.402276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:006655583df4b3a1cdb8cb108ae377f6f9e5fd3efad36604c957c823a2caef11

Observation 54dc86fa-7e99-4043-877c-ecfeb4b11693 · outbound

This paper cites Suppose that a+ (1/b) and b+ (z/a) are the roots of the equationx 2 −px+q= 0.

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs Suppose that a+ (1/b) and b+ (z/a) are the roots of the equationx 2 −px+q= 0

Reference 77

Resolution
unresolved
no resolver link, observed 2026-06-29T12:11:00.402276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:9d28d79092426d989713455d5209897732114405f78b27b5a5cedbd6f2beb955

Observation ec053849-7c0e-4b77-8036-b50f27b450f2 · outbound

This paper cites Suppose that a+ (1/b) and b+ (1/a) are the roots of the equation x2 −px+q= 0.

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs Suppose that a+ (1/b) and b+ (1/a) are the roots of the equation x2 −px+q= 0

Reference 78

Resolution
unresolved
no resolver link, observed 2026-06-29T12:11:00.402276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:fab7c8d772f8ab9106c21babfea9e122f873f979379e4ac0b0b500f4c95b604c

Observation 6a9ddd17-f0bf-43d6-aa68-e8129206045e · outbound

This paper cites Suppose that a+ (z/b) and b+ (1/a) are the roots of the equation x2 −px+q= 0.

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs Suppose that a+ (z/b) and b+ (1/a) are the roots of the equation x2 −px+q= 0

Reference 79

Resolution
unresolved
no resolver link, observed 2026-06-29T12:11:00.402276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:300f458dab8c522417e54f29babeedede5b9e6dc44c6be832f7082ce492912d5

Observation 03c22ac2-0321-4743-9fc4-0974f10285cf · outbound

This paper cites Suppose that a+ (1/b) and b+ (z/a) are the roots of the equation x2 −px+q= 0.

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs Suppose that a+ (1/b) and b+ (z/a) are the roots of the equation x2 −px+q= 0

Reference 80

Resolution
unresolved
no resolver link, observed 2026-06-29T12:11:00.402276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:594b0bbcf512949f65dd0b9272319a3f519c213de6850fbc7cc4bd6511257514

Pith citing papers

Observation 12ba30fa-be86-4e36-9321-482ee6e65a26 · inbound

DRIFT: Difficulty Routing Self-DIstillation with Rhythm-Gated Exploration and Success BuFfer Training cites this paper.

DRIFT: Difficulty Routing Self-DIstillation with Rhythm-Gated Exploration and Success BuFfer Training Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T07:24:21.461392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T07:19:00.012926Z digest=sha256:1d790e952b6806fa749e86d70d94c95eb58264ce8b4dfc9b2f4de0f5daca86a1