Pith. sign in

Paper Citation Record · LEDGER

Parameter Exploration for RLVR via Variational Learning

As of 18 August 2026, this Paper Citation Record lists 82 of 82 outbound references and 0 inbound Pith citation observations for arXiv:2608.09805.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.09805 v1

Coverage vector

measured 82 of 82 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T10:32:13.146209Z

measured 82 of 82 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

82 of 82 outbound references displayed

  • verified exact10
  • verified fuzzy28
  • unresolved43
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 463ddaed-3d44-4da6-9ce5-80f49d67d6bd · outbound

This paper cites A Survey of Exploration Methods in Reinforcement Learning.

Parameter Exploration for RLVR via Variational Learning A Survey of Exploration Methods in Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:12.730067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:12.730067Z digest=sha256:90cd4b490fcfa4cac1a68e3fd85c2feb05892ba0e1ddbff204adfd654efa3ea0

Observation 37c920ea-313f-4c46-8e5e-e92705d6dc25 · outbound

This paper cites Polaris: A post-training recipe for scaling reinforcement learning on advanced reasoning models, 2025.

Parameter Exploration for RLVR via Variational Learning Polaris: A post-training recipe for scaling reinforcement learning on advanced reasoning models, 2025

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:12.735104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:12.735104Z digest=sha256:660179c589ec84ee104612ad8ddce544653d3a2cf29684f061db998a539430f6

Observation 9f15f11d-376d-433e-a2d4-c7a09b942c1f · outbound

This paper cites Learning to explore with parameter-space noise: A deep dive into parameter-space noise for reinforcement learning with verifiable rewards.CoRR, abs/2602.02555, 2026.

Parameter Exploration for RLVR via Variational Learning Learning to explore with parameter-space noise: A deep dive into parameter-space noise for reinforcement learning with verifiable rewards.CoRR, abs/2602.02555, 2026

Reference 3

Resolution
verified exact
doi, observed 2026-08-11T10:32:14.059743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:12.739164Z digest=sha256:d272dcc417af6a2c1618d326453bfeebcb966e1ec891d3bd601bf2e922257d59

Observation ba81f022-c73f-4977-a899-1156cd2224c0 · outbound

This paper cites Llama-nemotron: Efficient reasoning models, 2025.

Parameter Exploration for RLVR via Variational Learning Llama-nemotron: Efficient reasoning models, 2025

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:12.743020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:12.743020Z digest=sha256:1d25c873d94d2fdae6fa48eb8d1bd96633aa0109a7be3a347f79b4d80c295653

Observation ae42e433-66d1-4662-a69d-e26076b8156a · outbound

This paper cites Weight un- certainty in neural network.

Parameter Exploration for RLVR via Variational Learning Weight un- certainty in neural network

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:12.747358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:12.747358Z digest=sha256:bd17ae2973aa21aa5e22ebae539c1778cb6c05201bcc376d6834a60bcf34a224

Observation 646564d1-5bb4-46a7-ac84-198d06b87cc4 · outbound

This paper cites FullStack Bench: Evaluating LLMs as Full Stack Coders.

Parameter Exploration for RLVR via Variational Learning FullStack Bench: Evaluating LLMs as Full Stack Coders

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:12.751253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:12.751253Z digest=sha256:6d5ad4ad78ae22b9ccac93e5926c7ed32ff88b8d18f09bcd3ba3b84464679ec5

Observation 06241385-ab5d-4963-b8db-254df635a039 · outbound

This paper cites Skyrl-v0: Train real-world long-horizon agents via reinforcement learning, 2025.

Parameter Exploration for RLVR via Variational Learning Skyrl-v0: Train real-world long-horizon agents via reinforcement learning, 2025

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:32:15.114998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:12.755926Z digest=sha256:6171a2365805b3382eda94c3006879addafd063528f1125e413ec96673007dc4

Observation 21a74956-aa23-430f-82c6-254cdabc43f2 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Parameter Exploration for RLVR via Variational Learning Evaluating Large Language Models Trained on Code

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:12.759973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:12.759973Z digest=sha256:da0e085f70233933f95fd5d67a426621b011ff2b023b6d996ea936fc35238540

Observation 51591005-14da-446a-84f1-1dfa858b03f0 · outbound

This paper cites Exploration vs exploitation: Rethinking RLVR through clipping, entropy, and spurious reward.

Parameter Exploration for RLVR via Variational Learning Exploration vs exploitation: Rethinking RLVR through clipping, entropy, and spurious reward

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:32:15.102504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:12.763941Z digest=sha256:40edaed6b1fc8bcf33ae2bb7817d11041b91eb83ec4217b978885098d90bc6a4

Observation 07287b1d-f253-4169-b518-6e5282df3d20 · outbound

This paper cites Improving LoRA with Variational Learning.

Parameter Exploration for RLVR via Variational Learning Improving LoRA with Variational Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:12.768570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:12.768570Z digest=sha256:a5496745f0deced26b316f1185266527a8f0f4c0bfc1c2fb201535ba3025e9d8

Observation 107d603d-cb86-4f0c-8ce4-f20e8e33bb98 · outbound

This paper cites The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models.

Parameter Exploration for RLVR via Variational Learning The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:12.774265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:12.774265Z digest=sha256:6985ac3d74a08ab80763a1959df6c8ee1fb08e6fa481426e5e343014767e6dd9

Observation b13e8654-d971-43e9-9564-14865009e741 · outbound

This paper cites Uncertainty-aware decoding with minimum bayes risk.

Parameter Exploration for RLVR via Variational Learning Uncertainty-aware decoding with minimum bayes risk

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:32:15.090956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:12.778746Z digest=sha256:c50cf33f27da2ea6ed3d38bd6f1eae863fc52b7b821282cd43ea800f90a1723a

Observation 8c5f775f-51d5-463a-9099-8ed57d78607f · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Parameter Exploration for RLVR via Variational Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:12.783229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:12.783229Z digest=sha256:efc7ace480fe7287d5d8cb2bb4ec3c2230122bf5f61e47f2a42461c1176321dc

Observation 123cc001-309e-4151-aecf-ccfc5a20a76b · outbound

This paper cites A survey on policy search for robotics.Found.

Parameter Exploration for RLVR via Variational Learning A survey on policy search for robotics.Found

Reference 14

Resolution
verified exact
doi, observed 2026-08-11T10:32:13.954304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:12.786642Z digest=sha256:1c78f1755498190bf9c456294598b51e3f2770590f5183abe4b38a27fb06c755

Observation 66d0116e-5f2b-4b80-975a-7f0c8792c0b5 · outbound

This paper cites Sharpness-aware min- imization for efficiently improving generalization.

Parameter Exploration for RLVR via Variational Learning Sharpness-aware min- imization for efficiently improving generalization

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:32:15.076200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:12.791126Z digest=sha256:447d778a1b6a4e895a8626c4a76c023d03d8ea00835eb7d6601d8c53e41d4725

Observation 0d56f047-1a95-4fb1-a99c-57e9f2decc1e · outbound

This paper cites Noisy networks for exploration.

Parameter Exploration for RLVR via Variational Learning Noisy networks for exploration

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:32:15.063344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:12.795514Z digest=sha256:f862b9007a9eb670cd9c18a1681756b1f86a1d3c426ad422e1f11de2f1d22919

Observation e18a1ea0-85d6-48b1-b172-5c81df055be6 · outbound

This paper cites Neural thickets: Diverse task experts are dense around pretrained weights.CoRR, abs/2603.12228, 2026.

Parameter Exploration for RLVR via Variational Learning Neural thickets: Diverse task experts are dense around pretrained weights.CoRR, abs/2603.12228, 2026

Reference 17

Resolution
verified exact
doi, observed 2026-08-11T10:32:13.938318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:12.805368Z digest=sha256:71e0d662455b8d1a1d2cc79a5d11550bd73559ce405d485c440041eec04cff72

Observation cc62d422-b240-4614-a421-cd1809e4a508 · outbound

This paper cites Practical variational inference for neural networks.

Parameter Exploration for RLVR via Variational Learning Practical variational inference for neural networks

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:32:15.039427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:12.810372Z digest=sha256:3ffe60f7322f2aa5aa50bf9ddb5028f6b7a332d68d2e810c92020127a7f970a9

Observation be16919a-7051-4fa5-b122-e5686d58ba26 · outbound

This paper cites Skywork Open Reasoner 1 Technical Report.

Parameter Exploration for RLVR via Variational Learning Skywork Open Reasoner 1 Technical Report

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:12.818963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:12.818963Z digest=sha256:190c45509a4fff2e06564ba21f14fbc5ddc2bf2851d0906a3100301d5db9b9d0

Observation 39dcb276-4953-4707-9f8b-bd6e0c39851a · outbound

This paper cites Measuring mathematical problem solving with the MATH dataset.

Parameter Exploration for RLVR via Variational Learning Measuring mathematical problem solving with the MATH dataset

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:12.823287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:12.823287Z digest=sha256:d36f34f96337ad992a2650c07e7e381a65c34c53bd381576c896f72346a8847a

Observation bbcfc14a-08ab-4b2f-8bbf-1ae24b45ff93 · outbound

This paper cites The curious case of neural text degeneration.

Parameter Exploration for RLVR via Variational Learning The curious case of neural text degeneration

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:12.827658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:12.827658Z digest=sha256:018bd770555ac85c7ec9087d8021d7f1a9209c5c235016bc6d721d99a52a906d

Observation aa45cfff-3eec-4ef9-b7ea-ce9b847dbacf · outbound

This paper cites Brorl: Scaling reinforcement learning via broadened exploration.CoRR, abs/2510.01180, 2025.

Parameter Exploration for RLVR via Variational Learning Brorl: Scaling reinforcement learning via broadened exploration.CoRR, abs/2510.01180, 2025

Reference 22

Resolution
verified exact
doi, observed 2026-08-11T10:32:13.827218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:12.832290Z digest=sha256:99d1b38548db7f7787a35621d30b17d25623b9c5b34ea9a9e49f13cfc5f7d0fd

Observation 45a2747d-aeee-4978-8f5f-d1e3705c635c · outbound

This paper cites Livecodebench: Holistic and contamination free evaluation of large language models for code.

Parameter Exploration for RLVR via Variational Learning Livecodebench: Holistic and contamination free evaluation of large language models for code

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:32:14.991288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:12.837036Z digest=sha256:56dfcaff89b08f2519f09e257c06c2e08892db12a142a0f56af087abe28d1fc2

Observation b29344e8-1bd0-4d1f-81af-47409f0aa669 · outbound

This paper cites Near-optimal regret bounds for reinforcement learning.J.

Parameter Exploration for RLVR via Variational Learning Near-optimal regret bounds for reinforcement learning.J

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:12.842280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:12.842280Z digest=sha256:b1b7efe16597d984608a39fd5891565cc203a65924e265338c765bb45beead4f

Observation b589645f-91ae-44e4-a097-1f32185282a3 · outbound

This paper cites Rethinking entropy regularization in large reasoning models.CoRR, abs/2509.25133, 2025.

Parameter Exploration for RLVR via Variational Learning Rethinking entropy regularization in large reasoning models.CoRR, abs/2509.25133, 2025

Reference 25

Resolution
verified exact
doi, observed 2026-08-11T10:32:13.757133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:12.846240Z digest=sha256:6d33ae431fcf3325e23fa025a751fbfb2f0c2ffa204f112f51cf657bcfe659ce

Observation 4fa8d917-90a4-44d0-8db7-8b6913d2551e · outbound

This paper cites The bayesian learning rule.Journal of Machine Learning Research, 24(281):1–46, 2023.

Parameter Exploration for RLVR via Variational Learning The bayesian learning rule.Journal of Machine Learning Research, 24(281):1–46, 2023

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:32:14.979814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:12.850986Z digest=sha256:48aff40d7ed7b3a11a87d8bc2d06a6000136b4cfce649d52692992aea8196028

Observation f3337831-493b-48d4-8920-4c952f4b3dd9 · outbound

This paper cites Fast and scalable bayesian deep learning by weight-perturbation in adam.

Parameter Exploration for RLVR via Variational Learning Fast and scalable bayesian deep learning by weight-perturbation in adam

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:32:14.967885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:12.854665Z digest=sha256:5aa7b08dd7dae4ec4e4e133b424190140ecee859e6e4701bef2203386022fecd

Observation 7af71b16-0a03-4670-a9b4-38097d635421 · outbound

This paper cites Generalized variational inference: Three arguments for deriving new posteriors.

Parameter Exploration for RLVR via Variational Learning Generalized variational inference: Three arguments for deriving new posteriors

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:32:14.953142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:12.858567Z digest=sha256:e47c1f5bdffefac9f5f3096746d7b127150a91b1781ea4b5bc2ba7692f25ab9b

Observation e15d98a9-9f10-44d7-9cb6-9be57a09ff38 · outbound

This paper cites The Case for Learning Application Behavior to Improve Hardware Energy Efficiency.

Parameter Exploration for RLVR via Variational Learning The Case for Learning Application Behavior to Improve Hardware Energy Efficiency

Reference 29

Resolution
metadata mismatch
local_arxiv, observed 2026-08-11T10:32:14.357333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:12.863076Z digest=sha256:11d1fdb6365df5745cd7f096abec3ff9da958161e1a4fb6c96e69987f2ad7850

Observation e0fefb07-c04c-48d0-aa5b-12d3b47a11b5 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

Parameter Exploration for RLVR via Variational Learning Efficient memory management for large language model serving with pagedattention

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:12.868294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:12.868294Z digest=sha256:898740bdb194e30e4a072faf49749d37a967b542cea7db608f8930a194f7f4f6

Observation 4f434f82-1798-4f91-af72-e455398bb5bd · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

Parameter Exploration for RLVR via Variational Learning Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:12.885592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:12.885592Z digest=sha256:a732b43c4ca487aae54638f2d1bed055179dbea7a4984b5ba43b69f2986b545d

Observation 32432725-071a-4e0f-afec-715aa718c91a · outbound

This paper cites Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo et al.

Parameter Exploration for RLVR via Variational Learning Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo et al

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:32:14.941961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:12.890555Z digest=sha256:a9049d6fecfa49ff8c1f662276e3fc5284402f3506b84634acd4d6b05685e6dd

Observation 0563aac8-04f5-4dc0-80f6-170c4957c00d · outbound

This paper cites Verified taco problems.

Parameter Exploration for RLVR via Variational Learning Verified taco problems

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:32:14.929410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:12.896238Z digest=sha256:e18ef6a49d970f5990891edabdaceddd21c695552c60dc99e72e84802d737966

Observation 09e3f26f-11c0-46cb-8543-b0b7b43bf160 · outbound

This paper cites Handling the positive-definite con- straint in the bayesian learning rule.

Parameter Exploration for RLVR via Variational Learning Handling the positive-definite con- straint in the bayesian learning rule

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:32:14.916744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:12.901478Z digest=sha256:d1b55ea1851398f72d1cdaeba00f2cda6c301ab5e4edbd232a828fe7d83d8c92

Observation bdc9df48-1918-4067-baef-34908d197424 · outbound

This paper cites When speed kills stability: Demystifying RL collapse from the training-inference mismatch.

Parameter Exploration for RLVR via Variational Learning When speed kills stability: Demystifying RL collapse from the training-inference mismatch

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:32:14.903579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:12.907014Z digest=sha256:2ce39bf36664749a647f5cc203eff143643a28a75719b16a567d28af8d55b5b8

Observation 48cddb81-4e20-40a9-b123-3e65c7504c61 · outbound

This paper cites Code-r1: Reproducing r1 for code with reliable rewards.

Parameter Exploration for RLVR via Variational Learning Code-r1: Reproducing r1 for code with reliable rewards

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:12.911912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:12.911912Z digest=sha256:85689b5e78b46808ed2f4925212e8bc9afd01659344eaec99e88438cdc2e7545

Observation 469e5597-9969-4364-bdd8-aefef327d7dd · outbound

This paper cites ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models.

Parameter Exploration for RLVR via Variational Learning ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:12.916702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:12.916702Z digest=sha256:2b144ed94054b000ef1532b8159436bede8871f058ce88fc9bcc167b011d546c

Observation c46c7011-3bd6-4da4-b5ac-95641628bd36 · outbound

This paper cites Regularization matters in policy optimization - an empirical study on continuous control.

Parameter Exploration for RLVR via Variational Learning Regularization matters in policy optimization - an empirical study on continuous control

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:32:14.881665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:12.920919Z digest=sha256:e1699fa4336dfc629c4009800a78bc2a34b779e7722ad668de3c9c045642aa0f

Observation 7d98b97b-749e-4c50-9625-8c85bb0645d3 · outbound

This paper cites Understanding r1-zero-like training: A critical perspective.

Parameter Exploration for RLVR via Variational Learning Understanding r1-zero-like training: A critical perspective

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:12.929977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:12.929977Z digest=sha256:1b5b1b3cba446b1a44a3e5b313068e075e1f85043380803c4fda5028c504778f

Observation d1efa133-d7be-4794-bd75-3f0a4a52fb6a · outbound

This paper cites Decoupled weight decay regularization.

Parameter Exploration for RLVR via Variational Learning Decoupled weight decay regularization

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:12.935853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:12.935853Z digest=sha256:89475a8c969a42c9f7d75d49549c620cdcd9801da8ae91811fe7ae0257da9eb9

Observation 7e5db798-f81c-4539-9cb9-43dfa506115c · outbound

This paper cites Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa et al.

Parameter Exploration for RLVR via Variational Learning Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa et al

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:32:14.834331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:12.940625Z digest=sha256:cc2be9fd7d27fb5f1f30d786531e60b6d9da60074cf672f4cdd0519cd04179f1

Observation 8eba8178-36aa-4337-9dff-9572eaa26618 · outbound

This paper cites American Invitational Mathematics Examination, 2026.

Parameter Exploration for RLVR via Variational Learning American Invitational Mathematics Examination, 2026

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:12.945458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:12.945458Z digest=sha256:d6889f1b6025485a32df6e7dc91cc2949ebf8fbeecd1af0303dbc134992b92e1

Observation fb3d2388-c503-41d7-8992-9b9b8033965a · outbound

This paper cites SOAP-Bubbles: Structured Weight Uncertainty for Neural Networks.

Parameter Exploration for RLVR via Variational Learning SOAP-Bubbles: Structured Weight Uncertainty for Neural Networks

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-08-11T10:32:14.270721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:12.952545Z digest=sha256:fee6e511268967a148fcc652003a619de8e30f904ef52e32acfa3517f2d6fd6b

Observation fdfe7acf-c09e-43e7-9b63-70e8b6eabbc9 · outbound

This paper cites SAM as an optimal relaxation of bayes.

Parameter Exploration for RLVR via Variational Learning SAM as an optimal relaxation of bayes

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:32:14.812635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:12.957266Z digest=sha256:10b7b4ca3169c5dede7749f0bd57e8d9833065d661152ecae9390d256f1e5836

Observation e36fba69-5a95-4a91-ac65-ca1b408713c6 · outbound

This paper cites Faster, more efficient RLHF through off-policy asynchronous learning.

Parameter Exploration for RLVR via Variational Learning Faster, more efficient RLHF through off-policy asynchronous learning

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:32:14.800601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:12.961285Z digest=sha256:4a91c884c4e19ae0ce56dbb722a6424ecb9bd15a2ba1a4e8108231170c5e31d3

Observation 31f11779-fdc0-4d42-8776-bacf1706d5bc · outbound

This paper cites Olmo 3.

Parameter Exploration for RLVR via Variational Learning Olmo 3

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:12.966627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:12.966627Z digest=sha256:111520f44bdc3915b410e0ca5f95639fd527dba12915e2bc37e4d1c05dedd196

Observation 73ad86e3-942c-475a-800f-ef103c20bef4 · outbound

This paper cites Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray et al.

Parameter Exploration for RLVR via Variational Learning Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray et al

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:32:14.787028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:12.973226Z digest=sha256:9b797c9a70e3fe14c9d039380a30b6893f325d5c2a2fc402bc293a0a5d51c4ec

Observation cef92906-a67a-47bc-82b4-1a8377b882dc · outbound

This paper cites Chen, Xi Chen, Tamim Asfour, Pieter Abbeel, and Marcin Andrychowicz.

Parameter Exploration for RLVR via Variational Learning Chen, Xi Chen, Tamim Asfour, Pieter Abbeel, and Marcin Andrychowicz

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:32:14.776599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:12.978337Z digest=sha256:faa2a616ac0941fa9f9975de86c09579d79cfabcf707ff23da6ed3ceff4dd8d7

Observation f75a85d1-56bd-4d9c-9930-7d709691dc09 · outbound

This paper cites Exploring parameter space in reinforcement learning.Paladyn J.

Parameter Exploration for RLVR via Variational Learning Exploring parameter space in reinforcement learning.Paladyn J

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:32:14.763936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:12.983370Z digest=sha256:0ec3be55010e550e3545153827190b0e6e4a14500a2f22dfe83882478d4d89ad

Observation cc300bdb-8913-4a54-b274-81d3787e8a95 · outbound

This paper cites Evolution Strategies as a Scalable Alternative to Reinforcement Learning.

Parameter Exploration for RLVR via Variational Learning Evolution Strategies as a Scalable Alternative to Reinforcement Learning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:13.001257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:13.001257Z digest=sha256:ba134141a0755c66bcdb09e378c442c6f872201798686b45051413d658e02772

Observation 2c1d7bd0-2199-4cbf-b47a-887012c3dd81 · outbound

This paper cites Parameter-exploring policy gradients.Neural Networks, 23(4):551–559, 2010.

Parameter Exploration for RLVR via Variational Learning Parameter-exploring policy gradients.Neural Networks, 23(4):551–559, 2010

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:13.006112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:13.006112Z digest=sha256:6ad7c2403e9248c39ae36bd7733662d31645c6a28ccbe6971290961ae15b7e43

Observation a38eab6a-6954-4594-88ba-7018c0acc17e · outbound

This paper cites e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs.

Parameter Exploration for RLVR via Variational Learning e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:13.013301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:13.013301Z digest=sha256:e57eb75a20a6317642a36f97d12c4ecaeb33fb22f60fbec8906e6f67658d5873

Observation 3c22fee3-d2a8-4aba-a5a5-480d4c63475d · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Parameter Exploration for RLVR via Variational Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:13.019444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:13.019444Z digest=sha256:74dd5578a004a11248b961aa611c45570414c67d5ec458efbad49d4daf845308

Observation cdc4d653-dc12-4ea8-b959-48296873be97 · outbound

This paper cites On entropy control in LLM-RL algorithms.CoRR, abs/2509.03493, 2025.

Parameter Exploration for RLVR via Variational Learning On entropy control in LLM-RL algorithms.CoRR, abs/2509.03493, 2025

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:13.024389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:13.024389Z digest=sha256:710813b8edd10d0c29c15272d211da1b13a68f0d0169f99ccfaffa18143f4c86

Observation fc207249-b380-469b-8164-1bfd76a837c6 · outbound

This paper cites Variational learning is effective for large deep networks.

Parameter Exploration for RLVR via Variational Learning Variational learning is effective for large deep networks

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:32:14.752260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:13.030530Z digest=sha256:6264b51c5b2367d99a410cd4715587a852cae3ab07f76ec9d02068f2fc57207c

Observation 5c5e5b2e-d02f-4f6e-a25a-efb5754ff767 · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

Parameter Exploration for RLVR via Variational Learning HybridFlow: A Flexible and Efficient RLHF Framework

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:13.039985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:13.039985Z digest=sha256:3470c2553802400b9e1103c6affc3aa478e7fd9f83556c0344399b7f1d94c76d

Observation 01429360-ac01-4efa-876a-3e51d1e403cb · outbound

This paper cites Strehl and Michael L.

Parameter Exploration for RLVR via Variational Learning Strehl and Michael L

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:13.044273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:13.044273Z digest=sha256:daddb2585d78f8f8a1a920c6828a5f49b13d1a1e8395328cc474573096a56115

Observation 0ed17521-795b-4680-9582-a16643fe2b55 · outbound

This paper cites Path integral policy improvement with covariance matrix adaptation.

Parameter Exploration for RLVR via Variational Learning Path integral policy improvement with covariance matrix adaptation

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:32:14.727527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:13.048926Z digest=sha256:1832a7e64dd6c80387405aab472eea353289ff40a2fbc1c70bcf2175ce1d6db1

Observation 2603b6f4-9c35-4596-81a3-4981b964bb16 · outbound

This paper cites RL grokking recipe: How does RL unlock and transfer new algorithms in LLMs? InThe Fourteenth International Conference on Learning Representations, 2026.

Parameter Exploration for RLVR via Variational Learning RL grokking recipe: How does RL unlock and transfer new algorithms in LLMs? InThe Fourteenth International Conference on Learning Representations, 2026

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:32:14.714895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:13.053355Z digest=sha256:656f6b7769d8898643a4831a2e9c70e7ccf91c4348ecee5dd8df3a6b60a2944a

Observation 2f6d0311-0935-4286-a635-9965396094f6 · outbound

This paper cites an unresolved cited work.

Parameter Exploration for RLVR via Variational Learning Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-11T10:32:14.703322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:13.059829Z digest=sha256:833516752f2452446953d5a7dbcb5ed5f2ed3ee9c5994c0d9896d7edc00d6498

Observation ae7c3df1-a809-46b8-bf67-5dce78e32c75 · outbound

This paper cites Sutton and Andrew G.

Parameter Exploration for RLVR via Variational Learning Sutton and Andrew G

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:32:14.691421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:13.066364Z digest=sha256:496e7b6005e47fa54d637b1455d4b712cd6ce57bea6c4021d987940c2cd0c252

Observation 907023fd-33b9-4a78-abcb-d1b231d7cdbc · outbound

This paper cites Theodorou, Jonas Buchli, and Stefan Schaal.

Parameter Exploration for RLVR via Variational Learning Theodorou, Jonas Buchli, and Stefan Schaal

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:13.071427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:13.071427Z digest=sha256:305d5210888991d13d9d47a655160b8c2b0bde50a125bf5c041bcfb479d972c0

Observation 90fa11a7-9efa-4a55-b408-9ac2e528f2aa · outbound

This paper cites Generalized exploration in policy search.

Parameter Exploration for RLVR via Variational Learning Generalized exploration in policy search

Reference 63

Resolution
verified exact
doi, observed 2026-08-11T10:32:13.571135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:13.076147Z digest=sha256:d9e9be27806d9c3cc19c26aa3a8818d5eff0e33eea37cb418981f7bbc071ff31

Observation b3f7816e-f0b6-4c34-bb47-255a26c11571 · outbound

This paper cites Aletheia: What Makes RLVR For Code Verifiers Tick?.

Parameter Exploration for RLVR via Variational Learning Aletheia: What Makes RLVR For Code Verifiers Tick?

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:13.081086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:13.081086Z digest=sha256:8f6416be5a3734685a2080e87b80c6551a76eecdfefdedd1ddeed89b81ae944d

Observation 6a678b06-4ac0-45a2-bfd0-706db61a7da8 · outbound

This paper cites Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning.

Parameter Exploration for RLVR via Variational Learning Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:13.087281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:13.087281Z digest=sha256:141b0e2b4026b3fce878c59d9ddf905047470b19bd0819cbc27e1155c04f33c9

Observation b1068912-5ebc-47e3-80c3-77ecc4ab946b · outbound

This paper cites Learning from delayed rewards.

Parameter Exploration for RLVR via Variational Learning Learning from delayed rewards

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:13.092467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:13.092467Z digest=sha256:88553e4e3561ff2938432b7c014dcf18a7954583716ff5b41fdcb641113246b1

Observation 87d8fd76-b25b-4726-88e9-bfa2efb01d1c · outbound

This paper cites Adversarial weight perturbation helps robust generalization.

Parameter Exploration for RLVR via Variational Learning Adversarial weight perturbation helps robust generalization

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:32:14.672871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:13.097808Z digest=sha256:b7c8833876f53e6048e243a11ee057143b322a10621d13626b10ac4470f8f88e

Observation 15766fb5-f584-47d9-9500-f27ae32303ae · outbound

This paper cites The invisible leash: Why RLVR may not escape its origin.CoRR, abs/2507.14843, 2025.

Parameter Exploration for RLVR via Variational Learning The invisible leash: Why RLVR may not escape its origin.CoRR, abs/2507.14843, 2025

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:13.102018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:13.102018Z digest=sha256:138f7c19c35772e15288fcf48c846f651e11e6bec5fa88ea215dd60ba3f1777b

Observation a2c6d4d7-5060-4d36-af4c-9dcc5e70c0d1 · outbound

This paper cites Reasoning or memorization? unreliable results of reinforcement learning due to data contamination.

Parameter Exploration for RLVR via Variational Learning Reasoning or memorization? unreliable results of reinforcement learning due to data contamination

Reference 69

Resolution
verified exact
doi, observed 2026-08-11T10:32:14.660615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:13.108060Z digest=sha256:189f82e81b89afc9a9fd8db61ad2d57f01702a9ee14a6b381fa72fdfe234c138

Observation 51e4aa03-2498-412f-89c4-f1a4775f3258 · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

Parameter Exploration for RLVR via Variational Learning Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:13.111958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:13.111958Z digest=sha256:5976adcabeb4593c2f65171c66595b327751a4ce1f68b186334152f2f30b5e50

Observation c678ef5a-4c70-40ef-8213-ed6ce0076ec3 · outbound

This paper cites Your efficient rl framework secretly brings you off-policy rl training, August 2025.

Parameter Exploration for RLVR via Variational Learning Your efficient rl framework secretly brings you off-policy rl training, August 2025

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:13.116229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:13.116229Z digest=sha256:2dba8fc1c0fd69efd2a898298c5492705411e50758c5c5b023903ff5692ce901

Observation 19cd3682-c4e9-46fa-8e09-ac3dbf307f84 · outbound

This paper cites The debate on RLVR reasoning capability boundary: Shrinkage, expansion, or both? A two-stage dynamic view.CoRR, abs/2510.04028, 2025.

Parameter Exploration for RLVR via Variational Learning The debate on RLVR reasoning capability boundary: Shrinkage, expansion, or both? A two-stage dynamic view.CoRR, abs/2510.04028, 2025

Reference 72

Resolution
verified exact
doi, observed 2026-08-11T10:32:13.447823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:13.120490Z digest=sha256:ff42b4d6097ac25a618ee554dbc6edfc5e1f863384284bd822f234bb8cbdb399

Observation 7fc9da07-b433-44a7-943c-b608181a1e43 · outbound

This paper cites DAPO: An open-source LLM reinforcement learning system at scale.

Parameter Exploration for RLVR via Variational Learning DAPO: An open-source LLM reinforcement learning system at scale

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:32:14.638208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:13.124865Z digest=sha256:55983ddacd114d7644a76b8c1b6f40b49883a0f78105ed110825fad9c051f67c

Observation 91df57ff-200d-4d8b-b1c3-ad18867ce443 · outbound

This paper cites Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining.

Parameter Exploration for RLVR via Variational Learning Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:13.129972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:13.129972Z digest=sha256:7b1305924aa54b0bd14adb270b4a8926d866d5b534f0d095bd758e10fd9a9d89

Observation 93db9b12-6f5e-4340-865a-1dbfa34e2bfa · outbound

This paper cites Group Sequence Policy Optimization.

Parameter Exploration for RLVR via Variational Learning Group Sequence Policy Optimization

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:13.137555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:13.137555Z digest=sha256:27b7a7d13ca97954c9ee7637eb617d5a6c30ad0310fc14abf23dcf6a414ddfb8

Observation 7d653e8f-31f8-4780-9f55-d034ef0ebb78 · outbound

This paper cites The surprising effectiveness of negative reinforcement in LLM reasoning.

Parameter Exploration for RLVR via Variational Learning The surprising effectiveness of negative reinforcement in LLM reasoning

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:32:14.622844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:13.142111Z digest=sha256:e87e2593dd8adbefca3f05a699e16528be6356d1f20ddb294f9121fb9e5385c9

Observation 053302a1-ebd4-4250-ac2e-12f4ab659af2 · outbound

This paper cites Exploring multi-temperature strategies for token- and rollout-level control in RLVR.

Parameter Exploration for RLVR via Variational Learning Exploring multi-temperature strategies for token- and rollout-level control in RLVR

Reference 77

Resolution
verified exact
doi, observed 2026-08-11T10:32:13.355286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:13.146209Z digest=sha256:25271a4b5b04c3771d7fd765e136d57cb65fdbbb3fd7c64ca1bb3db3b2aadc1d

Observation 0cbc3438-9d41-410a-8c07-9614d43308a1 · outbound

This paper cites URL https://doi.org/10.2478/s13230-010- 0002-4.

Parameter Exploration for RLVR via Variational Learning URL https://doi.org/10.2478/s13230-010- 0002-4

Reference 2010

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:12.988600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:12.988600Z digest=sha256:6263b6646358b5d2448f08b36321ff1fb4dc46b06ae935ddd828f7483923e2ff

Observation f8a980b0-da8d-4898-8282-325438106671 · outbound

This paper cites an unresolved cited work.

Parameter Exploration for RLVR via Variational Learning Unresolved cited work

Reference 2011

Resolution
unresolved
raw_fallback, observed 2026-08-11T10:32:15.027341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:12.814815Z digest=sha256:b1b53b3888098126de11f72e6d865a94fc5bdab7629a97a0528cadc6e16c1cba

Observation f282f0ce-7eb4-47dd-8d52-efd6e25bef35 · outbound

This paper cites an unresolved cited work.

Parameter Exploration for RLVR via Variational Learning Unresolved cited work

Reference 2018

Resolution
unresolved
raw_fallback, observed 2026-08-11T10:32:15.052051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:12.801070Z digest=sha256:9db0b9b768f11b74254ff6ed61865261801263665e670e911412bfc8e16b12e4

Observation 937248b9-6481-4813-a1fc-2729e8ad8412 · outbound

This paper cites an unresolved cited work.

Parameter Exploration for RLVR via Variational Learning Unresolved cited work

Reference 2021

Resolution
unresolved
raw_fallback, observed 2026-08-11T10:32:14.868197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:12.924363Z digest=sha256:75283a743fe5af6a3d51c71a586d14846581a1e2cf8b7f9569e2e263e65df131

Observation a210d05b-a4f2-4727-9d7d-3edfdf420c50 · outbound

This paper cites an unresolved cited work.

Parameter Exploration for RLVR via Variational Learning Unresolved cited work

Reference 2024

Resolution
unresolved
raw_fallback, observed 2026-08-11T10:32:14.739485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:13.035731Z digest=sha256:47cabbb6ef54ee3bae01318e6d42c55b86ef17d1ed7bfe4868d3e105fdafe097

Pith citing papers

No inbound Pith citation observations are available.