Pith. sign in

Paper Citation Record · LEDGER

How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors

As of 10 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 1 inbound Pith citation observation for arXiv:2605.08817.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.08817 v1

Coverage vector

measured 53 of 53 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-12T03:25:04.955816Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-29T12:58:13.585763Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-06-29T13:03:26.327885Z

Reference resolution

53 of 53 outbound references displayed

  • verified exact33
  • verified fuzzy17
  • unresolved1
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4f368f85-6f12-43d8-8bc5-79ba5c3f2624 · outbound

This paper cites Learning to explore with parameter-space noise: A deep dive into parameter-space noise for reinforcement learning with verifiable rewards.

How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors Learning to explore with parameter-space noise: A deep dive into parameter-space noise for reinforcement learning with verifiable rewards

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:26:19.278367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:25:04.955816Z digest=sha256:6f3ce81bf252bd77ddeeba111e28189c263422fb4896959c59cbbd50957aee8f

Observation 4b0df4f7-fd55-45b2-a108-e368da976f4a · outbound

This paper cites Blei, Alp Kucukelbir, and Jon D.

How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors Blei, Alp Kucukelbir, and Jon D

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T19:56:48.183326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:25:04.955816Z digest=sha256:a04d1ee24bf0f2519ae18fa7230b5d6132149a09b17cfd9a93fa1d3133c5c8e2

Observation e8051922-815f-4367-8c4b-300a29563b98 · outbound

This paper cites SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models.

How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:43:34.765363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:25:04.955816Z digest=sha256:98adafc12047fb100993375ad53073915788e9227b5d5005d59c67a634fad280

Observation 439ee1c2-8708-4a96-aed9-9632c154ff14 · outbound

This paper cites Seed-Prover: Deep and Broad Reasoning for Automated Theorem Proving.

How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors Seed-Prover: Deep and Broad Reasoning for Automated Theorem Proving

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:26:19.268538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:25:04.955816Z digest=sha256:3edcc9ec7991bbc85ddcc90d826d0aae58565586442e84c659ccb538a70ab043

Observation ea1edc64-3503-4d43-a1e1-a809d0918803 · outbound

This paper cites Infogan: Interpretable representation learning by information maximizing generative adversarial nets.

How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors Infogan: Interpretable representation learning by information maximizing generative adversarial nets

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T19:56:48.179838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:25:04.955816Z digest=sha256:ef92794a63f9c7d8489d64208df3c5c21d4a7972be2a657f5b778db8c0f20e05

Observation a9173ba0-c924-44b1-b186-7d8790d0b7f9 · outbound

This paper cites An Empirical Study on Eliciting and Improving R1-like Reasoning Models.

How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors An Empirical Study on Eliciting and Improving R1-like Reasoning Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:26:19.275213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:25:04.955816Z digest=sha256:08d5ad3d02d9658f58a281aa73f254da502900587f94f1eae5557e5d044293c6

Observation b2396714-c593-4147-b06c-da2776362bf8 · outbound

This paper cites Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models.

How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:26:19.271730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:25:04.955816Z digest=sha256:ff2ddda44724289ce0fd112a0827ec7871a347221f37d538e0c1b449fc1a2bc4

Observation 2b7d09e5-44fd-4fcd-9c04-15e1634f165e · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors Training Verifiers to Solve Math Word Problems

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-12T03:26:19.343068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:25:04.955816Z digest=sha256:26437c18301110b5357935f0bf4ab78571e00a36758ff5c0cff90b32c202f1f6

Observation f60b3ed7-fa3b-48dc-a8f2-dfa714c6a319 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-12T03:26:19.320307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:25:04.955816Z digest=sha256:1c6045b50807e3897aa6072e93ae95accb670708b25ebae7c904eff90b3f2b18

Observation 436461c4-71fe-48be-8068-dff99b3dd313 · outbound

This paper cites The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models.

How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-12T12:31:34.026635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:25:04.955816Z digest=sha256:8ebcbdb5afdb877441b73a27fc5a7cb839ef503027ed9edf37cc5956d4fc74b7

Observation 34e702ce-cb53-4b81-8f1e-ae601ef5d328 · outbound

This paper cites Xing, and Zhiting Hu.

How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors Xing, and Zhiting Hu

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T19:56:48.131519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:25:04.955816Z digest=sha256:fb6701f2594dd2d3e357d4d223c5f3e71bc7e9c10ad665ca012e822d55e7a40d

Observation b9bf5241-aeae-434b-b038-f053c7dd3e60 · outbound

This paper cites Differential smooth- ing mitigates sharpening and improves llm reasoning.arXiv preprint arXiv:2511.19942.

How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors Differential smooth- ing mitigates sharpening and improves llm reasoning.arXiv preprint arXiv:2511.19942

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:26:19.297284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:25:04.955816Z digest=sha256:e16bebd9f6f805d93c96763e3b67fe440d284250e36bf4675610993dfcf3bea4

Observation 6587639b-b977-462f-8ead-432998bb39a3 · outbound

This paper cites arXiv preprint arXiv:2505.17621 , year=.

How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors arXiv preprint arXiv:2505.17621 , year=

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:26:19.303759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:25:04.955816Z digest=sha256:f328b2dacc9fbfe489a6d5b975522736e9ef6e11becf18091c993cdbede6b7ee

Observation eb4c74fc-3128-4049-9902-bc49d27dd86f · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-12T03:26:19.354147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:25:04.955816Z digest=sha256:58a823f53e3203ef170bb7d9f22ccb15988a70ea7e2dfc6284d09af93f82f61d

Observation de2dd941-906a-4b27-be90-ee4af0ea2737 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors Measuring Mathematical Problem Solving With the MATH Dataset

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-12T03:26:19.325851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:25:04.955816Z digest=sha256:822ce6c814b518bcc4a4221a9f89074d250fc580c6042f10f9c384c899d9e9d9

Observation 2952e0bd-3a0d-4bcd-b856-5c71194cc6e1 · outbound

This paper cites Diversity-incentivized exploration for versatile reasoning.

How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors Diversity-incentivized exploration for versatile reasoning

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:26:19.337760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:25:04.955816Z digest=sha256:412e5de6b97c1fc8d02659c3d651af5f0c86cc013531dc26fc28ce2d402429da

Observation 3b62f61a-6b29-43d9-99e4-14b92613575c · outbound

This paper cites Low-probability tokens sustain exploration in reinforcement learning with verifiable reward.arXiv preprint arXiv:2510.03222.

How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors Low-probability tokens sustain exploration in reinforcement learning with verifiable reward.arXiv preprint arXiv:2510.03222

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:26:19.282521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:25:04.955816Z digest=sha256:f8125765a04e7911993a08d2026ccafd4408efb4535b555d7cb067eb5ae77afa

Observation 0059778b-b8b7-4a87-84cc-2619f3e89388 · outbound

This paper cites OpenAI o1 System Card.

How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors OpenAI o1 System Card

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-12T03:26:19.345462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:25:04.955816Z digest=sha256:0ad2b4699335bf3e5f1d8d1ea6cb3528f9033cb2e6e0c124b7d2bc1f971ddb80

Observation fbb45143-aa81-4026-92a9-848007ebcf15 · outbound

This paper cites Reasoning with Sampling: Your Base Model is Smarter Than You Think.

How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors Reasoning with Sampling: Your Base Model is Smarter Than You Think

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:16:05.617665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:25:04.955816Z digest=sha256:40a59a7ef27cb3691649fa531107cca959d2aa8536dccb760e771b0aa1786d1b

Observation 6f8b424e-ebc1-4de9-8e51-38621493438e · outbound

This paper cites LLM Post-Training: A Deep Dive into Reasoning Large Language Models.

How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors LLM Post-Training: A Deep Dive into Reasoning Large Language Models

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T03:26:19.317306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:25:04.955816Z digest=sha256:3b1c185f3013028ae167ceeb11d77d0105747741878fafc9c240516536aeef55

Observation 9800f4d2-e055-43cd-9a09-9cec2c746a0d · outbound

This paper cites The power of scale for parameter-efficient prompt tuning.

How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors The power of scale for parameter-efficient prompt tuning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T19:56:48.172842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:25:04.955816Z digest=sha256:a41661386cee6775664d4e4327f4f9668523c906c4151b1abdea5ffd3386d6b7

Observation 18f170f9-efa0-4dca-bb34-63685a49d6f2 · outbound

This paper cites Solving quan- titative reasoning problems with language models.Advances in neural information processing systems, 35:3843–3857.

How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors Solving quan- titative reasoning problems with language models.Advances in neural information processing systems, 35:3843–3857

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T19:56:48.176723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:25:04.955816Z digest=sha256:941dda242e97fd1fcc15e559af7596aafbcc915db4a284036de8930d7224779d

Observation a6a7c79f-606e-48bf-8ce5-e3ce59db90fc · outbound

This paper cites Prefix-tuning: Optimizing continuous prompts for generation.

How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors Prefix-tuning: Optimizing continuous prompts for generation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T19:56:48.158593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:25:04.955816Z digest=sha256:b4fd5e67d56f24ff4be6bdb395a3a5f45a4b3a709e72de49b0e627794e5528a9

Observation a5dd20e1-20cd-4935-9e33-fc17ad9bff44 · outbound

This paper cites Let’s verify step by step.

How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors Let’s verify step by step

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T19:56:48.145755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:25:04.955816Z digest=sha256:6458a71f34de8892e7ed9e95eefaf634607983ec41c4fe2da18477791f7f9a8a

Observation 8dceb4a6-4324-47e3-9e72-142660872234 · outbound

This paper cites P- Tuning: Prompt tuning can be comparable to fine-tuning across scales and tasks.

How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors P- Tuning: Prompt tuning can be comparable to fine-tuning across scales and tasks

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T19:56:48.138754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:25:04.955816Z digest=sha256:34b1411c38fdefc7f1cc98fb7364642f113d5d2754eb29a19717de7fa7e674b8

Observation 10e50bd9-0962-402d-b744-22d692e2bd45 · outbound

This paper cites an unresolved cited work.

How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-05-12T19:56:48.142150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:25:04.955816Z digest=sha256:4a9eaa9914a10b57f83fa5196d7733bbd8022b77049b2fd252aba29414bd5326

Observation 528185d7-192e-4ee3-b95c-9839e87b720e · outbound

This paper cites Learning how to ask: Querying LMs with mixtures of soft prompts.

How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors Learning how to ask: Querying LMs with mixtures of soft prompts

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T19:56:48.134962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:25:04.955816Z digest=sha256:a5e87192540c4edb1063374643090d80193173dbc5c4853148219456816cb883

Observation 38a28626-14ec-4270-8335-6447b26deba0 · outbound

This paper cites Toolformer: Language models can teach themselves to use tools.Advances in Neural Information Processing Systems, 36: 68539–68551.

How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors Toolformer: Language models can teach themselves to use tools.Advances in Neural Information Processing Systems, 36: 68539–68551

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T19:56:48.169795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:25:04.955816Z digest=sha256:14398cea51aaff543a380e34d9c337cd235f974ab419fe771a208fc45a9e6090

Observation e472d010-8778-496c-a67b-901562182a93 · outbound

This paper cites Upskill: Mutual information skill learning for structured response diversity in llms.

How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors Upskill: Mutual information skill learning for structured response diversity in llms

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:26:19.331003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:25:04.955816Z digest=sha256:17dbfc6ebfb21da87f3c3ce4f491c8f19ad20410360ed49f35569df82026f938

Observation 0fb8b061-f431-40c6-aa22-cc59c2bd523e · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-12T03:26:19.288149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:25:04.955816Z digest=sha256:fba5b0ab7f7706df2b8f4da6b79312b52b324baaacbb2c56b97e01e09ab4cd8a

Observation 78982944-9a84-4601-b6ac-87542b219e28 · outbound

This paper cites Logan IV , Eric Wallace, and Sameer Singh.

How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors Logan IV , Eric Wallace, and Sameer Singh

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T19:56:48.162074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:25:04.955816Z digest=sha256:f0c03adb7df3afbe5ac76454b48a8229dbdc11d2ea9631aaae0248dc7b580a91

Observation 036c8d27-ff8d-4ed0-b2ea-19671f01c92d · outbound

This paper cites OpenAI GPT-5 System Card.

How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors OpenAI GPT-5 System Card

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-12T03:26:19.340247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:25:04.955816Z digest=sha256:877ce51f4486c2f98b43c6c05106eb2f4326fe6f421089b9e018f7d94989baae

Observation e53a0f20-be19-42e0-9cdd-74e8d88b2ed4 · outbound

This paper cites Coda-prompt: Continual decomposed attention-based prompting for rehearsal-free continual learning.

How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors Coda-prompt: Continual decomposed attention-based prompting for rehearsal-free continual learning

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T19:56:48.165637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:25:04.955816Z digest=sha256:d21fad9c6dc9cc2cf4922f459c6ca5b585ce43329206bb865b79571e04fa318f

Observation 10e501fd-6580-4127-9e70-0ccc5cc4d645 · outbound

This paper cites Gtpo and grpo-s: Token and sequence-level reward shaping with policy entropy.arXiv preprint arXiv:2508.04349.

How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors Gtpo and grpo-s: Token and sequence-level reward shaping with policy entropy.arXiv preprint arXiv:2508.04349

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:26:19.306763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:25:04.955816Z digest=sha256:23f9dd9ec8abbee271ecc05e00e7f5067c62e5dbfd329ef828161c8a8aaf0c52

Observation 3bee837e-9f5a-4678-82f3-f62fb9e4e3aa · outbound

This paper cites SPoT: Better frozen model adaptation through soft prompt transfer.

How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors SPoT: Better frozen model adaptation through soft prompt transfer

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T19:56:48.148900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:25:04.955816Z digest=sha256:7266612f1565ffa900b1feb2fd26284fe1044e94744465441cdecf93961af320

Observation 735b28f4-7c1a-40d1-a853-961c06c0b121 · outbound

This paper cites Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning.

How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-12T12:12:09.004283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:25:04.955816Z digest=sha256:6f71d9c797fe1a8b48f659c7ec9ffef5ea8e2e4f1b4572083aaf7e56c0d70c56

Observation 5ed5422a-c503-448a-b535-65416d99516a · outbound

This paper cites Dualprompt: Complementary prompting for rehearsal-free continual learning.

How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors Dualprompt: Complementary prompting for rehearsal-free continual learning

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T19:56:48.123054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:25:04.955816Z digest=sha256:60f8fa11da1325af647f4579034cf46ce5c931d8b26fcf5c8d883c4884e8b47e

Observation 3f98af22-5d3b-4f52-8a2b-23023d695a6c · outbound

This paper cites Learning to prompt for continual learning.

How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors Learning to prompt for continual learning

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T19:56:48.155640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:25:04.955816Z digest=sha256:bff448e784204609c26332e325c957b4e220a715901c2842d9f0e8faad958678

Observation 0c2b2a42-7aaf-4eb7-99f0-620fab430eb2 · outbound

This paper cites The invisible leash: Why rlvr may or may not escape its origin.

How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors The invisible leash: Why rlvr may or may not escape its origin

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:26:19.313766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:25:04.955816Z digest=sha256:8cd00e85fc1667d002504277f5ef7656b81eda15b413e8b55323a60d0ee4ec30

Observation 86bcc709-06b1-4bc1-a6bc-c6e60cad6917 · outbound

This paper cites DeepSeek-Prover: Advancing Theorem Proving in LLMs through Large-Scale Synthetic Data.

How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors DeepSeek-Prover: Advancing Theorem Proving in LLMs through Large-Scale Synthetic Data

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:26:19.285607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:25:04.955816Z digest=sha256:a62b934eccce3e2a56656fa7029ff97bc22e332d7836724671d2ceaa2e2a926e

Observation 1dbef7fa-7842-42d6-8cde-207846b5c6e1 · outbound

This paper cites Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models.

How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-15T21:20:59.729674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:25:04.955816Z digest=sha256:3b826d33055750cf097d5e93ff20c11d023cb33bdcd405682a57d14b270efcb8

Observation d8e3e400-fa51-4e04-a6cd-08b47dfe0513 · outbound

This paper cites Qwen2.5 Technical Report.

How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors Qwen2.5 Technical Report

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-12T03:26:19.323364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:25:04.955816Z digest=sha256:15e66e505aa28cada08b9497057efbe3f49bca61054578dd0684e00dd1cbb29e

Observation 7a7e283b-8e73-441d-abf5-d8ece75c4355 · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-05-12T03:26:19.368149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:25:04.955816Z digest=sha256:018f7772c855d2dfad846fff9a9d61990dcf56d85ef7b1458daed6b78331faa9

Observation 4bb19d9d-4d05-492d-9152-6470b6eca44e · outbound

This paper cites Qwen3 Technical Report.

How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors Qwen3 Technical Report

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-12T03:26:19.356776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:25:04.955816Z digest=sha256:d6e112f36a51af6d703cd6d9de6ac1502deac21ef5c51ca71866c18f7810ff2c

Observation bd4aff42-9dd7-4a92-b89d-8f7968dfa5d4 · outbound

This paper cites Leandojo: Theorem proving with retrieval-augmented language models.Advances in Neural Information Processing Systems, 36: 21573–21612.

How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors Leandojo: Theorem proving with retrieval-augmented language models.Advances in Neural Information Processing Systems, 36: 21573–21612

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T19:56:48.152430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:25:04.955816Z digest=sha256:f7aaa0fde69c547357da0852f04116b9215258e5e240cf1e7724d3a42acb0a91

Observation e64c9f67-4163-4174-a1f1-d7aaa56ff31c · outbound

This paper cites React: Synergizing reasoning and acting in language models.

How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors React: Synergizing reasoning and acting in language models

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T19:56:48.128189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:25:04.955816Z digest=sha256:450ecb4bb75879227042774d59e918988330ddc869366c37912a40a9f4a9feb0

Observation 3bc9414c-7a0e-4ec7-900f-5ef56d81adfd · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-05-12T03:26:19.351483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:25:04.955816Z digest=sha256:f12a293925feaae89d8bc73f435ee46657e02d6abe55caae6231421cd3b15664

Observation 07d44e0c-edeb-4a08-a848-7e0629a8f0ed · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?.

How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-05-12T03:26:19.365465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:25:04.955816Z digest=sha256:d4e43421a1530e4e87ca8ffc60f182ab50d0205459af7178a034a31d834cbd79

Observation 0506aae2-4d04-4b22-9aaa-fb3297ab6b16 · outbound

This paper cites Count counts: Motivating exploration in llm reasoning with count-based intrinsic rewards.

How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors Count counts: Motivating exploration in llm reasoning with count-based intrinsic rewards

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:26:19.359601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:25:04.955816Z digest=sha256:1f1be3ab8ce9a717e4a1227c69577d51a403421fa0eb44ee77474aaaa1e71e81

Observation d4df3c85-90ef-440d-8b2f-bd916f831d3d · outbound

This paper cites Group Sequence Policy Optimization.

How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors Group Sequence Policy Optimization

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-05-12T03:26:19.333785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:25:04.955816Z digest=sha256:77376e90615b93abb3a94caf72e8f73525d73eb2be09760f4922f104cffbd79a

Observation f373a8cc-ce6f-4c16-8b6d-3b79152a151b · outbound

This paper cites First Return, Entropy-Eliciting Explore.

How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors First Return, Entropy-Eliciting Explore

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:26:19.362669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:25:04.955816Z digest=sha256:2e9f67bcd7e11838634fc968d6257eaaf2995963a858466974eb1136c1648c90

Observation f5287d11-4896-4d14-84da-9d1e1943270e · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors Instruction-Following Evaluation for Large Language Models

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-05-12T03:26:19.328220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:25:04.955816Z digest=sha256:cf4c0baecbe01a04ac3b09a1df2ecaa772519d4561e9669571737c60a61ac0c0

Observation b9f57432-8acc-44a6-87ad-c68a09f7a097 · outbound

This paper cites The surprising effectiveness of negative reinforcement in llm reasoning.arXiv preprint arXiv:2506.01347.

How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors The surprising effectiveness of negative reinforcement in llm reasoning.arXiv preprint arXiv:2506.01347

Reference 54

Resolution
malformed identifier
arxiv_id, observed 2026-05-12T03:26:19.310174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:25:04.955816Z digest=sha256:a2cb9335df12e4ab5da4c34f3e1008ae9ec521104937e0ba52d0cc607c0d92e8

Pith citing papers

Observation 79ed85e0-7648-49f4-9e6c-7af397234ec5 · inbound

EAPO: Entropy-Driven Adaptive Positive-Negative Sample Weighting for Policy Optimization in Open-Ended QA cites this paper.

EAPO: Entropy-Driven Adaptive Positive-Negative Sample Weighting for Policy Optimization in Open-Ended QA How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T13:03:26.329169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T12:58:13.585763Z digest=sha256:bd2b6c48849eb19d0b35e827be1215ce78d1e0d1a1374cbd29477e66d988398e