Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-12T03:25:04.955816Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 1 inbound Pith citation observation for arXiv:2605.08817.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-12T03:25:04.955816Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-29T12:58:13.585763Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-06-29T13:03:26.327885Z
53 of 53 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 4f368f85-6f12-43d8-8bc5-79ba5c3f2624 · outbound
How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors Learning to explore with parameter-space noise: A deep dive into parameter-space noise for reinforcement learning with verifiable rewards
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4b0df4f7-fd55-45b2-a108-e368da976f4a · outbound
How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors Blei, Alp Kucukelbir, and Jon D
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e8051922-815f-4367-8c4b-300a29563b98 · outbound
How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 439ee1c2-8708-4a96-aed9-9632c154ff14 · outbound
How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors Seed-Prover: Deep and Broad Reasoning for Automated Theorem Proving
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ea1edc64-3503-4d43-a1e1-a809d0918803 · outbound
How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors Infogan: Interpretable representation learning by information maximizing generative adversarial nets
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a9173ba0-c924-44b1-b186-7d8790d0b7f9 · outbound
How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors An Empirical Study on Eliciting and Improving R1-like Reasoning Models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b2396714-c593-4147-b06c-da2776362bf8 · outbound
How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2b7d09e5-44fd-4fcd-9c04-15e1634f165e · outbound
How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors Training Verifiers to Solve Math Word Problems
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f60b3ed7-fa3b-48dc-a8f2-dfa714c6a319 · outbound
How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 436461c4-71fe-48be-8068-dff99b3dd313 · outbound
How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 34e702ce-cb53-4b81-8f1e-ae601ef5d328 · outbound
How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors Xing, and Zhiting Hu
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b9bf5241-aeae-434b-b038-f053c7dd3e60 · outbound
How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors Differential smooth- ing mitigates sharpening and improves llm reasoning.arXiv preprint arXiv:2511.19942
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6587639b-b977-462f-8ead-432998bb39a3 · outbound
How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors arXiv preprint arXiv:2505.17621 , year=
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation eb4c74fc-3128-4049-9902-bc49d27dd86f · outbound
How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation de2dd941-906a-4b27-be90-ee4af0ea2737 · outbound
How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors Measuring Mathematical Problem Solving With the MATH Dataset
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2952e0bd-3a0d-4bcd-b856-5c71194cc6e1 · outbound
How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors Diversity-incentivized exploration for versatile reasoning
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3b62f61a-6b29-43d9-99e4-14b92613575c · outbound
How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors Low-probability tokens sustain exploration in reinforcement learning with verifiable reward.arXiv preprint arXiv:2510.03222
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0059778b-b8b7-4a87-84cc-2619f3e89388 · outbound
How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors OpenAI o1 System Card
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation fbb45143-aa81-4026-92a9-848007ebcf15 · outbound
How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors Reasoning with Sampling: Your Base Model is Smarter Than You Think
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6f8b424e-ebc1-4de9-8e51-38621493438e · outbound
How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors LLM Post-Training: A Deep Dive into Reasoning Large Language Models
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9800f4d2-e055-43cd-9a09-9cec2c746a0d · outbound
How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors The power of scale for parameter-efficient prompt tuning
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 18f170f9-efa0-4dca-bb34-63685a49d6f2 · outbound
How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors Solving quan- titative reasoning problems with language models.Advances in neural information processing systems, 35:3843–3857
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a6a7c79f-606e-48bf-8ce5-e3ce59db90fc · outbound
How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors Prefix-tuning: Optimizing continuous prompts for generation
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a5dd20e1-20cd-4935-9e33-fc17ad9bff44 · outbound
How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors Let’s verify step by step
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8dceb4a6-4324-47e3-9e72-142660872234 · outbound
How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors P- Tuning: Prompt tuning can be comparable to fine-tuning across scales and tasks
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 10e50bd9-0962-402d-b744-22d692e2bd45 · outbound
How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors Unresolved cited work
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 528185d7-192e-4ee3-b95c-9839e87b720e · outbound
How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors Learning how to ask: Querying LMs with mixtures of soft prompts
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 38a28626-14ec-4270-8335-6447b26deba0 · outbound
How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors Toolformer: Language models can teach themselves to use tools.Advances in Neural Information Processing Systems, 36: 68539–68551
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e472d010-8778-496c-a67b-901562182a93 · outbound
How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors Upskill: Mutual information skill learning for structured response diversity in llms
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0fb8b061-f431-40c6-aa22-cc59c2bd523e · outbound
How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 78982944-9a84-4601-b6ac-87542b219e28 · outbound
How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors Logan IV , Eric Wallace, and Sameer Singh
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 036c8d27-ff8d-4ed0-b2ea-19671f01c92d · outbound
How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors OpenAI GPT-5 System Card
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e53a0f20-be19-42e0-9cdd-74e8d88b2ed4 · outbound
How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors Coda-prompt: Continual decomposed attention-based prompting for rehearsal-free continual learning
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 10e501fd-6580-4127-9e70-0ccc5cc4d645 · outbound
How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors Gtpo and grpo-s: Token and sequence-level reward shaping with policy entropy.arXiv preprint arXiv:2508.04349
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3bee837e-9f5a-4678-82f3-f62fb9e4e3aa · outbound
How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors SPoT: Better frozen model adaptation through soft prompt transfer
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 735b28f4-7c1a-40d1-a853-961c06c0b121 · outbound
How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5ed5422a-c503-448a-b535-65416d99516a · outbound
How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors Dualprompt: Complementary prompting for rehearsal-free continual learning
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3f98af22-5d3b-4f52-8a2b-23023d695a6c · outbound
How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors Learning to prompt for continual learning
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0c2b2a42-7aaf-4eb7-99f0-620fab430eb2 · outbound
How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors The invisible leash: Why rlvr may or may not escape its origin
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 86bcc709-06b1-4bc1-a6bc-c6e60cad6917 · outbound
How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors DeepSeek-Prover: Advancing Theorem Proving in LLMs through Large-Scale Synthetic Data
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1dbef7fa-7842-42d6-8cde-207846b5c6e1 · outbound
How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d8e3e400-fa51-4e04-a6cd-08b47dfe0513 · outbound
How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors Qwen2.5 Technical Report
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7a7e283b-8e73-441d-abf5-d8ece75c4355 · outbound
How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4bb19d9d-4d05-492d-9152-6470b6eca44e · outbound
How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors Qwen3 Technical Report
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation bd4aff42-9dd7-4a92-b89d-8f7968dfa5d4 · outbound
How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors Leandojo: Theorem proving with retrieval-augmented language models.Advances in Neural Information Processing Systems, 36: 21573–21612
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e64c9f67-4163-4174-a1f1-d7aaa56ff31c · outbound
How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors React: Synergizing reasoning and acting in language models
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3bc9414c-7a0e-4ec7-900f-5ef56d81adfd · outbound
How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 07d44e0c-edeb-4a08-a848-7e0629a8f0ed · outbound
How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0506aae2-4d04-4b22-9aaa-fb3297ab6b16 · outbound
How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors Count counts: Motivating exploration in llm reasoning with count-based intrinsic rewards
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d4df3c85-90ef-440d-8b2f-bd916f831d3d · outbound
How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors Group Sequence Policy Optimization
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f373a8cc-ce6f-4c16-8b6d-3b79152a151b · outbound
How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors First Return, Entropy-Eliciting Explore
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f5287d11-4896-4d14-84da-9d1e1943270e · outbound
How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors Instruction-Following Evaluation for Large Language Models
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b9f57432-8acc-44a6-87ad-c68a09f7a097 · outbound
How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors The surprising effectiveness of negative reinforcement in llm reasoning.arXiv preprint arXiv:2506.01347
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 79ed85e0-7648-49f4-9e6c-7af397234ec5 · inbound
EAPO: Entropy-Driven Adaptive Positive-Negative Sample Weighting for Policy Optimization in Open-Ended QA How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.