Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-28T11:13:21.565082Z
Paper Citation Record · LEDGER
As of 4 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 0 inbound Pith citation observations for arXiv:2606.03234.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-28T11:13:21.565082Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
37 of 37 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 726fc06c-2251-462f-94f6-a29321e6aeb6 · outbound
Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning Probing classifiers: Promises, shortcomings, and advances.Computational Linguistics, 48(1): 207–219, 2022
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 250b1fc2-6150-4f56-9b67-d4e689fd5286 · outbound
Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning Evaluating Large Language Models Trained on Code
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 19c9a916-55fe-4077-8575-8634ae7dc557 · outbound
Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning Fapo: flawed-aware policy optimization for efficient and reliable reasoning.arXiv preprint arXiv:2510.22543, 2025
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 69e080db-36eb-454d-a052-5d12825e52b2 · outbound
Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 14f14886-e3b6-4760-a694-4d6fd2fc28e7 · outbound
Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning Reward inside the model: A lightweight hidden-state reward model for llm’s best-of-n sampling
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 099b8cf7-d684-4de6-97dd-e85f8e3ac597 · outbound
Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning Foundation Models for Semantic Novelty in Reinforcement Learning
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a675e5bc-513f-407a-b0a4-be8e59e389c7 · outbound
Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning HMMT february competition.https://www.hmmt.org/, 2025
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8fee68db-00a0-4eed-a189-405a5936fdd5 · outbound
Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning Rewarding the unlikely: Lifting grpo beyond distribution sharpening
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e4d5daa-c12f-4ff6-aaa4-0077cb4c55da · outbound
Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc77d68d-c518-4809-9aa7-c3db5c2fec50 · outbound
Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 039a75b8-36bc-4c93-831c-4cd08087cc24 · outbound
Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning Solving quantitative reasoning problems with language models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba8ccdcc-64d8-4c07-b1a1-472c9b830409 · outbound
Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning Leveraging Error Diversity in Group Rollouts for Reinforcement Learning
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 267d07a5-f394-44c3-ac86-4d9bd858a288 · outbound
Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning Understanding R1-Zero-Like Training: A Critical Perspective
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 918abf7b-a54a-41c8-90bd-b7f399d6f19a · outbound
Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning Contrastive Reasoning Alignment: Reinforcement Learning from Hidden Representations
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 793842e9-72be-44c2-a5f1-407c71554794 · outbound
Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning AIME problems and solutions.https://maa.org/, 2024
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fac7aaa3-596d-4943-9d20-4c24dc7bff4e · outbound
Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning AMC 10/12 problems and solutions.https://maa.org/, 2024
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1fa4859f-0667-46c4-976d-96d8ae349115 · outbound
Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning Ngrpo: Negative-enhanced group relative policy optimization
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 24f2fb66-db74-4aa9-9cf4-8eeeeb01ba83 · outbound
Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning Training language models to follow instructions with human feedback
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff243a49-9c3d-412f-b6a0-82d1339436cf · outbound
Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning Relational knowledge distillation
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c74f7ca1-cf00-4cef-9c71-7b6cc29f80e6 · outbound
Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning RIDE: Rewarding Impact-Driven Exploration for Procedurally-Generated Environments
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 902a469f-cc30-484f-b1e0-a483f8043394 · outbound
Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 459eabe6-40ea-4059-8e66-c0f25668fd16 · outbound
Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9704f948-8cf8-4ecc-bf8b-84d2697d3c85 · outbound
Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning arXiv preprint arXiv:2508.03772 , year=
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c235b19d-da3d-4111-83a3-00a5882c1576 · outbound
Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning LLM Reasoning as Trajectories: Step-Specific Representation Geometry and Correctness Signals
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3f29e08e-eefc-48db-9ad6-1781d50a28eb · outbound
Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning arXiv preprint arXiv:2511.00794 , year=
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 17239b58-90e8-4605-80a4-fe3b2a9c386e · outbound
Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning Contrastive Representation Distillation
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f4b768f0-1e6a-47a8-bdc8-7827472204c5 · outbound
Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning LLM-Empowered State Representation for Reinforcement Learning
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 46e0c3db-2e8d-4f1c-83aa-0f4ed8e81d79 · outbound
Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning Closing the Modality Reasoning Gap for Speech Large Language Models
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation fe052b4f-c69f-4a88-bf87-f8e3ce9a11ff · outbound
Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning Step-wise Rubric Rewards for LLM Reasoning
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0d39a0e1-4727-46e6-8ce0-d9c40a002928 · outbound
Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning Qwen3 Technical Report
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 89266322-9a9d-44f1-b6b3-a9eff2d05f24 · outbound
Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning Regularizing hidden states enables learning generalizable reward model for llms.Advancesin Neural Information Processing Systems, 37:62279–62309, 2024
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1fe08952-f9bf-4b2c-b525-e7640571654f · outbound
Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning Dapo: An open-source llm reinforcement learning system at scale.Advances in Neural Information Processing Systems, 38:113222–113244, 2026
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba1d5106-159f-4966-ad2f-234fd759a8b2 · outbound
Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 95de35a7-47dd-4860-bd5c-b3c0f5ed3b1f · outbound
Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a1bfd013-df4d-4e38-ba6b-10f73cbb447e · outbound
Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning Geometric-mean policy optimization
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1bb72bb5-4672-4b52-818b-dabbb4a86bc7 · outbound
Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning Group Sequence Policy Optimization
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ccaefb89-7afd-4951-bd42-546bec4ef861 · outbound
Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning Representation Engineering: A Top-Down Approach to AI Transparency
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
No inbound Pith citation observations are available.