Pith. sign in

Paper Citation Record · LEDGER

Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning

As of 4 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 0 inbound Pith citation observations for arXiv:2606.03234.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.03234 v1

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-28T11:13:21.565082Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

37 of 37 outbound references displayed

  • verified exact21
  • verified fuzzy0
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch4

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 726fc06c-2251-462f-94f6-a29321e6aeb6 · outbound

This paper cites Probing classifiers: Promises, shortcomings, and advances.Computational Linguistics, 48(1): 207–219, 2022.

Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning Probing classifiers: Promises, shortcomings, and advances.Computational Linguistics, 48(1): 207–219, 2022

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-28T11:13:21.565082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T11:13:21.565082Z digest=sha256:12028c0b122d2d5f08ba6e42f51514d911f791e089b8288ead591ab289638ec0

Observation 250b1fc2-6150-4f56-9b67-d4e689fd5286 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning Evaluating Large Language Models Trained on Code

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:06:27.261820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T11:13:21.565082Z digest=sha256:a6b873ac1d16c2f2439353d0da9e284f387559ebcdb1a9ef2c15ca2e1c4128a1

Observation 19c9a916-55fe-4077-8575-8634ae7dc557 · outbound

This paper cites Fapo: flawed-aware policy optimization for efficient and reliable reasoning.arXiv preprint arXiv:2510.22543, 2025.

Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning Fapo: flawed-aware policy optimization for efficient and reliable reasoning.arXiv preprint arXiv:2510.22543, 2025

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:06:27.258479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T11:13:21.565082Z digest=sha256:765c74bde7a0ff669a0e3b483ebec4c8709fdd226ae5cc7962b311acdd722bb4

Observation 69e080db-36eb-454d-a052-5d12825e52b2 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:06:27.277353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T11:13:21.565082Z digest=sha256:cefda7b1ce6c7b29cd9d2e8d380cd3d3773abc85d9171d81d23125bb53d309a6

Observation 14f14886-e3b6-4760-a694-4d6fd2fc28e7 · outbound

This paper cites Reward inside the model: A lightweight hidden-state reward model for llm’s best-of-n sampling.

Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning Reward inside the model: A lightweight hidden-state reward model for llm’s best-of-n sampling

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-28T11:13:21.565082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T11:13:21.565082Z digest=sha256:be8464c58869c12d6b89e1cf8214de7a4cb733fec4e1d25fd6d09e2ccfed24e7

Observation 099b8cf7-d684-4de6-97dd-e85f8e3ac597 · outbound

This paper cites Foundation Models for Semantic Novelty in Reinforcement Learning.

Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning Foundation Models for Semantic Novelty in Reinforcement Learning

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T02:06:27.283338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T11:13:21.565082Z digest=sha256:2225ee7b6cd6969f09711712e1fe16ae8a4a4d236fbc1a5b1a5ada7575dd3c47

Observation a675e5bc-513f-407a-b0a4-be8e59e389c7 · outbound

This paper cites HMMT february competition.https://www.hmmt.org/, 2025.

Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning HMMT february competition.https://www.hmmt.org/, 2025

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-28T11:13:21.565082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T11:13:21.565082Z digest=sha256:1bb565edc7b1ea76181dc320d63e3b1c804b9328bfd2657ab431fea1bdf832c7

Observation 8fee68db-00a0-4eed-a189-405a5936fdd5 · outbound

This paper cites Rewarding the unlikely: Lifting grpo beyond distribution sharpening.

Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning Rewarding the unlikely: Lifting grpo beyond distribution sharpening

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-28T11:13:21.565082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T11:13:21.565082Z digest=sha256:8911886fc1d6f5825e2dc64b11914116cbf5c8c9d59c94df0eb6d07f0c433ff6

Observation 5e4d5daa-c12f-4ff6-aaa4-0077cb4c55da · outbound

This paper cites Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems.

Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-28T11:13:21.565082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T11:13:21.565082Z digest=sha256:09ea61be2f48335285f75ccbff286e0783dbf5910c996fa624b269d1f348e64d

Observation bc77d68d-c518-4809-9aa7-c3db5c2fec50 · outbound

This paper cites APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization.

Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:06:27.261864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T11:13:21.565082Z digest=sha256:afa60d5004cb066975d0a65e35a8cde697dc7746a479aea027027874a8f65c58

Observation 039a75b8-36bc-4c93-831c-4cd08087cc24 · outbound

This paper cites Solving quantitative reasoning problems with language models.

Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning Solving quantitative reasoning problems with language models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-28T11:13:21.565082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T11:13:21.565082Z digest=sha256:55ba4adbc19eb84d1c530d0d6ceaac702d731144acb4cdf72eefe921c922f740

Observation ba8ccdcc-64d8-4c07-b1a1-472c9b830409 · outbound

This paper cites Leveraging Error Diversity in Group Rollouts for Reinforcement Learning.

Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning Leveraging Error Diversity in Group Rollouts for Reinforcement Learning

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:06:27.264764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T11:13:21.565082Z digest=sha256:4f6334b3fda8e3f7f431b56123f3a6f0c646b2a5eea4b538389256f35331285b

Observation 267d07a5-f394-44c3-ac86-4d9bd858a288 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning Understanding R1-Zero-Like Training: A Critical Perspective

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:06:27.258704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T11:13:21.565082Z digest=sha256:93fc2311f461c203ff0d964c361215f1d2cfb307afb450034728d73e6025d1f1

Observation 918abf7b-a54a-41c8-90bd-b7f399d6f19a · outbound

This paper cites Contrastive Reasoning Alignment: Reinforcement Learning from Hidden Representations.

Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning Contrastive Reasoning Alignment: Reinforcement Learning from Hidden Representations

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:06:27.267586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T11:13:21.565082Z digest=sha256:4e7a201c37053d0b34508cdf4c1893893218b0568f5953bc36dae04c539f2dd0

Observation 793842e9-72be-44c2-a5f1-407c71554794 · outbound

This paper cites AIME problems and solutions.https://maa.org/, 2024.

Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning AIME problems and solutions.https://maa.org/, 2024

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-28T11:13:21.565082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T11:13:21.565082Z digest=sha256:d37c7cdb11e93a41d242c7ee63e8c0388fba9e20abcb32f9c82e2aea4f12007f

Observation fac7aaa3-596d-4943-9d20-4c24dc7bff4e · outbound

This paper cites AMC 10/12 problems and solutions.https://maa.org/, 2024.

Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning AMC 10/12 problems and solutions.https://maa.org/, 2024

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-28T11:13:21.565082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T11:13:21.565082Z digest=sha256:1ecb65216cc4422c0389b783b75c68ed31e10b2501036ca137724e4adbd78b5c

Observation 1fa4859f-0667-46c4-976d-96d8ae349115 · outbound

This paper cites Ngrpo: Negative-enhanced group relative policy optimization.

Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning Ngrpo: Negative-enhanced group relative policy optimization

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:06:27.249056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T11:13:21.565082Z digest=sha256:5b6c2ca4245e24f3256e0821611051b1fc66a01da917fdaf42962507cc201623

Observation 24f2fb66-db74-4aa9-9cf4-8eeeeb01ba83 · outbound

This paper cites Training language models to follow instructions with human feedback.

Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning Training language models to follow instructions with human feedback

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-28T11:13:21.565082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T11:13:21.565082Z digest=sha256:48c34bad5388486cf4f555699b8e5330d7df6c9c2373117942d1ec99272b7b33

Observation ff243a49-9c3d-412f-b6a0-82d1339436cf · outbound

This paper cites Relational knowledge distillation.

Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning Relational knowledge distillation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-28T11:13:21.565082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T11:13:21.565082Z digest=sha256:fdbf6f66438b17fd5456b35df6e76b801c7c003edddfecc268a0bc387988d7a1

Observation c74f7ca1-cf00-4cef-9c71-7b6cc29f80e6 · outbound

This paper cites RIDE: Rewarding Impact-Driven Exploration for Procedurally-Generated Environments.

Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning RIDE: Rewarding Impact-Driven Exploration for Procedurally-Generated Environments

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T02:06:27.252391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T11:13:21.565082Z digest=sha256:d0c4bdeca0598636f7ff826f9db2d679eaef3e5d699e7c086f84d74039c29ef7

Observation 902a469f-cc30-484f-b1e0-a483f8043394 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:06:27.268188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T11:13:21.565082Z digest=sha256:f7ca7863d17f63137179bd6cec2e9ef7585276be0c48a9f05947a0a020f082ab

Observation 459eabe6-40ea-4059-8e66-c0f25668fd16 · outbound

This paper cites Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning.

Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:06:27.265290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T11:13:21.565082Z digest=sha256:e6c976715cf6a5891e37fd7be956b56b3a31a33fcad909674008d3a62ebce369

Observation 9704f948-8cf8-4ecc-bf8b-84d2697d3c85 · outbound

This paper cites arXiv preprint arXiv:2508.03772 , year=.

Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning arXiv preprint arXiv:2508.03772 , year=

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:06:27.274217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T11:13:21.565082Z digest=sha256:44dd823fa16775b368071748759484b6509674c368f7a4223005425fafea8dcb

Observation c235b19d-da3d-4111-83a3-00a5882c1576 · outbound

This paper cites LLM Reasoning as Trajectories: Step-Specific Representation Geometry and Correctness Signals.

Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning LLM Reasoning as Trajectories: Step-Specific Representation Geometry and Correctness Signals

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:06:27.280205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T11:13:21.565082Z digest=sha256:154f47bc325dff43366ded6bbc801925e88d5bfd84e43f2dcbf958093e249d40

Observation 3f29e08e-eefc-48db-9ad6-1781d50a28eb · outbound

This paper cites arXiv preprint arXiv:2511.00794 , year=.

Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning arXiv preprint arXiv:2511.00794 , year=

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:06:27.255280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T11:13:21.565082Z digest=sha256:2035f24899850496fa87ce97ecbc0cca4ea7fd721622be0242cc8fce69ffb12f

Observation 17239b58-90e8-4605-80a4-fe3b2a9c386e · outbound

This paper cites Contrastive Representation Distillation.

Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning Contrastive Representation Distillation

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T02:06:27.234120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T11:13:21.565082Z digest=sha256:15d41b81acfd147971ec5676435cd75aeccf492071cdb7ef926ca19072b29051

Observation f4b768f0-1e6a-47a8-bdc8-7827472204c5 · outbound

This paper cites LLM-Empowered State Representation for Reinforcement Learning.

Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning LLM-Empowered State Representation for Reinforcement Learning

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:06:27.233967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T11:13:21.565082Z digest=sha256:f08339bfed26fd8b4213cd5c87c98c4704a331c071d59344612c02e02b53d4f8

Observation 46e0c3db-2e8d-4f1c-83aa-0f4ed8e81d79 · outbound

This paper cites Closing the Modality Reasoning Gap for Speech Large Language Models.

Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning Closing the Modality Reasoning Gap for Speech Large Language Models

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:06:27.220432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T11:13:21.565082Z digest=sha256:68c05baa9d984e8ef6eac2f104f81a158e07ac708626f4f03b8c1d6f1bc018a7

Observation fe052b4f-c69f-4a88-bf87-f8e3ce9a11ff · outbound

This paper cites Step-wise Rubric Rewards for LLM Reasoning.

Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning Step-wise Rubric Rewards for LLM Reasoning

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:06:27.223751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T11:13:21.565082Z digest=sha256:c96e48361c0b4a2a866120981b4840874ecb0151e983b54e5315a0bde59158a2

Observation 0d39a0e1-4727-46e6-8ce0-d9c40a002928 · outbound

This paper cites Qwen3 Technical Report.

Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning Qwen3 Technical Report

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:06:27.226743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T11:13:21.565082Z digest=sha256:a7d648aa0c900a9ea47c2e16f93b5875cd84cda1580f600dd2762e37d5aa805d

Observation 89266322-9a9d-44f1-b6b3-a9eff2d05f24 · outbound

This paper cites Regularizing hidden states enables learning generalizable reward model for llms.Advancesin Neural Information Processing Systems, 37:62279–62309, 2024.

Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning Regularizing hidden states enables learning generalizable reward model for llms.Advancesin Neural Information Processing Systems, 37:62279–62309, 2024

Reference 31

Resolution
unresolved
no resolver link, observed 2026-06-28T11:13:21.565082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T11:13:21.565082Z digest=sha256:e0b05b3546afd64deab97ac97b3c7286b0da0299990fefeea6bb824248a53d28

Observation 1fe08952-f9bf-4b2c-b525-e7640571654f · outbound

This paper cites Dapo: An open-source llm reinforcement learning system at scale.Advances in Neural Information Processing Systems, 38:113222–113244, 2026.

Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning Dapo: An open-source llm reinforcement learning system at scale.Advances in Neural Information Processing Systems, 38:113222–113244, 2026

Reference 32

Resolution
unresolved
no resolver link, observed 2026-06-28T11:13:21.565082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T11:13:21.565082Z digest=sha256:9fe6877f6ffa9ccb8e46dd5e4e58582cdec8edcb1edfc6a3f61b3fbac734f424

Observation ba1d5106-159f-4966-ad2f-234fd759a8b2 · outbound

This paper cites Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think.

Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:06:27.241288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T11:13:21.565082Z digest=sha256:14d3a5b102544b971723867eaae1bc5ea0a0f5dea0e760b24f0974fb882ab1c6

Observation 95de35a7-47dd-4860-bd5c-b3c0f5ed3b1f · outbound

This paper cites VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks.

Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:06:27.270607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T11:13:21.565082Z digest=sha256:ecdc1b10f9095ba5b2c9f97141ec843136ca4d30c8a8345e8e88638e58c6805e

Observation a1bfd013-df4d-4e38-ba6b-10f73cbb447e · outbound

This paper cites Geometric-mean policy optimization.

Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning Geometric-mean policy optimization

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:06:27.226675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T11:13:21.565082Z digest=sha256:18cdf99a539b1b6ea5f880cd6338f81160c296cbc53af63ce4f5d959730427eb

Observation 1bb72bb5-4672-4b52-818b-dabbb4a86bc7 · outbound

This paper cites Group Sequence Policy Optimization.

Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning Group Sequence Policy Optimization

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:06:27.237419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T11:13:21.565082Z digest=sha256:a852aff9c741ea4af783471bcc752ef8c0674f6b246e3833465c984ee45f3e92

Observation ccaefb89-7afd-4951-bd42-546bec4ef861 · outbound

This paper cites Representation Engineering: A Top-Down Approach to AI Transparency.

Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning Representation Engineering: A Top-Down Approach to AI Transparency

Reference 37

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T02:06:27.198922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T11:13:21.565082Z digest=sha256:556696ba35a4dbf4be14334ec43cfce93efd99e38b9a59dfb288a448e90c6d33

Pith citing papers

No inbound Pith citation observations are available.