Pith. sign in

Paper Citation Record · LEDGER

RLPR: Extrapolating RLVR to General Domains without Verifiers

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 25 inbound Pith citation observations for arXiv:2506.18254.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.18254 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 25 of 25 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T14:53:04.646543Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T06:29:38.225670Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 79755e74-e6e5-458f-b5dd-2e89aacd48ca · inbound

Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning cites this paper.

Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning RLPR: Extrapolating RLVR to General Domains without Verifiers

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:04.646543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:04.646543Z digest=sha256:ac824f4de5b7b0834a6de4042a7dc523a8905c64db61e4a5d4098477df942688

Observation 8969b9a4-7393-42e5-95c1-4d046f3c9338 · inbound

StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models cites this paper.

StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models RLPR: Extrapolating RLVR to General Domains without Verifiers

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T23:29:17.022564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:29:17.022564Z digest=sha256:b70a66eb4dabbf7b01683a11a265d2be3ca59537577d8594e83075aba8ff4221

Observation 8c1750ea-d176-45cf-a88a-bc2dd7c10d68 · inbound

An Explainable Machine Learning Framework for Railway Predictive Maintenance using Data Streams from the Metro Operator of Portugal cites this paper.

An Explainable Machine Learning Framework for Railway Predictive Maintenance using Data Streams from the Metro Operator of Portugal RLPR: Extrapolating RLVR to General Domains without Verifiers

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T23:25:50.287093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:25:50.287093Z digest=sha256:e5cf9b0c333defcb1a3a45fbe2bcfca4f5ec968e0ebceacdff7de1f8067944fe

Observation 06e66c7c-7c58-4988-8903-f54b84e49500 · inbound

Do Ethical AI Principles Matter to Users? A Large-Scale Analysis of User Sentiment and Satisfaction cites this paper.

Do Ethical AI Principles Matter to Users? A Large-Scale Analysis of User Sentiment and Satisfaction RLPR: Extrapolating RLVR to General Domains without Verifiers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T23:07:12.004152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:07:12.004152Z digest=sha256:1e55dba2396cb6a1cbfef7d4a57985b93f963f4798abff939362a1cdb250277b

Observation 86079210-c35f-4451-aea5-043af72958f2 · inbound

Compute as Teacher: Turning Inference Compute Into Reference-Free Supervision cites this paper.

Compute as Teacher: Turning Inference Compute Into Reference-Free Supervision RLPR: Extrapolating RLVR to General Domains without Verifiers

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-18T15:46:34.115346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T15:45:09.730804Z digest=sha256:96f9a0909da31d6e4fd361ff0e610e7e31ea8c12119c721d88dac0640998d00e

Observation 054c9827-e6dc-4db5-8108-e7bc3b7f6fd8 · inbound

TRINITY: An Evolved LLM Coordinator cites this paper.

TRINITY: An Evolved LLM Coordinator RLPR: Extrapolating RLVR to General Domains without Verifiers

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:21:24.781976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T01:19:10.420266Z digest=sha256:8d4f8cfd688772cb1abf5f2017718bcec39ff181e71d22d50739a802e1a65988

Observation 2b790b15-16be-4b3e-9cfc-f1531b9f4188 · inbound

StaRPO: Stability-Augmented Reinforcement Policy Optimization cites this paper.

StaRPO: Stability-Augmented Reinforcement Policy Optimization RLPR: Extrapolating RLVR to General Domains without Verifiers

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:20:59.963670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T18:11:49.056805Z digest=sha256:4a60d4b41fef1de3a70378fe572d7f0785007431c5723d3ea1e1fd33a7843d78

Observation c47942a3-ba62-454a-aa9f-655e9224bdec · inbound

Trust Your Memory: Verifiable Control of Smart Homes through Reinforcement Learning with Multi-dimensional Rewards cites this paper.

Trust Your Memory: Verifiable Control of Smart Homes through Reinforcement Learning with Multi-dimensional Rewards RLPR: Extrapolating RLVR to General Domains without Verifiers

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:06:01.178591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T16:49:00.343580Z digest=sha256:a7c09c34626ed07a8574f094291963bcccf2d75b8109e292b19d44855bb7f3db

Observation 3a2be3a4-1eb2-4e55-bd07-8f2556bdeb5d · inbound

HEALing Entropy Collapse: Enhancing Exploration in Few-Shot RLVR via Hybrid-Domain Entropy Dynamics Alignment cites this paper.

HEALing Entropy Collapse: Enhancing Exploration in Few-Shot RLVR via Hybrid-Domain Entropy Dynamics Alignment RLPR: Extrapolating RLVR to General Domains without Verifiers

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:25:55.287022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T05:23:08.478393Z digest=sha256:3cb7fc52995d1c9f544eff97908adf31972933ea220f696975e7bdb7d9df055e

Observation edb6943d-e58d-4e84-8861-e882d01b06da · inbound

Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning cites this paper.

Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning RLPR: Extrapolating RLVR to General Domains without Verifiers

Reference 160

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:15:49.503384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T19:15:27.406778Z digest=sha256:76921ce9759f0f5255ea110949b257e4a9619ffe7222b8321dda0eabad7058ec

Observation 85edc4c9-1f3d-4cd6-af42-7efe2ab62f6f · inbound

G-Zero: Self-Play for Open-Ended Generation from Zero Data cites this paper.

G-Zero: Self-Play for Open-Ended Generation from Zero Data RLPR: Extrapolating RLVR to General Domains without Verifiers

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:11:24.558508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:a23dd878f5b23afcd8b284199a1b5ca9e1f7df5baa711cb0bc4af75144e7ab33

Observation 72cb19f3-c96d-4734-b47f-68a9b6dacd23 · inbound

Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities cites this paper.

Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities RLPR: Extrapolating RLVR to General Domains without Verifiers

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-19T17:22:41.998128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T17:20:39.969447Z digest=sha256:fa1f17e1b5e624cc534d3f5beb721ad44eb0a4a14136a4c7961fb03c92f3c9ae

Observation 901cfdcf-ce73-4be3-b531-742fd2ea79ef · inbound

Selective Off-Policy Reference Tuning with Plan Guidance cites this paper.

Selective Off-Policy Reference Tuning with Plan Guidance RLPR: Extrapolating RLVR to General Domains without Verifiers

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:47:05.185587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T01:28:18.615371Z digest=sha256:5ddf89d59aec69a0f5926700aac5992057cfd391883a89077b6adbb167f79ead

Observation de309758-8d6d-4c05-9cdc-a0fcca03b0d7 · inbound

Selective Off-Policy Reference Tuning with Plan Guidance cites this paper.

Selective Off-Policy Reference Tuning with Plan Guidance RLPR: Extrapolating RLVR to General Domains without Verifiers

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:22:59.347610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-14T21:20:24.066520Z digest=sha256:9ce212f8eef9900ce97dc846fb793b85d527f03d339a0148ba48980fb23aae24

Observation d5447ff0-8b0c-4b8d-ab76-5f6809095f65 · inbound

LamPO: A Lambda Style Policy Optimization for Reasoning Language Models cites this paper.

LamPO: A Lambda Style Policy Optimization for Reasoning Language Models RLPR: Extrapolating RLVR to General Domains without Verifiers

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:13:58.419624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T05:11:35.785227Z digest=sha256:795df087bc164fedaa394aac7e3fc4ea6aa2b622193ac7f6beb7da9594514d05

Observation 2571351c-abdd-42c0-88d9-539d5988e6cb · inbound

Combinatorial Synthesis: Scaling Code RLVR via Atomic Decomposition and Recombination cites this paper.

Combinatorial Synthesis: Scaling Code RLVR via Atomic Decomposition and Recombination RLPR: Extrapolating RLVR to General Domains without Verifiers

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-28T23:12:47.141445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T23:08:05.597810Z digest=sha256:6a11a85c0153836c82a58f3d277b8e33f8a8e37848ff9c77ffeede205e6f4a5f

Observation 01487990-7a6a-4622-869a-f7e4c6e99a89 · inbound

CARE-RL: Capability-Aware Reinforcement Learning for Mitigating Cross-Domain Conflicts cites this paper.

CARE-RL: Capability-Aware Reinforcement Learning for Mitigating Cross-Domain Conflicts RLPR: Extrapolating RLVR to General Domains without Verifiers

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:02:34.047073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T19:01:14.340754Z digest=sha256:d968dfbca33f48d4e1c3ae69f50c70c8f41bb280a31eabbee1d86f1bef1df90e

Observation 67d565f6-b0ee-485c-81a4-c78fa1535f7f · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation RLPR: Extrapolating RLVR to General Domains without Verifiers

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:56:13.380190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:e5177509c956b5c4e808d805ea120bbdf9c48b1249fb9784fc50c7ad93908c8a

Observation 7b35d5da-b98f-4acc-a2b8-543a0c429259 · inbound

GeoMin: Data-Efficient Semi-Supervised RLVR via Geometric Distribution Modeling cites this paper.

GeoMin: Data-Efficient Semi-Supervised RLVR via Geometric Distribution Modeling RLPR: Extrapolating RLVR to General Domains without Verifiers

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-02T06:06:40.808650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T07:45:43.320339Z digest=sha256:b0a8fedade726b21744c2a4f65a3ddcee371e9a63da21e8cd0e2a94f55ebc7a0

Observation 1e2ad807-18ad-4438-a044-3e8a0c8d781a · inbound

Reasoning Arena: Trace Tournaments When Verifiable Rewards Fall Short cites this paper.

Reasoning Arena: Trace Tournaments When Verifiable Rewards Fall Short RLPR: Extrapolating RLVR to General Domains without Verifiers

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:17:29.542547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T17:18:03.459289Z digest=sha256:e6edd8cfec8b6fa3523230ee8d1f78115a5cc3a477c2e38d83c252b23dab16c3

Observation b3e06ca5-3c98-46a5-a471-ab315d6f3d88 · inbound

CapRL++: Unified Reinforcement Learning with Verifiable Rewards for Dense Image and Video Captioning cites this paper.

CapRL++: Unified Reinforcement Learning with Verifiable Rewards for Dense Image and Video Captioning RLPR: Extrapolating RLVR to General Domains without Verifiers

Reference 93

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:17:29.006584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T17:21:38.543724Z digest=sha256:93e789eefc07caf85082cdbe1875239da0c32d68145802a726d3b65417ea6951

Observation 91a78fa8-cb7c-40a4-be49-3ff58b70e6b9 · inbound

Sakana Fugu Technical Report cites this paper.

Sakana Fugu Technical Report RLPR: Extrapolating RLVR to General Domains without Verifiers

Reference 243

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:29:38.227309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-26T14:22:37.596720Z digest=sha256:b9d27fbafe4646172e1d18010b6057cfe312285b294f6d661532f2cc6c194f9d

Observation 3ecd9f5c-fb34-47fb-93e2-e73d4881b216 · inbound

Copy Less, Ground More: Overcoming Repetitive Copying in Long-Context Reasoning via Evidence-Aware Reinforcement Learning cites this paper.

Copy Less, Ground More: Overcoming Repetitive Copying in Long-Context Reasoning via Evidence-Aware Reinforcement Learning RLPR: Extrapolating RLVR to General Domains without Verifiers

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T12:46:04.176578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:46:04.176578Z digest=sha256:099e0be5f6319a67a23138ac6b9ac6d0661c6fd24c551b4942a91cc82a016439

Observation b24dbb89-1a3f-4931-9a84-47e8a977714b · inbound

Copy Less, Ground More: Overcoming Repetitive Copying in Long-Context Reasoning via Evidence-Aware Reinforcement Learning cites this paper.

Copy Less, Ground More: Overcoming Repetitive Copying in Long-Context Reasoning via Evidence-Aware Reinforcement Learning RLPR: Extrapolating RLVR to General Domains without Verifiers

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T01:57:53.821375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:57:53.821375Z digest=sha256:40ab5516f14d1d8cef10347a330d355566413919477b00c5d3c9f486352e92df

Observation 4403b786-7262-4849-adda-33fb400c4398 · inbound

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design cites this paper.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design RLPR: Extrapolating RLVR to General Domains without Verifiers

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:11.704541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:11.704541Z digest=sha256:5760175eca69f02d49004660f9e6b6b1ceed50668161b29a07047923632d7dba