Pith. sign in

Paper Citation Record · LEDGER

Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks

As of 3 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 4 inbound Pith citation observations for arXiv:2506.13351.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.13351 v3

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-19T09:48:56.990745Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-03T06:30:56.289259+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T10:23:23.168854Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-02T23:07:27.269519Z

Reference resolution

37 of 37 outbound references displayed

  • verified exact30
  • verified fuzzy2
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch5

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1c1b743d-b5eb-483e-b48c-637472cb1aaf · outbound

This paper cites The pitfalls of next-token prediction.

Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks The pitfalls of next-token prediction

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:52:14.138276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-19T09:48:56.990745Z digest=sha256:44c379d8413bdc24a63bf7d0f7af947365d12fd54ee41186a2520042ca8846b5

Observation 179665dc-1fcf-454b-9c3c-6bf43aa56ab8 · outbound

This paper cites Language Models are Hidden Reasoners: Unlocking Latent Reasoning Capabilities via Self-Rewarding.

Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks Language Models are Hidden Reasoners: Unlocking Latent Reasoning Capabilities via Self-Rewarding

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T09:52:14.073815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-19T09:48:56.990745Z digest=sha256:7097132fc34180f0faef8fad19e867477efb0f82fd0e708144a5a03d91a60356

Observation 6d886974-da84-403d-9a03-a5b3cba3403c · outbound

This paper cites Enhancing uncertainty modeling with semantic graph for hallucination detection.

Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks Enhancing uncertainty modeling with semantic graph for hallucination detection

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:52:14.087180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-19T09:48:56.990745Z digest=sha256:c3c5ffa9f5248ea1f8b978538c1cde8204e1266df683e04d785731507ba12f90

Observation b71a2602-5e4e-42db-bc26-2cf26d7a6d0f · outbound

This paper cites Think, Prune, Train, Improve: Scaling Reasoning without Scaling Models.

Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks Think, Prune, Train, Improve: Scaling Reasoning without Scaling Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:52:14.119976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-19T09:48:56.990745Z digest=sha256:36ac9bafb65360cab7957f01852195e25343c540875f7e3d150128ef0008cee9

Observation 752909ee-b529-4540-8312-0ee5af3e777a · outbound

This paper cites KTO: Model Alignment as Prospect Theoretic Optimization.

Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks KTO: Model Alignment as Prospect Theoretic Optimization

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-19T09:52:14.105843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-19T09:48:56.990745Z digest=sha256:13381dd5b8f7d6dc9759fed692dcc32ce42d3adc91af6fd9df5a4e235563556a

Observation ad3a6eb2-fe66-45df-899e-6c975039dc0e · outbound

This paper cites A Survey on LLM-as-a-Judge.

Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks A Survey on LLM-as-a-Judge

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-19T09:52:14.103020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-19T09:48:56.990745Z digest=sha256:986c3be18703398178194bc46d0a52c7c182d6f2742a695216330c54bc5ceb14

Observation 73fb2433-34c9-418f-b41d-ea294dc85efb · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-19T09:52:14.170732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-19T09:48:56.990745Z digest=sha256:f9063af29b07dae9c0afaab08e902745c2f59d0e17b67649698a34e4ef2b3a9e

Observation 76a1baff-bdc8-4954-9157-54deaae169d3 · outbound

This paper cites Language Model Cascades: Token-level uncertainty and beyond.

Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks Language Model Cascades: Token-level uncertainty and beyond

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:52:14.114295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-19T09:48:56.990745Z digest=sha256:70a7b3792373cf923a0a09b4879bda0b64f57005343c57c02e641dc0a7414359

Observation cfe89eaf-3165-4c07-abf4-1bc669127b95 · outbound

This paper cites Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model.

Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-19T09:52:14.062147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-19T09:48:56.990745Z digest=sha256:11833735895d0d4cab364f4a141c905457535cfe76185e54553d428c217d69d2

Observation 27fc0b52-f5ff-4bdb-a168-adbcd38d5fa2 · outbound

This paper cites Self-Evolved Reward Learning for LLMs.

Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks Self-Evolved Reward Learning for LLMs

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:52:14.068048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-19T09:48:56.990745Z digest=sha256:5e6bcceee6d8d64095f6808ce4b90c812705d92626c88aef0f307ad7e28d12bd

Observation 474cd497-e9cc-4eb2-aaf9-b8898aceabd2 · outbound

This paper cites Towards Reasoning in Large Language Models: A Survey.

Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks Towards Reasoning in Large Language Models: A Survey

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-19T09:52:14.148806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-19T09:48:56.990745Z digest=sha256:347665679dc618bd04848978f8bb52ff4a9f9219b31ecec28500b2aa1b36e750

Observation 7202e720-1d9f-4df2-b612-4e6f98e094d0 · outbound

This paper cites OpenAI o1 System Card.

Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks OpenAI o1 System Card

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-19T09:52:14.117083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-19T09:48:56.990745Z digest=sha256:6db25b54f71e4b37d52dd03749e2815fa11976d3fe422cab21d7c7d7ccb8b913

Observation a8b57cdf-c553-4123-9ce6-413cc63ce350 · outbound

This paper cites s3: You don’t need that much data to train a search agent via rl.

Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks s3: You don’t need that much data to train a search agent via rl

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:52:14.181371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-19T09:48:56.990745Z digest=sha256:5c5f769639015219d882a13519d9a49b47ee5176280ad0a2e88debacdc935ecb

Observation 05bd0ce7-3127-4dea-9c8a-0f5d40bcb033 · outbound

This paper cites CASIMIR: A Corpus of Scientific Articles enhanced with Multiple Author-Integrated Revisions.

Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks CASIMIR: A Corpus of Scientific Articles enhanced with Multiple Author-Integrated Revisions

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:52:14.108516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-19T09:48:56.990745Z digest=sha256:662eb41be20a4efaea2b958d1c775deaa364659a2cffb6ba800911b95022c653

Observation 0cb9471e-feb2-4b79-9286-0d95976ef168 · outbound

This paper cites ParaRev: Building a dataset for Scientific Paragraph Revision annotated with revision instruction.

Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks ParaRev: Building a dataset for Scientific Paragraph Revision annotated with revision instruction

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:52:14.065438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-19T09:48:56.990745Z digest=sha256:9e19a46f7ca4676724a9577ca80c5569e1ce4e2c3cc023bd9374fc9bd921e411

Observation 99011ba1-ab7a-4c81-8572-1e6628cf102d · outbound

This paper cites Log Probabilities Are a Reliable Estimate of Semantic Plausibility in Base and Instruction-Tuned Language Models.

Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks Log Probabilities Are a Reliable Estimate of Semantic Plausibility in Base and Instruction-Tuned Language Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:52:14.129471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-19T09:48:56.990745Z digest=sha256:63fb908677a9de148bf3734aff2e6605afa8ced38e81f6f46e524e111797bf82

Observation 103b764c-d556-4b2b-9bc6-2338fcdc6286 · outbound

This paper cites Gemini 2.5: Our most intelligent AI model.

Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks Gemini 2.5: Our most intelligent AI model

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T10:03:02.362976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-19T09:48:56.990745Z digest=sha256:cf1ff1794639c598c8172164fb8518809e0b19b74b114949284385d56f78c911

Observation 707dcb62-1026-45f1-b86b-11462a399b4e · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-19T09:52:14.135246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-19T09:48:56.990745Z digest=sha256:bebfb8d7292fa7b0a04253aa7ee5ec952302ae01dce675f21887ea02acc926b4

Observation 2786c23c-680d-489f-9c4f-6d2f8954ec77 · outbound

This paper cites Pan and H.

Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks Pan and H

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T09:52:14.080320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-19T09:48:56.990745Z digest=sha256:eb0bd89f0b94ae0e787d4ca4df1e2ad82dc19f8da141b44b5e29880fa1b73f15

Observation 4aec18e5-fffe-4e03-8174-e28aba8cd4d8 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks Understanding R1-Zero-Like Training: A Critical Perspective

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-19T09:52:14.077117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-19T09:48:56.990745Z digest=sha256:3fcf7346a7318c7f78845865495cb98cf4f765ebf438293bb39c78060ea65dc1

Observation e0278dc4-8130-4281-990a-5a781151c114 · outbound

This paper cites Microsoft.

Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks Microsoft

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T10:03:02.361109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-19T09:48:56.990745Z digest=sha256:343dfd60225687211abe6a0c62645abeb2585a4e6ccf04b1eb3b7409e9f923f0

Observation 09a83d25-af58-4e89-a882-da1e0dc94b45 · outbound

This paper cites s1: Simple test-time scaling.

Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks s1: Simple test-time scaling

Reference 22

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T09:52:14.090011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-19T09:48:56.990745Z digest=sha256:9b9524bac797437cf1cfe93e67ae3ab3fb019ee9344d7e5d1f965ddd78a56b2b

Observation ce90225b-c8f8-479b-bf1c-8504c24fbff9 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks Proximal Policy Optimization Algorithms

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-19T09:52:14.132222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-19T09:48:56.990745Z digest=sha256:6ce96a5025629edf0c44358520562b36518a91d5df2476d4edcc9f676ca054e4

Observation 7a06cae3-a6b7-4b61-87ee-2ee1a1e5f28d · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-19T09:52:14.174956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-19T09:48:56.990745Z digest=sha256:5b131d6dddc5652e74bdee89b2ca1790edb6ca2ddd9e0408f7550e6be35fffeb

Observation c0442bea-760a-4676-97e7-2fc13f8b4d84 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-19T09:52:14.145232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-19T09:48:56.990745Z digest=sha256:909acbe1efe8e0859d3bae664e59797d9e2d55504e478ea4a71e8d10499a512b

Observation 7745fa08-d56a-4cc0-aa9c-f06996f5d17d · outbound

This paper cites Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains.

Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:52:14.093197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-19T09:48:56.990745Z digest=sha256:33ca5edffacd554f4e965cb43314f8f87adadad7f6ab61052bf528d5e7b5fc07

Observation ed179084-8be9-4cb2-abeb-95217cf7a310 · outbound

This paper cites Beyond Verifiable Rewards: Scaling Reinforcement Learning for Language Models to Unverifiable Data.

Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks Beyond Verifiable Rewards: Scaling Reinforcement Learning for Language Models to Unverifiable Data

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T09:52:14.141843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-19T09:48:56.990745Z digest=sha256:b93844b595c7bde9b4e2f01593d49740d983d2c443fc1114d76b9a825acb9fdd

Observation 4e99e11f-e55e-4aa3-80d5-5f162e0ee8da · outbound

This paper cites A Stitch in Time Saves Nine: Detecting and Mitigating Hallucinations of LLMs by Validating Low-Confidence Generation.

Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks A Stitch in Time Saves Nine: Detecting and Mitigating Hallucinations of LLMs by Validating Low-Confidence Generation

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T09:52:14.126420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-19T09:48:56.990745Z digest=sha256:66ff1ccccbf65ec9cc63fb5e31ff6ffa5fe4dd00aa09e7fb5e22e786bf73ea43

Observation 415cfe12-9a22-4b4a-8b22-8b65b53e2caf · outbound

This paper cites Genius: A Generalizable and Purely Unsupervised Self-Training Framework For Advanced Reasoning.

Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks Genius: A Generalizable and Purely Unsupervised Self-Training Framework For Advanced Reasoning

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:52:14.122863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-19T09:48:56.990745Z digest=sha256:185eef9960ae87c6dd2a206a7e671a59e856fb496c2444d327fd3725bdf7b1f8

Observation 95300ed3-3406-401a-a6fc-5af3c0cdc81e · outbound

This paper cites LIMO: Less is More for Reasoning.

Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks LIMO: Less is More for Reasoning

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-19T09:52:14.059223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-19T09:48:56.990745Z digest=sha256:96bd8257ce6bb59e0ffbcd49f9911646f27acf390f8f1a32c19c8ecc22985ffb

Observation c92f8f91-3a7a-4a5d-94f8-4032ea539520 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-19T09:52:14.083201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-19T09:48:56.990745Z digest=sha256:22a844ccd70189341bac6edc31a26fa953ee027fb7d9c33370fddf59b79057d8

Observation fa98ac6d-2f5b-4c26-8992-8370e5586ff0 · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?.

Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-19T09:52:14.097098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-19T09:48:56.990745Z digest=sha256:5e4df905a2ce96cbe17cd42956f818c300347b714c3be0c027a7b4e55e643117

Observation e9b1f88b-5d58-4037-921a-64db6211d813 · outbound

This paper cites Scaling of Search and Learning: A Roadmap to Reproduce o1 from Reinforcement Learning Perspective.

Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks Scaling of Search and Learning: A Roadmap to Reproduce o1 from Reinforcement Learning Perspective

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:52:14.111450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-19T09:48:56.990745Z digest=sha256:c0ae019420db44859f1b0916c75391413866b8abeb3523006b752881dc1e3814

Observation 74be46e2-347f-41b3-9412-e3c4e10abec1 · outbound

This paper cites Automatic Chain of Thought Prompting in Large Language Models.

Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks Automatic Chain of Thought Prompting in Large Language Models

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-19T09:52:14.178050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-19T09:48:56.990745Z digest=sha256:e98cdf71d4472e09d2800bc2235e2d271bdc18cafed48da0420207e519b2958e

Observation 7200d265-0d91-40f9-b422-a5fac556602d · outbound

This paper cites Absolute Zero: Reinforced Self-play Reasoning with Zero Data.

Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks Absolute Zero: Reinforced Self-play Reasoning with Zero Data

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-19T09:52:14.070782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-19T09:48:56.990745Z digest=sha256:8fb425ef4c7981a998d09a05fe0403e2a8fbfaf0d99127ffbcfc030a83c8582b

Observation 09f1d75b-3d1f-411a-81d3-dc06e696881d · outbound

This paper cites Reinforcing General Reasoning without Verifiers.

Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks Reinforcing General Reasoning without Verifiers

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:52:14.184491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-19T09:48:56.990745Z digest=sha256:e1864df50a29db26fcaedaae141b12d4bf1a06566cd0ec11168fb7bb94de3bf5

Observation cb697b38-80ee-4fda-8a36-4d881b70c9f2 · outbound

This paper cites TTRL: Test-Time Reinforcement Learning.

Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks TTRL: Test-Time Reinforcement Learning

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-19T09:52:14.100204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-19T09:48:56.990745Z digest=sha256:359ed1c43dba2d5f4fa06c44989d215522ae11730f2eca3062cde64350778a60

Pith citing papers

Observation 3984fa88-075b-41d5-be28-a20913e35529 · inbound

Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities cites this paper.

Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-19T17:22:41.983841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-19T17:20:39.969447Z digest=sha256:49102d3e9ccd73b497cdf87e1e126d7540e6b028f61dde1f5a8bea6102787ea2

Observation 8ee4c914-181f-481e-bcdc-8889895a1c0f · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks

Reference 169

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T20:56:13.576795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:df6a0f726ea2a8d66a5950df5d79492e6e5512d6c5a29e6505738de230947c8f

Observation 8d635831-158b-4d2e-9f7c-e682eb7be96e · inbound

Momentum for Reasoning: Dense Intrinsic Signals in Policy Optimization cites this paper.

Momentum for Reasoning: Dense Intrinsic Signals in Policy Optimization Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T23:07:27.270961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-06-27T18:26:22.701076Z digest=sha256:82ca358bcc07de3b538d2a08266ab4256b90dc193d31f2bf7f052f8b794bb098

Observation f2777606-e344-481d-91d4-74dea9cbc06a · inbound

Open Security Benchmark: Towards Autonomous Enterprise Cyber Defense cites this paper.

Open Security Benchmark: Towards Autonomous Enterprise Cyber Defense Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-01T10:23:23.168854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:23:23.168854Z digest=sha256:660f7a482381b68c40cc92402a1119bf0ec4c2734587a6eb0cb245dd4c2e9ed9