Pith. sign in

Paper Citation Record · LEDGER

Reinforcing General Reasoning without Verifiers

As of 19 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 18 inbound Pith citation observations for arXiv:2505.21493.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.21493 v1

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:35:56.253150Z

measured 78 of 78 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T19:01:06.330755Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T20:56:13.505929Z

Reference resolution

60 of 60 outbound references displayed

  • verified exact0
  • verified fuzzy12
  • unresolved47
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a93d6290-e684-40aa-84ad-0648f6d7895e · outbound

This paper cites Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs.

Reinforcing General Reasoning without Verifiers Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:49.676520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:49.676520Z digest=sha256:5d4628d6217c9574f2e3e26ab2aa9e2c8a50cf9eb8ffebe9c11f8dc399d58c2f

Observation 6b8912b9-44d6-4993-8c75-365c99efb6f4 · outbound

This paper cites an unresolved cited work.

Reinforcing General Reasoning without Verifiers Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:36:00.810288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:35:49.783935Z digest=sha256:f9af9efecae7614d8103b4f1d3754d6bba5dc3bdf88f0e125e2498c31c00711a

Observation da40036f-4d63-460f-9951-a80959496b15 · outbound

This paper cites Bootstrapping language models with dpo implicit rewards.

Reinforcing General Reasoning without Verifiers Bootstrapping language models with dpo implicit rewards

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:36:00.555733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:35:49.895146Z digest=sha256:6e77715baadef5d371ebeb7479457a2b5e660dde3f7e3f5896f4d43c1f755e1e

Observation 52ebb6e6-cb33-47bf-9fd6-91affe7684c7 · outbound

This paper cites Language Models are Hidden Reasoners: Unlocking Latent Reasoning Capabilities via Self-Rewarding.

Reinforcing General Reasoning without Verifiers Language Models are Hidden Reasoners: Unlocking Latent Reasoning Capabilities via Self-Rewarding

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:50.020849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:50.020849Z digest=sha256:0bc59c66f375bfa9b3d0f014059767a7a9a7ce22b9232634a589caf0c0947b86

Observation 81e3becc-9596-4c67-8111-196dba531996 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Reinforcing General Reasoning without Verifiers Training Verifiers to Solve Math Word Problems

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:50.157115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:50.157115Z digest=sha256:785296093738912542c292c699f1f8137e05f831041944c818dc70bd875681a9

Observation 2d669c5b-0d3a-4ba5-9a8e-f0cf67637d61 · outbound

This paper cites SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines.

Reinforcing General Reasoning without Verifiers SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:50.307619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:50.307619Z digest=sha256:d0e551ed60e0b2f75d486589df68021d52cbe309dd2f8b5bb41a6dd1fbd8295a

Observation 7a802705-31d5-4dfe-9f52-c59d5cddffa1 · outbound

This paper cites Scaling laws for reward model overoptimization.

Reinforcing General Reasoning without Verifiers Scaling laws for reward model overoptimization

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:50.451904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:50.451904Z digest=sha256:ba7b98bcdfd570cd28edd7bc61bfb00f3d51f451f93bc4fadd4537a584e73efe

Observation 41cd47ac-0822-4c76-8982-68592cfe2bbc · outbound

This paper cites RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning.

Reinforcing General Reasoning without Verifiers RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:50.613928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:50.613928Z digest=sha256:389e38e76b11ffa93ff27a964ead576cd558bdadea4571bb611c86539cd340d4

Observation 5dade0db-61c1-4c40-8094-7951010b67e5 · outbound

This paper cites The Llama 3 Herd of Models.

Reinforcing General Reasoning without Verifiers The Llama 3 Herd of Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:50.694745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:50.694745Z digest=sha256:9ce38e29112fd83461526a81879344c1d35313b3bdc7252624f0558c3db7a495

Observation f51e8249-efdf-4f50-a80c-83e7ace68b80 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Reinforcing General Reasoning without Verifiers DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:50.841396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:50.841396Z digest=sha256:e4afc099073c99e914fad1fde85d8685873a941c551945a9584ec3bc22034b26

Observation 532f803e-3d01-46a4-aeea-6d6fa3efeaf9 · outbound

This paper cites Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems, 2024.

Reinforcing General Reasoning without Verifiers Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems, 2024

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:50.986072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:50.986072Z digest=sha256:d2dae98a102f1fb969fb7159d73041e56800c2e68a49830cc6d75d6830fc4915

Observation 8ab29f53-b30b-494d-88ee-a43a1f872e9f · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Reinforcing General Reasoning without Verifiers Measuring Mathematical Problem Solving With the MATH Dataset

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:51.098076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:51.098076Z digest=sha256:86873e24d99ec658f366bdbab6f816478d426ae60cc25dfb3d03a62d8556100e

Observation 3317d73c-bec0-4a93-9f00-2ebfaa75998e · outbound

This paper cites Measuring mathematical problem solving with the math dataset.

Reinforcing General Reasoning without Verifiers Measuring mathematical problem solving with the math dataset

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:51.205483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:51.205483Z digest=sha256:e0e832028e963f0b271d3e5878e980cd571217f914695e8b8a22064f27073fab

Observation b2bbc0d8-6d7c-4e9a-8559-ef1948d36c91 · outbound

This paper cites Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model.

Reinforcing General Reasoning without Verifiers Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:51.324740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:51.324740Z digest=sha256:4a702f13e0048e2110bd7c472d96e56fe8aeba064d2ebaeecaf6460566bbc3ce

Observation b479d45d-7c4b-4ae8-800e-17e7f52fb058 · outbound

This paper cites Self-Improvement in Language Models: The Sharpening Mechanism.

Reinforcing General Reasoning without Verifiers Self-Improvement in Language Models: The Sharpening Mechanism

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:51.425938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:51.425938Z digest=sha256:3f05b53cb680bd9a2f30e41c73d6a66e7124bfbb3d9054c3473dcc5f68af9faf

Observation 81cf8b30-cfc3-4b32-a192-89c7731e1297 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

Reinforcing General Reasoning without Verifiers Gonzalez, Hao Zhang, and Ion Stoica

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:51.526601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:51.526601Z digest=sha256:7bb43c41671f196eb228009321aafee9c5acc4f30f84a93713edef7d325a19f6

Observation 7d2439f9-9c29-4fb7-b7ca-c8cb2be5b697 · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

Reinforcing General Reasoning without Verifiers Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:51.600499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:51.600499Z digest=sha256:f9831caaca44170d2d8b703df7a151efd2b61e46bab620cb4e17dbe4c029a912

Observation 59922537-1b14-47ce-a2d4-346a1010e5b7 · outbound

This paper cites Solving quantitative reasoning problems with language models.Advances in Neural Information Processing Systems, 35:3843–3857, 2022.

Reinforcing General Reasoning without Verifiers Solving quantitative reasoning problems with language models.Advances in Neural Information Processing Systems, 35:3843–3857, 2022

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:51.708819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:51.708819Z digest=sha256:26245706e51f1d4cf75c30923f75ec0f8a3890b8ee3fb4d348fc0f79f2ff7105

Observation cadd1958-83e1-4260-a2fe-7ebce56bb0ee · outbound

This paper cites Numinamath: The largest public dataset in ai4maths with 860k pairs of competition math problems and solutions.Hugging Face repository, 13:9, 2024.

Reinforcing General Reasoning without Verifiers Numinamath: The largest public dataset in ai4maths with 860k pairs of competition math problems and solutions.Hugging Face repository, 13:9, 2024

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:51.819498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:51.819498Z digest=sha256:728a54e3af2c8b420dc78476e4713bed492c8d81d5d130d4e06f0cafbbe57ee6

Observation fb322e58-6b54-4677-94c9-26d5f8f67426 · outbound

This paper cites Let’s verify step by step.

Reinforcing General Reasoning without Verifiers Let’s verify step by step

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:51.952732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:51.952732Z digest=sha256:3672cf25a86aa30862eac2858e45d99a1957fd9eac9623cbf391a35509b57e4e

Observation 23d49f70-44e6-486d-8234-6972ef41ec39 · outbound

This paper cites X-Reasoner: Towards Generalizable Reasoning Across Modalities and Domains.

Reinforcing General Reasoning without Verifiers X-Reasoner: Towards Generalizable Reasoning Across Modalities and Domains

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:52.104790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:52.104790Z digest=sha256:b0d32e08723c0360a54c4acae67af26b0b10e523eeada5ae5993786500d17aef

Observation 8f616da6-ee5f-48fa-a183-eb4a4e5ecd88 · outbound

This paper cites Oat: A research-friendly framework for llm online alignment.https://github.com/sail-sg/oat, 2024.

Reinforcing General Reasoning without Verifiers Oat: A research-friendly framework for llm online alignment.https://github.com/sail-sg/oat, 2024

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:36:00.206502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:35:52.211512Z digest=sha256:b21891535634eb2f259e7ea67c12ec7805a0da780bf8fda3e93eb9f66c0cdd56

Observation 99224617-7326-4c1a-85f0-2dd010ea96bc · outbound

This paper cites There may not be aha moment in r1-zero-like training — a pilot study.

Reinforcing General Reasoning without Verifiers There may not be aha moment in r1-zero-like training — a pilot study

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:52.284816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:52.284816Z digest=sha256:f0f016b8339b7532414c79f867a4f603e60558ca1abcffff4a3e2243b45e076b

Observation 8fd9a995-3dc6-479b-9102-e66be0b39631 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

Reinforcing General Reasoning without Verifiers Understanding R1-Zero-Like Training: A Critical Perspective

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:52.414743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:52.414743Z digest=sha256:9d2137a7be76d5a81ac0901c70c09c614a9b3c873c4f367e596f43827fb244a9

Observation 00842862-6da4-4cb5-947f-450fb3d521b8 · outbound

This paper cites Deepcoder: A fully open-source 14b coder at o3-mini level, 2025.

Reinforcing General Reasoning without Verifiers Deepcoder: A fully open-source 14b coder at o3-mini level, 2025

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:52.527377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:52.527377Z digest=sha256:d040fc72e5b6221f8d5bb48b2924da5670ba803d1e82caf3b9e5340497c3b3df

Observation f99e1cdc-3791-47c3-88d4-7fe27993c55d · outbound

This paper cites Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica.

Reinforcing General Reasoning without Verifiers Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:52.600474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:52.600474Z digest=sha256:8622fe26c65d19a761798b095ef31176625ecad1cc74d923e725f3811c67623c

Observation e8604899-443b-4c3e-8189-b8cef03b172b · outbound

This paper cites General-Reasoner: Advancing LLM Reasoning Across All Domains.

Reinforcing General Reasoning without Verifiers General-Reasoner: Advancing LLM Reasoning Across All Domains

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:52.681204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:52.681204Z digest=sha256:461d627631eaed69cb20955776e8ec03ac9bd222f17b04d07b8348bcdaae67ec

Observation a64d249a-1c61-4104-b764-d954622e262e · outbound

This paper cites Ng, Daishi Harada, and Stuart J.

Reinforcing General Reasoning without Verifiers Ng, Daishi Harada, and Stuart J

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:59.955749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:35:52.844277Z digest=sha256:a48eb8c37385b777e358b638c4b0950398f773e0ab562769187eee6206db95f0

Observation 4a747535-adc6-4e76-9340-bbe4f3af9954 · outbound

This paper cites Learning to reason with llms, 2024.

Reinforcing General Reasoning without Verifiers Learning to reason with llms, 2024

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:59.437638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:35:53.099446Z digest=sha256:57a716c6e3a4d57ebfa13a8737b04712264bd06ed99eac64a7e8d84c0acc29fc

Observation fedd8f05-30b9-4aac-8644-585576047986 · outbound

This paper cites Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022.

Reinforcing General Reasoning without Verifiers Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:53.235958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:53.235958Z digest=sha256:6ee8b763c71006333a9b8728b2a065d11bbdef5d0c677b6f3cd305fb9f5c912b

Observation 6de7d42f-7f10-4ba6-967c-3f7b9cc6ec85 · outbound

This paper cites Tinyzero.

Reinforcing General Reasoning without Verifiers Tinyzero

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:53.355905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:53.355905Z digest=sha256:2c66979bef454c2a7d4481c222783b309e88c468f47a70eeb53c7abe9df566d9

Observation a8008712-a0b1-452c-a993-0fa78b48d8e1 · outbound

This paper cites Training chain-of- thought via latent-variable inference.Advances in Neural Information Processing Systems, 36: 72819–72841, 2023.

Reinforcing General Reasoning without Verifiers Training chain-of- thought via latent-variable inference.Advances in Neural Information Processing Systems, 36: 72819–72841, 2023

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:59.181728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:35:53.438954Z digest=sha256:5f79571ebb4c88959802d3e7c903b4692643d446c2f2796f07d7caa989ef0208

Observation 80afe6a7-a471-41af-9798-d8c60a771a76 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Reinforcing General Reasoning without Verifiers Direct preference optimization: Your language model is secretly a reward model

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:53.528942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:53.528942Z digest=sha256:68cbbd56494e6e2e12a2ac8f3cbe2fd77e8216f68ceea66564ca5a51e449794f

Observation a1dd8e9e-70b0-4602-9844-d1cab760b508 · outbound

This paper cites Learning to drive a bicycle using reinforcement learning and shaping.

Reinforcing General Reasoning without Verifiers Learning to drive a bicycle using reinforcement learning and shaping

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:58.957200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:35:53.652657Z digest=sha256:e7f1583e49421690ba8f8c571d7e7d69d66e99a6f8d9f2981aedb699936cd78d

Observation bff29ac9-4452-474e-a5a8-662cfbb18d31 · outbound

This paper cites Gpqa: A graduate-level google-proof q&a benchmark.

Reinforcing General Reasoning without Verifiers Gpqa: A graduate-level google-proof q&a benchmark

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:53.729825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:53.729825Z digest=sha256:a034367a1ca2ef3ea819fe2da91f0d2265a3f8f95499886447f4979c26ff362e

Observation ca43c1f2-a018-4109-b4ff-3389791b2fca · outbound

This paper cites Proximal Policy Optimization Algorithms.

Reinforcing General Reasoning without Verifiers Proximal Policy Optimization Algorithms

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:53.862659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:53.862659Z digest=sha256:ab2378c5c4df59050e38b883a4f785f0786a6557efcaa3e30d322f98e55b8d5d

Observation a7849d0c-5308-4572-aae7-00b1eb99e966 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Reinforcing General Reasoning without Verifiers DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:53.962521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:53.962521Z digest=sha256:26389e4f2ddb1d40dde2d49ab0d71045ebd0035513b17d5c8fc04d31bcca9661

Observation facf22d4-454e-4613-b858-c10da0a74038 · outbound

This paper cites Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains.

Reinforcing General Reasoning without Verifiers Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:54.067551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:54.067551Z digest=sha256:0bdb5f5512b6a9850998dad746729179524aa992d6e3931330f09ff235e78903

Observation 13680724-420a-4471-9655-96475f45143a · outbound

This paper cites Sutton and Andrew G.

Reinforcing General Reasoning without Verifiers Sutton and Andrew G

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:54.172581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:54.172581Z digest=sha256:d20b2e1030854525aa4f174001d71ce92380bc806977bc1bec61d4cd0674e3e8

Observation c5e30f52-8d0b-4008-a8f2-fc41b25e47ad · outbound

This paper cites Beyond Verifiable Rewards: Scaling Reinforcement Learning for Language Models to Unverifiable Data.

Reinforcing General Reasoning without Verifiers Beyond Verifiable Rewards: Scaling Reinforcement Learning for Language Models to Unverifiable Data

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:54.275667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:54.275667Z digest=sha256:8aee329bf1270aa644cbb458da80a01e44ee9ba787790d05ca8c13d34ce906f4

Observation 23fbb93f-8030-4c13-9649-781b8f2dca25 · outbound

This paper cites Qwen3, April 2025.

Reinforcing General Reasoning without Verifiers Qwen3, April 2025

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:54.354782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:54.354782Z digest=sha256:a97ef41d987bb8376ce04918a4e31d120090751e8619a38c7ddbee4c8393391d

Observation a661f906-bd43-41e7-96cd-89f67894319d · outbound

This paper cites Mmlu-pro: A more robust and challenging multi-task language understanding benchmark.

Reinforcing General Reasoning without Verifiers Mmlu-pro: A more robust and challenging multi-task language understanding benchmark

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:54.447387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:54.447387Z digest=sha256:27d2fd1bfa5832105088cb2bd601a3b18470204d18105c2d99ed94b70c91b657

Observation 8ae8c311-f20e-43d8-ad0d-9983c0c8b323 · outbound

This paper cites Qwen2.5 Technical Report.

Reinforcing General Reasoning without Verifiers Qwen2.5 Technical Report

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:54.550113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:54.550113Z digest=sha256:705f811774c080fbbe2ceaf25c7bd08c90f6f5b1b2f12617df32a562a548b447

Observation 6ab5a380-d554-4d25-a2ea-5ba8c30cf2dd · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

Reinforcing General Reasoning without Verifiers Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:54.659939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:54.659939Z digest=sha256:2ae2f2d988548c91346b90b84634af1190f50a83f7a60c2bc50aebca2d7e28fa

Observation dfbbdad7-3e5b-4f76-a2f8-96a7f26b2e10 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Reinforcing General Reasoning without Verifiers DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:54.837737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:54.837737Z digest=sha256:06c6195ee35d3ad71194546efe720488ecda85023cdc6e900857d3756f8b45dd

Observation 0b46341a-8a36-4346-a992-69c3bac38ab0 · outbound

This paper cites Self-rewarding language models.International Conference on Machine Learning,.

Reinforcing General Reasoning without Verifiers Self-rewarding language models.International Conference on Machine Learning,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:58.735265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:35:54.952680Z digest=sha256:361e846038c0cba01e9c027c7ee048aebe0f4b0c9c681b009b21adb0bced3598

Observation ea9dad09-2bb7-4512-9ae8-730ae885cb81 · outbound

This paper cites Naturalreasoning: Reasoning in the wild with 2.8 m challenging questions.arXiv preprint arXiv:2502.13124, 2025.

Reinforcing General Reasoning without Verifiers Naturalreasoning: Reasoning in the wild with 2.8 m challenging questions.arXiv preprint arXiv:2502.13124, 2025

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:55.134741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:55.134741Z digest=sha256:75c94dc21a1e694d0a72efe6453f92dad101a1b59913b08a85f7a42cffee9387

Observation eb9e4194-f947-421d-8802-662ec50072e4 · outbound

This paper cites Scaling Relationship on Learning Mathematical Reasoning with Large Language Models.

Reinforcing General Reasoning without Verifiers Scaling Relationship on Learning Mathematical Reasoning with Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:55.254650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:55.254650Z digest=sha256:eb84f4064214d35f8cb1d4631381f3628c62c68578fa5ea62f66594f2ade58d1

Observation 9cb71e8a-757a-459b-a8f0-6bced49c0046 · outbound

This paper cites Mammoth2: Scaling instructions from the web.Advances in Neural Information Processing Systems, 37:90629–90660, 2024.

Reinforcing General Reasoning without Verifiers Mammoth2: Scaling instructions from the web.Advances in Neural Information Processing Systems, 37:90629–90660, 2024

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:55.410399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:55.410399Z digest=sha256:c5975af41fb327a7ccb1ecf4a1da6f8cb8c058e24ede0330cf1a80b3cfa5423d

Observation fd3fc11e-4ffb-4e99-b255-fabad4fbd632 · outbound

This paper cites 7b model and 8k examples: Emerging reasoning with reinforcement learning is both effective and efficient.https://hkust-nlp.notion.site/simplerl-reason, 2025.

Reinforcing General Reasoning without Verifiers 7b model and 8k examples: Emerging reasoning with reinforcement learning is both effective and efficient.https://hkust-nlp.notion.site/simplerl-reason, 2025

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:58.609689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:35:55.503236Z digest=sha256:e367513b8c05bee219d4f13952f8091a4ee235a8678da6fe9c61e964679be682

Observation db5b202b-4c9f-4cf3-8395-4e90c75576e0 · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Reinforcing General Reasoning without Verifiers Fine-Tuning Language Models from Human Preferences

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:55.613828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:55.613828Z digest=sha256:c714ef49a37ee4c3a2f52aed9143e478416fcab10e4695a3ede66b381887b775

Observation ec4bb7de-61fa-4edc-990c-b6107dad8896 · outbound

This paper cites TTRL: Test-Time Reinforcement Learning.

Reinforcing General Reasoning without Verifiers TTRL: Test-Time Reinforcement Learning

Reference 52

Resolution
malformed identifier
no resolver link, observed 2026-08-07T13:35:55.747382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:55.747382Z digest=sha256:e67911593eca32a4e9dfde73a3e2bbf62001e76f3b5257eef883983b59b0477c

Observation acce323f-2ddd-4261-a551-78492765c929 · outbound

This paper cites an unresolved cited work.

Reinforcing General Reasoning without Verifiers Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:35:58.331998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:35:55.845493Z digest=sha256:f61cc6ef2c41a1b860a7798096963ebd65e7a0d968c23bcb3109b57bd0082b4b

Observation e6763064-39f5-48b3-b01b-87fda9d95806 · outbound

This paper cites This relation is usually given in a form where a graph or a formula relates period to absolute magnitude (a measure of intrinsic brightness).

Reinforcing General Reasoning without Verifiers This relation is usually given in a form where a graph or a formula relates period to absolute magnitude (a measure of intrinsic brightness)

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:58.140035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:35:55.913862Z digest=sha256:564927cf1a1932198866f3fdfb59b44488541fd9f847576808a088e313cc8291

Observation 120c5616-51b9-4211-9664-27fd8f35cbdc · outbound

This paper cites an unresolved cited work.

Reinforcing General Reasoning without Verifiers Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:35:57.928563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:35:55.981091Z digest=sha256:76699d89d8ba7f6e4ccd29c258abfcfc957282f67385639e46a9404e1b360de1

Observation 0edea18e-4c38-4d9f-974d-6f07835b4b5f · outbound

This paper cites ""everyone else is doing it.

Reinforcing General Reasoning without Verifiers ""everyone else is doing it

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:57.686510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:35:56.073560Z digest=sha256:ae79133dc0cf471280c7fefa76806e073651d659ea538ae6e3128d8a885ae15e

Observation ecf44972-dfd3-463b-97e4-8ff3ab1de69d · outbound

This paper cites Their reasoning is fear-based, and they view rules as set by authority figures.

Reinforcing General Reasoning without Verifiers Their reasoning is fear-based, and they view rules as set by authority figures

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:57.494309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:35:56.170121Z digest=sha256:5ba4754b66da1b6da7a2a8be7dd10af67bdb9bd37c67a9169b041840566f5d81

Observation ef0107df-7832-4415-b647-576f0b7fd712 · outbound

This paper cites what’s in it for me?.

Reinforcing General Reasoning without Verifiers what’s in it for me?

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:57.168558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:35:56.253150Z digest=sha256:9c5cc3908cb2a48b589cb8333504bcda33ba4b656f3856641c59eba51634b431

Observation 0e3fc1d9-feba-4c02-b9d5-adcd9d901f15 · outbound

This paper cites an unresolved cited work.

Reinforcing General Reasoning without Verifiers Unresolved cited work

Reference 1999

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:35:59.715432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:35:52.950970Z digest=sha256:1fa25a92f3d683330a9f2e28c62ff453601b47fc114798e3d185be4481109b3e

Observation 2c3364b5-f502-45b4-9dda-64bae6f72e75 · outbound

This paper cites Self-Rewarding Language Models.

Reinforcing General Reasoning without Verifiers Self-Rewarding Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:55.039840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:55.039840Z digest=sha256:18f3f35dda80f73f788a555263ed3180964aa398a79b927f66602b21f0ec78af

Pith citing papers

Observation 5011bd04-bd69-40bb-b9b5-63b1463ceb1d · inbound

Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems cites this paper.

Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems Reinforcing General Reasoning without Verifiers

Reference 167

Resolution
verified exact
arxiv_id, observed 2026-05-22T21:42:10.812398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T21:39:49.832151Z digest=sha256:64d6c3a02f510e23b314c530c028427e6279d118b2682e20657a35f06f680ec6

Observation 881d37f7-a003-4bc1-a8de-289095639b6c · inbound

Reinforcement Pre-Training cites this paper.

Reinforcement Pre-Training Reinforcing General Reasoning without Verifiers

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:22.158083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:26:22.158083Z digest=sha256:a00af2c675703b771ac3af8aa12f97b486b4bda3ec87e2b612ffb66d8f936a84

Observation 09f1d75b-3d1f-411a-81d3-dc06e696881d · inbound

Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks cites this paper.

Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks Reinforcing General Reasoning without Verifiers

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:52:14.184491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T09:48:56.990745Z digest=sha256:0335ac5d55b43d66442c0deb8c218ee9f7ed1609701fa661f690406681cd8699

Observation cf019ec0-544f-4ace-a61c-286476ddd018 · inbound

RLPR: Extrapolating RLVR to General Domains without Verifiers cites this paper.

RLPR: Extrapolating RLVR to General Domains without Verifiers Reinforcing General Reasoning without Verifiers

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:06.330755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:01:06.330755Z digest=sha256:b01062d9f368408c4f569049a2964e2cbb5fd02562e28d7d6f030f3cb2974f44

Observation d9cb06d2-5012-4061-985a-434b37d473ca · inbound

URPO: A Unified Reward & Policy Optimization Framework for Large Language Models cites this paper.

URPO: A Unified Reward & Policy Optimization Framework for Large Language Models Reinforcing General Reasoning without Verifiers

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T14:51:25.447032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:51:25.447032Z digest=sha256:3da23f8b2e5e3ce7ce8783cc0f1569157a66c2093bdaa9c4a9a646768709ca2d

Observation c801d62d-bb02-4a36-ab60-326a1d65e71a · inbound

UQ: Assessing Language Models on Unsolved Questions cites this paper.

UQ: Assessing Language Models on Unsolved Questions Reinforcing General Reasoning without Verifiers

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.893401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.893401Z digest=sha256:fb81abba2333d6d09d347bc6d77786b54f77461c96243d928969cf634233b805

Observation 85b3a33d-241e-4e2f-ab1e-ef958c464377 · inbound

Reverse-Engineered Reasoning for Open-Ended Generation cites this paper.

Reverse-Engineered Reasoning for Open-Ended Generation Reinforcing General Reasoning without Verifiers

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T00:06:08.567934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T00:06:08.567934Z digest=sha256:da6076d8fb5a71566820aa9e6c1d520a3789950d6576ac99db79c69138c42a6e

Observation f8b2fe70-730b-4ddc-8020-c762a6a2adf3 · inbound

Compute as Teacher: Turning Inference Compute Into Reference-Free Supervision cites this paper.

Compute as Teacher: Turning Inference Compute Into Reference-Free Supervision Reinforcing General Reasoning without Verifiers

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T15:46:34.124175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T15:45:09.730804Z digest=sha256:ce3c3a076db9ac52cede711067fffede9389fefdbecc280500c9969b3f87f53e

Observation adc3cf64-8802-4479-ba29-f2d3a969ed65 · inbound

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle cites this paper.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Reinforcing General Reasoning without Verifiers

Reference 257

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:48.740176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:48.740176Z digest=sha256:926b4e3b24f5ce074c60557fdbc1f11a31d4e9b9f60fd8cc62b34c1871319c93

Observation 57b2310b-f1db-4c8c-8d96-305438cda973 · inbound

Learning to Pose Problems: Reasoning-Driven and Solver-Adaptive Data Synthesis cites this paper.

Learning to Pose Problems: Reasoning-Driven and Solver-Adaptive Data Synthesis Reinforcing General Reasoning without Verifiers

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-17T23:00:25.487127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-17T22:59:48.260121Z digest=sha256:4a6b8db0ed56dee48b184d524d10f7ed6efa40e900f6a75240041c996e0713ec

Observation b607286b-d83d-43df-9474-bd80a6ef4580 · inbound

VI-CuRL: Stabilizing Verifier-Independent RL Reasoning via Confidence-Guided Variance Reduction cites this paper.

VI-CuRL: Stabilizing Verifier-Independent RL Reasoning via Confidence-Guided Variance Reduction Reinforcing General Reasoning without Verifiers

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-25T07:26:41.714012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-25T07:26:35.767179Z digest=sha256:a919b77e0069374a0baded810a8e1ce84ded7f87c5aa8302207c3b9894348751

Observation 52f3cb1b-099f-4c53-9e14-867cb07c84b5 · inbound

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data cites this paper.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Reinforcing General Reasoning without Verifiers

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T05:36:02.036290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:20fdc7987d5d2c76116af06d6a035d94ce12886b9e71691ef53cc88a816056d2

Observation 9c6b2e4e-3b78-4ece-bb46-3f61bff623b2 · inbound

Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning cites this paper.

Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning Reinforcing General Reasoning without Verifiers

Reference 187

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:15:49.560996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-10T19:15:27.406778Z digest=sha256:dc1c55d7b90d4e1f14a3acc43cf411470100de6c8668e1026e703300c4f37525

Observation e01d6ce9-375f-44d7-b0c2-6248af681050 · inbound

G-Zero: Self-Play for Open-Ended Generation from Zero Data cites this paper.

G-Zero: Self-Play for Open-Ended Generation from Zero Data Reinforcing General Reasoning without Verifiers

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:11:24.522825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:f774dcb281053b9eecaa925c8c0619d09d65ec744bfecfc4bbf1cb5ccdb8a681

Observation 8727f605-db0e-467b-8e17-1e63a94ffd2b · inbound

Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities cites this paper.

Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities Reinforcing General Reasoning without Verifiers

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-19T17:22:41.894869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T17:20:39.969447Z digest=sha256:615a5ef120507b2642e2df06d64757755f0727efa88fc62bc55cc4873ce7f4dd

Observation a24852dd-a7ee-4c3c-9469-f35df85cafb7 · inbound

CARE-RL: Capability-Aware Reinforcement Learning for Mitigating Cross-Domain Conflicts cites this paper.

CARE-RL: Capability-Aware Reinforcement Learning for Mitigating Cross-Domain Conflicts Reinforcing General Reasoning without Verifiers

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T19:02:33.978221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-28T19:01:14.340754Z digest=sha256:feab142d2263c15c2e880cc8e7fd4d32ff820ea66d38bca71bc666bae5405074

Observation d86298e8-0466-4040-9e5a-1d394911d33f · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation Reinforcing General Reasoning without Verifiers

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T20:56:13.507223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:538b7b148098adc006acf8f8271253f8b4c6f5ca28bc69014f0989e21d557c78

Observation 1c040be7-4644-4a71-ad7a-08c77ae88e6e · inbound

Predictive Divergence Masks for LLM RL cites this paper.

Predictive Divergence Masks for LLM RL Reinforcing General Reasoning without Verifiers

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-14T08:49:37.615924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T08:49:37.615924Z digest=sha256:1d0ffc8ba003cb74109200b9a607c0c871813a2769b32ed70a6d3d27923b72a5