Pith. sign in

Paper Citation Record · LEDGER

Reinforcing General Reasoning without Verifiers

As of 8 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 16 inbound Pith citation observations for arXiv:2505.21493.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.21493 v1

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:35:56.253150Z

measured 76 of 76 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:26:22.158083Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T20:56:13.505929Z

Reference resolution

60 of 60 outbound references displayed

  • verified exact0
  • verified fuzzy12
  • unresolved47
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a93d6290-e684-40aa-84ad-0648f6d7895e · outbound

This paper cites Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs.

Reinforcing General Reasoning without Verifiers Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:49.676520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:49.676520Z digest=sha256:cdc230d0c0177b39d529b3dd6d1240868dc488dea6388c5a45e9220d3c72ebf3

Observation 6b8912b9-44d6-4993-8c75-365c99efb6f4 · outbound

This paper cites an unresolved cited work.

Reinforcing General Reasoning without Verifiers Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:36:00.810288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:35:49.783935Z digest=sha256:25407c56e7f4425e8a8583445dad6e1df696cc7f1ed4bb3529960caede8f5e3d

Observation da40036f-4d63-460f-9951-a80959496b15 · outbound

This paper cites Bootstrapping language models with dpo implicit rewards.

Reinforcing General Reasoning without Verifiers Bootstrapping language models with dpo implicit rewards

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:36:00.555733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:35:49.895146Z digest=sha256:741c9da86a769ff04ffdeafae43905f093182dd8b863f973af69b29dcd47237a

Observation 52ebb6e6-cb33-47bf-9fd6-91affe7684c7 · outbound

This paper cites Language Models are Hidden Reasoners: Unlocking Latent Reasoning Capabilities via Self-Rewarding.

Reinforcing General Reasoning without Verifiers Language Models are Hidden Reasoners: Unlocking Latent Reasoning Capabilities via Self-Rewarding

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:50.020849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:50.020849Z digest=sha256:b67a4e063c9e2242656d46ffb74fcf6396f9cb37d1623b3b91c01b504f13f5c4

Observation 81e3becc-9596-4c67-8111-196dba531996 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Reinforcing General Reasoning without Verifiers Training Verifiers to Solve Math Word Problems

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:50.157115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:50.157115Z digest=sha256:6c967afbe4bc25463aa03117ef662d7b3c8b390b1eef039cc0ab3dd04c42326a

Observation 2d669c5b-0d3a-4ba5-9a8e-f0cf67637d61 · outbound

This paper cites SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines.

Reinforcing General Reasoning without Verifiers SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:50.307619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:50.307619Z digest=sha256:0b67cc7ed407911b2ddce4e77d25cae8ba14d8740ec43e0a0930dd8f17b92f32

Observation 7a802705-31d5-4dfe-9f52-c59d5cddffa1 · outbound

This paper cites Scaling laws for reward model overoptimization.

Reinforcing General Reasoning without Verifiers Scaling laws for reward model overoptimization

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:50.451904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:50.451904Z digest=sha256:6019fe2a0aa5001c5afc85c3bac42ec0131d0de1b76f08a18745ff25da14c82d

Observation 41cd47ac-0822-4c76-8982-68592cfe2bbc · outbound

This paper cites RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning.

Reinforcing General Reasoning without Verifiers RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:50.613928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:50.613928Z digest=sha256:a10937519ccab7d34716bad6bf4ce8f8abff2bb71827685adb1b023ab5998aa1

Observation 5dade0db-61c1-4c40-8094-7951010b67e5 · outbound

This paper cites The Llama 3 Herd of Models.

Reinforcing General Reasoning without Verifiers The Llama 3 Herd of Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:50.694745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:50.694745Z digest=sha256:02337f7ad69ef5c64ae999a353cfc108fad4e5ee0d4ba646fc106a1645df0f68

Observation f51e8249-efdf-4f50-a80c-83e7ace68b80 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Reinforcing General Reasoning without Verifiers DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:50.841396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:50.841396Z digest=sha256:ece9460c529a16af247a226f63b39c8b89bf05aeaa4338cadf6ebd5a36308103

Observation 532f803e-3d01-46a4-aeea-6d6fa3efeaf9 · outbound

This paper cites Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems, 2024.

Reinforcing General Reasoning without Verifiers Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems, 2024

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:50.986072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:50.986072Z digest=sha256:6c44d237158e16cfff0ad3393d7d210ff59516b26e69089a216a1f3b4d1d7e70

Observation 8ab29f53-b30b-494d-88ee-a43a1f872e9f · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Reinforcing General Reasoning without Verifiers Measuring Mathematical Problem Solving With the MATH Dataset

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:51.098076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:51.098076Z digest=sha256:60266b8bcc274e983c8961095ac580e8879c325f805263fba91dc77a3ed2266f

Observation 3317d73c-bec0-4a93-9f00-2ebfaa75998e · outbound

This paper cites Measuring mathematical problem solving with the math dataset.

Reinforcing General Reasoning without Verifiers Measuring mathematical problem solving with the math dataset

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:51.205483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:51.205483Z digest=sha256:8b1c69ca3493f68c18d60630c2ebc7d2f5d6b2313c31fd394a5c9ccc8edf7a93

Observation b2bbc0d8-6d7c-4e9a-8559-ef1948d36c91 · outbound

This paper cites Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model.

Reinforcing General Reasoning without Verifiers Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:51.324740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:51.324740Z digest=sha256:650ec66a07c97241d14693a1cfaf33bec0b40b37c35a3796c03d04e5b872fe38

Observation b479d45d-7c4b-4ae8-800e-17e7f52fb058 · outbound

This paper cites Self-Improvement in Language Models: The Sharpening Mechanism.

Reinforcing General Reasoning without Verifiers Self-Improvement in Language Models: The Sharpening Mechanism

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:51.425938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:51.425938Z digest=sha256:74603cd4bc802745c286124b2238ea7be5f1e60be07abd7b3fa89b70b7dbeaba

Observation 81cf8b30-cfc3-4b32-a192-89c7731e1297 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

Reinforcing General Reasoning without Verifiers Gonzalez, Hao Zhang, and Ion Stoica

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:51.526601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:51.526601Z digest=sha256:45c31aadf7347c15a7b0cfcb51c30b3b3e5dd57e643c5c5bc6ff5196659387cb

Observation 7d2439f9-9c29-4fb7-b7ca-c8cb2be5b697 · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

Reinforcing General Reasoning without Verifiers Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:51.600499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:51.600499Z digest=sha256:d39f4b3d61448e1ec001c3d687ec0ad1ca2be82b2a3b55eb0df77e3fe10fd524

Observation 59922537-1b14-47ce-a2d4-346a1010e5b7 · outbound

This paper cites Solving quantitative reasoning problems with language models.Advances in Neural Information Processing Systems, 35:3843–3857, 2022.

Reinforcing General Reasoning without Verifiers Solving quantitative reasoning problems with language models.Advances in Neural Information Processing Systems, 35:3843–3857, 2022

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:51.708819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:51.708819Z digest=sha256:10740b304c808b7ed61b2d66f1239f3f11c728a04f82ebe972929990d97aa1bd

Observation cadd1958-83e1-4260-a2fe-7ebce56bb0ee · outbound

This paper cites Numinamath: The largest public dataset in ai4maths with 860k pairs of competition math problems and solutions.Hugging Face repository, 13:9, 2024.

Reinforcing General Reasoning without Verifiers Numinamath: The largest public dataset in ai4maths with 860k pairs of competition math problems and solutions.Hugging Face repository, 13:9, 2024

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:51.819498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:51.819498Z digest=sha256:84e454a44942c02d15b010501eec8ebce04b1ebdcbd3572db619dfbb4e5d6d1e

Observation fb322e58-6b54-4677-94c9-26d5f8f67426 · outbound

This paper cites Let’s verify step by step.

Reinforcing General Reasoning without Verifiers Let’s verify step by step

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:51.952732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:51.952732Z digest=sha256:54a2e7137135ac658486d6f19057307ce57966c4089f7ac2b80956a80fd55157

Observation 23d49f70-44e6-486d-8234-6972ef41ec39 · outbound

This paper cites X-Reasoner: Towards Generalizable Reasoning Across Modalities and Domains.

Reinforcing General Reasoning without Verifiers X-Reasoner: Towards Generalizable Reasoning Across Modalities and Domains

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:52.104790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:52.104790Z digest=sha256:42c458e71b5a70627709c8f8f785b86480ecb40770108b94ad1dce2c8d0c07b6

Observation 8f616da6-ee5f-48fa-a183-eb4a4e5ecd88 · outbound

This paper cites Oat: A research-friendly framework for llm online alignment.https://github.com/sail-sg/oat, 2024.

Reinforcing General Reasoning without Verifiers Oat: A research-friendly framework for llm online alignment.https://github.com/sail-sg/oat, 2024

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:36:00.206502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:35:52.211512Z digest=sha256:648b92b7e20f287f1b2fedbb82e0a74b9258dd6a0813050556fe42fd143a47c2

Observation 99224617-7326-4c1a-85f0-2dd010ea96bc · outbound

This paper cites There may not be aha moment in r1-zero-like training — a pilot study.

Reinforcing General Reasoning without Verifiers There may not be aha moment in r1-zero-like training — a pilot study

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:52.284816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:52.284816Z digest=sha256:5db344fa6df33ef288d319fc94b4f6d5c0890f1e0b582d1cacbbeeb26a0d46be

Observation 8fd9a995-3dc6-479b-9102-e66be0b39631 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

Reinforcing General Reasoning without Verifiers Understanding R1-Zero-Like Training: A Critical Perspective

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:52.414743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:52.414743Z digest=sha256:37eb0802f4d2f2900ecec4fee6d36c3f3c88d15e1a97df6fe7ce729f57bd872a

Observation 00842862-6da4-4cb5-947f-450fb3d521b8 · outbound

This paper cites Deepcoder: A fully open-source 14b coder at o3-mini level, 2025.

Reinforcing General Reasoning without Verifiers Deepcoder: A fully open-source 14b coder at o3-mini level, 2025

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:52.527377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:52.527377Z digest=sha256:1c5f1ac2e257c7678f8a8fa5dbd0a09951e4750f2019f5277c9a0df6d527da47

Observation f99e1cdc-3791-47c3-88d4-7fe27993c55d · outbound

This paper cites Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica.

Reinforcing General Reasoning without Verifiers Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:52.600474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:52.600474Z digest=sha256:11d4a4c10d65d2762c15a077302884549056b45e27b26bf5d8d7b39c4fe0688b

Observation e8604899-443b-4c3e-8189-b8cef03b172b · outbound

This paper cites General-Reasoner: Advancing LLM Reasoning Across All Domains.

Reinforcing General Reasoning without Verifiers General-Reasoner: Advancing LLM Reasoning Across All Domains

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:52.681204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:52.681204Z digest=sha256:02c8484c4abd0a1d5cca925a244f9dd3d309836f1931f44e690eaa57a68ab608

Observation a64d249a-1c61-4104-b764-d954622e262e · outbound

This paper cites Ng, Daishi Harada, and Stuart J.

Reinforcing General Reasoning without Verifiers Ng, Daishi Harada, and Stuart J

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:59.955749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:35:52.844277Z digest=sha256:bcfb5c4995c6b4880708b2f0374b7494fcd6a145a514502d35c633e41e08fe12

Observation 4a747535-adc6-4e76-9340-bbe4f3af9954 · outbound

This paper cites Learning to reason with llms, 2024.

Reinforcing General Reasoning without Verifiers Learning to reason with llms, 2024

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:59.437638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:35:53.099446Z digest=sha256:34e7f8faf01f3eb2590d092bfd2c20e9138b0b3559391626c4715b3e1c797207

Observation fedd8f05-30b9-4aac-8644-585576047986 · outbound

This paper cites Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022.

Reinforcing General Reasoning without Verifiers Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:53.235958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:53.235958Z digest=sha256:bb4a7a4e38f6e0d93da9f89ecbaaef890aa8c864df26e91d3fae4eeea1ef0ef4

Observation 6de7d42f-7f10-4ba6-967c-3f7b9cc6ec85 · outbound

This paper cites Tinyzero.

Reinforcing General Reasoning without Verifiers Tinyzero

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:53.355905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:53.355905Z digest=sha256:5f6a6870675c621e338aea03712a34ab7ec2bad83bcad5c478458628ffac2cd3

Observation a8008712-a0b1-452c-a993-0fa78b48d8e1 · outbound

This paper cites Training chain-of- thought via latent-variable inference.Advances in Neural Information Processing Systems, 36: 72819–72841, 2023.

Reinforcing General Reasoning without Verifiers Training chain-of- thought via latent-variable inference.Advances in Neural Information Processing Systems, 36: 72819–72841, 2023

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:59.181728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:35:53.438954Z digest=sha256:214c440e148e19df121c7b98dad77f40a4dd16b853aa8b392d8a7977e0bc2da5

Observation 80afe6a7-a471-41af-9798-d8c60a771a76 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Reinforcing General Reasoning without Verifiers Direct preference optimization: Your language model is secretly a reward model

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:53.528942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:53.528942Z digest=sha256:f54bebe4ebde6562e83abd702bf853fe48557116e014b116fc3b42a5c99dd781

Observation a1dd8e9e-70b0-4602-9844-d1cab760b508 · outbound

This paper cites Learning to drive a bicycle using reinforcement learning and shaping.

Reinforcing General Reasoning without Verifiers Learning to drive a bicycle using reinforcement learning and shaping

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:58.957200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:35:53.652657Z digest=sha256:f8bc001f32ba6a5effa33ebf62c8b5afbd359a873308a0c1e9a7970f6c931643

Observation bff29ac9-4452-474e-a5a8-662cfbb18d31 · outbound

This paper cites Gpqa: A graduate-level google-proof q&a benchmark.

Reinforcing General Reasoning without Verifiers Gpqa: A graduate-level google-proof q&a benchmark

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:53.729825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:53.729825Z digest=sha256:731e9c0db9673341b930baef865f74b55d55e9a409442e7a1f42373b055b73ce

Observation ca43c1f2-a018-4109-b4ff-3389791b2fca · outbound

This paper cites Proximal Policy Optimization Algorithms.

Reinforcing General Reasoning without Verifiers Proximal Policy Optimization Algorithms

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:53.862659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:53.862659Z digest=sha256:2882eccc96fa9422adeaed40b87feaa3a3e980ef491f57b7fdc306dc5b8b8bb7

Observation a7849d0c-5308-4572-aae7-00b1eb99e966 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Reinforcing General Reasoning without Verifiers DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:53.962521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:53.962521Z digest=sha256:bf83fb24a8a33c3a7ae199bec19becd80c6bd0d0ece9b2740c6b92b4ae0ebd03

Observation facf22d4-454e-4613-b858-c10da0a74038 · outbound

This paper cites Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains.

Reinforcing General Reasoning without Verifiers Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:54.067551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:54.067551Z digest=sha256:dc101ad30baf714ca8abbf115acacd421af6a7e26fb7486a2cc4d3c0a8c6e9d3

Observation 13680724-420a-4471-9655-96475f45143a · outbound

This paper cites Sutton and Andrew G.

Reinforcing General Reasoning without Verifiers Sutton and Andrew G

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:54.172581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:54.172581Z digest=sha256:9ef2c9410dd59f1c7dcd591f5bd3988fca75ecab51b99061df6aae8d0f18455c

Observation c5e30f52-8d0b-4008-a8f2-fc41b25e47ad · outbound

This paper cites Beyond Verifiable Rewards: Scaling Reinforcement Learning for Language Models to Unverifiable Data.

Reinforcing General Reasoning without Verifiers Beyond Verifiable Rewards: Scaling Reinforcement Learning for Language Models to Unverifiable Data

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:54.275667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:54.275667Z digest=sha256:07cb2fea3e9d17a69ef3e3f3320c2395f7d0da8863e8c95de07ff207b59b6582

Observation 23fbb93f-8030-4c13-9649-781b8f2dca25 · outbound

This paper cites Qwen3, April 2025.

Reinforcing General Reasoning without Verifiers Qwen3, April 2025

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:54.354782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:54.354782Z digest=sha256:3cf9f7f8536a1256bf3a12edc0d0399105b97291c70ca0fdb19f9a64c1785738

Observation a661f906-bd43-41e7-96cd-89f67894319d · outbound

This paper cites Mmlu-pro: A more robust and challenging multi-task language understanding benchmark.

Reinforcing General Reasoning without Verifiers Mmlu-pro: A more robust and challenging multi-task language understanding benchmark

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:54.447387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:54.447387Z digest=sha256:7755062b7a44916130f389577ffa868df0c05985ea20ce4a5858ef5d6c4ce636

Observation 8ae8c311-f20e-43d8-ad0d-9983c0c8b323 · outbound

This paper cites Qwen2.5 Technical Report.

Reinforcing General Reasoning without Verifiers Qwen2.5 Technical Report

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:54.550113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:54.550113Z digest=sha256:587b0f5b5512e5ead76e4af3b14f8d6b8372503c1068c09bbd93aa43949c36c1

Observation 6ab5a380-d554-4d25-a2ea-5ba8c30cf2dd · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

Reinforcing General Reasoning without Verifiers Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:54.659939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:54.659939Z digest=sha256:0ce26f117b4b240a26b2d0ca65354f8a89d05d1cbb5c342d7c7919303b96bec7

Observation dfbbdad7-3e5b-4f76-a2f8-96a7f26b2e10 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Reinforcing General Reasoning without Verifiers DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:54.837737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:54.837737Z digest=sha256:e3cc61de2d98c0f49a7a47e45a6fcea0029f1cdd98260851790fcb772ab8967c

Observation 0b46341a-8a36-4346-a992-69c3bac38ab0 · outbound

This paper cites Self-rewarding language models.International Conference on Machine Learning,.

Reinforcing General Reasoning without Verifiers Self-rewarding language models.International Conference on Machine Learning,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:58.735265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:35:54.952680Z digest=sha256:f49d58f69863cfa0ec7def3af86d731af6282261494c47059b83420b0a6ca041

Observation ea9dad09-2bb7-4512-9ae8-730ae885cb81 · outbound

This paper cites Naturalreasoning: Reasoning in the wild with 2.8 m challenging questions.arXiv preprint arXiv:2502.13124, 2025.

Reinforcing General Reasoning without Verifiers Naturalreasoning: Reasoning in the wild with 2.8 m challenging questions.arXiv preprint arXiv:2502.13124, 2025

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:55.134741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:55.134741Z digest=sha256:f93227bf4e130331f4e141d445dba352d11c38df5cc38fdd39c2329a1208990a

Observation eb9e4194-f947-421d-8802-662ec50072e4 · outbound

This paper cites Scaling Relationship on Learning Mathematical Reasoning with Large Language Models.

Reinforcing General Reasoning without Verifiers Scaling Relationship on Learning Mathematical Reasoning with Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:55.254650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:55.254650Z digest=sha256:11c7998ea13402b400da1ff9d8a40829ffa6779f896c801a1538b9eb752ad4df

Observation 9cb71e8a-757a-459b-a8f0-6bced49c0046 · outbound

This paper cites Mammoth2: Scaling instructions from the web.Advances in Neural Information Processing Systems, 37:90629–90660, 2024.

Reinforcing General Reasoning without Verifiers Mammoth2: Scaling instructions from the web.Advances in Neural Information Processing Systems, 37:90629–90660, 2024

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:55.410399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:55.410399Z digest=sha256:bbba4a3d666c5848c11c1978bb28c148b8267a8f1acf39c61b76baf3414b6e0d

Observation fd3fc11e-4ffb-4e99-b255-fabad4fbd632 · outbound

This paper cites 7b model and 8k examples: Emerging reasoning with reinforcement learning is both effective and efficient.https://hkust-nlp.notion.site/simplerl-reason, 2025.

Reinforcing General Reasoning without Verifiers 7b model and 8k examples: Emerging reasoning with reinforcement learning is both effective and efficient.https://hkust-nlp.notion.site/simplerl-reason, 2025

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:58.609689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:35:55.503236Z digest=sha256:fe7bfe222339d8f186614dccf3d7fa1a16a54d91218dbe71162753adcf322fa7

Observation db5b202b-4c9f-4cf3-8395-4e90c75576e0 · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Reinforcing General Reasoning without Verifiers Fine-Tuning Language Models from Human Preferences

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:55.613828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:55.613828Z digest=sha256:f79525ebc53a30f577f59ec3ee4581ca92fab8f7724308bb93b3b5e950ba7cd6

Observation ec4bb7de-61fa-4edc-990c-b6107dad8896 · outbound

This paper cites TTRL: Test-Time Reinforcement Learning.

Reinforcing General Reasoning without Verifiers TTRL: Test-Time Reinforcement Learning

Reference 52

Resolution
malformed identifier
no resolver link, observed 2026-08-07T13:35:55.747382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:55.747382Z digest=sha256:052db66399d535091178fae276aa72f622b5cbf2c8e866a639946af3cc2e05db

Observation acce323f-2ddd-4261-a551-78492765c929 · outbound

This paper cites an unresolved cited work.

Reinforcing General Reasoning without Verifiers Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:35:58.331998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:35:55.845493Z digest=sha256:81d323b7ba99814a839dc58aea3244b474ddb662e916f81d4afcc4313324658a

Observation e6763064-39f5-48b3-b01b-87fda9d95806 · outbound

This paper cites This relation is usually given in a form where a graph or a formula relates period to absolute magnitude (a measure of intrinsic brightness).

Reinforcing General Reasoning without Verifiers This relation is usually given in a form where a graph or a formula relates period to absolute magnitude (a measure of intrinsic brightness)

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:58.140035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:35:55.913862Z digest=sha256:3134b2292ec3f716d179e20a1f24a315fadc5083a54fce5102ae1aab036dc3e4

Observation 120c5616-51b9-4211-9664-27fd8f35cbdc · outbound

This paper cites an unresolved cited work.

Reinforcing General Reasoning without Verifiers Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:35:57.928563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:35:55.981091Z digest=sha256:66416091f24c67265f79eecacf44ed7c8a431cc052dfc13f1a3514dc327253c4

Observation 0edea18e-4c38-4d9f-974d-6f07835b4b5f · outbound

This paper cites ""everyone else is doing it.

Reinforcing General Reasoning without Verifiers ""everyone else is doing it

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:57.686510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:35:56.073560Z digest=sha256:8faef6d5fa3528edf330d0750406a96ff703a45a3edfdca3368c015fbb0bf55b

Observation ecf44972-dfd3-463b-97e4-8ff3ab1de69d · outbound

This paper cites Their reasoning is fear-based, and they view rules as set by authority figures.

Reinforcing General Reasoning without Verifiers Their reasoning is fear-based, and they view rules as set by authority figures

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:57.494309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:35:56.170121Z digest=sha256:cf2c13250f8177e8b41f1f92457fab1210d4302d04bee7cf1f3143c30a20d78c

Observation ef0107df-7832-4415-b647-576f0b7fd712 · outbound

This paper cites what’s in it for me?.

Reinforcing General Reasoning without Verifiers what’s in it for me?

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:57.168558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:35:56.253150Z digest=sha256:1345453371599b4a9e3433ced9172ab7bdfca760e79ab2108b3e7c88470451a3

Observation 0e3fc1d9-feba-4c02-b9d5-adcd9d901f15 · outbound

This paper cites an unresolved cited work.

Reinforcing General Reasoning without Verifiers Unresolved cited work

Reference 1999

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:35:59.715432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:35:52.950970Z digest=sha256:e50910957c6ec9953e0cc0c3f707f0d2116c73c0422f43a1c0a9ca8e3d969d65

Observation 2c3364b5-f502-45b4-9dda-64bae6f72e75 · outbound

This paper cites Self-Rewarding Language Models.

Reinforcing General Reasoning without Verifiers Self-Rewarding Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:55.039840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:55.039840Z digest=sha256:4ad89cdb6b52eeb9b6d41652d7c0acc7bbbf623cdbc0fa434434c2177cc1b697

Pith citing papers

Observation 5011bd04-bd69-40bb-b9b5-63b1463ceb1d · inbound

Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems cites this paper.

Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems Reinforcing General Reasoning without Verifiers

Reference 167

Resolution
verified exact
arxiv_id, observed 2026-05-22T21:42:10.812398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T21:39:49.832151Z digest=sha256:f146702cde53425c27efd4e5f6e4fdbbbfbfb9f621439057ca9cd70320fc16f6

Observation 881d37f7-a003-4bc1-a8de-289095639b6c · inbound

Reinforcement Pre-Training cites this paper.

Reinforcement Pre-Training Reinforcing General Reasoning without Verifiers

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:22.158083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:26:22.158083Z digest=sha256:4f5bc93fc9f969ec97eb39b06eb273851c96ef631791d8f16f4c3de2bee7387b

Observation 09f1d75b-3d1f-411a-81d3-dc06e696881d · inbound

Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks cites this paper.

Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks Reinforcing General Reasoning without Verifiers

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:52:14.184491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T09:48:56.990745Z digest=sha256:91439f63f2750703f33989e1d46fde6d9bfd7a0166a6b580e91edb989b00b8bd

Observation d9cb06d2-5012-4061-985a-434b37d473ca · inbound

URPO: A Unified Reward & Policy Optimization Framework for Large Language Models cites this paper.

URPO: A Unified Reward & Policy Optimization Framework for Large Language Models Reinforcing General Reasoning without Verifiers

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T14:51:25.447032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:51:25.447032Z digest=sha256:c60b98075b8cfa10806e1bedc1646b036b4d3757103602add4c326ee72aabacd

Observation 85b3a33d-241e-4e2f-ab1e-ef958c464377 · inbound

Reverse-Engineered Reasoning for Open-Ended Generation cites this paper.

Reverse-Engineered Reasoning for Open-Ended Generation Reinforcing General Reasoning without Verifiers

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T00:06:08.567934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T00:06:08.567934Z digest=sha256:7cb660ac7fa7737c980d461029ef65039cb3d1b76363cd5839dd6976b281410d

Observation f8b2fe70-730b-4ddc-8020-c762a6a2adf3 · inbound

Compute as Teacher: Turning Inference Compute Into Reference-Free Supervision cites this paper.

Compute as Teacher: Turning Inference Compute Into Reference-Free Supervision Reinforcing General Reasoning without Verifiers

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T15:46:34.124175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T15:45:09.730804Z digest=sha256:a29731ad6872ecd970a76dfc35beab3d05021f39d711ec87364ba7316adf7fee

Observation adc3cf64-8802-4479-ba29-f2d3a969ed65 · inbound

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle cites this paper.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Reinforcing General Reasoning without Verifiers

Reference 257

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:48.740176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:48.740176Z digest=sha256:3ef9c9c5c475d33cef7fbc64355230868cd0af9d92321a6a6c68c855d63f5f20

Observation 57b2310b-f1db-4c8c-8d96-305438cda973 · inbound

Learning to Pose Problems: Reasoning-Driven and Solver-Adaptive Data Synthesis cites this paper.

Learning to Pose Problems: Reasoning-Driven and Solver-Adaptive Data Synthesis Reinforcing General Reasoning without Verifiers

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-17T23:00:25.487127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T22:59:48.260121Z digest=sha256:eea6a828100784064eaba130cb1168bd4d487ba6a3320afa3cbd0ee56458247a

Observation b607286b-d83d-43df-9474-bd80a6ef4580 · inbound

VI-CuRL: Stabilizing Verifier-Independent RL Reasoning via Confidence-Guided Variance Reduction cites this paper.

VI-CuRL: Stabilizing Verifier-Independent RL Reasoning via Confidence-Guided Variance Reduction Reinforcing General Reasoning without Verifiers

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-25T07:26:41.714012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T07:26:35.767179Z digest=sha256:1d805a5b1d392dab9886cb9a5cb43da763404bc21244a5639f901f2145c5fb69

Observation 52f3cb1b-099f-4c53-9e14-867cb07c84b5 · inbound

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data cites this paper.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Reinforcing General Reasoning without Verifiers

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T05:36:02.036290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:7b25a0deeecb15a513813206240c85acc4348f5bda48a78b55527ef81e4aec65

Observation 9c6b2e4e-3b78-4ece-bb46-3f61bff623b2 · inbound

Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning cites this paper.

Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning Reinforcing General Reasoning without Verifiers

Reference 187

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:15:49.560996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T19:15:27.406778Z digest=sha256:f3269f25ff1514390ce91c03550b4449cb31ccadd72c42668cf861e05508fbc7

Observation e01d6ce9-375f-44d7-b0c2-6248af681050 · inbound

G-Zero: Self-Play for Open-Ended Generation from Zero Data cites this paper.

G-Zero: Self-Play for Open-Ended Generation from Zero Data Reinforcing General Reasoning without Verifiers

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:11:24.522825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:2d639f73136c7f5a41824c437e271a8bdaf561b8a36d84cd3938ac975bc952d5

Observation 8727f605-db0e-467b-8e17-1e63a94ffd2b · inbound

Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities cites this paper.

Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities Reinforcing General Reasoning without Verifiers

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-19T17:22:41.894869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T17:20:39.969447Z digest=sha256:25387a911210a0cc8570e1b68b393135e4db8c10ad4e9dd358b7348358115545

Observation a24852dd-a7ee-4c3c-9469-f35df85cafb7 · inbound

CARE-RL: Capability-Aware Reinforcement Learning for Mitigating Cross-Domain Conflicts cites this paper.

CARE-RL: Capability-Aware Reinforcement Learning for Mitigating Cross-Domain Conflicts Reinforcing General Reasoning without Verifiers

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T19:02:33.978221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T19:01:14.340754Z digest=sha256:41c1d19bdb4fea3ccd835f6c828cf96fdfbd984a8a8c8a40764b724b403fafdd

Observation d86298e8-0466-4040-9e5a-1d394911d33f · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation Reinforcing General Reasoning without Verifiers

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T20:56:13.507223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:665a37bd961b967d1abd9ac3cbf6d18e2bc841c1a1dd3ff203f18777271071f6

Observation 1c040be7-4644-4a71-ad7a-08c77ae88e6e · inbound

Predictive Divergence Masks for LLM RL cites this paper.

Predictive Divergence Masks for LLM RL Reinforcing General Reasoning without Verifiers

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-14T08:49:37.615924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T08:49:37.615924Z digest=sha256:488df3dfec0f4c9683d06c67f418ced5774f6db53c0231d982f84d934fc8ed10