Pith. sign in

Paper Citation Record · LEDGER

Exploring the Potential of Offline RL for Reasoning in LLMs: A Preliminary Study

As of 19 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 1 inbound Pith citation observation for arXiv:2505.02142.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.02142 v1

Coverage vector

measured 64 of 64 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T01:04:38.905912Z

measured 65 of 65 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-26T12:15:08.304150Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T08:09:40.573829Z

Reference resolution

64 of 64 outbound references displayed

  • verified exact0
  • verified fuzzy7
  • unresolved56
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cfc6ffb7-03ba-4eba-9fd6-71af80d16ba8 · outbound

This paper cites Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models.

Exploring the Potential of Offline RL for Reasoning in LLMs: A Preliminary Study Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:38.684571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:38.684571Z digest=sha256:1e45e35d9347ae6573daf312d4ca4c9e0f01a5da81655c6a9f51fc12e1f101d2

Observation 21aaef41-7ec3-4365-960b-e84b79d8d4a6 · outbound

This paper cites 2023 AMC 8 Problems/Problem 10.

Exploring the Potential of Offline RL for Reasoning in LLMs: A Preliminary Study 2023 AMC 8 Problems/Problem 10

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:38.688977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:38.688977Z digest=sha256:ca61f7963025d13b9c45ef50a3bf76e8c9ffe06f4149b40662649c689199be41

Observation e60a095f-4414-472f-a3bc-0f3766cc8b3e · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Exploring the Potential of Offline RL for Reasoning in LLMs: A Preliminary Study DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:38.696928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:38.696928Z digest=sha256:2d971bf4add277a46582d3eb4127b1c406ea7edce6aeb5bb3b459f1779270cfc

Observation dfb5a672-30d7-486a-87c1-6d3faeed708a · outbound

This paper cites Aime_1983_2024 (revision 6283828), 2025.

Exploring the Potential of Offline RL for Reasoning in LLMs: A Preliminary Study Aime_1983_2024 (revision 6283828), 2025

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T01:04:39.509068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T01:04:38.701302Z digest=sha256:2f9e1c501cacdc5975f2aadded934a24a1cf360334ee9587ce2a0f521f18b644

Observation b981d4f7-0535-4099-aa09-7de0a0976236 · outbound

This paper cites KTO: Model Alignment as Prospect Theoretic Optimization.

Exploring the Potential of Offline RL for Reasoning in LLMs: A Preliminary Study KTO: Model Alignment as Prospect Theoretic Optimization

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:38.704305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:38.704305Z digest=sha256:862e1195541e5c9f97bc5b47b4ea988b91c4bfdf9458e549e19aa60e7a0f371a

Observation c03598fa-80d8-423c-80a7-12c5a4bc109c · outbound

This paper cites Open r1: A fully open reproduction of deepseek-r1, January 2025.

Exploring the Potential of Offline RL for Reasoning in LLMs: A Preliminary Study Open r1: A fully open reproduction of deepseek-r1, January 2025

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:38.707916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:38.707916Z digest=sha256:620a216f9ff78f10699ff058fb53a1adc2c3d9bfbbbe4d1204e3c3d52d17f575

Observation bfd2a199-21d7-4047-b36c-e29037e716e8 · outbound

This paper cites IPO: Your Language Model is Secretly a Preference Classifier.

Exploring the Potential of Offline RL for Reasoning in LLMs: A Preliminary Study IPO: Your Language Model is Secretly a Preference Classifier

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:38.711041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:38.711041Z digest=sha256:9834cbcd0cdd7015e0a66cc2f58cc6bae6a29b4bdd3fd106ec6ddc161527be60

Observation b043f8c1-2096-43e6-9b3c-ded5ad26ccf3 · outbound

This paper cites Fine flan: Seqio to parquet so you don’t have to.

Exploring the Potential of Offline RL for Reasoning in LLMs: A Preliminary Study Fine flan: Seqio to parquet so you don’t have to

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:38.714253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:38.714253Z digest=sha256:209bbd497c33f680e0f182edb1e3b7165ffbef36c251956aa9072b805e30c846

Observation 35edec94-921d-417f-8268-b5b7d7e4a3df · outbound

This paper cites Deepmath-103k: A large-scale, challenging, decontaminated, and verifiable mathematical dataset for advancing reasoning.

Exploring the Potential of Offline RL for Reasoning in LLMs: A Preliminary Study Deepmath-103k: A large-scale, challenging, decontaminated, and verifiable mathematical dataset for advancing reasoning

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T01:04:39.485870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T01:04:38.718052Z digest=sha256:190ef1ffe55f07f772c08a043ef2af700aba442071c38582c44f9a33114af5a9

Observation 36ba94a6-3d6c-401e-8be6-f05dc26f9909 · outbound

This paper cites Logic-701: A benchmark dataset for logical reasoning in english and russian.

Exploring the Potential of Offline RL for Reasoning in LLMs: A Preliminary Study Logic-701: A benchmark dataset for logical reasoning in english and russian

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:38.720852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:38.720852Z digest=sha256:e94a2c1fd0d1852043c86c95027a990afe7835aec851ae05e8431f634f1b4dd8

Observation 8acd92c1-cf25-489f-b27b-1220df8292fa · outbound

This paper cites OpenCoder: The Open Cookbook for Top-Tier Code Large Language Models.

Exploring the Potential of Offline RL for Reasoning in LLMs: A Preliminary Study OpenCoder: The Open Cookbook for Top-Tier Code Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:38.723914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:38.723914Z digest=sha256:3fac1ae635c052575fdf24f124b3d2630cdba76475f9ff014a82173577749f07

Observation d021b717-b790-4671-a5ca-0bb263ddb1a9 · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

Exploring the Potential of Offline RL for Reasoning in LLMs: A Preliminary Study LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:38.727591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:38.727591Z digest=sha256:8cf1020b8913f5d058a9ae8ebb20ecea865b8730908cff4c9f01a49d1cd010d5

Observation 42d522cd-ce98-4ee4-b609-2adaff4b37ad · outbound

This paper cites Ncert biology 11th dataset.

Exploring the Potential of Offline RL for Reasoning in LLMs: A Preliminary Study Ncert biology 11th dataset

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:38.731892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:38.731892Z digest=sha256:1fa7d3d667d92fecaf0459ae83384b55201ad76f19eea7f7c08bc119bf147a62

Observation c94e6c7e-6def-4fae-854b-0f554c700c60 · outbound

This paper cites Ncert biology 12th dataset.

Exploring the Potential of Offline RL for Reasoning in LLMs: A Preliminary Study Ncert biology 12th dataset

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:38.735710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:38.735710Z digest=sha256:e34e1ac86a7ae59fa9ed9c1cb4190e88f27214919936fddcd09642f983e2f282

Observation e7ef0b56-65a4-4d23-b65e-da49396f755d · outbound

This paper cites Ncert chemistry 11th dataset.

Exploring the Potential of Offline RL for Reasoning in LLMs: A Preliminary Study Ncert chemistry 11th dataset

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:38.738934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:38.738934Z digest=sha256:9cf62047beb0c161be5ee49c4e15a78475e494ac9de741d5cec9006c79442309

Observation c6bead1f-4387-4d28-bc3e-50f346df2916 · outbound

This paper cites Ncert chemistry 12th dataset.

Exploring the Potential of Offline RL for Reasoning in LLMs: A Preliminary Study Ncert chemistry 12th dataset

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:38.742197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:38.742197Z digest=sha256:81bab477b9993577f400d27250aec9a1fa38f1f755f223c9b032fd4854ed0f40

Observation 8cde3bd2-2c7e-4634-8363-63f554372d81 · outbound

This paper cites Ncert physics 11th dataset.

Exploring the Potential of Offline RL for Reasoning in LLMs: A Preliminary Study Ncert physics 11th dataset

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:38.746718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:38.746718Z digest=sha256:88cea0ab877e9a21b1e6ac7cdda32f04b7114cb14a607015066a9555ebf2da60

Observation c0bbfb90-8d5d-4a05-98d5-fa4f063b1e26 · outbound

This paper cites Ncert physics 12th dataset.

Exploring the Potential of Offline RL for Reasoning in LLMs: A Preliminary Study Ncert physics 12th dataset

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T01:04:39.439637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T01:04:38.750319Z digest=sha256:70bfbe0a6c1cc66a1514d78a0a1be9ef5844155c464fc4721c89bf981c020aad

Observation e9709543-80d9-4672-8f4b-5cca0aa9c277 · outbound

This paper cites Miranda, Alisa Liu, Nouha Dziri, Shane Lyu, Yuling Gu, Saumya Malik, Victoria Graf, Jena D.

Exploring the Potential of Offline RL for Reasoning in LLMs: A Preliminary Study Miranda, Alisa Liu, Nouha Dziri, Shane Lyu, Yuling Gu, Saumya Malik, Victoria Graf, Jena D

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:38.753815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:38.753815Z digest=sha256:fda446381eb97a9f1538575cfb6ae175372eeb42925ff1629269d0bf8846c9d5

Observation 52f0d26b-1f39-4ea6-8f0c-0cafd8564693 · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

Exploring the Potential of Offline RL for Reasoning in LLMs: A Preliminary Study Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:38.756644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:38.756644Z digest=sha256:a0e7af9144b4ab9431aefd5eef0ffcf7aea2b163817e991b0daf3b4694237403

Observation f43b0eef-f44c-4906-9d5a-7d44d721057b · outbound

This paper cites Numinamath.

Exploring the Potential of Offline RL for Reasoning in LLMs: A Preliminary Study Numinamath

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:38.759853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:38.759853Z digest=sha256:800183ae2472f96a29732560451884b185352c95be2d3b93bd0bbab907f1c9c6

Observation 2f863bc1-17ab-45ff-af56-419ba43e37e8 · outbound

This paper cites From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline.

Exploring the Potential of Offline RL for Reasoning in LLMs: A Preliminary Study From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:38.762722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:38.762722Z digest=sha256:51980b985558994a7dd7af54485ee6e24312a2f132faf012512602f8b9316ebb

Observation ea88edfd-e382-4f14-a47f-eafaec48804a · outbound

This paper cites Gonzalez, and Ion Stoica.

Exploring the Potential of Offline RL for Reasoning in LLMs: A Preliminary Study Gonzalez, and Ion Stoica

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T01:04:39.416887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T01:04:38.765848Z digest=sha256:be20d14bf4bd632e7a18eb3cfcdd5fb1e6a9f7a7662f18594e112da457eeba8c

Observation b7ce3c49-d795-499a-a456-2ee1ea860d6c · outbound

This paper cites Openorca: An open dataset of gpt augmented flan reasoning traces.

Exploring the Potential of Offline RL for Reasoning in LLMs: A Preliminary Study Openorca: An open dataset of gpt augmented flan reasoning traces

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:38.769300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:38.769300Z digest=sha256:c681133f3375c1e09eac691a3a56c755085d5fd7b5c48ed1e14fcad7ed7bb4a3

Observation 64020ec3-470c-4727-8ea9-87132461dc8b · outbound

This paper cites Kw-r1: A simple implementation of the grpo algorithm.

Exploring the Potential of Offline RL for Reasoning in LLMs: A Preliminary Study Kw-r1: A simple implementation of the grpo algorithm

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T01:04:39.400427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T01:04:38.774225Z digest=sha256:43e32460dd8587ee02d961b722a0d9ef20ac412cc2edae4d40e31a5bb5ec20f9

Observation facd71f4-229c-44e6-a287-0e06394fdd04 · outbound

This paper cites Length Desensitization in Direct Preference Optimization.

Exploring the Potential of Offline RL for Reasoning in LLMs: A Preliminary Study Length Desensitization in Direct Preference Optimization

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:38.778559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:38.778559Z digest=sha256:1162edccc6865dca3960e031c30f6cbf26182c03a04d36b10da5b8caa811dc83

Observation 0670d9b8-4713-456d-be9f-fabe8423b899 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

Exploring the Potential of Offline RL for Reasoning in LLMs: A Preliminary Study Understanding R1-Zero-Like Training: A Critical Perspective

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:38.781933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:38.781933Z digest=sha256:e2c4c9851f7affe4252bd4795e9a45bdfbf1721df103b68ef563be45ba4a7bcd

Observation d431f0b3-a664-4330-8609-ed74c031d357 · outbound

This paper cites an unresolved cited work.

Exploring the Potential of Offline RL for Reasoning in LLMs: A Preliminary Study Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-16T01:04:39.390630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T01:04:38.785530Z digest=sha256:f54e8dc03e9e6f81f6dd6cab3c36a77f86a9e938fad673381b67a93b0fcf7b7d

Observation 0577ce0a-83a9-4fd4-8c11-5b7329c71616 · outbound

This paper cites Le, Barret Zoph, Jason Wei, and Adam Roberts.

Exploring the Potential of Offline RL for Reasoning in LLMs: A Preliminary Study Le, Barret Zoph, Jason Wei, and Adam Roberts

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:38.789797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:38.789797Z digest=sha256:48455e3da0ddc9f7cb41a7dbc660f94229033de850586ac55e9734f4de924621

Observation 4b1d1f10-5720-4b46-91d9-c527e7e721a9 · outbound

This paper cites Deepcoder: A fully open-source 14b coder at o3-mini level.

Exploring the Potential of Offline RL for Reasoning in LLMs: A Preliminary Study Deepcoder: A fully open-source 14b coder at o3-mini level

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:38.793190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:38.793190Z digest=sha256:5a4072da8641f6688414953bfb0c6f27eeaa7f80fc86700983f74289d47e9cb6

Observation 54d018df-6ac2-4545-910f-3acbebc95abd · outbound

This paper cites Chemistry-qa.

Exploring the Potential of Offline RL for Reasoning in LLMs: A Preliminary Study Chemistry-qa

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:38.796500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:38.796500Z digest=sha256:36138c06d6e947610e152cdbfb2a59dea93ab8e61f04fbe5e85f8ed628d83fd7

Observation cd907a48-355e-4704-a676-b64958b1a11b · outbound

This paper cites s1: Simple test-time scaling.

Exploring the Potential of Offline RL for Reasoning in LLMs: A Preliminary Study s1: Simple test-time scaling

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:38.799528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:38.799528Z digest=sha256:e2434c01a5462ffdf57554d0aebbd2f9e75ea69f2a304d738d6ab835f76c06e7

Observation 9cc0aef8-6577-47a1-b6ba-dcb9d7ec331b · outbound

This paper cites Llama-nemotron-post-training-dataset.

Exploring the Potential of Offline RL for Reasoning in LLMs: A Preliminary Study Llama-nemotron-post-training-dataset

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T01:04:39.361345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T01:04:38.803057Z digest=sha256:0267c1538357d044bc60066c1f567e9ae342f2f503e9f8230b4310417b0697d1

Observation bb9b5037-b839-4a3c-af70-4eeb96a5a5f1 · outbound

This paper cites Llama-3_1-nemotron-ultra-253b-v1.

Exploring the Potential of Offline RL for Reasoning in LLMs: A Preliminary Study Llama-3_1-nemotron-ultra-253b-v1

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:38.806310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:38.806310Z digest=sha256:be88c8bc59892ebe3d94254033eae2256b1df3a1d44ca90d4cf59f0fb821ebcb

Observation 0b60eb25-de3b-4dd7-86bc-e3098757d1da · outbound

This paper cites Infinity instruct.

Exploring the Potential of Offline RL for Reasoning in LLMs: A Preliminary Study Infinity instruct

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:38.809759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:38.809759Z digest=sha256:24a59897134b3cbec240fd2b42449e4a2bd0c57dea0c44b410a6279d0e1ecc40

Observation 65cd8540-6e1e-410f-becc-2038e4d44abf · outbound

This paper cites Verifiable coding problems (python).

Exploring the Potential of Offline RL for Reasoning in LLMs: A Preliminary Study Verifiable coding problems (python)

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:38.813326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:38.813326Z digest=sha256:7f3a1f338236fd7c1d8c6b5372fc22b3e45ae8c48d63c2e2d113fafa9b1d5453

Observation 5be56665-1c51-4312-bc76-20fb0449a1ad · outbound

This paper cites Learning to reason with llms, 2024.

Exploring the Potential of Offline RL for Reasoning in LLMs: A Preliminary Study Learning to reason with llms, 2024

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:38.816799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:38.816799Z digest=sha256:cc4843e55498018f141152dc28af7ccd0e380fe0c885078d5336bc8abefc95b4

Observation 2e5fa7e0-01f8-4933-8485-2132de1b6538 · outbound

This paper cites Training language models to follow instructions with human feedback.

Exploring the Potential of Offline RL for Reasoning in LLMs: A Preliminary Study Training language models to follow instructions with human feedback

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:38.820093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:38.820093Z digest=sha256:f582408dd5be2e466c571effa0df9902e62105787cefd4074bf60c7c70b8ad3d

Observation 4b7aec5e-ab2b-41de-9601-c70744877383 · outbound

This paper cites Codeforces cots.

Exploring the Potential of Offline RL for Reasoning in LLMs: A Preliminary Study Codeforces cots

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:38.823918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:38.823918Z digest=sha256:530ac7ad81ff54e9b43e2cbc9bc6e48753e90b5175233e505d01d779502f37ed

Observation c50da048-1071-44d1-9016-a30d6181b3a2 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Exploring the Potential of Offline RL for Reasoning in LLMs: A Preliminary Study Direct preference optimization: Your language model is secretly a reward model

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:38.827452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:38.827452Z digest=sha256:f2bf0815683fc920991b2a5683b5965e219721c7f24ca0263994cba432aec2b3

Observation 753f0b64-6e9e-460c-bbda-031399d893e8 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

Exploring the Potential of Offline RL for Reasoning in LLMs: A Preliminary Study GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:38.830986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:38.830986Z digest=sha256:9fe9d66d01ae1980d4db37cff1af045b65537748c14bd336758a17d70a6a9b07

Observation 92d7cc52-07a7-4321-8c8e-0a0355e26dd9 · outbound

This paper cites an unresolved cited work.

Exploring the Potential of Offline RL for Reasoning in LLMs: A Preliminary Study Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:38.834346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:38.834346Z digest=sha256:7e00c065dfb4d066a032a1e0a1b8d56bc2559eac8192bd0b3dad6c2f94e13402

Observation bd0549cc-770a-4436-92f5-61b169cc9406 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Exploring the Potential of Offline RL for Reasoning in LLMs: A Preliminary Study Proximal Policy Optimization Algorithms

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:38.837931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:38.837931Z digest=sha256:29ec2ee6298bdc2ebd9c300a8ea61f889ddd8909af625105fd9e4428e781c2b4

Observation f086508c-8213-4813-b586-6bda26b8ad29 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Exploring the Potential of Offline RL for Reasoning in LLMs: A Preliminary Study DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:38.841121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:38.841121Z digest=sha256:6da8a2fbd6916ab2851bb9fb702b77e133c821521e024ae2a365f2cef2eaeb7e

Observation 3debcc52-d762-4f34-b563-ec1ec39d6de5 · outbound

This paper cites Qwen2.5: A party of foundation models, September 2024.

Exploring the Potential of Offline RL for Reasoning in LLMs: A Preliminary Study Qwen2.5: A party of foundation models, September 2024

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:38.844397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:38.844397Z digest=sha256:c0a971a7609c39309ec09ac94c2eed7c7d6e72775aa2ccbb718d3ff696844c55

Observation 0348ce26-7209-4680-9726-41fb111acac4 · outbound

This paper cites Wizardlm evol-instruct 70k dataset.

Exploring the Potential of Offline RL for Reasoning in LLMs: A Preliminary Study Wizardlm evol-instruct 70k dataset

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:38.847460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:38.847460Z digest=sha256:15e0a594aa8346a14cc1d3827f148c7b95f5696e968f960f9e9ab41966c88dfb

Observation 47c34423-45bb-4bc5-a98d-66f46ff80514 · outbound

This paper cites Openhermes 2.5: An open dataset of synthetic data for generalist llm assistants, 2023.

Exploring the Potential of Offline RL for Reasoning in LLMs: A Preliminary Study Openhermes 2.5: An open dataset of synthetic data for generalist llm assistants, 2023

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:38.850481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:38.850481Z digest=sha256:6f4703284f917f72a17ef9dc67bf9f779d0ae72fef44cd08b62a1ea8f12e278a

Observation eedb6bee-05fb-4a92-a0c7-7f105e290fa1 · outbound

This paper cites DeepDistill: Enhancing LLM Reasoning Capabilities via Large-Scale Difficulty-Graded Data Training.

Exploring the Potential of Offline RL for Reasoning in LLMs: A Preliminary Study DeepDistill: Enhancing LLM Reasoning Capabilities via Large-Scale Difficulty-Graded Data Training

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:38.853466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:38.853466Z digest=sha256:9efd66edaeb535fbd96688d8ab3cf2c4ed04a82443d9fae131e4123351e2251f

Observation ba424970-494b-490e-8b2f-b9ed22ca704a · outbound

This paper cites Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks.

Exploring the Potential of Offline RL for Reasoning in LLMs: A Preliminary Study Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:38.856568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:38.856568Z digest=sha256:3dcc0d6fdf911500a5f68772004a161e28e5de04af18a3cda4b5ce38f523f2da

Observation 91904807-967f-4db1-86cd-37e533561d43 · outbound

This paper cites Smith, Hannaneh Hajishirzi, and Daniel Khashabi.

Exploring the Potential of Offline RL for Reasoning in LLMs: A Preliminary Study Smith, Hannaneh Hajishirzi, and Daniel Khashabi

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:38.859743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:38.859743Z digest=sha256:d190a432b5587ea8b71b7ab9f32aeb747f00efb1a25c9a10ec2ded9e98b617c5

Observation ec7a51fd-cdcc-4664-bd13-e3da0b781e75 · outbound

This paper cites Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M.

Exploring the Potential of Offline RL for Reasoning in LLMs: A Preliminary Study Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:38.862534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:38.862534Z digest=sha256:a9d48642804c14ad73e95af250b1433283aefc189f159b266526267a4631be1c

Observation d1257ba9-7fda-4218-8b8a-3b8827da053e · outbound

This paper cites KodCode: A Diverse, Challenging, and Verifiable Synthetic Dataset for Coding.

Exploring the Potential of Offline RL for Reasoning in LLMs: A Preliminary Study KodCode: A Diverse, Challenging, and Verifiable Synthetic Dataset for Coding

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:38.865499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:38.865499Z digest=sha256:6987c34ccd64812b3a694a51c236648d449d215d20ea8ab927a1b04b32f8b8e2

Observation cc608c9c-c12e-4716-ad8f-1e5d2a0f5d7b · outbound

This paper cites Qwen2 Technical Report.

Exploring the Potential of Offline RL for Reasoning in LLMs: A Preliminary Study Qwen2 Technical Report

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:38.868464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:38.868464Z digest=sha256:59cd81e130df10e4b4398b9b5935faa9c6dbec4d7e385a3621c516166f23ffb1

Observation 4daa4ef2-7916-475e-a215-45761ca69b20 · outbound

This paper cites MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models.

Exploring the Potential of Offline RL for Reasoning in LLMs: A Preliminary Study MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:38.872002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:38.872002Z digest=sha256:4c776fbd4c1aff733718d7f0b2fae34e36e17b42c06bf64394c977ea63a1cf8e

Observation c0ff78bd-e467-4fe8-89fa-d8e224d68732 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Exploring the Potential of Offline RL for Reasoning in LLMs: A Preliminary Study DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:38.875723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:38.875723Z digest=sha256:65f1b3d71737d847843ce9d7616ac2d9957c79be28fc82bd9aed7d925143acea

Observation 965369bd-dfbd-4177-bbed-3666517ad2ce · outbound

This paper cites Free Process Rewards without Process Labels.

Exploring the Potential of Offline RL for Reasoning in LLMs: A Preliminary Study Free Process Rewards without Process Labels

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:38.879469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:38.879469Z digest=sha256:16b8f5c063b74506e5cd797c6d234608b1b74d19ffe45fffe472a3e5610c7c60

Observation 88222009-4be8-4abc-9e76-f8086dd15721 · outbound

This paper cites Naturalreasoning: Rea- soning in the wild with 2.8m challenging questions, 2025.

Exploring the Potential of Offline RL for Reasoning in LLMs: A Preliminary Study Naturalreasoning: Rea- soning in the wild with 2.8m challenging questions, 2025

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:38.883074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:38.883074Z digest=sha256:c53cf965c275aa5e5cc4c8b6832768ef9bb4048f6354675ccc1712d532a91837

Observation 58731850-bb23-41c3-8764-5e480bd74fcd · outbound

This paper cites VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks.

Exploring the Potential of Offline RL for Reasoning in LLMs: A Preliminary Study VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:38.886277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:38.886277Z digest=sha256:23a2a1cb40197ebf9db3be7d15b0a2c83df398ba5fe1b3a9e64e818017a75d35

Observation 0dab8b0b-d2ba-43d3-8881-a4ecc0b15ec3 · outbound

This paper cites CodeRL: Mastering Code Generation through Pretrained Models and Deep Reinforcement Learning.

Exploring the Potential of Offline RL for Reasoning in LLMs: A Preliminary Study CodeRL: Mastering Code Generation through Pretrained Models and Deep Reinforcement Learning

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:38.890403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:38.890403Z digest=sha256:01e21687c9169469ea7f0cfe4b18284efd6031964f67849b4268c063266c3694

Observation fe5866d2-194b-4ebb-9a3c-9d3b314333ff · outbound

This paper cites InfinityMATH: A Scalable Instruction Tuning Dataset in Programmatic Mathematical Reasoning.

Exploring the Potential of Offline RL for Reasoning in LLMs: A Preliminary Study InfinityMATH: A Scalable Instruction Tuning Dataset in Programmatic Mathematical Reasoning

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:38.893490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:38.893490Z digest=sha256:6ababd4f536858c3ff54622c07cdb4d8745de492516c66ebe824460766f1165e

Observation 7e271105-5461-4e7d-a69e-1c64742f4812 · outbound

This paper cites Beyond IID: Optimizing Instruction Learning from the Perspective of Instruction Interaction and Dependency.

Exploring the Potential of Offline RL for Reasoning in LLMs: A Preliminary Study Beyond IID: Optimizing Instruction Learning from the Perspective of Instruction Interaction and Dependency

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:38.896787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:38.896787Z digest=sha256:9143e90debfdfdd5f97b512afa0a1b20b0d81616bb76f3097480ededa0e700c2

Observation b05ddc4c-e6e0-44d5-8d31-9137448fea7e · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

Exploring the Potential of Offline RL for Reasoning in LLMs: A Preliminary Study Instruction-Following Evaluation for Large Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:38.900467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:38.900467Z digest=sha256:76ad36ecb3415ce6960f583287ade397fce89dd0c5c0c6e4514599c4de85817b

Observation 5dbfd169-9ebf-4bf4-937a-9d09a617fa74 · outbound

This paper cites Hey Siri.

Exploring the Potential of Offline RL for Reasoning in LLMs: A Preliminary Study Hey Siri

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T01:04:39.275023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T01:04:38.905912Z digest=sha256:94f139e11d827abc4ccf3d637a63dab8262dedb71835a38a8d8cc3a7b70af024

Observation 07956d67-68b3-43fc-9354-a4ca7e89ccda · outbound

This paper cites an unresolved cited work.

Exploring the Potential of Offline RL for Reasoning in LLMs: A Preliminary Study Unresolved cited work

Reference 2023

Resolution
parse uncertain
no resolver link, observed 2026-08-16T01:04:38.693266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:38.693266Z digest=sha256:d6b07408cec8d73cd4602166b03bfb4fed24f278e7a8b68d314d96d4925f193e

Pith citing papers

Observation 10c0d3fa-2308-45cb-9039-12172632c9c6 · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning Exploring the Potential of Offline RL for Reasoning in LLMs: A Preliminary Study

Reference 204

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:09:40.575141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:ac892979ae9cf6ea6152fa7ca296bcf7fb521536b0ec3b9d293a420477ce5e3c