Pith. sign in

Paper Citation Record · LEDGER

Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning

As of 9 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 2 inbound Pith citation observations for arXiv:2508.04848.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.04848 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T23:49:05.708180Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-11T01:49:15.136031Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T16:01:22.618150Z

Reference resolution

21 of 21 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4cd5b0c6-2b61-4a8b-802e-9cd4f7868183 · outbound

This paper cites Evaluating LLMs and Prompting Strategies for Automated Hardware Diagnosis from Textual User-Reports.

Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning Evaluating LLMs and Prompting Strategies for Automated Hardware Diagnosis from Textual User-Reports

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T23:49:07.511703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T23:49:02.285130Z digest=sha256:25f6459e6738c5961517fe9be154db54b3e2fc430ce46229d8baa8970d8b77f6

Observation 26e2ce09-01a4-4690-ac40-280bd2d77a25 · outbound

This paper cites an unresolved cited work.

Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-05T23:49:07.820130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T23:49:02.550966Z digest=sha256:1e92cf800f029bd94e6e064aece5736c2162d1f2c91563f06203e0a8efe2ec73

Observation 3a6443b6-1c35-4fb7-a39b-93fcaba8ab0b · outbound

This paper cites arXiv preprint arXiv:2503.12434.

Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning arXiv preprint arXiv:2503.12434

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T23:49:02.715209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:49:02.715209Z digest=sha256:24c73d318f30c85024558a0fdba8799617e28688aea4e718000e1147cdbb360a

Observation cc53431a-943a-4af7-afb6-ffb268577a7a · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T23:49:02.929694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:49:02.929694Z digest=sha256:5d05dd16df4eb895011d691c42c51bfdc784d4515523496af15b0be4f958ae4e

Observation eed7fb36-b81f-4e42-afde-45576bac2ccc · outbound

This paper cites VinePPO: Refining Credit Assignment in RL Training of LLMs.

Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T23:49:03.216906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:49:03.216906Z digest=sha256:baf1abf8d61ea3d95eef6f820cbc819fa02e382b70ad3018056009f836ec5bb8

Observation c089ee1b-f9ea-4421-8d1d-ce90170f4779 · outbound

This paper cites LLM Post-Training: A Deep Dive into Reasoning Large Language Models.

Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning LLM Post-Training: A Deep Dive into Reasoning Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T23:49:03.474747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:49:03.474747Z digest=sha256:ac3b2b372764ab52254385f331349d9e9a777655fdc0d8c2e7451e2adea9e38c

Observation 5069f2cf-b853-4e39-a927-395baf2b1c31 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T23:49:04.041085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:49:04.041085Z digest=sha256:a4c6defdc470fcecd63ee55c2fd363f27263b5d576a8cb3cc1510d1cb49b100b

Observation 2a099c1a-ee77-4545-bebb-118df17d7438 · outbound

This paper cites Using Causality for Enhanced Prediction of Web Traffic Time Series.

Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning Using Causality for Enhanced Prediction of Web Traffic Time Series

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T23:49:06.723730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T23:49:04.374494Z digest=sha256:72e6fdcc94a7820ae124e4c2b878a6f5c7060bab0aff9193f050685a531d87ce

Observation 9453a80a-4f89-45ac-adb3-85b33cfdf8d6 · outbound

This paper cites What is the Alignment Objective of GRPO?.

Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning What is the Alignment Objective of GRPO?

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T23:49:04.813342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:49:04.813342Z digest=sha256:ea86ef4d9e55623e6fe96ced250473b8f2c3c00faee35f67452c384c7be3f506

Observation da8eb2be-757c-408d-9572-55982b3900b1 · outbound

This paper cites Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning.

Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T23:49:05.025670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:49:05.025670Z digest=sha256:ef58bb43647b6496c9d6885ecb447ab5ed38e85474221ab54122042720cdec41

Observation 4dd15b6d-ed00-430b-b9bf-739dc07bb09e · outbound

This paper cites Training Large Language Models to Reason via EM Policy Gradient.

Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning Training Large Language Models to Reason via EM Policy Gradient

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T23:49:05.199743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:49:05.199743Z digest=sha256:8f9605e34b637ae6aa9ffaf6549790ec2a6db71e18efc704a6d830fe193e0f73

Observation b51dbaea-cabf-458e-bd33-21a885642e6d · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T23:49:05.395955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:49:05.395955Z digest=sha256:6f032691047b50af630ca5595ca09be01499533a9b50b4f7a1484ed0cac75c40

Observation 8197227e-d83e-46e8-b5bb-5485649366a8 · outbound

This paper cites arXiv preprint arXiv:2505.17508.

Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning arXiv preprint arXiv:2505.17508

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T23:49:05.541952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:49:05.541952Z digest=sha256:35d57618861171ce3a11a078de81d7972181a9750ff2d37bcd8cf091b29aa29c

Observation 8a413b4d-c7e9-4a91-b607-8b6cf0a66171 · outbound

This paper cites Monte Carlo Tree Search for Comprehensive Exploration in LLM-Based Automatic Heuristic Design.

Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning Monte Carlo Tree Search for Comprehensive Exploration in LLM-Based Automatic Heuristic Design

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T23:49:05.708180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:49:05.708180Z digest=sha256:b5c17256d059907ec7c620fb796d446be6e7abb329de5ae0a16465615a694ffc

Observation c5a07c9f-a069-4938-9dbe-c3d61df449e2 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning Proximal Policy Optimization Algorithms

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-05T23:49:03.820612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:49:03.820612Z digest=sha256:638aaf558e1ec753686fdb3cc4e6a555bad5e4b84667532234577f3a71538f74

Observation e9fbee50-7381-4d9e-9880-7e86a8620486 · outbound

This paper cites A Generic Method for Fine-grained Category Discovery in Natural Language Texts.

Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning A Generic Method for Fine-grained Category Discovery in Natural Language Texts

Reference 2019

Resolution
verified exact
local_arxiv, observed 2026-08-05T23:49:07.013350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T23:49:04.213823Z digest=sha256:ea7eaab9095c034408c98a1267c41a7e80e4446972aa95c262c96bd9060c774b

Observation 10b799cd-fa19-409b-8e5e-13ff2688b4a9 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning Measuring Massive Multitask Language Understanding

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-05T23:49:03.062950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:49:03.062950Z digest=sha256:ed0e8e4cf93051beee79ec8720fb5011aded78723d4b3e176cce17107296753d

Observation 9fdfde4a-26c7-4548-9b83-561e8375ee4e · outbound

This paper cites Paint4Poem: A Dataset for Artistic Visualization of Classical Chinese Poems.

Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning Paint4Poem: A Dataset for Artistic Visualization of Classical Chinese Poems

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-05T23:49:03.658162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:49:03.658162Z digest=sha256:50a2e9320c2b5fcd542a27c78bb9f77b32c1a17b9063c3477b9ada108aaebce7

Observation f8585f12-a48c-4d6d-b420-834437982951 · outbound

This paper cites Anti-Overestimation Dialogue Policy Learning for Task-Completion Dialogue System.

Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning Anti-Overestimation Dialogue Policy Learning for Task-Completion Dialogue System

Reference 2022

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T23:49:06.307243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T23:49:04.536418Z digest=sha256:8e5ad31efb8f0f763d389bd1721b5c3216ce8809870910529fe9cc7cdedb4a55

Observation 22c06452-16ca-4c32-b0ab-d55ec42d5471 · outbound

This paper cites Meta-Models: An Architecture for Decoding LLM Behaviors Through Interpreted Embeddings and Natural Language.

Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning Meta-Models: An Architecture for Decoding LLM Behaviors Through Interpreted Embeddings and Natural Language

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T23:49:02.409134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:49:02.409134Z digest=sha256:e4c03fa2db0fe6b2dab2b1860be8542057dee424bf6d64b74ed4f4dc2f079bf5

Observation 6c9329a9-57f5-468f-aeec-0c66a4b91d01 · outbound

This paper cites Qwen2.5-VL Technical Report.

Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning Qwen2.5-VL Technical Report

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-05T23:49:02.173108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:49:02.173108Z digest=sha256:53501ecc11f6418a120cc9f95932fea432ded99d0e2dc353e4da7adb76583a69

Pith citing papers

Observation b426df3d-129f-4e23-a706-9cb26d330ce0 · inbound

Persistent Visual Memory: Sustaining Perception for Deep Generation in LVLMs cites this paper.

Persistent Visual Memory: Sustaining Perception for Deep Generation in LVLMs Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:01:22.621931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T18:53:06.494640Z digest=sha256:213c50fedba784ca573cf5bb18693f6799e06e6724cd0acd9c29036bfb72d91c

Observation b60805a6-808a-4981-b738-8bea31069308 · inbound

Persistent Visual Memory: Sustaining Perception for Deep Generation in LVLMs cites this paper.

Persistent Visual Memory: Sustaining Perception for Deep Generation in LVLMs Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:50:51.662264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T01:49:15.136031Z digest=sha256:1d9ed73269043601b149b37e58aafff23f7391d7dcadbdeb115e14f5d50c5a82