Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T23:49:05.708180Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 2 inbound Pith citation observations for arXiv:2508.04848.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T23:49:05.708180Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-11T01:49:15.136031Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-11T16:01:22.618150Z
21 of 21 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 4cd5b0c6-2b61-4a8b-802e-9cd4f7868183 · outbound
Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning Evaluating LLMs and Prompting Strategies for Automated Hardware Diagnosis from Textual User-Reports
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 26e2ce09-01a4-4690-ac40-280bd2d77a25 · outbound
Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3a6443b6-1c35-4fb7-a39b-93fcaba8ab0b · outbound
Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning arXiv preprint arXiv:2503.12434
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc53431a-943a-4af7-afb6-ffb268577a7a · outbound
Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eed7fb36-b81f-4e42-afde-45576bac2ccc · outbound
Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c089ee1b-f9ea-4421-8d1d-ce90170f4779 · outbound
Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning LLM Post-Training: A Deep Dive into Reasoning Large Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5069f2cf-b853-4e39-a927-395baf2b1c31 · outbound
Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a099c1a-ee77-4545-bebb-118df17d7438 · outbound
Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning Using Causality for Enhanced Prediction of Web Traffic Time Series
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9453a80a-4f89-45ac-adb3-85b33cfdf8d6 · outbound
Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning What is the Alignment Objective of GRPO?
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da8eb2be-757c-408d-9572-55982b3900b1 · outbound
Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4dd15b6d-ed00-430b-b9bf-739dc07bb09e · outbound
Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning Training Large Language Models to Reason via EM Policy Gradient
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b51dbaea-cabf-458e-bd33-21a885642e6d · outbound
Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8197227e-d83e-46e8-b5bb-5485649366a8 · outbound
Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning arXiv preprint arXiv:2505.17508
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a413b4d-c7e9-4a91-b607-8b6cf0a66171 · outbound
Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning Monte Carlo Tree Search for Comprehensive Exploration in LLM-Based Automatic Heuristic Design
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5a07c9f-a069-4938-9dbe-c3d61df449e2 · outbound
Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning Proximal Policy Optimization Algorithms
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9fbee50-7381-4d9e-9880-7e86a8620486 · outbound
Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning A Generic Method for Fine-grained Category Discovery in Natural Language Texts
Reference 2019
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 10b799cd-fa19-409b-8e5e-13ff2688b4a9 · outbound
Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning Measuring Massive Multitask Language Understanding
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9fdfde4a-26c7-4548-9b83-561e8375ee4e · outbound
Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning Paint4Poem: A Dataset for Artistic Visualization of Classical Chinese Poems
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8585f12-a48c-4d6d-b420-834437982951 · outbound
Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning Anti-Overestimation Dialogue Policy Learning for Task-Completion Dialogue System
Reference 2022
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 22c06452-16ca-4c32-b0ab-d55ec42d5471 · outbound
Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning Meta-Models: An Architecture for Decoding LLM Behaviors Through Interpreted Embeddings and Natural Language
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c9329a9-57f5-468f-aeec-0c66a4b91d01 · outbound
Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning Qwen2.5-VL Technical Report
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b426df3d-129f-4e23-a706-9cb26d330ce0 · inbound
Persistent Visual Memory: Sustaining Perception for Deep Generation in LVLMs Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b60805a6-808a-4981-b738-8bea31069308 · inbound
Persistent Visual Memory: Sustaining Perception for Deep Generation in LVLMs Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.