Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T15:17:13.821920Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 0 inbound Pith citation observations for arXiv:2509.03537.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T15:17:13.821920Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
18 of 18 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 5a7a71da-eeb0-4562-b917-d3a509c04373 · outbound
AR$^2$: Adversarial Reinforcement Learning for Abstract Reasoning in Large Language Models Evaluating Large Language Models Trained on Code
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8e3db9b-6914-4760-826b-fecc9e6486ec · outbound
AR$^2$: Adversarial Reinforcement Learning for Abstract Reasoning in Large Language Models SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b73d0ffa-457b-4b1d-85c5-b595229fe28f · outbound
AR$^2$: Adversarial Reinforcement Learning for Abstract Reasoning in Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbe70dc4-8e78-4d4e-88a8-b072a15b3a7e · outbound
AR$^2$: Adversarial Reinforcement Learning for Abstract Reasoning in Large Language Models StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8753fe9-fef7-4f22-9d23-1e4dc5afc420 · outbound
AR$^2$: Adversarial Reinforcement Learning for Abstract Reasoning in Large Language Models Competitive Programming with Large Reasoning Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 397166e6-a4a4-44d1-8f37-0ffa783c02d6 · outbound
AR$^2$: Adversarial Reinforcement Learning for Abstract Reasoning in Large Language Models Unresolved cited work
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 16e4a717-7845-47c8-b931-7bb989f39dba · outbound
AR$^2$: Adversarial Reinforcement Learning for Abstract Reasoning in Large Language Models Language Models Can Teach Themselves to Program Better
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f12c963a-9083-4579-8128-fc2233fb8adc · outbound
AR$^2$: Adversarial Reinforcement Learning for Abstract Reasoning in Large Language Models Qwen2.5-Coder Technical Report
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e56cf67-6e4d-4b7c-92ac-30340b4a9d38 · outbound
AR$^2$: Adversarial Reinforcement Learning for Abstract Reasoning in Large Language Models LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 944c58a2-a88b-4d2a-9dbe-d3b2c24a94d1 · outbound
AR$^2$: Adversarial Reinforcement Learning for Abstract Reasoning in Large Language Models Understanding R1-Zero-Like Training: A Critical Perspective
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf84de14-32ec-4fa4-bdbf-4020e10daa64 · outbound
AR$^2$: Adversarial Reinforcement Learning for Abstract Reasoning in Large Language Models CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5759bce-d896-4232-b2df-88f3d6ad35f4 · outbound
AR$^2$: Adversarial Reinforcement Learning for Abstract Reasoning in Large Language Models Qwen2.5 Technical Report
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4cf570fe-e19a-413a-98c7-b137748981b8 · outbound
AR$^2$: Adversarial Reinforcement Learning for Abstract Reasoning in Large Language Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 826bf715-dc69-4193-92a9-815a966ded12 · outbound
AR$^2$: Adversarial Reinforcement Learning for Abstract Reasoning in Large Language Models Execution-based Code Generation using Deep Reinforcement Learning
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44eb70b4-099c-4dfc-a3ce-bf381b2e2cd2 · outbound
AR$^2$: Adversarial Reinforcement Learning for Abstract Reasoning in Large Language Models Sutton and Andrew G
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4337c025-72a6-4375-9146-b13ac84b2cc5 · outbound
AR$^2$: Adversarial Reinforcement Learning for Abstract Reasoning in Large Language Models Reinforcement Learning Enhanced LLMs: A Survey
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e33c6417-4a03-41ba-be60-dfa7156370dc · outbound
AR$^2$: Adversarial Reinforcement Learning for Abstract Reasoning in Large Language Models the word is a subset of the puzzle’s characters)
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3ece2b84-8724-4fa4-a6b2-6345d1d91dd0 · outbound
AR$^2$: Adversarial Reinforcement Learning for Abstract Reasoning in Large Language Models includes the first character of the puzzle and contains all the characters from the puzzle
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
No inbound Pith citation observations are available.