Pith. sign in

Paper Citation Record · LEDGER

AR$^2$: Adversarial Reinforcement Learning for Abstract Reasoning in Large Language Models

As of 8 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 0 inbound Pith citation observations for arXiv:2509.03537.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.03537 v1

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T15:17:13.821920Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

18 of 18 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5a7a71da-eeb0-4562-b917-d3a509c04373 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

AR$^2$: Adversarial Reinforcement Learning for Abstract Reasoning in Large Language Models Evaluating Large Language Models Trained on Code

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T15:17:13.732880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:17:13.732880Z digest=sha256:5b9f14f8f6fe7f25776ec513cb45f0bc082b61e0d4d525ccfbc625b656adf1c8

Observation a8e3db9b-6914-4760-826b-fecc9e6486ec · outbound

This paper cites SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training.

AR$^2$: Adversarial Reinforcement Learning for Abstract Reasoning in Large Language Models SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T15:17:13.738630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:17:13.738630Z digest=sha256:6782d18d9b3fc871939c57ab23277db11434f55f939b01b464f4192025de8db1

Observation b73d0ffa-457b-4b1d-85c5-b595229fe28f · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

AR$^2$: Adversarial Reinforcement Learning for Abstract Reasoning in Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T15:17:13.743935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:17:13.743935Z digest=sha256:2885e277fa9e3c0dce683182dc95f29ecd06eb8905c3745ef32ff71d0c10784b

Observation bbe70dc4-8e78-4d4e-88a8-b072a15b3a7e · outbound

This paper cites StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback.

AR$^2$: Adversarial Reinforcement Learning for Abstract Reasoning in Large Language Models StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T15:17:13.749999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:17:13.749999Z digest=sha256:350387ee894206ec4644b55c095c95ee68d3c3eac5f276a88dc37059d1907e0b

Observation b8753fe9-fef7-4f22-9d23-1e4dc5afc420 · outbound

This paper cites Competitive Programming with Large Reasoning Models.

AR$^2$: Adversarial Reinforcement Learning for Abstract Reasoning in Large Language Models Competitive Programming with Large Reasoning Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T15:17:13.755616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:17:13.755616Z digest=sha256:80075a6bb3356e2f8f5d3a929d04eb7b518ae4b4c58b06cea2dd41cc64cba744

Observation 397166e6-a4a4-44d1-8f37-0ffa783c02d6 · outbound

This paper cites an unresolved cited work.

AR$^2$: Adversarial Reinforcement Learning for Abstract Reasoning in Large Language Models Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-05T15:17:14.122287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T15:17:13.761076Z digest=sha256:19f516fb154e02ea092d91d37af4c756fac0e9fbedb0625ff30f3b2dcf0595ec

Observation 16e4a717-7845-47c8-b931-7bb989f39dba · outbound

This paper cites Language Models Can Teach Themselves to Program Better.

AR$^2$: Adversarial Reinforcement Learning for Abstract Reasoning in Large Language Models Language Models Can Teach Themselves to Program Better

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T15:17:13.766596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:17:13.766596Z digest=sha256:652deff31f6adce96753e13e784629e041f838643b2baae58a95f6740536b92f

Observation f12c963a-9083-4579-8128-fc2233fb8adc · outbound

This paper cites Qwen2.5-Coder Technical Report.

AR$^2$: Adversarial Reinforcement Learning for Abstract Reasoning in Large Language Models Qwen2.5-Coder Technical Report

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T15:17:13.771367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:17:13.771367Z digest=sha256:8d39239775d4a1d1b897700d43d64c6d68b30085d390db614a4455ed53be165e

Observation 5e56cf67-6e4d-4b7c-92ac-30340b4a9d38 · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

AR$^2$: Adversarial Reinforcement Learning for Abstract Reasoning in Large Language Models LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T15:17:13.776157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:17:13.776157Z digest=sha256:90c089a7b2629088e458248f315062a52b0d6701942d2c0aa10f19010b2f1c2b

Observation 944c58a2-a88b-4d2a-9dbe-d3b2c24a94d1 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

AR$^2$: Adversarial Reinforcement Learning for Abstract Reasoning in Large Language Models Understanding R1-Zero-Like Training: A Critical Perspective

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T15:17:13.781115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:17:13.781115Z digest=sha256:8a2ae6cdd68e2e00e01a16bd73e3c325bb364e6009e4a844aa722038ce16fad4

Observation bf84de14-32ec-4fa4-bdbf-4020e10daa64 · outbound

This paper cites CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings.

AR$^2$: Adversarial Reinforcement Learning for Abstract Reasoning in Large Language Models CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T15:17:13.786271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:17:13.786271Z digest=sha256:e98cba196186268a0cd0fbadd7f65307987dc8f7a9e67fda0584b95bba4d5014

Observation c5759bce-d896-4232-b2df-88f3d6ad35f4 · outbound

This paper cites Qwen2.5 Technical Report.

AR$^2$: Adversarial Reinforcement Learning for Abstract Reasoning in Large Language Models Qwen2.5 Technical Report

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T15:17:13.791467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:17:13.791467Z digest=sha256:c05f7b891e5fb093778d92cccae788cd874a8e49a64bd81749eb50835f0ea785

Observation 4cf570fe-e19a-413a-98c7-b137748981b8 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

AR$^2$: Adversarial Reinforcement Learning for Abstract Reasoning in Large Language Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T15:17:13.797081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:17:13.797081Z digest=sha256:337ec9dcbc9ea8ab16309fd9bd5aa5362946bc264c544f60b529e5e4abcfe84f

Observation 826bf715-dc69-4193-92a9-815a966ded12 · outbound

This paper cites Execution-based Code Generation using Deep Reinforcement Learning.

AR$^2$: Adversarial Reinforcement Learning for Abstract Reasoning in Large Language Models Execution-based Code Generation using Deep Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T15:17:13.801965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:17:13.801965Z digest=sha256:47ccf43f3e8a0361a2572337e6f4126ecdf4747dc9d9c87b59d5d54b3841a6ee

Observation 44eb70b4-099c-4dfc-a3ce-bf381b2e2cd2 · outbound

This paper cites Sutton and Andrew G.

AR$^2$: Adversarial Reinforcement Learning for Abstract Reasoning in Large Language Models Sutton and Andrew G

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:17:14.107674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T15:17:13.807366Z digest=sha256:c2d30ff1e1ffb58a52b29314bb458c647cb0ae23eea47b9081c6c8eb257a30c8

Observation 4337c025-72a6-4375-9146-b13ac84b2cc5 · outbound

This paper cites Reinforcement Learning Enhanced LLMs: A Survey.

AR$^2$: Adversarial Reinforcement Learning for Abstract Reasoning in Large Language Models Reinforcement Learning Enhanced LLMs: A Survey

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T15:17:13.811953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:17:13.811953Z digest=sha256:424eda0bcc930cd40e54dc07dfc335c693f5371db6bb73764c42e20a32c53ae0

Observation e33c6417-4a03-41ba-be60-dfa7156370dc · outbound

This paper cites the word is a subset of the puzzle’s characters).

AR$^2$: Adversarial Reinforcement Learning for Abstract Reasoning in Large Language Models the word is a subset of the puzzle’s characters)

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:17:14.093333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T15:17:13.817245Z digest=sha256:2a1fc71112384cabd4cc724a530831123ecbe5b9e04d74e3e3784b677d09b4f5

Observation 3ece2b84-8724-4fa4-a6b2-6345d1d91dd0 · outbound

This paper cites includes the first character of the puzzle and contains all the characters from the puzzle.

AR$^2$: Adversarial Reinforcement Learning for Abstract Reasoning in Large Language Models includes the first character of the puzzle and contains all the characters from the puzzle

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:17:14.077781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T15:17:13.821920Z digest=sha256:a578fcc6141f42e231b93c8d6c437bb6c76f128d4e93c6bb932a1db83f4d6d3b

Pith citing papers

No inbound Pith citation observations are available.