Pith. sign in

Paper Citation Record · LEDGER

Domain-Adaptable Reinforcement Learning for Code Generation with Dense Rewards

As of 4 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 0 inbound Pith citation observations for arXiv:2605.21180.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.21180 v1

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-21T05:42:15.582971Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

28 of 28 outbound references displayed

  • verified exact10
  • verified fuzzy10
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch6

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b9b86ec6-f301-4765-84e6-46a2427b5ad1 · outbound

This paper cites Ruff : An extremely fast Python linter and code formatter, written in Rust.

Domain-Adaptable Reinforcement Learning for Code Generation with Dense Rewards Ruff : An extremely fast Python linter and code formatter, written in Rust

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T05:43:59.552239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-21T05:42:15.582971Z digest=sha256:8b3212b598c3092a65e12b5b48656f936fdcb6a72f63de3fc9ae0d4219a23f7f

Observation 8133df44-3410-4c24-8cd5-20a42d67030f · outbound

This paper cites Program Synthesis with Large Language Models.

Domain-Adaptable Reinforcement Learning for Code Generation with Dense Rewards Program Synthesis with Large Language Models

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T05:43:58.961120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-21T05:42:15.582971Z digest=sha256:0fd6941f68f6627abb111f36f3708bb45243ff489c319903c862fe519bacdd87

Observation c85392c3-e681-4761-aa0e-26133ce612f6 · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

Domain-Adaptable Reinforcement Learning for Code Generation with Dense Rewards RT-1: Robotics Transformer for Real-World Control at Scale

Reference 3

Resolution
verified exact
doi, observed 2026-05-21T05:43:58.604489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-21T05:42:15.582971Z digest=sha256:4abedae9e1d2eabecdee1f654a4996347710ab12f007c0c652d8534662a0f8ac

Observation b0453ae8-0315-4425-b6dd-b4d646c1c963 · outbound

This paper cites RoboScript: Code Generation for Free-Form Manipulation Tasks across Real and Simulation.

Domain-Adaptable Reinforcement Learning for Code Generation with Dense Rewards RoboScript: Code Generation for Free-Form Manipulation Tasks across Real and Simulation

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:43:58.592880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-21T05:42:15.582971Z digest=sha256:da70f604ea032492bf4fdc9835b5bc72826c9dbe35ec4ad7da11d815350153b3

Observation 90d25cf0-a675-45f8-b4b4-1aa453c9726b · outbound

This paper cites An llm-powered natural-to-robotic language translation framework with correctness guarantees.

Domain-Adaptable Reinforcement Learning for Code Generation with Dense Rewards An llm-powered natural-to-robotic language translation framework with correctness guarantees

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T05:43:58.558988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-21T05:42:15.582971Z digest=sha256:478d43c487644ecbb331af3fe77d4e7a747bf50d2e97d345b1725f3674a78bc0

Observation 48fead79-5725-40a4-b4b0-8393bc014b0f · outbound

This paper cites an unresolved cited work.

Domain-Adaptable Reinforcement Learning for Code Generation with Dense Rewards Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-05-21T05:43:59.554039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-21T05:42:15.582971Z digest=sha256:ba753dce8dd4b4518a0160dbd96a825760cfc83aad4116fc7e621e34e2dde7bb

Observation c911c28f-26cb-495c-9a47-87da6f5bb977 · outbound

This paper cites DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence.

Domain-Adaptable Reinforcement Learning for Code Generation with Dense Rewards DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T05:43:58.577731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-21T05:42:15.582971Z digest=sha256:91a3270ab0686763cab2e1398e637e0593d633199d2907bceef48cbf63348402

Observation 4180046e-0274-4442-a0a3-537f45b7e63c · outbound

This paper cites Available: https://doi.org/10.1145/3695988.

Domain-Adaptable Reinforcement Learning for Code Generation with Dense Rewards Available: https://doi.org/10.1145/3695988

Reference 8

Resolution
metadata mismatch
doi, observed 2026-05-21T05:43:58.595877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-21T05:42:15.582971Z digest=sha256:5311aca6b9a414f99e459aa35645232e2b35c92769e988a9195c573f94c75c64

Observation 2561cf79-7fda-468f-97eb-2b81e3935ecc · outbound

This paper cites J., Shen, Y., Wallis, P., Allen - Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W.

Domain-Adaptable Reinforcement Learning for Code Generation with Dense Rewards J., Shen, Y., Wallis, P., Allen - Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T05:43:59.555728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-21T05:42:15.582971Z digest=sha256:eb6c43405134695edb649b7129a32a401a2b12f3215e6a98f89fc3759949cc3f

Observation 90cfef53-dc5b-4a35-b95b-523bd3e78b56 · outbound

This paper cites robo-instruct.

Domain-Adaptable Reinforcement Learning for Code Generation with Dense Rewards robo-instruct

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T05:43:59.572126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-21T05:42:15.582971Z digest=sha256:88a6107fad88383808f1d7ad540718f1191aa8e886b088f5a04e42736edcfff1

Observation 4a10d21f-e333-4984-b3b5-0852e90d21fa · outbound

This paper cites J., Guha, A., and Biswas, J.

Domain-Adaptable Reinforcement Learning for Code Generation with Dense Rewards J., Guha, A., and Biswas, J

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:43:58.589351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-21T05:42:15.582971Z digest=sha256:2026ac2b311f06d5495b5e87b0c57935fffc0b4a29cf4a35a8b5f133d8807680

Observation c4ecf552-f470-48b2-ae30-9f877ef0fd36 · outbound

This paper cites Deploying and evaluating llms to program service mobile robots.

Domain-Adaptable Reinforcement Learning for Code Generation with Dense Rewards Deploying and evaluating llms to program service mobile robots

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:43:58.584888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-21T05:42:15.582971Z digest=sha256:0ea0e98e99d7ee91d1e25efe692af7938691c28c4f4ef944c05f81572ce756a6

Observation 26418a78-aa14-4c70-8073-9deb430cc186 · outbound

This paper cites Inner monologue: Embodied reasoning through planning with language models.

Domain-Adaptable Reinforcement Learning for Code Generation with Dense Rewards Inner monologue: Embodied reasoning through planning with language models

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T05:43:59.569815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-21T05:42:15.582971Z digest=sha256:c73020d6e637485e65216ae8b91774bf4b5f1e99b609763e110b0663b17c694a

Observation 0c9f0919-8f01-4f5c-afc8-f013dc199857 · outbound

This paper cites Reinforcement Learning via Self-Distillation.

Domain-Adaptable Reinforcement Learning for Code Generation with Dense Rewards Reinforcement Learning via Self-Distillation

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-21T05:43:58.581066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-21T05:42:15.582971Z digest=sha256:ba78b364879a372e305a54e4e090e31a30d7a030ac3510dbba73d02d30c40608

Observation b121e478-b222-412f-8c46-e9a34cdded99 · outbound

This paper cites TRL - transformers reinforcement learning.

Domain-Adaptable Reinforcement Learning for Code Generation with Dense Rewards TRL - transformers reinforcement learning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T05:43:59.568175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-21T05:42:15.582971Z digest=sha256:dde835127f785b1e97f98009db170343077c804a097a5a5d75a39ee259944e09

Observation 687b79d3-96df-4bce-b32c-34747e3168a7 · outbound

This paper cites Qwen2.5-Coder Technical Report.

Domain-Adaptable Reinforcement Learning for Code Generation with Dense Rewards Qwen2.5-Coder Technical Report

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-21T05:43:58.554962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-21T05:42:15.582971Z digest=sha256:d15b8dc82a486af66b8708e0ab1ecdd8573b71db8e7900bbe09bfe200ca3a4f4

Observation 49bde5ed-0633-4773-8440-183bda1de084 · outbound

This paper cites Cotran: An llm-based code translator using reinforcement learning with feedback from compiler and symbolic execution.

Domain-Adaptable Reinforcement Learning for Code Generation with Dense Rewards Cotran: An llm-based code translator using reinforcement learning with feedback from compiler and symbolic execution

Reference 17

Resolution
verified exact
doi, observed 2026-05-21T05:43:58.573280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-21T05:42:15.582971Z digest=sha256:7006e8fbd447a79d44e506f2185271860c39fc9c62938872c4ca3698bee229ee

Observation baf7c9a0-e0aa-49c7-847b-6af586c7f422 · outbound

This paper cites D., Savarese, S., and Hoi, S.

Domain-Adaptable Reinforcement Learning for Code Generation with Dense Rewards D., Savarese, S., and Hoi, S

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:43:58.966059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-21T05:42:15.582971Z digest=sha256:405f999416665bca5f8944b040fabd36708ecb17f1e25544225115b0e93afc78

Observation 88826f8c-b54f-44d7-b95d-5a97c2767e66 · outbound

This paper cites ImmFusion: Robust mmWave-RGB Fusion for 3D Human Body Reconstruction in All Weather Conditions.

Domain-Adaptable Reinforcement Learning for Code Generation with Dense Rewards ImmFusion: Robust mmWave-RGB Fusion for 3D Human Body Reconstruction in All Weather Conditions

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T05:43:58.564641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-21T05:42:15.582971Z digest=sha256:9022ace1fd722380e512f2f2d7b88ccad0e9f342a36e348591496b2d99480579

Observation 2789a8c6-be80-4ab5-96e8-0c0dd213f445 · outbound

This paper cites S., Wang, Y., and Zhang, L.

Domain-Adaptable Reinforcement Learning for Code Generation with Dense Rewards S., Wang, Y., and Zhang, L

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T05:43:59.564897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-21T05:42:15.582971Z digest=sha256:befee8cf272c697e26fc01a92a317c7fbbe52649f68c412ef2107b67b98fd76f

Observation ed8cc369-24dd-4e18-92f0-078452450abc · outbound

This paper cites OpenCodeInstruct : A large-scale instruction tuning dataset for code LLMs.

Domain-Adaptable Reinforcement Learning for Code Generation with Dense Rewards OpenCodeInstruct : A large-scale instruction tuning dataset for code LLMs

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T05:43:59.566605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-21T05:42:15.582971Z digest=sha256:7d9260853340c987019fa21e8ea7839d3028e6b57006594e286fdaa3da6562b1

Observation 462069f8-15a1-4d51-8ead-e450391a2677 · outbound

This paper cites Qwen2.5-Coder-1.5B-Instruct.

Domain-Adaptable Reinforcement Learning for Code Generation with Dense Rewards Qwen2.5-Coder-1.5B-Instruct

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T05:43:59.561511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-21T05:42:15.582971Z digest=sha256:cf5ab2e951c6c24883186a73aa68f1bcfbf9435b1b1c6eb4a00ec30231b29c04

Observation 2dd069ff-415c-4118-a287-1324911b9522 · outbound

This paper cites D., and S \" u nderhauf, N.

Domain-Adaptable Reinforcement Learning for Code Generation with Dense Rewards D., and S \" u nderhauf, N

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T05:43:59.557910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-21T05:42:15.582971Z digest=sha256:8e133b90163c9de67de16d6e053a31eeb1250248d8180b78132b64ab719fa4a5

Observation 99a164ea-f509-4b89-a16e-14ca7db09dea · outbound

This paper cites Proximal Policy Optimization Algorithms.

Domain-Adaptable Reinforcement Learning for Code Generation with Dense Rewards Proximal Policy Optimization Algorithms

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-21T05:43:58.963505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-21T05:42:15.582971Z digest=sha256:d700d60991723b2df422654b0f7c3a4ebbc37ba7dc8fdfe55dd5ecd7735f2f29

Observation 9695d68d-569a-4743-a7af-3446727cce58 · outbound

This paper cites an unresolved cited work.

Domain-Adaptable Reinforcement Learning for Code Generation with Dense Rewards Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-05-21T05:43:59.559812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-21T05:42:15.582971Z digest=sha256:dd5d4f36ec487f169b82ca6897e18bf0742765720c95f78d47216d16d53c0ea5

Observation ae303f31-147c-4409-ad13-0a701048101a · outbound

This paper cites LLMs for Coding and Robotics Education.

Domain-Adaptable Reinforcement Learning for Code Generation with Dense Rewards LLMs for Coding and Robotics Education

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:43:58.570058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-21T05:42:15.582971Z digest=sha256:ba9ca849145ed5c786b88e2157de47caa603d82453c3af849943cc2e9f608053

Observation 1202e8be-7bde-4ebf-bc11-1a811ea9bdd5 · outbound

This paper cites In: 2023 IEEE/CVF International Conference on Computer Vision (ICCV).

Domain-Adaptable Reinforcement Learning for Code Generation with Dense Rewards In: 2023 IEEE/CVF International Conference on Computer Vision (ICCV)

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T05:43:58.600246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-21T05:42:15.582971Z digest=sha256:4731c81f6df2a14aa318544deff2052b68f5610849fadfbbd5dc822e5226a558

Observation 1e044c54-3d31-4686-adee-f685e173ed99 · outbound

This paper cites Syncode: LLM generation with grammar augmentation.

Domain-Adaptable Reinforcement Learning for Code Generation with Dense Rewards Syncode: LLM generation with grammar augmentation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T05:43:59.563174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-21T05:42:15.582971Z digest=sha256:2acc20747a6f8ffcb00509e80136278c437b3bd8fc214df87734dc3b6eb55c2e

Pith citing papers

No inbound Pith citation observations are available.