Pith. sign in

Paper Citation Record · LEDGER

OpenCodeInstruct: A Large-scale Instruction Tuning Dataset for Code LLMs

As of 5 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:2504.04030.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.04030 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T21:18:25.994910Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 5114ecfe-fb52-4e44-8466-be759840ff37 · inbound

AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators cites this paper.

AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators OpenCodeInstruct: A Large-scale Instruction Tuning Dataset for Code LLMs

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-05T21:18:25.994910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:18:25.994910Z digest=sha256:99f628e2fefd7be4bbbfa3203cc8be0bee359ef9f6a1362c4e3ab97feec61087

Observation baa35bcd-0c52-41a8-85fa-08b8d0f41694 · inbound

NVIDIA Nemotron 3: Efficient and Open Intelligence cites this paper.

NVIDIA Nemotron 3: Efficient and Open Intelligence OpenCodeInstruct: A Large-scale Instruction Tuning Dataset for Code LLMs

Reference 146

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T01:40:42.670697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T01:40:42.190369Z digest=sha256:1f271b9308814495a05ff4a1ed03d58f0dc7e488a59496f261a2f1414ddbf494

Observation c8316051-861b-4dcf-a6c8-52b676137502 · inbound

A Survey on Evaluating Quality and Trustworthiness in LLM-Generated Data cites this paper.

A Survey on Evaluating Quality and Trustworthiness in LLM-Generated Data OpenCodeInstruct: A Large-scale Instruction Tuning Dataset for Code LLMs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T08:15:12.744350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T08:15:12.744350Z digest=sha256:d15fa62bcb664f7b15f85fd674a4590748392513e6a81b8f81bf7a68d540154a

Observation fa6aefeb-02bf-462d-9177-0c714affd228 · inbound

Robust Policy Optimization to Prevent Catastrophic Forgetting cites this paper.

Robust Policy Optimization to Prevent Catastrophic Forgetting OpenCodeInstruct: A Large-scale Instruction Tuning Dataset for Code LLMs

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:37:24.407770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T05:33:42.965249Z digest=sha256:8110d888309758d04d223a5dd3dcb25dfd542d39e450d8c86c20030ec7de6426

Observation 1b03bbc1-ba20-4d30-b821-32e88e5efb3d · inbound

Yet Even Less Is Even Better For Agentic, Reasoning, and Coding LLMs cites this paper.

Yet Even Less Is Even Better For Agentic, Reasoning, and Coding LLMs OpenCodeInstruct: A Large-scale Instruction Tuning Dataset for Code LLMs

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:38:22.655467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T22:33:48.548362Z digest=sha256:de879d4db82e336d2882656a214ecfe148f2ace814e6a2e8a341a5e1737252bf

Observation c3373941-bd2d-4ad9-910b-324dc696c669 · inbound

DMax: Aggressive Parallel Decoding for dLLMs cites this paper.

DMax: Aggressive Parallel Decoding for dLLMs OpenCodeInstruct: A Large-scale Instruction Tuning Dataset for Code LLMs

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:41:04.786360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T17:58:17.880199Z digest=sha256:544849be2a6e5a5488d5e44ae19c3e8ad983de68f51a6a37f0bccdf4fa3209e3

Observation f9d0ae2f-014c-4052-8f73-2f37f5475b6e · inbound

DMax: Aggressive Parallel Decoding for dLLMs cites this paper.

DMax: Aggressive Parallel Decoding for dLLMs OpenCodeInstruct: A Large-scale Instruction Tuning Dataset for Code LLMs

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:47:40.144387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T16:46:56.743268Z digest=sha256:103c0f5900e2165faff02ff1c2ae19a8825cdf55a0ad44f3ff80cccf510211e1

Observation 38fc9fbf-4a98-466c-8ec7-f2325af5656c · inbound

SimCT: Recovering Lost Supervision for Cross-Tokenizer On-Policy Distillation cites this paper.

SimCT: Recovering Lost Supervision for Cross-Tokenizer On-Policy Distillation OpenCodeInstruct: A Large-scale Instruction Tuning Dataset for Code LLMs

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:05:53.589428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T03:01:13.168529Z digest=sha256:3a7857e4825defd65a7d78c8d059e463cbeead5714806a291b3ee3f47f9ae195

Observation 50a37ed7-eed9-421c-861d-4b004b9c2bc6 · inbound

SimCT: Recovering Lost Supervision for Cross-Tokenizer On-Policy Distillation cites this paper.

SimCT: Recovering Lost Supervision for Cross-Tokenizer On-Policy Distillation OpenCodeInstruct: A Large-scale Instruction Tuning Dataset for Code LLMs

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-22T10:31:25.025236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T10:30:11.309910Z digest=sha256:fd972b3e8d0f6654ce8c0ea4f72a484b34b64be9b2de6ef4530e3a1097b57f3b

Observation f5bc92ad-ee08-4f21-9546-ef67792c1d72 · inbound

Subjective Code Preferences in Experts and Large Language Models cites this paper.

Subjective Code Preferences in Experts and Large Language Models OpenCodeInstruct: A Large-scale Instruction Tuning Dataset for Code LLMs

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T23:24:01.795920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T23:18:27.778928Z digest=sha256:c9bb13055264821315ca61ff9ec0ba628d202e9e63028fe0c3d96b499129390b

Observation db3acef2-28ee-4e85-939f-981f5c37e986 · inbound

PrivCode++: Latent-Conditioned Differentially Private Code Generation for Comprehensive Guarantees cites this paper.

PrivCode++: Latent-Conditioned Differentially Private Code Generation for Comprehensive Guarantees OpenCodeInstruct: A Large-scale Instruction Tuning Dataset for Code LLMs

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T01:27:31.726851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-27T16:28:06.351376Z digest=sha256:93f3b0fc142cd6e8d24855db7eabe5734eef287e6a67ed1020143261a069ac58

Observation d35718cc-4719-4ccb-88b9-3b7cd3862511 · inbound

Grammar-Constrained Decoding Can Jailbreak LLMs into Generating Malicious Code cites this paper.

Grammar-Constrained Decoding Can Jailbreak LLMs into Generating Malicious Code OpenCodeInstruct: A Large-scale Instruction Tuning Dataset for Code LLMs

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-07-03T11:58:06.215545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T09:14:50.233629Z digest=sha256:ef27f50e0bf53189d960c4041b5284f75e0848326676ecfd3142dbcbb534c5f4

Observation f294068d-c7d5-4bac-8940-64061ec80329 · inbound

CODEBLOCK: Learning to Supervise Code at the Right Granularity cites this paper.

CODEBLOCK: Learning to Supervise Code at the Right Granularity OpenCodeInstruct: A Large-scale Instruction Tuning Dataset for Code LLMs

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T08:17:45.490427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T10:52:00.147391Z digest=sha256:3a53a8e0009d1ea2bb0a826c84f17ba05c5d41f1e97e5e19545f16229f36ecdc

Observation 21429693-07b6-4bb0-8fa8-601b497a6285 · inbound

Self-CTRL: Self-Consistency Training with Reinforcement Learning cites this paper.

Self-CTRL: Self-Consistency Training with Reinforcement Learning OpenCodeInstruct: A Large-scale Instruction Tuning Dataset for Code LLMs

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:08:56.021233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T01:38:48.296421Z digest=sha256:9f3c56777ef250f59b57346936e3a055979b2ac4196c4ea588f7551166d5f770

Observation 4f2cd127-bba2-4ff2-9545-9c174dd75a88 · inbound

Masked Language Flow Models cites this paper.

Masked Language Flow Models OpenCodeInstruct: A Large-scale Instruction Tuning Dataset for Code LLMs

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-01T19:06:02.657077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T01:17:56.122002Z digest=sha256:5409aebb95d4c1d417065c33529308d18f0ffe9c0bae47e2f2f1161cce219a30

Observation eeb760af-1fa6-4b34-b7bd-508008d233cf · inbound

Reward-Free Code Alignment from Pretrained or Fine-Tuned LLM: Unpacking the Trade-offs for Code Generation cites this paper.

Reward-Free Code Alignment from Pretrained or Fine-Tuned LLM: Unpacking the Trade-offs for Code Generation OpenCodeInstruct: A Large-scale Instruction Tuning Dataset for Code LLMs

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-06-30T08:34:26.753153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T08:30:37.016856Z digest=sha256:befa36e6903059aa4d73997c114bfdcef81e64490f90bc0f305c05cb6f66372e

Observation 4ea2e034-07cd-4d37-b120-681d3dd472aa · inbound

AdaFlash: Adaptive Speculative Decoding via On-Policy Distilled Diffusion Drafters cites this paper.

AdaFlash: Adaptive Speculative Decoding via On-Policy Distilled Diffusion Drafters OpenCodeInstruct: A Large-scale Instruction Tuning Dataset for Code LLMs

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-01T13:10:37.439253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:10:37.439253Z digest=sha256:2550155f2908503244583b2e85b806cac58d88db6ca453ae4fcf27d47a6814c4

Observation d3763eb7-31d1-48d1-a294-64f0373fbae2 · inbound

AngelSpec: Towards Real-World High Performance Inference with Speculative Decoding cites this paper.

AngelSpec: Towards Real-World High Performance Inference with Speculative Decoding OpenCodeInstruct: A Large-scale Instruction Tuning Dataset for Code LLMs

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T01:22:12.369265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T01:22:12.369265Z digest=sha256:5e3cd55cb4deddc73bfde3bd6ff587d01d3804367ea7eb6d685d3805c942b84b