Pith. sign in

Paper Citation Record · LEDGER

FullStack Bench: Evaluating LLMs as Full Stack Coders

As of 21 August 2026, this Paper Citation Record lists 76 of 76 outbound references and 28 inbound Pith citation observations for arXiv:2412.00535.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.00535 v6

Coverage vector

measured 76 of 76 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T05:20:05.505030Z

measured 104 of 104 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 28 of 28 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:02:43.399252Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

76 of 76 outbound references displayed

  • verified exact1
  • verified fuzzy10
  • unresolved65
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

2
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 6f3cd70f-8664-42c6-8cec-212d7817d145 · outbound

This paper cites Abadi, A.

FullStack Bench: Evaluating LLMs as Full Stack Coders Abadi, A

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:20:06.249592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T05:20:05.278848Z digest=sha256:defab91de106134a67aefeed7b7f8da06ba054a7bbba31228c098086ceb51506

Observation 87e018a8-fe43-4271-88da-66d8d22a0e9d · outbound

This paper cites GPT-4 Technical Report.

FullStack Bench: Evaluating LLMs as Full Stack Coders GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T05:20:05.283095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:20:05.283095Z digest=sha256:9de09603b5bdd81fa24a57a973b11a13f59caa7258b85b62fe18c97970182d81

Observation b5308a0e-e859-47c6-9e38-1cbb0158b1c0 · outbound

This paper cites Guiding Language Models of Code with Global Context using Monitors.

FullStack Bench: Evaluating LLMs as Full Stack Coders Guiding Language Models of Code with Global Context using Monitors

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T05:20:05.286611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:20:05.286611Z digest=sha256:12abfdd7dd69f32ded2ec5100084e0ca8dd58ea32db85e36aecb4237a66affd2

Observation 3533f3b0-85bb-4607-a2b4-37d4c309b370 · outbound

This paper cites SantaCoder: don't reach for the stars!.

FullStack Bench: Evaluating LLMs as Full Stack Coders SantaCoder: don't reach for the stars!

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T05:20:05.289718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:20:05.289718Z digest=sha256:c623d54d1bfe26e39f781dc61af15a01c87fd5bc267fd12f2f8d08db1924fdbd

Observation da587a18-9d0c-47e8-a1ca-e701a62f438c · outbound

This paper cites Athiwaratkun, S.

FullStack Bench: Evaluating LLMs as Full Stack Coders Athiwaratkun, S

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T05:20:05.293097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:20:05.293097Z digest=sha256:94a8c3906debb4f52fe38c30960fa25d6c78251007431207394eb7ac7d810c0d

Observation e71362df-0543-4c8f-bbfd-e805624a7f06 · outbound

This paper cites Program Synthesis with Large Language Models.

FullStack Bench: Evaluating LLMs as Full Stack Coders Program Synthesis with Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T05:20:05.299398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:20:05.299398Z digest=sha256:8dbe2a7a0c4f8fbae95b102d9beefc46a71e45f4b556eb6d555927afd74edd89

Observation 57609d92-1dcc-489f-81fa-3ab3350b3def · outbound

This paper cites Qwen Technical Report.

FullStack Bench: Evaluating LLMs as Full Stack Coders Qwen Technical Report

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T05:20:05.302038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:20:05.302038Z digest=sha256:cac759259804cb5fcda882cd8e05a3b7a6105a0d7fe1b10972fe54abd697e3b4

Observation d0f108f9-07fa-404a-aa87-a9c6e4b804a3 · outbound

This paper cites CodePlan: Repository-level Coding using LLMs and Planning.

FullStack Bench: Evaluating LLMs as Full Stack Coders CodePlan: Repository-level Coding using LLMs and Planning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T05:20:05.307436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:20:05.307436Z digest=sha256:d11596ac8ba7450cd50e90653b9016503c3732a4115822ca6cde76019921520e

Observation 2e116869-6856-4c68-b5b6-c5b16f1a8de9 · outbound

This paper cites Black, L.

FullStack Bench: Evaluating LLMs as Full Stack Coders Black, L

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T05:20:05.310636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:20:05.310636Z digest=sha256:446f84ab34e977cebdb5799fc2553fb00128e3ca10c3245322124c548b368585

Observation 6aab817e-0457-4ba5-9539-484de413b8e7 · outbound

This paper cites Black, S.

FullStack Bench: Evaluating LLMs as Full Stack Coders Black, S

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T05:20:05.313377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:20:05.313377Z digest=sha256:d200f0258480717e0744f67d56147945ae6d849d15bde167852d73cb5f0c9088

Observation ab0d37de-0ee3-4c0c-8500-07ad65129277 · outbound

This paper cites MultiPL-E: A Scalable and Extensible Approach to Benchmarking Neural Code Generation.

FullStack Bench: Evaluating LLMs as Full Stack Coders MultiPL-E: A Scalable and Extensible Approach to Benchmarking Neural Code Generation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T05:20:05.316083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:20:05.316083Z digest=sha256:bd4487545df64965ba0698bf3637f8b034fb340532fa4f6f115396111dc447af

Observation 1a724412-aec9-49c0-8066-5f4cc04908aa · outbound

This paper cites Cassano, J.

FullStack Bench: Evaluating LLMs as Full Stack Coders Cassano, J

Reference 13

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-12T05:20:06.003163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T05:20:05.319047Z digest=sha256:5aaf3d908ee69d5bcb6aea8bc7132848ab483bcdab8b0bf132db5770a7b48faf

Observation 16aa5753-9cdf-41ab-8d51-e0033ba29954 · outbound

This paper cites McEval: Massively Multilingual Code Evaluation.

FullStack Bench: Evaluating LLMs as Full Stack Coders McEval: Massively Multilingual Code Evaluation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T05:20:05.321813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:20:05.321813Z digest=sha256:f25c8a41d26b2d68a9571b788ef8be44cd538e11db41f16a249666f19feefbad

Observation ce27147a-62e6-44fe-a44a-8ff938306970 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

FullStack Bench: Evaluating LLMs as Full Stack Coders Evaluating Large Language Models Trained on Code

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T05:20:05.327644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:20:05.327644Z digest=sha256:3fb6c810c4d28b3ff71ff24e41225c300ed6a7eee6de10babbcb6eadb956f729

Observation ca876e82-cdc2-4bf9-8226-99af66c92d8e · outbound

This paper cites Chowdhery, S.

FullStack Bench: Evaluating LLMs as Full Stack Coders Chowdhery, S

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:20:06.236654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T05:20:05.330423Z digest=sha256:5a46e45d456bd9598a9068067f2f4c67996a2b72336b1b9cc2689dba22db9ee7

Observation 8928e6d0-8246-4d0a-819b-6573295cc81c · outbound

This paper cites MHPP: Exploring the Capabilities and Limitations of Language Models Beyond Basic Code Generation.

FullStack Bench: Evaluating LLMs as Full Stack Coders MHPP: Exploring the Capabilities and Limitations of Language Models Beyond Basic Code Generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T05:20:05.333354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:20:05.333354Z digest=sha256:6aed051d100e39872a91be9650ad387dd750834d30a282675c81b650a6267211

Observation 6d6a7549-73e3-44aa-a48f-4888e03485ca · outbound

This paper cites R2C2-Coder: Enhancing and Benchmarking Real-world Repository-level Code Completion Abilities of Code Large Language Models.

FullStack Bench: Evaluating LLMs as Full Stack Coders R2C2-Coder: Enhancing and Benchmarking Real-world Repository-level Code Completion Abilities of Code Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T05:20:05.335928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:20:05.335928Z digest=sha256:49b47a984ff9df7af0d05cba17eee7a82191bfefdbd6a1c5768de2abedb10f59

Observation fe6e42b1-3176-4ff7-bed5-fefe82cfbbc2 · outbound

This paper cites CoCoMIC: Code Completion By Jointly Modeling In-file and Cross-file Context.

FullStack Bench: Evaluating LLMs as Full Stack Coders CoCoMIC: Code Completion By Jointly Modeling In-file and Cross-file Context

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T05:20:05.338700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:20:05.338700Z digest=sha256:dfed5923370ad80d9bedaaf0ca5eb35332ffe91196664630afb1d3b2c76fc708

Observation 14a9c9bb-32e5-4c17-82cd-baedec22879d · outbound

This paper cites Multi-Programming Language Sandbox for LLMs.

FullStack Bench: Evaluating LLMs as Full Stack Coders Multi-Programming Language Sandbox for LLMs

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T05:20:05.341415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:20:05.341415Z digest=sha256:62b32963e8cfb9448584d4b5e40fdf85cadcdb1efacb2e78001883a6ca9bbc44

Observation f215d007-4a18-4c0c-b685-d4481799b9a6 · outbound

This paper cites Fried, A.

FullStack Bench: Evaluating LLMs as Full Stack Coders Fried, A

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:20:06.228933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T05:20:05.344193Z digest=sha256:c84e66ece5361b36a5731f6e86bb232db8da6829d70c7c451a76dd4c2b5444de

Observation a0b96630-b583-4b10-b252-9d3c719ffa3d · outbound

This paper cites an unresolved cited work.

FullStack Bench: Evaluating LLMs as Full Stack Coders Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-12T05:20:06.219799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T05:20:05.346740Z digest=sha256:7f056237fe80d672e910c31c80e09af51dd7868a69fbf70237bc94d27da38967

Observation a795a31c-3557-4b7d-bec2-4d9b5759129a · outbound

This paper cites CRUXEval: A Benchmark for Code Reasoning, Understanding and Execution.

FullStack Bench: Evaluating LLMs as Full Stack Coders CRUXEval: A Benchmark for Code Reasoning, Understanding and Execution

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T05:20:05.349369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:20:05.349369Z digest=sha256:d0d406c1faf7fe58383daf820b1a3d7daf23397c29a837feaf8946fb121b8b7d

Observation d70c94e6-37ca-489c-adfb-4fbcf9e14e7c · outbound

This paper cites DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence.

FullStack Bench: Evaluating LLMs as Full Stack Coders DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T05:20:05.355918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:20:05.355918Z digest=sha256:a732a31ad21035b3cc8739135866927aa48dd2806ed6eb36a736c75208386165

Observation e6b2b047-b222-4786-a92d-ac446ed07b64 · outbound

This paper cites Hendrycks, S.

FullStack Bench: Evaluating LLMs as Full Stack Coders Hendrycks, S

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:20:06.212099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T05:20:05.358298Z digest=sha256:34e8c8a448c890578d85272bf708c383bdddae24a7d26f04b7545e96ade7d26f

Observation 4df8d9f2-c08e-4fe4-8c73-bfa4177341d0 · outbound

This paper cites CoSQA: 20,000+ Web Queries for Code Search and Question Answering.

FullStack Bench: Evaluating LLMs as Full Stack Coders CoSQA: 20,000+ Web Queries for Code Search and Question Answering

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T05:20:05.360941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:20:05.360941Z digest=sha256:52390b288c45dcfbfe2831a56f6c9428b3c5dd1e9fb8d5cf00061424e17a8074

Observation 6d518335-d4fd-464a-a5bf-7e22daca8d11 · outbound

This paper cites Huang, T.

FullStack Bench: Evaluating LLMs as Full Stack Coders Huang, T

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:20:06.204976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T05:20:05.363640Z digest=sha256:6ab690f12810471aa2499eda0f12d8f04592c08275f5451bf1f7d8df605d696f

Observation c23d30ef-92ff-4444-8cc4-fd3b89c7e560 · outbound

This paper cites Huang, T.

FullStack Bench: Evaluating LLMs as Full Stack Coders Huang, T

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:20:06.196841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T05:20:05.366181Z digest=sha256:57737a784bd1e52516b3710791f33dbd5c5615765c3c447e8f48935df4325131

Observation d642dcd8-ed4d-4d21-b078-f321dfe0cfe4 · outbound

This paper cites Qwen2.5-Coder Technical Report.

FullStack Bench: Evaluating LLMs as Full Stack Coders Qwen2.5-Coder Technical Report

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T05:20:05.369428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:20:05.369428Z digest=sha256:38f3661a957139bf5ddc6658535cc45854b36b06de91dc9ffd38d2ecc970c670

Observation a361ae43-9b96-41db-a933-133240f8d471 · outbound

This paper cites an unresolved cited work.

FullStack Bench: Evaluating LLMs as Full Stack Coders Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-12T05:20:06.188109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T05:20:05.372215Z digest=sha256:579da42af58d0fea06c5671bbed07d0fd998aaa98aec104cafa6fb796641d486

Observation 2300580e-4c56-46b8-856b-71617e7558d4 · outbound

This paper cites CodeSearchNet Challenge: Evaluating the State of Semantic Code Search.

FullStack Bench: Evaluating LLMs as Full Stack Coders CodeSearchNet Challenge: Evaluating the State of Semantic Code Search

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T05:20:05.374768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:20:05.374768Z digest=sha256:3d2a36ab33f2cdab9cbf1be35192b6ebf6a69bbb0fcdbdb010be36cfe459f2f2

Observation c394e215-1a5e-4a0b-9fb0-6ed32cb7dca0 · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

FullStack Bench: Evaluating LLMs as Full Stack Coders LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T05:20:05.377628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:20:05.377628Z digest=sha256:96d1c53d90cd1544d95c540fb5e7222c32947231038b7ef8f6dd75314ba63613

Observation 12773f51-5145-4653-8da2-41da8f65f4a0 · outbound

This paper cites an unresolved cited work.

FullStack Bench: Evaluating LLMs as Full Stack Coders Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-12T05:20:06.179630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T05:20:05.380401Z digest=sha256:d5baa63724d9cdcb048c73a6fe865369c0596a9c1c8531c9ae171ec95f73a6d9

Observation 35154dbe-9b69-4865-9c9c-abf9f558b0f9 · outbound

This paper cites an unresolved cited work.

FullStack Bench: Evaluating LLMs as Full Stack Coders Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-12T05:20:06.169630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T05:20:05.382914Z digest=sha256:e0f4b56e13dd8073c0ee367a9daaa6e742490059c7293f9e9abf2e062ad82b93

Observation adc2739d-7e84-46c6-a933-3c229eb54a7e · outbound

This paper cites an unresolved cited work.

FullStack Bench: Evaluating LLMs as Full Stack Coders Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-12T05:20:06.160031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T05:20:05.385563Z digest=sha256:57cdd417339c683ed8e28359420d05a6f628288e4b4197ad534c4eb3f67b7431

Observation 9ca28e6c-d273-4255-91b1-522918a407fa · outbound

This paper cites DS-1000: A Natural and Reliable Benchmark for Data Science Code Generation.

FullStack Bench: Evaluating LLMs as Full Stack Coders DS-1000: A Natural and Reliable Benchmark for Data Science Code Generation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T05:20:05.388102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:20:05.388102Z digest=sha256:4b8f6c34353c1efc6389ddb991534ee0005cdb16e4ce6f7fe84e4ed7d5ef0904

Observation 483d1c58-33e1-4c33-9a90-af4c8e5016a9 · outbound

This paper cites an unresolved cited work.

FullStack Bench: Evaluating LLMs as Full Stack Coders Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-12T05:20:06.150520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T05:20:05.390822Z digest=sha256:9f1327b3c6e1f911efb26c2d519b06c208cc302a568658733179700af9e8ffb6

Observation fd255c38-c4c1-492b-a455-5ff3c739b810 · outbound

This paper cites Difysandbox.

FullStack Bench: Evaluating LLMs as Full Stack Coders Difysandbox

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:20:06.140969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T05:20:05.393332Z digest=sha256:1bb8c7a2b8e93618330355a756259fd8cb6ed3e5c8e76f5b44235bfb489b3582

Observation 886c4559-7c33-47c4-8f1f-d6340945ce18 · outbound

This paper cites CodeRL: Mastering Code Generation through Pretrained Models and Deep Reinforcement Learning.

FullStack Bench: Evaluating LLMs as Full Stack Coders CodeRL: Mastering Code Generation through Pretrained Models and Deep Reinforcement Learning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T05:20:05.396000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:20:05.396000Z digest=sha256:69a6472065cdfe8792abd784991c45380fd05c15a623edd8c7e3e5fe0b1c7961

Observation 93710c02-f040-4f29-812e-a474a5fdea57 · outbound

This paper cites StarCoder: may the source be with you!.

FullStack Bench: Evaluating LLMs as Full Stack Coders StarCoder: may the source be with you!

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T05:20:05.398710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:20:05.398710Z digest=sha256:20097572f2b137564c1e60a3627af7404f6ced17b324c48d9f0726de6d379f02

Observation 9c861291-32d0-4613-9598-42bb5b7fa98b · outbound

This paper cites Competition-Level Code Generation with AlphaCode.

FullStack Bench: Evaluating LLMs as Full Stack Coders Competition-Level Code Generation with AlphaCode

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T05:20:05.401614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:20:05.401614Z digest=sha256:06dfcbe77ed38181510379ae25c3593e88ac256db84c6d91a8a074decf79513d

Observation fdbad7e8-58a2-4de9-87ce-d8eb3a56be29 · outbound

This paper cites ProCQA: A Large-scale Community-based Programming Question Answering Dataset for Code Search.

FullStack Bench: Evaluating LLMs as Full Stack Coders ProCQA: A Large-scale Community-based Programming Question Answering Dataset for Code Search

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T05:20:05.404456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:20:05.404456Z digest=sha256:5cbcda351eecd602757b5f5499bdc056c624e91fd8dba07f248bee9d0a495efd

Observation 4d84d6ef-7222-42f5-b66f-d46a1f1d27fd · outbound

This paper cites an unresolved cited work.

FullStack Bench: Evaluating LLMs as Full Stack Coders Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-12T05:20:06.131991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T05:20:05.407238Z digest=sha256:296dfb5fc071f2df0d547787d8afce398fa3a69d26d670f7848444137de6c1c9

Observation c47d663f-257f-4754-b4e9-f5aa3ecb5280 · outbound

This paper cites an unresolved cited work.

FullStack Bench: Evaluating LLMs as Full Stack Coders Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-12T05:20:06.122995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T05:20:05.409886Z digest=sha256:bf2705e458e70f0d17bd62c655c83618130c5fc221b8a2ea90341b6611fee313

Observation 19bfa075-9ae0-4a0b-94d6-0291a3da05c4 · outbound

This paper cites VerilogEval: Evaluating Large Language Models for Verilog Code Generation.

FullStack Bench: Evaluating LLMs as Full Stack Coders VerilogEval: Evaluating Large Language Models for Verilog Code Generation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T05:20:05.412832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:20:05.412832Z digest=sha256:90bb0dc7770477d990c3d32412d2ecdb98428bcda2d5743e5ed3f85c76ea71ab

Observation f907edef-f7cd-469d-b360-2e911f74fb32 · outbound

This paper cites MdEval: Massively Multilingual Code Debugging.

FullStack Bench: Evaluating LLMs as Full Stack Coders MdEval: Massively Multilingual Code Debugging

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T05:20:05.415542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:20:05.415542Z digest=sha256:3e59085b9158983f83f8b9a671399840a3f51012314abd7590dc3a1e52836cf1

Observation 53623dae-89c5-4a26-abcc-1151ff97177d · outbound

This paper cites RepoBench: Benchmarking Repository-Level Code Auto-Completion Systems.

FullStack Bench: Evaluating LLMs as Full Stack Coders RepoBench: Benchmarking Repository-Level Code Auto-Completion Systems

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T05:20:05.418526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:20:05.418526Z digest=sha256:2770e202b4bba17feaa0d5788cfb956eb07126cc90f914a2a170d1dc60a7ea9f

Observation ba143727-baac-48d7-b607-986f25759fd0 · outbound

This paper cites StarCoder 2 and The Stack v2: The Next Generation.

FullStack Bench: Evaluating LLMs as Full Stack Coders StarCoder 2 and The Stack v2: The Next Generation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T05:20:05.421369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:20:05.421369Z digest=sha256:b60bf9b7277e3ebe8a70384a908eab89ef52da2b877001915f0c33ef3178fcf4

Observation e30fdd54-472a-4392-9d9b-d850fe3a749c · outbound

This paper cites an unresolved cited work.

FullStack Bench: Evaluating LLMs as Full Stack Coders Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-12T05:20:06.113393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T05:20:05.424082Z digest=sha256:82f0e8d25aada5ce6229017145d021c9838a061ba578cefa1c6f19b2a16b63ba

Observation 5b06595a-52dc-4c30-a6f6-3d115d6a6d75 · outbound

This paper cites an unresolved cited work.

FullStack Bench: Evaluating LLMs as Full Stack Coders Unresolved cited work

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T05:20:05.426936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:20:05.426936Z digest=sha256:29952894b2ff66981e607cb630f5a9311b6135a6843df4ec1bd56cfefb2e32e0

Observation 8312e3f0-9e47-4281-824e-18f7fe827f81 · outbound

This paper cites Madaan, N.

FullStack Bench: Evaluating LLMs as Full Stack Coders Madaan, N

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T05:20:05.429513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:20:05.429513Z digest=sha256:b0f5fa2314a759c8156ccfd52206efa83434352e4ad9a8e870ae374ad6c80f14

Observation d488b940-e973-4d75-8591-d9a8dee41716 · outbound

This paper cites Nijkamp, B.

FullStack Bench: Evaluating LLMs as Full Stack Coders Nijkamp, B

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:20:06.099638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T05:20:05.432049Z digest=sha256:d87c1b183767e82154d346a6e755a9c0f19d8fd8d22ae5ccd6dce442899c555d

Observation 75e11b82-682d-428f-abfc-2c6599cb1037 · outbound

This paper cites PyTorch: An Imperative Style, High-Performance Deep Learning Library.

FullStack Bench: Evaluating LLMs as Full Stack Coders PyTorch: An Imperative Style, High-Performance Deep Learning Library

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T05:20:05.434437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:20:05.434437Z digest=sha256:9fcc5c7fb955972b888dabc1d06bfa999f7d78295b06027a29492e374652c0a7

Observation 7dfa2a46-1536-4b0d-b0e2-0c5753288b76 · outbound

This paper cites an unresolved cited work.

FullStack Bench: Evaluating LLMs as Full Stack Coders Unresolved cited work

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T05:20:05.437271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:20:05.437271Z digest=sha256:0a9a0d23f919ba916a9ed08734d8159cdc65b7e31ed427bc8b10111d2c8f5178

Observation 58109c6c-f2fa-43ed-ada5-2e77ff980481 · outbound

This paper cites RunBugRun -- An Executable Dataset for Automated Program Repair.

FullStack Bench: Evaluating LLMs as Full Stack Coders RunBugRun -- An Executable Dataset for Automated Program Repair

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T05:20:05.440540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:20:05.440540Z digest=sha256:a3bdbcb974ab6498872409c017f6826a652647ec682f1b81672bd3da6cd53abf

Observation 5647213e-ef1d-4e65-837f-270eeb6df02f · outbound

This paper cites Richter and H.

FullStack Bench: Evaluating LLMs as Full Stack Coders Richter and H

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:20:06.090404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T05:20:05.443378Z digest=sha256:00d5a8604b6208acebbad70361a9565b46ce94b2ac9e3f921ed2221cb176b22a

Observation 4dfbd0ff-da1e-4e76-842c-6277063e3abe · outbound

This paper cites Code Llama: Open Foundation Models for Code.

FullStack Bench: Evaluating LLMs as Full Stack Coders Code Llama: Open Foundation Models for Code

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-12T05:20:05.446045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:20:05.446045Z digest=sha256:192ac1e8a4d3b7d9f8b855dac829bda2cfb25efd1e7d1da33241652c6e46fd96

Observation bc5a08fd-4022-41bf-a647-fd4bbc5fded2 · outbound

This paper cites RepoFusion: Training Code Models to Understand Your Repository.

FullStack Bench: Evaluating LLMs as Full Stack Coders RepoFusion: Training Code Models to Understand Your Repository

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-12T05:20:05.449223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:20:05.449223Z digest=sha256:cb64b1b518b85512e54f9f502ba7436ef4131e7fbc624f0a6ec7086eb1609700

Observation d160a3e9-60a5-4d22-95c5-ce3af963b220 · outbound

This paper cites Shrivastava, H.

FullStack Bench: Evaluating LLMs as Full Stack Coders Shrivastava, H

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:20:06.082608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T05:20:05.451863Z digest=sha256:7324c62830c585037fe29dd88b68ff3de4fef010e494bfa31b85c9e1b271e0fe

Observation 2b3a5e4f-c510-46aa-892c-90776783d379 · outbound

This paper cites TableGPT2: A Large Multimodal Model with Tabular Data Integration.

FullStack Bench: Evaluating LLMs as Full Stack Coders TableGPT2: A Large Multimodal Model with Tabular Data Integration

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-12T05:20:05.454496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:20:05.454496Z digest=sha256:ca53f6335785a80a965450cdf250b61e7ea7e7f6b2fcb7de9a872dbb41a50a4f

Observation 7ee3b980-09b5-4a3a-a8e7-598f0c81faf2 · outbound

This paper cites an unresolved cited work.

FullStack Bench: Evaluating LLMs as Full Stack Coders Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-12T05:20:06.075004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T05:20:05.457236Z digest=sha256:e7a33facff73f72cd424dcbec0ad37ca133cbf6bdce03e5bc591f505406447f4

Observation a8a63f71-e51a-44e7-b731-ca8f9e22eb05 · outbound

This paper cites The Llama 3 Herd of Models.

FullStack Bench: Evaluating LLMs as Full Stack Coders The Llama 3 Herd of Models

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-12T05:20:05.459692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:20:05.459692Z digest=sha256:6f8d04e56417f333d94f6c7e27994abdfb313599847432c4a34ec2a953f79f26

Observation e5f77d3f-dd67-475d-8b9b-828fdaee41b1 · outbound

This paper cites DebugBench: Evaluating Debugging Capability of Large Language Models.

FullStack Bench: Evaluating LLMs as Full Stack Coders DebugBench: Evaluating Debugging Capability of Large Language Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-12T05:20:05.462687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:20:05.462687Z digest=sha256:03cae82ed65f842342d17433d07359f1b752c020343e824e9a386b359fba54b1

Observation d3fb5fc5-02c8-4dba-bcf3-1c56d9ef7985 · outbound

This paper cites Skywork: A More Open Bilingual Foundation Model.

FullStack Bench: Evaluating LLMs as Full Stack Coders Skywork: A More Open Bilingual Foundation Model

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-12T05:20:05.466346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:20:05.466346Z digest=sha256:d525318bfd7c5a638b3035f830658b4a22239bc54df3b1ca31e3acda92fc9162

Observation d174db6d-19a8-4ec6-9725-01e83adc936b · outbound

This paper cites TableBench: A Comprehensive and Complex Benchmark for Table Question Answering.

FullStack Bench: Evaluating LLMs as Full Stack Coders TableBench: A Comprehensive and Complex Benchmark for Table Question Answering

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-12T05:20:05.469752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:20:05.469752Z digest=sha256:8e8198d9b26660668344f1b1ae527aedcc04cdca74ed5008fa21aeb6f4c59362

Observation 27cb9275-0f0d-4130-ab85-9fb389b14df0 · outbound

This paper cites an unresolved cited work.

FullStack Bench: Evaluating LLMs as Full Stack Coders Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-08-12T05:20:06.065464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T05:20:05.472547Z digest=sha256:b9b61bc1dd4134eee41a2f0aca33086e467405d85fe5c02d134d5d2b116ff5c0

Observation 677e37ac-a539-471c-a199-f5b014d4fdbe · outbound

This paper cites CodeTransOcean: A Comprehensive Multilingual Benchmark for Code Translation.

FullStack Bench: Evaluating LLMs as Full Stack Coders CodeTransOcean: A Comprehensive Multilingual Benchmark for Code Translation

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-12T05:20:05.475790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:20:05.475790Z digest=sha256:28598ba97154a847571aa4978c8e5b2a8d535cf793c854a2f7bcb99cf2337a55

Observation f13ae7a9-78b3-4431-bb82-45e06f408143 · outbound

This paper cites Yi: Open Foundation Models by 01.AI.

FullStack Bench: Evaluating LLMs as Full Stack Coders Yi: Open Foundation Models by 01.AI

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-12T05:20:05.478461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:20:05.478461Z digest=sha256:2971a5f2cd9eddc3324ff20e4604e30c1611d527616338d937428004c85b02ec

Observation 33313f8e-31f0-42db-948e-45ff093e4c35 · outbound

This paper cites an unresolved cited work.

FullStack Bench: Evaluating LLMs as Full Stack Coders Unresolved cited work

Reference 71

Resolution
unresolved
raw_fallback, observed 2026-08-12T05:20:06.056737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T05:20:05.481547Z digest=sha256:f38f85d72522af326f720333fe6c4388b3383627866b7b64d27366f9cc72d471

Observation 2da2f8ae-c6aa-45b0-84ae-66cb35d2fc09 · outbound

This paper cites RepoCoder: Repository-Level Code Completion Through Iterative Retrieval and Generation.

FullStack Bench: Evaluating LLMs as Full Stack Coders RepoCoder: Repository-Level Code Completion Through Iterative Retrieval and Generation

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-12T05:20:05.483902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:20:05.483902Z digest=sha256:7a5026c8337b9e76ae24618d4d4b9a651c13ebf0e8668b8471b9f1526073deac

Observation d500d700-a55a-404c-91b9-0dafc3f4708b · outbound

This paper cites NaturalCodeBench: Examining Coding Performance Mismatch on HumanEval and Natural User Prompts.

FullStack Bench: Evaluating LLMs as Full Stack Coders NaturalCodeBench: Examining Coding Performance Mismatch on HumanEval and Natural User Prompts

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-12T05:20:05.487084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:20:05.487084Z digest=sha256:1d28da350807667d4fd5a5780307d17a3e5f81ee2125f13e0bdb96fc7acc7391

Observation f784b43a-7e0d-4d6a-b3da-3d0663f19525 · outbound

This paper cites CodeGemma: Open Code Models Based on Gemma.

FullStack Bench: Evaluating LLMs as Full Stack Coders CodeGemma: Open Code Models Based on Gemma

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-12T05:20:05.489981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:20:05.489981Z digest=sha256:b7e8d3f008512c066df9771991a1c8f749eeec749a93cb42685cc06dc263fc6d

Observation 85473a2b-318e-49f6-b9bb-ff1429be8b21 · outbound

This paper cites MiniF2F: a cross-system benchmark for formal Olympiad-level mathematics.

FullStack Bench: Evaluating LLMs as Full Stack Coders MiniF2F: a cross-system benchmark for formal Olympiad-level mathematics

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-12T05:20:05.492679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:20:05.492679Z digest=sha256:449d45a653bb9d0bcdaab3aa0d9c3e91f2fad3889c4521f0f5c8b3a51b9d2c3f

Observation bd0abebd-5eed-463a-ac0b-158055f1c777 · outbound

This paper cites CodeGeeX: A Pre-Trained Model for Code Generation with Multilingual Benchmarking on HumanEval-X.

FullStack Bench: Evaluating LLMs as Full Stack Coders CodeGeeX: A Pre-Trained Model for Code Generation with Multilingual Benchmarking on HumanEval-X

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-12T05:20:05.495686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:20:05.495686Z digest=sha256:6ebc122fa91c65ce146c8038524ba45aea621a7725916f0a5f6a4b7ecece2286

Observation 58dd1094-d60b-43b8-9688-657c5a9a0f84 · outbound

This paper cites LIME: Less Is More for MLLM Evaluation.

FullStack Bench: Evaluating LLMs as Full Stack Coders LIME: Less Is More for MLLM Evaluation

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-12T05:20:05.498485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:20:05.498485Z digest=sha256:634b366e8de89173287123df2dd1b119fdae6049071f9196166267d29e1f17b2

Observation 55277862-3be4-489b-a37e-60d14bc1fe2b · outbound

This paper cites XLCoST: A Benchmark Dataset for Cross-lingual Code Intelligence.

FullStack Bench: Evaluating LLMs as Full Stack Coders XLCoST: A Benchmark Dataset for Cross-lingual Code Intelligence

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-12T05:20:05.501444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:20:05.501444Z digest=sha256:5f3f3fbc5ac6c90105db47f1cf887d5fa377da0055fd6faf77a62752141dd863

Observation e2d9b971-c799-4fe1-9f17-2cf21051d1ce · outbound

This paper cites DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence.

FullStack Bench: Evaluating LLMs as Full Stack Coders DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-12T05:20:05.505030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:20:05.505030Z digest=sha256:30b26fe59b2bc8a238fe7ac80adf06851db5e710830f25df9df162d8d9efba11

Pith citing papers

Observation 91dc54a7-2761-4458-be49-bf4c913b6e6c · inbound

ExecRepoBench: Multi-level Executable Code Completion Evaluation cites this paper.

ExecRepoBench: Multi-level Executable Code Completion Evaluation FullStack Bench: Evaluating LLMs as Full Stack Coders

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T14:27:05.655691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:27:05.655691Z digest=sha256:715d86027ddf19fb4393615a8d72f6dd16a141af70bc585dad9ca71e74b903c1

Observation b0ce753c-5281-4844-97cd-82d98d4d30fb · inbound

BitsAI-CR: Automated Code Review via LLM in Practice cites this paper.

BitsAI-CR: Automated Code Review via LLM in Practice FullStack Bench: Evaluating LLMs as Full Stack Coders

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T14:40:07.177530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:40:07.177530Z digest=sha256:22532c229a5c301110af25e3507652717173fa782ab3b6123a1572a1d7084e5c

Observation 36a40ee4-029d-4743-aa5e-5805644c6875 · inbound

Multi-Agent Collaboration for Multilingual Code Instruction Tuning cites this paper.

Multi-Agent Collaboration for Multilingual Code Instruction Tuning FullStack Bench: Evaluating LLMs as Full Stack Coders

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T12:37:25.734725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:37:25.734725Z digest=sha256:85e0af33ae64d71fe8d9e05318692e96e420c51c8a95c55858e8a4111cc2c38d

Observation 63ddeeb5-2b43-4b99-a28f-8582a873b955 · inbound

Generative AI Act II: Test Time Scaling Drives Cognition Engineering cites this paper.

Generative AI Act II: Test Time Scaling Drives Cognition Engineering FullStack Bench: Evaluating LLMs as Full Stack Coders

Reference 200

Resolution
unresolved
no resolver link, observed 2026-08-16T12:02:43.399252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:02:43.399252Z digest=sha256:9808fd89696c638521d2f15c1ac4db03585b7052cecb4ed6a2aa997a43481b75

Observation 4e36ea01-a816-479b-8e55-e5e83f090ed0 · inbound

Seed-Coder: Let the Code Model Curate Data for Itself cites this paper.

Seed-Coder: Let the Code Model Curate Data for Itself FullStack Bench: Evaluating LLMs as Full Stack Coders

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:56.816308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:56.816308Z digest=sha256:324c663ef0685f4de9279f4c9dd4c22059a58fb8d69a331a9eff9f9c136006b7

Observation 3685dfb9-cbbf-440f-81cb-00da98b0e11c · inbound

CodeContests+: High-Quality Test Case Generation for Competitive Programming cites this paper.

CodeContests+: High-Quality Test Case Generation for Competitive Programming FullStack Bench: Evaluating LLMs as Full Stack Coders

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T10:16:52.892806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:16:52.892806Z digest=sha256:d2661ca66d2c2cb0dff8822851b5fe9b8a08e569e3d2550b60aa3a79d603ffed

Observation 4e68e30d-277c-49e1-95bf-b374a7ea367d · inbound

AdaptiveLLM: A Framework for Selecting Optimal Cost-Efficient LLM for Code-Generation Based on CoT Length cites this paper.

AdaptiveLLM: A Framework for Selecting Optimal Cost-Efficient LLM for Code-Generation Based on CoT Length FullStack Bench: Evaluating LLMs as Full Stack Coders

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T04:30:43.658947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:30:43.658947Z digest=sha256:aad6b6dac77c22a10f848081997a44604dddb2aca89cab28436179a39edf1244

Observation c3a115a9-4e4a-4d45-8caa-b5e6e63d6d05 · inbound

RLPR: Extrapolating RLVR to General Domains without Verifiers cites this paper.

RLPR: Extrapolating RLVR to General Domains without Verifiers FullStack Bench: Evaluating LLMs as Full Stack Coders

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:06.248686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:01:06.248686Z digest=sha256:fb1139c1d1360870037e5ae3bbc13c0508d0233814f99c3da6becc19871aaead

Observation bc990494-3d29-4615-ba53-57ab17dee1af · inbound

Turning the Tide: Repository-based Code Reflection cites this paper.

Turning the Tide: Repository-based Code Reflection FullStack Bench: Evaluating LLMs as Full Stack Coders

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T17:51:13.610216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:51:13.610216Z digest=sha256:c86698db4d6d124b71f5f8a4fe679c3c78a3166a3ca04e7a9514d4f75bfda1d3

Observation cc06240c-ddf5-42b5-bc80-efc8a971c01d · inbound

IFEvalCode: Controlled Code Generation cites this paper.

IFEvalCode: Controlled Code Generation FullStack Bench: Evaluating LLMs as Full Stack Coders

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T11:44:34.021098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:44:34.021098Z digest=sha256:9c42aed2e94cb27e7dccf8059422ff4dd0df16e71313e755b73049c1f3a88eea

Observation 06a31577-b3b6-41d8-a96b-58f6ee127d9c · inbound

WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent cites this paper.

WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent FullStack Bench: Evaluating LLMs as Full Stack Coders

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:56:23.902839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-15T18:56:23.817544Z digest=sha256:438c20bd166ffec2a161b0c903d1b123487fa4d52387cf0c2a7f867d45acb811

Observation 691aaab4-f958-4149-b143-d5c85afa298b · inbound

Vibration-Based Energy Metric for Restoring Needle Alignment in Autonomous Robotic Ultrasound cites this paper.

Vibration-Based Energy Metric for Restoring Needle Alignment in Autonomous Robotic Ultrasound FullStack Bench: Evaluating LLMs as Full Stack Coders

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T22:28:15.505496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:28:15.505496Z digest=sha256:3ba641e416f6243f788ec17784fd13f7eee1e7b1559c202b9a7c671e64a4de9d

Observation 84d01275-a476-432f-adc0-7b9059005a7f · inbound

Dream-Coder 7B: An Open Diffusion Language Model for Code cites this paper.

Dream-Coder 7B: An Open Diffusion Language Model for Code FullStack Bench: Evaluating LLMs as Full Stack Coders

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T12:56:35.116898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:56:35.116898Z digest=sha256:284d6c227a03efe6cc2b348e70f9a326757bb19e71c9c80843b6c395c2624ac7

Observation 2078eff1-5f64-449f-ac54-1769f309f2dc · inbound

UniCode: Augmenting Evaluation for Code Reasoning cites this paper.

UniCode: Augmenting Evaluation for Code Reasoning FullStack Bench: Evaluating LLMs as Full Stack Coders

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T09:39:47.399871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:39:47.399871Z digest=sha256:965839a45a5fbc29535ca7031deb8a669464792d160088e5fc631be137f1351c

Observation 6d7713cf-1c96-4680-9b4b-8991009a3ea7 · inbound

SCALER:Synthetic Scalable Adaptive Learning Environment for Reasoning cites this paper.

SCALER:Synthetic Scalable Adaptive Learning Environment for Reasoning FullStack Bench: Evaluating LLMs as Full Stack Coders

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-16T16:28:05.743083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-16T16:24:04.132572Z digest=sha256:c647be3590ee5d91b9b2e190b3c745d2878d035b0a840a30fededc6971436f79

Observation 3ea0ae09-9371-4ef4-aeb6-37c6c3c4553a · inbound

CharTool: Tool-Integrated Visual Reasoning for Chart Understanding cites this paper.

CharTool: Tool-Integrated Visual Reasoning for Chart Understanding FullStack Bench: Evaluating LLMs as Full Stack Coders

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:08:12.681617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-13T20:07:23.153064Z digest=sha256:1988f8ffa01dc3b0f9a602bc5e83b8c7e5652c253249a30402433f476df2c39b

Observation 1d632ded-40aa-4766-a353-f0309625ad76 · inbound

InCoder-32B-Thinking: Industrial Code World Model for Thinking cites this paper.

InCoder-32B-Thinking: Industrial Code World Model for Thinking FullStack Bench: Evaluating LLMs as Full Stack Coders

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T18:58:08.705577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-13T18:56:31.975626Z digest=sha256:6e8304a3fc4ffe9061f3bbeb26a335f48a2b31fb5db065c94b53931f97c3cc36

Observation d20db14b-b003-45d1-a46d-1714cf89a392 · inbound

GeoBrowse: A Geolocation Benchmark for Agentic Tool Use with Expert-Annotated Reasoning Traces cites this paper.

GeoBrowse: A Geolocation Benchmark for Agentic Tool Use with Expert-Annotated Reasoning Traces FullStack Bench: Evaluating LLMs as Full Stack Coders

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T17:33:02.359854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-13T17:31:08.575993Z digest=sha256:e8a3466523d312119df1feea57eb4c3df8f7f6026b0bbb00ad3f09ea28034c50

Observation 6ea69460-01e2-4c41-a164-69ea5768da40 · inbound

SWE-WebDevBench: Evaluating Coding Agent Application Platforms as Virtual Software Agencies cites this paper.

SWE-WebDevBench: Evaluating Coding Agent Application Platforms as Virtual Software Agencies FullStack Bench: Evaluating LLMs as Full Stack Coders

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:16:08.964076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-08T16:24:49.710882Z digest=sha256:e79b509e5d1688c6fd3d0c62c561a50492054c082c88d3db29e243d287e8addb

Observation 711bdc2f-7f5b-4f60-a1e0-2f4d27955f21 · inbound

Multi-Objective and Mixed-Reward Reinforcement Learning via Reward-Decorrelated Policy Optimization cites this paper.

Multi-Objective and Mixed-Reward Reinforcement Learning via Reward-Decorrelated Policy Optimization FullStack Bench: Evaluating LLMs as Full Stack Coders

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:27:52.265134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T19:24:05.375951Z digest=sha256:69af3e19d7f7dfcbeecddd838265ab8f93fd9cc6a5a8f768c214e16e2ad33b33

Observation 45947a59-3b54-4b42-8481-4ef33ceb4fbe · inbound

Multi-Objective and Mixed-Reward Reinforcement Learning via Reward-Decorrelated Policy Optimization cites this paper.

Multi-Objective and Mixed-Reward Reinforcement Learning via Reward-Decorrelated Policy Optimization FullStack Bench: Evaluating LLMs as Full Stack Coders

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T19:27:51.248601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T19:24:05.375951Z digest=sha256:cfb915092c74555637f6ff209535b8f52be5be415b21cdf801aae60b25f8dc83

Observation 2b27823f-fdf5-4492-9a77-6ad80320d633 · inbound

Efficient Agentic Reasoning Through Self-Regulated Simulative Planning cites this paper.

Efficient Agentic Reasoning Through Self-Regulated Simulative Planning FullStack Bench: Evaluating LLMs as Full Stack Coders

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-22T06:34:40.962219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-22T06:33:36.846345Z digest=sha256:17a88bfee9c35a9693f8b60aa074dd57629e43c1ae016b2a2ec6ac9d15d3107f

Observation 76f0cbe1-a7f7-4592-a315-fb39d240f6d6 · inbound

Spreadsheet-RL: Advancing Large Language Model Agents on Realistic Spreadsheet Tasks via Reinforcement Learning cites this paper.

Spreadsheet-RL: Advancing Large Language Model Agents on Realistic Spreadsheet Tasks via Reinforcement Learning FullStack Bench: Evaluating LLMs as Full Stack Coders

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:44:38.479911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-22T05:44:12.807291Z digest=sha256:3b5e9edc2624d33019dd69c0280173b8479c19f6c43372360901b2a6c0e3eb33

Observation 7bb56558-e906-442b-a332-b3fcc8b46601 · inbound

MONA: Muon Optimizer with Nesterov Acceleration for Scalable Language Model Training cites this paper.

MONA: Muon Optimizer with Nesterov Acceleration for Scalable Language Model Training FullStack Bench: Evaluating LLMs as Full Stack Coders

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-06-29T19:23:54.161758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-29T19:15:49.229099Z digest=sha256:9ef62fe4df775537269201bf94d7beb48d133eac34bf42a6f8003bafa4addb0a

Observation ffa1d5e4-05c8-44ec-80ad-3116e8cad077 · inbound

Ling and Ring 2.6 Technical Report: Efficient and Instant Agentic Intelligence at Trillion-Parameter Scale cites this paper.

Ling and Ring 2.6 Technical Report: Efficient and Instant Agentic Intelligence at Trillion-Parameter Scale FullStack Bench: Evaluating LLMs as Full Stack Coders

Reference 188

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T22:17:25.583378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-02T22:10:59.568675Z digest=sha256:ed927958c9f29b246540190f211c5fa5d2b43fc4c7ef078fa5630f0656a3b5c0

Observation 84725537-852c-47a6-b27d-5ac3754f5fe3 · inbound

Towards Reliable C-to-Rust Translation with Rule-Guided Reasoning and Reinforcement Learning cites this paper.

Towards Reliable C-to-Rust Translation with Rule-Guided Reasoning and Reinforcement Learning FullStack Bench: Evaluating LLMs as Full Stack Coders

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T11:14:59.621613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:14:59.621613Z digest=sha256:ca117b19ceffb0ea7a5410662d190fbc839bf36320914c0f798e7e25d4c4de76

Observation 00da7b69-8be1-4464-ae4f-fdaf423434fb · inbound

SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD cites this paper.

SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD FullStack Bench: Evaluating LLMs as Full Stack Coders

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T10:42:39.936078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T10:42:39.936078Z digest=sha256:87cbb1268603df1dbc55df59af1810e94c17759fc8d2ff4446ae16dd57f956ca

Observation 646564d1-5bb4-46a7-ac84-198d06b87cc4 · inbound

Parameter Exploration for RLVR via Variational Learning cites this paper.

Parameter Exploration for RLVR via Variational Learning FullStack Bench: Evaluating LLMs as Full Stack Coders

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:12.751253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:12.751253Z digest=sha256:6d5ad4ad78ae22b9ccac93e5926c7ed32ff88b8d18f09bcd3ba3b84464679ec5