Pith. sign in

Paper Citation Record · LEDGER

NoCode-bench: A Benchmark for Evaluating Natural Language-Driven Feature Addition

As of 11 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 6 inbound Pith citation observations for arXiv:2507.18130.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.18130 v3

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T14:43:14.383453Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T16:27:41.491607Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T22:44:02.092336Z

Reference resolution

47 of 47 outbound references displayed

  • verified exact0
  • verified fuzzy4
  • unresolved43
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 733c37ce-153f-4b08-817d-950d81e2ae7b · outbound

This paper cites an unresolved cited work.

NoCode-bench: A Benchmark for Evaluating Natural Language-Driven Feature Addition Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:43:14.893845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T14:43:14.262991Z digest=sha256:87e96722f9fb40480dbfd55d29b5b3a7003e5d273a8dfe79152845130a20ef92

Observation bebbfaac-7e62-4b8b-874d-37cc30faf89e · outbound

This paper cites an unresolved cited work.

NoCode-bench: A Benchmark for Evaluating Natural Language-Driven Feature Addition Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:43:14.885765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T14:43:14.266264Z digest=sha256:dfcc90e53d2ce63252caa36bd1823a984b0a023e989e8b1da50f6a0a017742ae

Observation 03a8c34a-b913-4f01-9709-56be5d71bfd9 · outbound

This paper cites an unresolved cited work.

NoCode-bench: A Benchmark for Evaluating Natural Language-Driven Feature Addition Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:43:14.877792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T14:43:14.269276Z digest=sha256:bd871c0519c56a3833bc27e7b9298007c5cc710d06bb9faaa83eec82bea85dc5

Observation b8f866f0-9607-428c-8a92-6e5aebf4efa3 · outbound

This paper cites an unresolved cited work.

NoCode-bench: A Benchmark for Evaluating Natural Language-Driven Feature Addition Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:43:14.868958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T14:43:14.271948Z digest=sha256:10eb693d8bde8df8c9b69923463a5de6b83eeeeac9fbd2c88be59aff23dd1b28

Observation 4964010a-830a-4784-9731-37e7bf017f55 · outbound

This paper cites an unresolved cited work.

NoCode-bench: A Benchmark for Evaluating Natural Language-Driven Feature Addition Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:43:14.859921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T14:43:14.274534Z digest=sha256:5be96e3d977addff4805ef0a65ecd6d95885d79bbb8c685ac98e6f29ff4ad204

Observation 5d258ce8-d459-4ccc-9760-80404074deb9 · outbound

This paper cites an unresolved cited work.

NoCode-bench: A Benchmark for Evaluating Natural Language-Driven Feature Addition Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:43:14.850015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T14:43:14.277804Z digest=sha256:c086f181423f753a017c861daed454240ef749f05f62c427cf4396c9cf5fc96d

Observation 10db8e0a-c7f4-454b-b307-7d490531d042 · outbound

This paper cites an unresolved cited work.

NoCode-bench: A Benchmark for Evaluating Natural Language-Driven Feature Addition Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:43:14.841140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T14:43:14.280520Z digest=sha256:9a5fa0025167cd9fe88c29e0d21bd4c0e78bc8228003352f95852655ee2466df

Observation 440bde2c-6ca4-47de-ae52-a372a0e47903 · outbound

This paper cites an unresolved cited work.

NoCode-bench: A Benchmark for Evaluating Natural Language-Driven Feature Addition Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:43:14.832385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T14:43:14.282964Z digest=sha256:540cb64fd343b273f7887c9756e1edd99a7d736fd89ccf75d391bc03f00c8ea6

Observation a5b97720-58b2-4b57-9c8b-d165f48ca220 · outbound

This paper cites an unresolved cited work.

NoCode-bench: A Benchmark for Evaluating Natural Language-Driven Feature Addition Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:43:14.823477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T14:43:14.285407Z digest=sha256:4840f62ecdf4c95ff3370aecb701d6b9c25edac5209443672cc70e3e4ba52221

Observation 68481218-e7dd-4b5e-8105-0ae4bb0ad055 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

NoCode-bench: A Benchmark for Evaluating Natural Language-Driven Feature Addition Evaluating Large Language Models Trained on Code

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T14:43:14.287777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:43:14.287777Z digest=sha256:7b4e327a7cadc1beff3d5d84774cfc59b853d48dbd425349d96658481fb687c2

Observation 485b04a1-c8a1-444d-94ba-f111840a68c1 · outbound

This paper cites Jimenez, John Yang, Leyton Ho, Tejal Patwardhan, Kevin Liu, and Aleksander Madry.

NoCode-bench: A Benchmark for Evaluating Natural Language-Driven Feature Addition Jimenez, John Yang, Leyton Ho, Tejal Patwardhan, Kevin Liu, and Aleksander Madry

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:43:14.812635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T14:43:14.290538Z digest=sha256:30b7e96f1dd39a52a63ae91a73905d234aa68915cf153191617774c6a9f12d66

Observation 64461c92-eb9a-4e56-989d-e58a0146b5d2 · outbound

This paper cites an unresolved cited work.

NoCode-bench: A Benchmark for Evaluating Natural Language-Driven Feature Addition Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:43:14.802275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T14:43:14.293016Z digest=sha256:84c1aa8a6377ad265cfd96619d036f41fe97c80de75c9e523ce6c8f5e3e20d89

Observation 54c3eb89-88e0-4f4b-bcf3-718f78a8dc94 · outbound

This paper cites an unresolved cited work.

NoCode-bench: A Benchmark for Evaluating Natural Language-Driven Feature Addition Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:43:14.792205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T14:43:14.295380Z digest=sha256:d43bcee362ecfdfb5e1df17434d384878fd3947cff1fa527895b8a9a71a6c876

Observation bc34194d-c587-4232-810a-babaf5a139c7 · outbound

This paper cites an unresolved cited work.

NoCode-bench: A Benchmark for Evaluating Natural Language-Driven Feature Addition Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:43:14.781828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T14:43:14.297579Z digest=sha256:ba55a70080811d94f863e24a31011b3263ee5df216a1f825c44cda602de77368

Observation ad7b1351-1708-4be8-a507-a0e05ddee227 · outbound

This paper cites an unresolved cited work.

NoCode-bench: A Benchmark for Evaluating Natural Language-Driven Feature Addition Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:43:14.772113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T14:43:14.299839Z digest=sha256:4b0b3f66509c1c7781099123fc49cd234d8b0368f4c94ee15d5dd16c376ca37b

Observation d6641181-2442-4712-914e-1183641e21d4 · outbound

This paper cites an unresolved cited work.

NoCode-bench: A Benchmark for Evaluating Natural Language-Driven Feature Addition Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T14:43:14.302175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:43:14.302175Z digest=sha256:4675f60f9f43fba792834a141b7f2560cc1f754390a283a390a697cb91fb343a

Observation 37f60e10-1686-4a98-b436-c97efba55b5e · outbound

This paper cites an unresolved cited work.

NoCode-bench: A Benchmark for Evaluating Natural Language-Driven Feature Addition Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:43:14.757485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T14:43:14.304757Z digest=sha256:73639bd8b1dc0418e92888a29cee66a76b9ea703d252968bb4ccd1d1b456a12e

Observation 2f06833c-0bf2-4cb8-b758-61b7536f29ed · outbound

This paper cites an unresolved cited work.

NoCode-bench: A Benchmark for Evaluating Natural Language-Driven Feature Addition Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:43:14.748745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T14:43:14.307311Z digest=sha256:9f800ae05a936185638cbabb5d206cccbcdadf1c386d2f4063477b38a55072ae

Observation 9cf92708-93f2-40e8-8a85-a2c6df496ffb · outbound

This paper cites FEA-Bench: A Benchmark for Evaluating Repository-Level Code Generation for Feature Implementation.

NoCode-bench: A Benchmark for Evaluating Natural Language-Driven Feature Addition FEA-Bench: A Benchmark for Evaluating Repository-Level Code Generation for Feature Implementation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T14:43:14.309928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:43:14.309928Z digest=sha256:6aa395fc74317a5cbbea5e99b68639686427d8dd18ae210aa118c019a5a9fa0b

Observation 270c441a-3e03-49f4-b4cd-a20492db771c · outbound

This paper cites an unresolved cited work.

NoCode-bench: A Benchmark for Evaluating Natural Language-Driven Feature Addition Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:43:14.738598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T14:43:14.312861Z digest=sha256:a565bb9b0c6f5da2d946be385e2e37c78b96b3892cb37ddb6512d84b049ba7cc

Observation 6af1df62-334d-450e-b42d-979615867bf5 · outbound

This paper cites an unresolved cited work.

NoCode-bench: A Benchmark for Evaluating Natural Language-Driven Feature Addition Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T14:43:14.315412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:43:14.315412Z digest=sha256:ecf219e493273920cb19decdb09b156e307d7060e3cd3a34680a138edf79b092

Observation 2493913b-39b1-4a01-b21b-371b79fac0b0 · outbound

This paper cites GPT-4o System Card.

NoCode-bench: A Benchmark for Evaluating Natural Language-Driven Feature Addition GPT-4o System Card

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T14:43:14.317738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:43:14.317738Z digest=sha256:ef42bd43f90fc634238f37f342be07c7b983b663a24ab718ae69021d5f6337bd

Observation 5f04209c-a2a2-4536-8354-e8451f568687 · outbound

This paper cites an unresolved cited work.

NoCode-bench: A Benchmark for Evaluating Natural Language-Driven Feature Addition Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:43:14.729163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T14:43:14.320122Z digest=sha256:8c297b94326ef1b07d84395863da7825b305b8a6efeaaa8fd5fe80511f6b7761

Observation f9aac29c-d0d9-4621-9e85-6bbf893b4aff · outbound

This paper cites an unresolved cited work.

NoCode-bench: A Benchmark for Evaluating Natural Language-Driven Feature Addition Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:43:14.720124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T14:43:14.323007Z digest=sha256:7d91bee75403957bb60c57bc7be3cd050a5d5c40d448949a8fe265e00177c4d5

Observation 0b5e861b-bf35-4e54-ac1f-0f49fda45e15 · outbound

This paper cites an unresolved cited work.

NoCode-bench: A Benchmark for Evaluating Natural Language-Driven Feature Addition Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:43:14.711179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T14:43:14.326655Z digest=sha256:189d387f66e57950c8f131343e288170ffdfb021f49a3672142f789c4d801683

Observation 322fea1a-391a-4329-ab43-fc2bbd3d5a4c · outbound

This paper cites an unresolved cited work.

NoCode-bench: A Benchmark for Evaluating Natural Language-Driven Feature Addition Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:43:14.701045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T14:43:14.329136Z digest=sha256:f513153ae7d20bbdae636d116ad8ce95ca33b2237063269ba335a6bec1cc0a8f

Observation e1a90696-cf96-40f1-80f2-cd99dfa1be90 · outbound

This paper cites an unresolved cited work.

NoCode-bench: A Benchmark for Evaluating Natural Language-Driven Feature Addition Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T14:43:14.331844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:43:14.331844Z digest=sha256:0e8168d66cf6b33a171911cc6dbf865523b6dc892697a05752a3ab67ed1fb0d8

Observation 9db39d24-7605-4fab-ab05-23178030fac3 · outbound

This paper cites SWE-PolyBench: A multi-language benchmark for repository level evaluation of coding agents.

NoCode-bench: A Benchmark for Evaluating Natural Language-Driven Feature Addition SWE-PolyBench: A multi-language benchmark for repository level evaluation of coding agents

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T14:43:14.337616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:43:14.337616Z digest=sha256:bad3b3ba338c614455cdd66d35c55021115b8feeafd52b428ca429b7b1467d95

Observation bd6c01a8-7067-4071-a951-b66c888fa0d9 · outbound

This paper cites Evaluating Agent-based Program Repair at Google.

NoCode-bench: A Benchmark for Evaluating Natural Language-Driven Feature Addition Evaluating Agent-based Program Repair at Google

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T14:43:14.340240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:43:14.340240Z digest=sha256:58c4eb3735c79918e8402bf5a1fcfa502879c07b21b655e5220c3cd53243f9ec

Observation b55bcfa7-9faf-4e34-9566-c49349fde958 · outbound

This paper cites SpecRover: Code Intent Extraction via LLMs.

NoCode-bench: A Benchmark for Evaluating Natural Language-Driven Feature Addition SpecRover: Code Intent Extraction via LLMs

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T14:43:14.343532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:43:14.343532Z digest=sha256:c82600d127d3012fdfab7d7dc1ee3df1cfc8dfc70ed74b066165a40de499d180

Observation b2957699-7d72-4d9b-ac49-49280ae48909 · outbound

This paper cites an unresolved cited work.

NoCode-bench: A Benchmark for Evaluating Natural Language-Driven Feature Addition Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:43:14.677931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T14:43:14.346525Z digest=sha256:f90b0bcf9b79e5f722e40919ead42ab963f4769db5d595329038f86becabbaee

Observation 1669ad58-ac92-46bc-9656-158f66f16662 · outbound

This paper cites an unresolved cited work.

NoCode-bench: A Benchmark for Evaluating Natural Language-Driven Feature Addition Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:43:14.669410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T14:43:14.348970Z digest=sha256:f0ba87995fb6a59e0c27c6d18c74cd0e9c65e268eb5327f87de988ab9a4d65ab

Observation a1432d15-69e0-4923-b7f6-bbf8bda87d25 · outbound

This paper cites Xu, Xiangru Tang, Mingchen Zhuge, Jiayi Pan, Yueqi Song, Bowen Li, Jaskirat Singh, Hoang H.

NoCode-bench: A Benchmark for Evaluating Natural Language-Driven Feature Addition Xu, Xiangru Tang, Mingchen Zhuge, Jiayi Pan, Yueqi Song, Bowen Li, Jaskirat Singh, Hoang H

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:43:14.660245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T14:43:14.351356Z digest=sha256:7d2374743e4bffad1b972b0ed8edcd8bf5f6577793a0b0b92bc5a6f68cd42a44

Observation db9d1249-d622-49c4-9bfe-2a2cf1fb15f3 · outbound

This paper cites an unresolved cited work.

NoCode-bench: A Benchmark for Evaluating Natural Language-Driven Feature Addition Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:43:14.650626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T14:43:14.353838Z digest=sha256:0ad344439316471e7fc4ae409c19b89fa1f31760d411769d72cfe9db5a514db4

Observation 73f451ca-e7b1-4a5d-88d0-9cad2f3648f7 · outbound

This paper cites SWE-Fixer: Training Open-Source LLMs for Effective and Efficient GitHub Issue Resolution.

NoCode-bench: A Benchmark for Evaluating Natural Language-Driven Feature Addition SWE-Fixer: Training Open-Source LLMs for Effective and Efficient GitHub Issue Resolution

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T14:43:14.356628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:43:14.356628Z digest=sha256:50203df1de2364259d0ea67955f04193345e059fb68823464d6cf8a9a74ee1ba

Observation 210bc757-2461-4a61-b0a6-4d252afad372 · outbound

This paper cites an unresolved cited work.

NoCode-bench: A Benchmark for Evaluating Natural Language-Driven Feature Addition Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:43:14.641649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T14:43:14.359152Z digest=sha256:7dce61e586ef68276ebcbf9d2b4b01c29ecedb27d6e2fbf83c409a8199be7e63

Observation e6e58472-c8a4-4290-9fff-ecf0c5e4c6bc · outbound

This paper cites Hallucination is Inevitable: An Innate Limitation of Large Language Models.

NoCode-bench: A Benchmark for Evaluating Natural Language-Driven Feature Addition Hallucination is Inevitable: An Innate Limitation of Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T14:43:14.361465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:43:14.361465Z digest=sha256:3df126ae3a394492fc222bd971361af6b6216af72a2997a08c8eb9511d0bb67a

Observation 4cfed32c-e5fd-4c60-864e-3bf3519ed302 · outbound

This paper cites an unresolved cited work.

NoCode-bench: A Benchmark for Evaluating Natural Language-Driven Feature Addition Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:43:14.631898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T14:43:14.363955Z digest=sha256:e0a6affd99ac011d98312fad05ae657e08ed986af09d09df334f191f39c12a19

Observation 14353f41-7176-4593-b81a-5b9b4db3a407 · outbound

This paper cites Jimenez, Alex L.

NoCode-bench: A Benchmark for Evaluating Natural Language-Driven Feature Addition Jimenez, Alex L

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:43:14.622860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T14:43:14.366281Z digest=sha256:4cafbaa783692ce1505f2796eb76d89d83591ea7eef20a15a2963856c0e5cad6

Observation fece8940-7317-4167-8288-4c30d5e4ba22 · outbound

This paper cites SWE-smith: Scaling Data for Software Engineering Agents.

NoCode-bench: A Benchmark for Evaluating Natural Language-Driven Feature Addition SWE-smith: Scaling Data for Software Engineering Agents

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T14:43:14.368798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:43:14.368798Z digest=sha256:09b5a37c6fcc8df87dcabfe09b515dc28509ce1f1d4e0e21baed901739c7229a

Observation ecae2225-2c4f-4d45-9c2e-e786af9b3f88 · outbound

This paper cites an unresolved cited work.

NoCode-bench: A Benchmark for Evaluating Natural Language-Driven Feature Addition Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:43:14.613735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T14:43:14.371251Z digest=sha256:4a1bdd9e3df225bd9b6e81425e572d85dd308f84784bff937ca27f02ff3f1f79

Observation e3544be7-c3e1-4be7-b1de-138886a7b9ee · outbound

This paper cites an unresolved cited work.

NoCode-bench: A Benchmark for Evaluating Natural Language-Driven Feature Addition Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:43:14.605333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T14:43:14.373555Z digest=sha256:98341b836ca78d36dcbd19cac61008b2df50b6a2d29f3a48db4edef004918f32

Observation 263468f8-5862-4328-befc-d9ebd881a351 · outbound

This paper cites Multi-SWE-bench: A Multilingual Benchmark for Issue Resolving.

NoCode-bench: A Benchmark for Evaluating Natural Language-Driven Feature Addition Multi-SWE-bench: A Multilingual Benchmark for Issue Resolving

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T14:43:14.376031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:43:14.376031Z digest=sha256:6fb4ce240666b39d19c9ffe34596c49e28f4ec0f3689a732eac23134254fdab3

Observation e0ca7024-b7ec-4390-bc36-c3b0778c7dea · outbound

This paper cites SWE-bench-java: A GitHub Issue Resolving Benchmark for Java.

NoCode-bench: A Benchmark for Evaluating Natural Language-Driven Feature Addition SWE-bench-java: A GitHub Issue Resolving Benchmark for Java

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T14:43:14.378522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:43:14.378522Z digest=sha256:d09afc543c3e25975d9bd8cedda3f4286677286c73a9d9c93e853d3eb41c0771

Observation 6077e96b-8dbd-40f8-94f5-e2f0dd52f71c · outbound

This paper cites Siren's Song in the AI Ocean: A Survey on Hallucination in Large Language Models.

NoCode-bench: A Benchmark for Evaluating Natural Language-Driven Feature Addition Siren's Song in the AI Ocean: A Survey on Hallucination in Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T14:43:14.380986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:43:14.380986Z digest=sha256:a3331f0cdecbf97db252d80f8f66f0a0d2082658d7ec1c180ec5ad715a052245

Observation 5f1dadeb-d6c6-48f7-954e-6512f906a516 · outbound

This paper cites an unresolved cited work.

NoCode-bench: A Benchmark for Evaluating Natural Language-Driven Feature Addition Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:43:14.597165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T14:43:14.383453Z digest=sha256:8efd804002a895c195d45c7ed775bfa535b5effff8965af3c4cc5a5f14575435

Observation 4d1f2ff6-5169-46ec-98a8-f83355354aab · outbound

This paper cites In Proceedings of the 29th International Conference on Intelligent User Interfaces.

NoCode-bench: A Benchmark for Evaluating Natural Language-Driven Feature Addition In Proceedings of the 29th International Conference on Intelligent User Interfaces

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:43:14.686753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T14:43:14.334990Z digest=sha256:413d0e967952ac71d23e1b5c01cfda17f8a73af4053a7bf51a8e233b1334c36f

Pith citing papers

Observation 6334d273-62ac-4d2e-91b4-c23e2f895269 · inbound

Evaluating LLM Agents on Automated Software Analysis Tasks cites this paper.

Evaluating LLM Agents on Automated Software Analysis Tasks NoCode-bench: A Benchmark for Evaluating Natural Language-Driven Feature Addition

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-08-04T01:57:46.600006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T16:38:12.144252Z digest=sha256:0ac7a80482a497c19fa1138aa106da7389ca412a1dbeb0ac91791400cf7d808e

Observation 7ac59ec0-b3f7-423a-9513-8b21ce1d0c3b · inbound

Evaluating LLM Agents on Automated Software Analysis Tasks cites this paper.

Evaluating LLM Agents on Automated Software Analysis Tasks NoCode-bench: A Benchmark for Evaluating Natural Language-Driven Feature Addition

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T16:27:41.491607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:27:41.491607Z digest=sha256:ff3850f03fe8bc1e225653e4fab2c9e60492214a11ff8099f874b681bc85538c

Observation 9c2bfa89-9269-4709-a3e0-f35539c07c12 · inbound

Neurosymbolic Repo-level Code Localization cites this paper.

Neurosymbolic Repo-level Code Localization NoCode-bench: A Benchmark for Evaluating Natural Language-Driven Feature Addition

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-08-04T01:57:46.600006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T08:26:02.980824Z digest=sha256:2e0d82d76054ad5accbb310b55719f1f8d70a46ca391f27919fd3adcf9ab5982

Observation bb8d263a-4b43-4f1d-83ce-369018640cd0 · inbound

Reproduction Test Generation for Java SWE Issues cites this paper.

Reproduction Test Generation for Java SWE Issues NoCode-bench: A Benchmark for Evaluating Natural Language-Driven Feature Addition

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-08-04T01:57:46.600006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T16:56:33.346935Z digest=sha256:619b3bea0cbe0b9a22ceb5505013a2ef1b341c565871e5471c27e930651ef0cf

Observation 8b93b7b0-67e4-4050-85dd-184aaa9c23ac · inbound

Reproduction Test Generation for Java SWE Issues cites this paper.

Reproduction Test Generation for Java SWE Issues NoCode-bench: A Benchmark for Evaluating Natural Language-Driven Feature Addition

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-08-04T01:57:46.600006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-11T00:48:28.794195Z digest=sha256:c5e1b9c18e9fc6da0f525b086aae20b1ce289ee94456491d57d71ee2cb88447b

Observation 0dab70a1-1e79-422f-acb6-1289b7dec24e · inbound

Names Are All You Need: Effective and Safe Regression Test Selection for Python cites this paper.

Names Are All You Need: Effective and Safe Regression Test Selection for Python NoCode-bench: A Benchmark for Evaluating Natural Language-Driven Feature Addition

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-08-04T01:57:46.600006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-29T21:07:10.225387Z digest=sha256:6d37a2fa098760c33892f3318325d686b128eaae7e6bcf15031233b2eb5fdaba