Pith. sign in

Paper Citation Record · LEDGER

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions

As of 18 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 11 inbound Pith citation observations for arXiv:2507.20439.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.20439 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T13:38:53.742442Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T14:46:30.803463Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

43 of 43 outbound references displayed

  • verified exact3
  • verified fuzzy1
  • unresolved38
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 9df98d93-dda8-4a14-b4d0-f058be2027d3 · outbound

This paper cites IEEE Recommended Practice for Software Requirements Specifications.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions IEEE Recommended Practice for Software Requirements Specifications

Reference 1

Resolution
verified exact
raw_fallback, observed 2026-08-06T13:38:55.205137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:38:47.649609Z digest=sha256:3179e0875895a382da7c7f65daa5d587fc5772baf11ebb08c86ddf2b9a83b8d2

Observation a27cb455-ca01-4845-932d-4b37bcdd76f2 · outbound

This paper cites DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T13:38:47.768131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:38:47.768131Z digest=sha256:caa7436b530e02d824a32e0af97b7b30e93adc6af1f21cacb285a38f533f839c

Observation 23de4ca7-f7e8-4054-8996-c4466765f738 · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:39:00.084889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:38:47.886462Z digest=sha256:2df5cf9391dd3ad7424e50f45d7b9aab4f667f720e38b2bdea3a084a2666a80a

Observation c6a564b6-62bf-444e-bc17-df086f55d3ee · outbound

This paper cites Program Synthesis with Large Language Models.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Program Synthesis with Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T13:38:48.144648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:38:48.144648Z digest=sha256:241b8a50dd32dfaed79174d854913f6ba86bd7a314570a9649de3505e48912b3

Observation 0fac6cd1-51e4-41fa-9b8b-0bb56846416b · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:38:59.929664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:38:48.298871Z digest=sha256:da8fccbcf1488ccfa3d2ac33357fe3fd4c8d53e27d3be7a06cdde7281b256dda

Observation 87fd0afb-5e94-4171-9253-850d8306a7ec · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:38:59.673488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:38:48.420423Z digest=sha256:91ea18b69b1cee4116269406ae6d1e418c214049ffa2a9e6b6fc090b3a4f9266

Observation eff87302-791c-471f-ba3b-c2a38e176e51 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Evaluating Large Language Models Trained on Code

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T13:38:48.655287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:38:48.655287Z digest=sha256:8dcd27f8e0f71591bd4e5b3221c44a6bcb8816eba16920ede4337d9b7fca6eea

Observation b13455d5-6388-4554-8959-e2e884814991 · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T13:38:48.816979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:38:48.816979Z digest=sha256:ceb6e1110123bc97e1f79954fb927ee2986db9e9ba93791bad2d69160fb7edae

Observation b6a20980-cb7d-4729-a40e-20159448581b · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T13:38:49.100177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:38:49.100177Z digest=sha256:e46da99d3f4400271d6da08ada4c7cc1304477629f92f0f4af693d69659eef84

Observation 365d18e0-4cd8-49f8-9f39-dde997ede50a · outbound

This paper cites Systematic Evaluation of GPT-3 for Zero-Shot Personality Estimation.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Systematic Evaluation of GPT-3 for Zero-Shot Personality Estimation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T13:38:49.217874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:38:49.217874Z digest=sha256:a03622a191caf8ae2e2e2551765cd67e3e8d69e35ba6b1193adcf7e3b9255789

Observation 161d3724-ede3-4f83-94cc-e2353ac70e94 · outbound

This paper cites Measuring Coding Challenge Competence With APPS.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Measuring Coding Challenge Competence With APPS

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T13:38:49.326645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:38:49.326645Z digest=sha256:fc0bd59f0d94fd22ac448b9a6279265340b570d3879585c31fdf35d77895fce0

Observation 725e18f3-d124-42a1-99bc-91ff4d39a635 · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T13:38:49.436262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:38:49.436262Z digest=sha256:6d16446783ce64e7692b22a1ae1bb550734cdd21990f894b6a9eff643a5786b5

Observation 8c056224-3794-45f4-8d08-3fad5c1eee98 · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T13:38:49.547770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:38:49.547770Z digest=sha256:54338fa014ad908276d38973a481d8dfc44763240a57897dca0485831a8ebe1c

Observation ca011c1e-cbf9-4062-8c5e-770820f1c436 · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:38:59.481548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:38:49.766965Z digest=sha256:98624f990386787f6b08eac46dc56a53fdfa345108c86f5ce62b70155c194d1e

Observation 2fed75e1-c953-40cd-b241-b8c68755a4c1 · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:38:59.337232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:38:49.914221Z digest=sha256:124eb6d32671c1df5636158bc207c1f4b0a7e81dcd97c7177495482975ca0a97

Observation f85e91c9-e344-4da7-833b-a4e2fd9a7e11 · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:38:59.122903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:38:50.005883Z digest=sha256:5e574c5ec06237d1e774743da5d529a4d88a99c8dea3dbbf1ef25e1be2f79f34

Observation ee94a9d0-b71b-42cf-8498-dc7de55822a9 · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:38:58.949258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:38:50.063162Z digest=sha256:7277fc015c37a3d1a0595e3ae0b08416f68984311c38a73e7f33f102fb002587

Observation 5358355d-f85f-46c2-ab56-11a97fdb9406 · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T13:38:50.268739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:38:50.268739Z digest=sha256:617ca3dd13efaa76040d400a9ae25b12fd4f3249c7cb7b608e29497d35c95168

Observation 20e2a215-9354-4920-b9b5-6e3907489b99 · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 22

Resolution
metadata mismatch
raw_fallback, observed 2026-08-06T13:38:54.736511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:38:50.390088Z digest=sha256:7f9bf5677e4f8b3128a56ea30430b70a68d78bfd415ab4ea6929d7d8799ec3d2

Observation a18c74fa-41c2-4b64-9e60-c3939ed13364 · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T13:38:50.555980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:38:50.555980Z digest=sha256:35c7160abdb61c0559b04f61b8a893fac91183b17503aead9c3eaf785b6d8fb1

Observation aface219-411c-4fbe-9a10-cfbfea79d036 · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:38:58.616747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:38:50.824677Z digest=sha256:50d1ef64bf8aecb9c17162d41555a4b24b67286441d0642cdd1d79cc5c2d8c85

Observation 982fa63d-cedf-4c65-9275-22cf2d04f7b1 · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 25

Resolution
verified exact
doi, observed 2026-08-06T13:38:54.019782Z

Source-reported events for the cited work

correction dated 2022-04-04. Source: crossref record 10.1007/s00766-022-00378-4->10.1007/s00766-021-00367-z:correction, observed 2026-07-11T02:57:50.434999+00:00. This notice travels one citation hop only.

source=pdf_text observed=2026-08-06T13:38:50.698361Z digest=sha256:db0ff7f0011bb70532414e658b6d14346ae4978c1d5a4519a7403073c697ec40

Observation fa8f5d74-5fa1-4823-8f1e-083426e837c7 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T13:38:51.181674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:38:51.181674Z digest=sha256:dfc24b304c5f46c0b03095c1658024f58d438435bd2f471681f9948df5771bc7

Observation 39c13da2-363a-4edd-8061-d9b3ceb02ba4 · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:38:58.387181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:38:51.043563Z digest=sha256:57b99cb0b79f4d4b9540c8d07f06b65e6333d0ae457da5dc9bc9914f5377abe7

Observation 45e8c2d1-c011-4765-8ee6-54bf5031f8ad · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:38:57.675628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:38:51.425304Z digest=sha256:6f84e41f40eb91e2283d2e4e3c7a1633219a812ac6741c0c75268f20117d71b3

Observation 74c4d79d-8833-4016-9175-2a3328b65748 · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:38:58.007097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:38:51.285484Z digest=sha256:2c00f9455f24556b199a13a219174de183047f3a73b1f73fd95cd36c202a242a

Observation e34977af-ddde-4c41-b37f-5331bc09f3cb · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:38:57.302392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:38:51.687667Z digest=sha256:d24a233c6d6898af5df0ca6100770489b5a2d52987675c2308977fe5c5a41819

Observation dc21c2dd-ad0a-4abe-bfc5-21ccb202e113 · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:38:57.436680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:38:51.583331Z digest=sha256:d86b444541db29981b30bf653f6e03f60407d69e1df08bd29a991a1f61ae9fd0

Observation 7f9f76a7-6b47-4e1d-8575-5cde5af1b13c · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:38:57.053010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:38:51.934242Z digest=sha256:5413436037e57cf121b7df87456e541246fc764f1d8f4d6e01d1c76cb90aa5f6

Observation 7362a76e-f3f8-4a6d-b97a-5d3d60c1cc65 · outbound

This paper cites Code Llama: Open Foundation Models for Code.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Code Llama: Open Foundation Models for Code

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T13:38:51.815907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:38:51.815907Z digest=sha256:c338e9b3e316ee8713a6bce6a58ce17f9f804f0890d5ed742ced5a7220beff4a

Observation 61318ea6-a0f4-41c4-8f1b-7d14faacdcbc · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:38:56.559707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:38:52.256132Z digest=sha256:a74a5a6b482627ae9298b56fa1e519fa485c6191355c31c37a30b6925ecd9a9d

Observation 2542d470-b572-492f-803d-c90b569f308b · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:38:56.839141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:38:52.141096Z digest=sha256:5964e99798adf05fa21db821b4378019250062d802466b61be5391eddef54896

Observation 353dce04-6766-4e73-870b-f95226b2934c · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:38:56.312009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:38:52.526338Z digest=sha256:47571ad55506f8f2949be0dbbfa61449afab9a821daa1bfcb57d428dc248c97e

Observation e37dd0df-15a1-48b8-85a3-2fda37ee3973 · outbound

This paper cites Recent Advances in Software Effort Estimation using Machine Learning.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Recent Advances in Software Effort Estimation using Machine Learning

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-08-06T13:38:54.305710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:38:52.374590Z digest=sha256:0717c7cd813778f2fdb0740197184e41ed7ab5ebfd8ca778898f59a9ac3d0956

Observation 0cfa7f88-6cea-4a69-8188-34069b47c5d3 · outbound

This paper cites Is Your AI-Generated Code Really Safe? Evaluating Large Language Models on Secure Code Generation with CodeSecEval.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Is Your AI-Generated Code Really Safe? Evaluating Large Language Models on Secure Code Generation with CodeSecEval

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T13:38:52.828980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:38:52.828980Z digest=sha256:615cd5076e91a17516fcaaad2609b1a5f9190c81d2c825d12604b99f21b3a815

Observation 8d00f9a7-7715-4037-9396-90903b31b54f · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:38:56.058687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:38:52.695032Z digest=sha256:c7dd0218c001ef544664395c7602aa0317874a74f71fec4dbd820188e440548e

Observation cbc77a68-d9ca-4d9d-a33d-29c36136f676 · outbound

This paper cites DeceptPrompt: Exploiting LLM-driven Code Generation via Adversarial Natural Language Instructions.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions DeceptPrompt: Exploiting LLM-driven Code Generation via Adversarial Natural Language Instructions

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T13:38:53.175435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:38:53.175435Z digest=sha256:d343fbaa20e960d2af27eeee1f8f7c34f62570e5a7f1016e90807959b4829b14

Observation 8a3ba20e-0ce5-4251-8cc3-30bc68fb2477 · outbound

This paper cites ReCode: Robustness Evaluation of Code Generation Models.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions ReCode: Robustness Evaluation of Code Generation Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T13:38:53.002137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:38:53.002137Z digest=sha256:d75ef3384e5086db1e8e2a1061268e023c0453afa1b95fdff95a2d5c97fec471

Observation 71430e1e-2df9-4de9-8f9f-1a75c545da80 · outbound

This paper cites LiveCodeBench Pro: How Do Olympiad Medalists Judge LLMs in Competitive Programming?.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions LiveCodeBench Pro: How Do Olympiad Medalists Judge LLMs in Competitive Programming?

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T13:38:53.452180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:38:53.452180Z digest=sha256:9a523b2cb0eddcc1fdea980695f3a4dfa39bdf48fca3b5c51d6d606eed01ce6b

Observation f86c9d45-f912-4324-9e17-cf9f78ff88d5 · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:38:55.836809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:38:53.396905Z digest=sha256:53ca351fdd838ab93d6b3bb25ea15b196aefa4897cc506b765ff8f1e0502633c

Observation a1570c58-f6a5-43b8-b3bc-45a685e623fc · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:38:55.486547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:38:53.742442Z digest=sha256:03f5685025164107c0abfeca052f0770d716e418e56ef711faf78a36685e48ed

Observation 365c216c-8b0a-48cd-a867-06d1e6cf07da · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:38:55.691351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:38:53.611065Z digest=sha256:f2711e37dd30461823c05abd0ed0f2db0e94a6471ca111419c1d0026834f6cb3

Observation f0cc826d-8542-49ee-b570-66ae34b1b1de · outbound

This paper cites Software: Practice and experience 52, 1 (2022), 39–65.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Software: Practice and experience 52, 1 (2022), 39–65

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:38:58.753242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:38:50.181171Z digest=sha256:89b01c3791651ec9febc0950a04343d69d7829fcf532ccd1956c55d430d5b81c

Pith citing papers

Observation 198ab81f-943e-426d-98b9-64b4c293c359 · inbound

Defective Task Descriptions in LLM-Based Code Generation: Detection and Analysis cites this paper.

Defective Task Descriptions in LLM-Based Code Generation: Detection and Analysis When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:16:34.336735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T03:07:06.924437Z digest=sha256:6f1dbbbe693b63c7bb0f82bec9dd2e71096826f774ba7401771fd7899378cabb

Observation 70807c01-e5e1-4017-aeca-80197e2fbaef · inbound

When Prompt Under-Specification Improves Code Correctness: An Exploratory Study of Prompt Wording and Structure Effects on LLM-Based Code Generation cites this paper.

When Prompt Under-Specification Improves Code Correctness: An Exploratory Study of Prompt Wording and Structure Effects on LLM-Based Code Generation When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:17:08.166144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T03:00:26.137401Z digest=sha256:8429dcea438c1a0b295e391ff3ea119ed21377b0889f9fed2be8c9a48aa514c4

Observation 77d8fa42-aa49-4696-b4b2-df3eb3d40d82 · inbound

ClarifyCodeBench: Evaluating LLMs on Clarifying Ambiguous Requirements for Code Generation cites this paper.

ClarifyCodeBench: Evaluating LLMs on Clarifying Ambiguous Requirements for Code Generation When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-02T08:46:48.648730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-02T08:39:01.987915Z digest=sha256:e3888be3e4c6a22e37ba0dc11917c765766dc177b03e7541fa8a28e8f927ec21

Observation fe9cb4ff-ad6e-4b67-b6d8-139a902c8018 · inbound

Underspecification does not imply Incoherence: The Risks of Semantic Collapse in Coding Models cites this paper.

Underspecification does not imply Incoherence: The Risks of Semantic Collapse in Coding Models When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-03T09:07:47.017340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-03T09:01:48.501745Z digest=sha256:6199552a4f74f64ee22689b5f83cebe97abd71ee3d94238dd8bb327e5ca54126

Observation b9f16e6d-8e35-411a-8ed3-6728eba63bfd · inbound

Guiding Human Validation of LLM-Generated Code via Verifiable Literate Programming cites this paper.

Guiding Human Validation of LLM-Generated Code via Verifiable Literate Programming When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-07-03T08:47:49.713356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-03T08:42:04.804042Z digest=sha256:4c5a69a279f4da898377418bd760dcdb6f38413c8888deac50cc92aed51fe92a

Observation bb9764c9-2676-429e-8063-7ebddb3de03f · inbound

Obey, Diverge, Collapse: Blind Obedience to Incorrect Instructions Drives Code LLMs to Irrecoverable Code Semantic Collapse cites this paper.

Obey, Diverge, Collapse: Blind Obedience to Incorrect Instructions Drives Code LLMs to Irrecoverable Code Semantic Collapse When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-11T17:45:46.872944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T17:45:46.872944Z digest=sha256:1071765ffc189654a874b74730306e843f0814422b5c0b78fed6d1a99d0b11d8

Observation 158d4e41-4d94-43b7-a6fc-b6a5bb766ec5 · inbound

From Failing to Passing: Evolving Natural Language Prompt Optimization Rules for LLM Code Generation cites this paper.

From Failing to Passing: Evolving Natural Language Prompt Optimization Rules for LLM Code Generation When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-11T08:31:24.850506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T08:31:24.850506Z digest=sha256:67409e010b5d81cedd5483052b0acdaa00449c0b629bc62355f1ec9a670ff970

Observation 1bd8e22d-9906-4da1-90b9-62a267294cc7 · inbound

On the risk of coding before testing: An empirical study on LLM-based test generation workflow cites this paper.

On the risk of coding before testing: An empirical study on LLM-based test generation workflow When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions

Reference 51

Resolution
unresolved
no resolver link, observed 2026-07-11T08:12:23.512737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T08:12:23.512737Z digest=sha256:ef86017711755e32540e6befd2aa168178dc1b47f6fd7116de76dd41d4f9f087

Observation 4d2e8f86-3a11-4e4b-bc0e-b731365ec346 · inbound

Automatically Evolving Prompt Guidelines for Task-Specific Optimization cites this paper.

Automatically Evolving Prompt Guidelines for Task-Specific Optimization When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T14:46:30.803463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:46:30.803463Z digest=sha256:0944583deee6edd54452ce5b6c23e955a105db4bbeee6895258ff3526a346205

Observation c853259f-c273-4ff6-b6a2-3bac8e2e1e11 · inbound

AssumptionMiner: Extracting, Tracing, and Revising Implicit Assumptions in LLM Code Generation cites this paper.

AssumptionMiner: Extracting, Tracing, and Revising Implicit Assumptions in LLM Code Generation When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T04:15:29.179347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:15:29.179347Z digest=sha256:ea48c51d8ab3a6091d9f66c8e36bc80e64bc74fc9f3eb65061b5349b3d6fb8d5

Observation 2b7984f4-f449-4cc1-9e8d-bd32b8930620 · inbound

VClare: Resolving Imperfect Specifications in LLM-Based Verilog Generation cites this paper.

VClare: Resolving Imperfect Specifications in LLM-Based Verilog Generation When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-30T23:35:19.791382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T23:35:19.791382Z digest=sha256:24f3979d2ac3364cad04575275ecc1e2e4e4fa95396f453476ea9b69ac8f8151