Pith. sign in

Paper Citation Record · LEDGER

What Did I Do Wrong? Quantifying LLMs' Sensitivity and Consistency to Prompt Engineering

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2406.12334.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.12334 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:02:21.562642Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T22:00:41.512733Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b40e3e17-40fd-43cc-8829-2a2244769ae2 · inbound

Agents Require Metacognitive and Strategic Reasoning to Succeed in the Coming Labor Markets cites this paper.

Agents Require Metacognitive and Strategic Reasoning to Succeed in the Coming Labor Markets What Did I Do Wrong? Quantifying LLMs' Sensitivity and Consistency to Prompt Engineering

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:21.562642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:21.562642Z digest=sha256:bf2af3e0c928dc1579b72d3295943751d4d494e338869b3b26946febeb0b00b3

Observation 4f02dd20-f97d-4038-8f17-d0972a1d0d27 · inbound

Evaluating Prompt Engineering Techniques for Accuracy and Confidence Elicitation in Medical LLMs cites this paper.

Evaluating Prompt Engineering Techniques for Accuracy and Confidence Elicitation in Medical LLMs What Did I Do Wrong? Quantifying LLMs' Sensitivity and Consistency to Prompt Engineering

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:14.027058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:14.027058Z digest=sha256:87309078c88462be769b373e2dd8fc46c3325a57d80516a1d254d160b5bd2275

Observation 8619f564-aa14-42b3-8b44-b34998980077 · inbound

CEA-LIST at CheckThat! 2025: Evaluating LLMs as Detectors of Bias and Opinion in Text cites this paper.

CEA-LIST at CheckThat! 2025: Evaluating LLMs as Detectors of Bias and Opinion in Text What Did I Do Wrong? Quantifying LLMs' Sensitivity and Consistency to Prompt Engineering

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T18:42:30.735057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:42:30.735057Z digest=sha256:a09183e1daeecf3b1fbe054837dded6f8ec222ff1d61efc597497ad0a308386e

Observation 69b5a0ef-14d0-4054-87cb-d1499b3b3125 · inbound

A Conceptual Framework for Requirements Engineering of Pretrained-Model-Enabled Systems cites this paper.

A Conceptual Framework for Requirements Engineering of Pretrained-Model-Enabled Systems What Did I Do Wrong? Quantifying LLMs' Sensitivity and Consistency to Prompt Engineering

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T16:34:09.624223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:34:09.624223Z digest=sha256:e36343e28567e35d9af55ba638a7495b7a7c79f63324e2eb31fc3bc2ec538e4c

Observation 4f199e73-cd56-421b-91c0-3e0df4fdbbd3 · inbound

Position: Intelligent Coding Systems Should Write Programs with Justifications cites this paper.

Position: Intelligent Coding Systems Should Write Programs with Justifications What Did I Do Wrong? Quantifying LLMs' Sensitivity and Consistency to Prompt Engineering

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T23:01:59.849579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:01:59.849579Z digest=sha256:4d8738b20214e5588b3eaf461f921ff22a5bf657dc96feae900b44a5b3b9b02d

Observation f040e31f-271f-434e-905f-a6f0fe38d6ab · inbound

Reinforcement Learning for Machine Learning Engineering Agents cites this paper.

Reinforcement Learning for Machine Learning Engineering Agents What Did I Do Wrong? Quantifying LLMs' Sensitivity and Consistency to Prompt Engineering

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T12:24:02.310282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:24:02.310282Z digest=sha256:e0a952dd4d09f6b15f4fa3e613e65856be10d916b83262777ccaa9a93165cb33

Observation ce73f51e-41f1-457e-a692-db9a25da401d · inbound

From scratch to silver: Creating trustworthy training data for patent-SDG classification using Large Language Models cites this paper.

From scratch to silver: Creating trustworthy training data for patent-SDG classification using Large Language Models What Did I Do Wrong? Quantifying LLMs' Sensitivity and Consistency to Prompt Engineering

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-04T19:27:04.900393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T19:27:04.900393Z digest=sha256:3256379e88fe9480799cb6c58d35e541a6c2fe687999996c2822cc08b5f19f4e

Observation 658d5516-63eb-40d5-8f06-36bb65bf298a · inbound

When the Code Autopilot Breaks: Why LLMs Falter in Embedded Machine Learning cites this paper.

When the Code Autopilot Breaks: Why LLMs Falter in Embedded Machine Learning What Did I Do Wrong? Quantifying LLMs' Sensitivity and Consistency to Prompt Engineering

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T17:23:21.474317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:23:21.474317Z digest=sha256:9338292292741529c6a25664991ece84d47b7a250ec8695128e2f6a6ae99d916

Observation 0e4f38a2-3ed7-4f93-a100-f8d063abd32e · inbound

Position: AI Evaluations Should be Grounded on a Theory of Capability cites this paper.

Position: AI Evaluations Should be Grounded on a Theory of Capability What Did I Do Wrong? Quantifying LLMs' Sensitivity and Consistency to Prompt Engineering

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-21T22:00:41.515755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-21T21:57:55.834632Z digest=sha256:7347ec14ccec043bfa8f448b3b3f1c3e97e12a8adf87e60db4aaa34dbcd1e524

Observation 6dbf4a2b-625a-4177-bcbd-c9427ae46443 · inbound

PIAST: Rapid Prompting with In-context Augmentation for Scarce Training data cites this paper.

PIAST: Rapid Prompting with In-context Augmentation for Scarce Training data What Did I Do Wrong? Quantifying LLMs' Sensitivity and Consistency to Prompt Engineering

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-16T23:18:40.050451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T23:15:52.444217Z digest=sha256:8cb3fc272975a9fee6f4f985cc6936ea6691abbca8827ea97a4eb53b2a3b1270

Observation 9ef101a1-d5e1-4a5b-911a-080263920c69 · inbound

Information-Consistent Language Model Recommendations through Group Relative Policy Optimization cites this paper.

Information-Consistent Language Model Recommendations through Group Relative Policy Optimization What Did I Do Wrong? Quantifying LLMs' Sensitivity and Consistency to Prompt Engineering

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:08:36.166115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T22:05:15.926436Z digest=sha256:35c628800de3e890fd5aab28c3d60c6c909dc983dd5efcc690bad5f4d1ff430d

Observation b8fb136c-760a-4ab8-af63-520cecba3002 · inbound

Similar Pattern Annotation via Retrieval Knowledge for LLM-Based Test Code Fault Localization cites this paper.

Similar Pattern Annotation via Retrieval Knowledge for LLM-Based Test Code Fault Localization What Did I Do Wrong? Quantifying LLMs' Sensitivity and Consistency to Prompt Engineering

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:00:56.656182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T02:58:17.585369Z digest=sha256:0044c28d9d744f1e54cfd9e889c946ed142d033a090dcff3cf093499cd55696c