Pith. sign in

Paper Citation Record · LEDGER

Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

As of 14 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 96 inbound Pith citation observations for arXiv:2310.11324.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.11324 v2

Coverage vector

measured 64 of 64 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-17T01:59:51.326493Z

measured 160 of 160 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 96 of 96 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:32:52.510712Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

64 of 64 outbound references displayed

  • verified exact10
  • verified fuzzy43
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch9

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 23c5f13b-058f-4f8b-97de-a2411d423f31 · outbound

This paper cites Tweet: Susan & I found MMLU performance jump 6-10 points in the 40s by formatting multiple choice as (A) not A in MMLU (for internal model).

Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Tweet: Susan & I found MMLU performance jump 6-10 points in the 40s by formatting multiple choice as (A) not A in MMLU (for internal model)

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:59:51.394625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T01:59:51.326493Z digest=sha256:4cbdd1563b6d4469fdabe85cf0cc2a1bf2f66113eaf2fa55d8844bda3551447e

Observation 956be0df-66bb-4914-a0db-8cd4e7b5039f · outbound

This paper cites Falcon-40B : an open large language model with state-of-the-art performance.

Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Falcon-40B : an open large language model with state-of-the-art performance

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:59:51.501347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T01:59:51.326493Z digest=sha256:e448eb9d9cd0880443df2ea8cea3a6e0ec77f9b875554134c8a2ddab88a4068b

Observation 925bbcdb-d037-43c3-b065-71c87e4f4038 · outbound

This paper cites An empirical evaluation of thompson sampling.

Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting An empirical evaluation of thompson sampling

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:59:51.503979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T01:59:51.326493Z digest=sha256:2ef8681fd095dfc9d5dba19f5a4baa2ba07fad481eee2fae2f0c43cff9112168

Observation ba4e1fcf-f107-4b51-b6b3-062d01b12596 · outbound

This paper cites Better hypothesis testing for statistical machine translation: Controlling for optimizer instability.

Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Better hypothesis testing for statistical machine translation: Controlling for optimizer instability

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:59:51.506685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T01:59:51.326493Z digest=sha256:3490df45856896ebc0264e0d2898c89163ac6c5ff1571d38bff0bd5cfa30d6da

Observation 88b0f953-5c8b-4220-9bcd-651eb1f61fe3 · outbound

This paper cites GPT 3.int8(): 8-bit matrix multiplication for transformers at scale.

Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting GPT 3.int8(): 8-bit matrix multiplication for transformers at scale

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:59:51.509487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T01:59:51.326493Z digest=sha256:a5c3746f7f7bcc72a7d01332ef69cd2d6ea545fd6230a0a63392231f65b91c2f

Observation 35fd58a4-464d-4701-b1f4-c35e5b1aa193 · outbound

This paper cites Openprompt: An open-source framework for prompt-learning.

Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Openprompt: An open-source framework for prompt-learning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:59:51.512258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T01:59:51.326493Z digest=sha256:7647af87f9b81695f0739e81ccb48b12d36a48044184e519dd0401c2cf55d973

Observation ef7a012f-8209-4e1e-935a-00baa84995a4 · outbound

This paper cites Measuring and improving consistency in pretrained language models.

Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Measuring and improving consistency in pretrained language models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:59:51.515516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T01:59:51.326493Z digest=sha256:45f7967d8bfa1a43aade3612244b48c8b06491069a85b46b6f16df845aaecd7a

Observation c5b09f3e-1956-450c-8160-8942eb53edba · outbound

This paper cites Deep reinforcement learning that matters.

Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Deep reinforcement learning that matters

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:59:51.519366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T01:59:51.326493Z digest=sha256:365b117106b7a299d2e7ee758750d4c2d052e9f4278eb380aa4254e2e7de9c94

Observation 6c7f1cdf-d166-42c1-980b-8e8b26b7cfe1 · outbound

This paper cites Editing models with task arithmetic.

Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Editing models with task arithmetic

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:59:51.523190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T01:59:51.326493Z digest=sha256:5d690900750cc8c31b9a51e0d1173ad465a30f8ffe2ef739b55dba76bfd972c2

Observation efb1adda-a288-4065-a733-9ed2f4c65e3c · outbound

This paper cites How can we know what language models know? Transactions of the Association for Computational Linguistics, 8: 0 423--438.

Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting How can we know what language models know? Transactions of the Association for Computational Linguistics, 8: 0 423--438

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:59:51.526558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T01:59:51.326493Z digest=sha256:84485debe9de3fbdf64d5bc0bb2ca66cc3d38bdbfb733863729b643a6eb8d453

Observation c17940ac-8712-49eb-9a48-df739b684ee4 · outbound

This paper cites Asymptotically efficient adaptive allocation rules.

Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Asymptotically efficient adaptive allocation rules

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:59:51.529807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T01:59:51.326493Z digest=sha256:c7b11d4ec2e4220360e339ecd2a442c64a94cafbc88a5dea8dbabf282009c729

Observation 69a1127a-7e13-4c0a-badc-649939be276a · outbound

This paper cites The power of scale for parameter-efficient prompt tuning.

Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting The power of scale for parameter-efficient prompt tuning

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:59:51.533654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T01:59:51.326493Z digest=sha256:85d318e0cb3955ae37c5702c4f3ec9c0e9e44baf5cdfcbe627fee28177b91e48

Observation 1b610d35-481b-4300-8abb-3812474c5d8d · outbound

This paper cites Rouge: A package for automatic evaluation of summaries.

Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Rouge: A package for automatic evaluation of summaries

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:59:51.537693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T01:59:51.326493Z digest=sha256:6dd0c6321ae5dd783f7805436508841fcbee16841ba5b35a65554503ecb94244

Observation 1f5438eb-dfa8-414f-86d3-4f75dddc46a1 · outbound

This paper cites What makes chain-of-thought prompting effective? a counterfactual study.

Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting What makes chain-of-thought prompting effective? a counterfactual study

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:59:51.541594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T01:59:51.326493Z digest=sha256:decc7466d9fca7c9d71f3333dd4bf6d9057f27c6d816139c8e58718167a33376

Observation 9e5056de-a469-4e7b-802e-28445a228263 · outbound

This paper cites Stereoset: Measuring stereotypical bias in pretrained language models.

Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Stereoset: Measuring stereotypical bias in pretrained language models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:59:51.544935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T01:59:51.326493Z digest=sha256:628cfb87e56be3dd091bff5b38dccaffaaed14a18fe9e662ea6732b1ce0ce5e9

Observation 67abd06c-7927-4565-bdc4-01521636eed4 · outbound

This paper cites Learning how to ask: Querying lms with mixtures of soft prompts.

Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Learning how to ask: Querying lms with mixtures of soft prompts

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:59:51.547957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T01:59:51.326493Z digest=sha256:b1be6d316910e1b7bf33e180ce29fb62950da46ff39d07d3076e888e17fc4bfb

Observation 87eabb56-a08b-463f-84ba-6b8cc280e469 · outbound

This paper cites Self-diagnosis and self-debiasing: A proposal for reducing corpus-based bias in nlp.

Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Self-diagnosis and self-debiasing: A proposal for reducing corpus-based bias in nlp

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:59:51.551285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T01:59:51.326493Z digest=sha256:65488a0c86cdcf42603ab54eba891cf643303760fcd68b3b604e5f6bf803ce99

Observation 6bcac65a-229f-407b-a575-735fcb99ab1b · outbound

This paper cites Chatgpt: Optimizing language models for dialogue.

Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Chatgpt: Optimizing language models for dialogue

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:59:51.554500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T01:59:51.326493Z digest=sha256:c27a99d65b042ebe3a86f6d63360002704657d25ae9205bfbf39a74182a3a4fc

Observation ce4cb433-8914-4664-9e4b-0cd8553c6402 · outbound

This paper cites Bertscore: Evaluating text generation with bert.

Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Bertscore: Evaluating text generation with bert

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:59:51.557702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T01:59:51.326493Z digest=sha256:071ccb36cd3ce20127cd14be4fc0b533e342ee016371978b01428093d4320f08

Observation 363ad71d-9efc-4221-a71f-71beafb4a411 · outbound

This paper cites Large language models are human-level prompt engineers.

Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Large language models are human-level prompt engineers

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:59:51.428661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T01:59:51.326493Z digest=sha256:32acf20f995fc237ff551813f29c4c1328ac84ac1832690f7aa98f5afb4b48d8

Observation 303e6923-59d2-4188-9afa-ecef67cd933b · outbound

This paper cites Findings of the Association for Computational Linguistics: EMNLP 2023 , pages=.

Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Findings of the Association for Computational Linguistics: EMNLP 2023 , pages=

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:59:51.431448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T01:59:51.326493Z digest=sha256:b5b6b8e048d7a5d2ffe22cfcb708abccad3585db2f8349701dd52ee6ee55da3e

Observation ce95afe9-39b8-41aa-bf58-377ab12a4698 · outbound

This paper cites Scaling Learning Algorithms Towards.

Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Scaling Learning Algorithms Towards

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:59:51.434692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T01:59:51.326493Z digest=sha256:4e5f793a8a9bcc0c02869cb1ad8ce611ee4ddf39226073243331512ca5889ec6

Observation 3fdc71a9-eb1c-48dd-9eaf-2490eeefec66 · outbound

This paper cites and Osindero, Simon and Teh, Yee Whye , journal =.

Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting and Osindero, Simon and Teh, Yee Whye , journal =

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:59:51.437652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T01:59:51.326493Z digest=sha256:db31182b0ea83208b054c6d696eaf9ece95bb0e4604107d3680802fbb8a88236

Observation 7cb0a65f-d2bb-4230-aa33-fabbc7b8dbdd · outbound

This paper cites 2016 , publisher=.

Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting 2016 , publisher=

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:59:51.440521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T01:59:51.326493Z digest=sha256:d586f0dc4025a45c1fc29ea5229f8645567a4b7589f6f2a898f440bda0aa494a

Observation a894853f-7809-492e-8cc1-6e3d66184bc2 · outbound

This paper cites Transactions of the Association for Computational Linguistics , volume=.

Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Transactions of the Association for Computational Linguistics , volume=

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:59:51.443192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T01:59:51.326493Z digest=sha256:737079108f33c9d931a42bce8f82de8e230d325486cf0ac55a4f32c43268cb78

Observation adddd939-ed1f-43d7-af5e-356bd620117d · outbound

This paper cites Super- N atural I nstructions: Generalization via declarative instructions on 1600+ NLP tasks.

Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Super- N atural I nstructions: Generalization via declarative instructions on 1600+ NLP tasks

Reference 43

Resolution
verified exact
doi, observed 2026-05-17T01:59:51.384220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T01:59:51.326493Z digest=sha256:468297befad7345a71a41b0237cbd2467e75bc5bf018cdd3c6aca9bedf569975

Observation 6d3acaeb-17d6-4dd3-89c7-43c915e926ea · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 44

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T01:59:51.411132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T01:59:51.326493Z digest=sha256:5cadf3c739cb973cf6e58fe72668acf7bda47c6b5ebfcb3d89bc61ca36dbb709

Observation e7963705-ba5a-4820-af5b-668668322802 · outbound

This paper cites an unresolved cited work.

Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-05-17T01:59:51.445968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T01:59:51.326493Z digest=sha256:240a2914906db4f60850258d151ee21b296d8315fabc0acf39a371a79b8ad85b

Observation 0fdbc677-f88c-4be8-8d5e-7376cb06b0cb · outbound

This paper cites Reframing Instructional Prompts to GPTk's Language.

Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Reframing Instructional Prompts to GPTk's Language

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T01:59:51.407613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T01:59:51.326493Z digest=sha256:1e323da5fd6b02d0b0a0f914f7e5a8e58e9e6cf243fbe156eca39cd18ef698c3

Observation 0b22b42e-7e46-489c-b79e-0ab25d634b23 · outbound

This paper cites Transactions of the Association for Computational Linguistics , volume=.

Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Transactions of the Association for Computational Linguistics , volume=

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:59:51.449067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T01:59:51.326493Z digest=sha256:1f39556b28aa86b02bece3e915b52048ebf1297f60390cd5702363bdfe9d3f07

Observation 4e1ed639-7176-459f-a20c-b2f7e260e0a3 · outbound

This paper cites and Wallace, Eric and Singh, Sameer.

Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting and Wallace, Eric and Singh, Sameer

Reference 48

Resolution
verified exact
doi, observed 2026-05-17T01:59:51.386958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T01:59:51.326493Z digest=sha256:f405ab1d1faf8c1776896a714fe0f7111bb7a230cd9de3d3cc77989913963fa3

Observation 9332d508-fd90-48aa-8f5c-5c72f609a9de · outbound

This paper cites RLP rompt: Optimizing Discrete Text Prompts with Reinforcement Learning.

Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting RLP rompt: Optimizing Discrete Text Prompts with Reinforcement Learning

Reference 49

Resolution
verified exact
doi, observed 2026-05-17T01:59:51.361456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T01:59:51.326493Z digest=sha256:ec1a9c55a2284574525ca354dd2c9e874f65af4b246a98a1348bf0b0af45d924

Observation 2d1f930c-6029-4fd1-abd4-8bb73c703fcf · outbound

This paper cites G r IPS : Gradient-free, Edit-based Instruction Search for Prompting Large Language Models.

Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting G r IPS : Gradient-free, Edit-based Instruction Search for Prompting Large Language Models

Reference 50

Resolution
verified exact
doi, observed 2026-05-17T01:59:51.367718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T01:59:51.326493Z digest=sha256:c04fe375c962867db8f67fc364da8e659ba1228b375f786cc57abc65fdca4d61

Observation 684c532d-f3c3-412c-aee9-67edfa5e69c0 · outbound

This paper cites Toward Human Readable Prompt Tuning: Kubrick's The Shining is a good movie, and a good prompt too?.

Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Toward Human Readable Prompt Tuning: Kubrick's The Shining is a good movie, and a good prompt too?

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:59:51.425681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T01:59:51.326493Z digest=sha256:29462bcc223bf1b219fe20bcc1f542574f493e09be6ea8521c33f71f0cab025a

Observation 7802663b-0457-45e6-aaf6-30dfdac66826 · outbound

This paper cites The Eleventh International Conference on Learning Representations , year=.

Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting The Eleventh International Conference on Learning Representations , year=

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:59:51.452037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T01:59:51.326493Z digest=sha256:9141029173a4425b84f157fc420edfc34a4188a10781342e8425603bb1b998c2

Observation 5ac2f298-efaf-462f-9ad2-c7d11eb7e95e · outbound

This paper cites Automatic Prompt Optimization with "Gradient Descent" and Beam Search.

Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Automatic Prompt Optimization with "Gradient Descent" and Beam Search

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:59:51.390802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T01:59:51.326493Z digest=sha256:50e8dd8700157342554aea81cb4fa43417a04417a7dacb9d5dd1c5513cb9fefc

Observation cdd0a501-f844-4ed7-95e0-796f3a5ed4c0 · outbound

This paper cites Text summarization branches out , pages=.

Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Text summarization branches out , pages=

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:59:51.455527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T01:59:51.326493Z digest=sha256:d9d9306794424434ec93f3e742c7c29ba460e3bcdba51d98982185d7a7700a7e

Observation 0db92237-9bae-4670-a5f5-a1837037446d · outbound

This paper cites International Conference on Learning Representations , year=.

Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting International Conference on Learning Representations , year=

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:59:51.458580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T01:59:51.326493Z digest=sha256:ce322559fb55af7389f1466ab7b9ed979676ede5c3aebded18cb0a57fcaf8833

Observation 34689bb0-5418-4fde-b316-dc369ef0ed53 · outbound

This paper cites The Eleventh International Conference on Learning Representations , year=.

Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting The Eleventh International Conference on Learning Representations , year=

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:59:51.461095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T01:59:51.326493Z digest=sha256:d96a96903f0b036e028db8be9b14b88884db0875dff8be821ff49cbbef223b11

Observation 7acd681a-a8bc-45a2-b6ed-a7c4ae282724 · outbound

This paper cites 2023 , url =.

Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting 2023 , url =

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:59:51.463495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T01:59:51.326493Z digest=sha256:5bf981ccd0c18c8f22a0f93165d2adae301b05526aea1946971ef1e9bdde4cf2

Observation 3597a8de-b7d7-4221-bad5-b1edc0fab879 · outbound

This paper cites 2023 , eprint=.

Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting 2023 , eprint=

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:59:51.465742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T01:59:51.326493Z digest=sha256:6308519e278b18c3999739c79308c05a049af1840686400ba6d797ce6e5cde28

Observation 9079ed20-f698-4c86-9b07-489d1b1d3578 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 59

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T01:59:51.414513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T01:59:51.326493Z digest=sha256:f91c00bbd2b02139dc5687a0ca65407af2a9c4b2c6fc3ae009e53e4f0a15edcc

Observation 4b9b33c4-422a-41a7-9277-514c541c5a89 · outbound

This paper cites Jailbroken: How Does LLM Safety Training Fail?.

Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Jailbroken: How Does LLM Safety Training Fail?

Reference 60

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T01:59:51.421696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T01:59:51.326493Z digest=sha256:182d590f566b785d2350be2a53424578bbeb6623608d618969a2300c7bd81e8d

Observation e0f1cf4c-2221-4853-9611-d9eee7638c10 · outbound

This paper cites 2022 , url=.

Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting 2022 , url=

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:59:51.468076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T01:59:51.326493Z digest=sha256:ce6c247cf1eaef1140e27953bfeceb18276ee98b79196dd13127ed4a97a98c84

Observation b62c89cf-8186-4d10-b394-f6b6d82c0e14 · outbound

This paper cites Advances in neural information processing systems , volume=.

Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Advances in neural information processing systems , volume=

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:59:51.470604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T01:59:51.326493Z digest=sha256:c33e69aed1d0167d973f35d704f28a3151d50668023ffebb4fdd5431b2572254

Observation ed0459fa-53df-4574-84c7-53b74c5b28bc · outbound

This paper cites Advances in applied mathematics , volume=.

Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Advances in applied mathematics , volume=

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:59:51.473063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T01:59:51.326493Z digest=sha256:174dffc1ded91702e1a499011fb8c12b099a0a1e6fceb444f2d78ddb7e0bc202

Observation 48b41ed1-527d-43bd-bf2c-c133a2a03a5d · outbound

This paper cites A Prompt Pattern Catalog to Enhance Prompt Engineering with ChatGPT.

Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting A Prompt Pattern Catalog to Enhance Prompt Engineering with ChatGPT

Reference 64

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T01:59:51.398030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T01:59:51.326493Z digest=sha256:22dcb129b32f4cc85fc42b237395781d5c793645bd470ba9cca137ae9fcf51cc

Observation 79ef4430-c6ac-46b7-b90f-63670cb02eee · outbound

This paper cites Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity.

Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity

Reference 65

Resolution
verified exact
doi, observed 2026-05-17T01:59:51.377945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T01:59:51.326493Z digest=sha256:c1a490c660c3bad63b0da92258d34f6acdaec53644b7855f522d62c2eb3c71ec

Observation fa491dbd-8986-483e-8de1-53b530d9f23e · outbound

This paper cites Demystifying Prompts in Language Models via Perplexity Estimation.

Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Demystifying Prompts in Language Models via Perplexity Estimation

Reference 66

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T01:59:51.403617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T01:59:51.326493Z digest=sha256:3216e78df9e4928878840d1cbdd9d5f529edea168f993eb498c86cf727bafdae

Observation c09083b8-870b-4ffb-9aee-d26dcf3e37d7 · outbound

This paper cites doi: 10.18653/v1/2021.acl-long.295.

Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting doi: 10.18653/v1/2021.acl-long.295

Reference 67

Resolution
metadata mismatch
doi, observed 2026-05-17T01:59:51.381003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T01:59:51.326493Z digest=sha256:403f2f676c0ae30a7fdaf334116f15b2aa82ab30ac524d5539355324246f8275

Observation f0463af4-f301-44b1-8e27-58c4ebc9fd6b · outbound

This paper cites and Levy, Omer.

Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting and Levy, Omer

Reference 68

Resolution
verified exact
doi, observed 2026-05-17T01:59:51.370516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T01:59:51.326493Z digest=sha256:097ae4372e3402dfd2c9a94c9e139bdf19abd2434f73346edeb97c3f3ccc5363

Observation d7bd6c92-6451-4ae4-85c8-20ad0c8e3e11 · outbound

This paper cites Chen and C.

Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Chen and C

Reference 69

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T01:59:51.374277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T01:59:51.326493Z digest=sha256:a8e537616a5c417f507a5e7bd95629252cec4c7b02439b4c981cac612a04888d

Observation 277fac3e-bce9-4b1f-bc44-e224a5b0c1be · outbound

This paper cites an unresolved cited work.

Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Unresolved cited work

Reference 70

Resolution
unresolved
raw_fallback, observed 2026-05-17T01:59:51.475332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T01:59:51.326493Z digest=sha256:a4ff82bf6ff09b9c0d7a507955960e530dd3b7d44e500fc827ddfd294a0b19cc

Observation 6315a095-6442-4f59-8434-5c95ad579103 · outbound

This paper cites Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT) , year=.

Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT) , year=

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:59:51.477813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T01:59:51.326493Z digest=sha256:26a9567d909cf857f90f7a8e22461cf48cdec99dae739b2b94eb2dbd10fd5bdb

Observation 63a03461-8ef6-4f3b-9801-72fd56b21afb · outbound

This paper cites Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics: System Demonstrations , pages=.

Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics: System Demonstrations , pages=

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:59:51.480294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T01:59:51.326493Z digest=sha256:c0e3f57f153218b48436edc52d31cd5de42c1bfd8fde6b0fdf16ecd4c9fcdbb8

Observation f0fa1666-343a-4a81-94c1-afc124fdc093 · outbound

This paper cites Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing , pages=.

Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing , pages=

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:59:51.483065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T01:59:51.326493Z digest=sha256:2267f632c518f84ba69f06022be8fd65774310807e07d8a0f859b955dab5546b

Observation 06807b0d-c111-442f-9c49-bfc1f1dca216 · outbound

This paper cites Prompt Waywardness: The Curious Case of Discretized Interpretation of Continuous Prompts.

Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Prompt Waywardness: The Curious Case of Discretized Interpretation of Continuous Prompts

Reference 74

Resolution
verified exact
doi, observed 2026-05-17T01:59:51.365010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T01:59:51.326493Z digest=sha256:f8a892e01491866f379c7495b117ae56d905c7342a17d4c0b53a79f44c0ba5c4

Observation ac4ccb65-6b8b-41a3-83d2-722b6f0c8d2b · outbound

This paper cites Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing , pages=.

Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing , pages=

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:59:51.485685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T01:59:51.326493Z digest=sha256:6685f8ca3f3e40fc237b1e725d239d2080ed25197335dbd95dec4bc9f7a98c9f

Observation 5cda8762-5e00-4b05-bfa4-73a7a1c30a49 · outbound

This paper cites Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies , pages=.

Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies , pages=

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:59:51.488168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T01:59:51.326493Z digest=sha256:97a87fe586f3812edcfa9fa791a5a703d75b1d22f590dc9cb3caaaee4a237ba2

Observation 62ee2e47-b803-4397-ae12-dfa2aa8d5498 · outbound

This paper cites Reproducibility of Benchmarked Deep Reinforcement Learning Tasks for Continuous Control.

Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Reproducibility of Benchmarked Deep Reinforcement Learning Tasks for Continuous Control

Reference 77

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T01:59:51.418148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T01:59:51.326493Z digest=sha256:68a494500fff1ccfcec6fde41486ec52a5b3111946dd9dd163443c380f9e2a76

Observation 90da363f-d2d8-4b37-b585-cfe12678bc2a · outbound

This paper cites OpenAI blog , year=.

Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting OpenAI blog , year=

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:59:51.490490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T01:59:51.326493Z digest=sha256:420deed4538c1be66028efbc8e5da652519001fbd480d80f3c4c05c96759f12c

Observation 79f94082-5c5e-4783-a0f5-42fabe4a2263 · outbound

This paper cites Transactions of the Association for Computational Linguistics , volume=.

Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Transactions of the Association for Computational Linguistics , volume=

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:59:51.492948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T01:59:51.326493Z digest=sha256:6932fd358c63164a2b45278aa26e487ac91355d9157d3c472f36b194f643aaa0

Observation 1d185803-b062-4568-92b2-ee65b70b3a8a · outbound

This paper cites Proceedings of the AAAI conference on artificial intelligence , volume=.

Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Proceedings of the AAAI conference on artificial intelligence , volume=

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:59:51.495994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T01:59:51.326493Z digest=sha256:93222c976558365cb39b0b1c0ba102ce0ec60ebeab71fe630fa24a1cbfda7ee5

Observation 017a4e09-2529-4a05-a318-df7e76e82b77 · outbound

This paper cites 2023 , booktitle=.

Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting 2023 , booktitle=

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:59:51.498435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T01:59:51.326493Z digest=sha256:7156ce0fee599f20c181b5b6a30df2778c5d21edf9f599ee167756198dbcea2b

Pith citing papers

Observation 67d1c560-d29b-4c70-b960-e933eb8b09ac · inbound

Lessons from the Trenches on Reproducible Evaluation of Language Models cites this paper.

Lessons from the Trenches on Reproducible Evaluation of Language Models Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 104

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:59:51.558892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-16T18:44:49.519995Z digest=sha256:c81f53b1be95984abfc9fef16386a8c031b19027fdcb9a0b3ceeada49524cb03

Observation afc502ba-3a02-4d8f-978b-05972e225180 · inbound

Leveraging LLMs for Predictive Insights in Food Policy and Behavioral Interventions cites this paper.

Leveraging LLMs for Predictive Insights in Food Policy and Behavioral Interventions Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 2013

Resolution
unresolved
no resolver link, observed 2026-08-12T21:35:15.564134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:35:15.564134Z digest=sha256:292b440d8a9fa59f55b7d9de89d13c693e192f3a0c7c902ee6de6de3fb36e1b7

Observation cf5aaaec-a456-484e-9538-44747de18621 · inbound

Does Prompt Formatting Have Any Impact on LLM Performance? cites this paper.

Does Prompt Formatting Have Any Impact on LLM Performance? Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T19:38:26.789374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:38:26.789374Z digest=sha256:56feff897459fc07abc7f5dfe46b6384cdd69522780d83618b7cb06b05e9ba05

Observation b23bf9eb-c0aa-4a52-96a4-9573e21b26b9 · inbound

BetterBench: Assessing AI Benchmarks, Uncovering Issues, and Establishing Best Practices cites this paper.

BetterBench: Assessing AI Benchmarks, Uncovering Issues, and Establishing Best Practices Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-12T17:01:31.658676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:01:31.658676Z digest=sha256:0e333d8884738f00cce213e14e31cbd7db566e0540f3e3ae418326b6a33cef79

Observation 0a4fd0ef-9e60-46d5-a808-f148bfbb1e3f · inbound

AdaptAgent: Adapting Multimodal Web Agents with Few-Shot Learning from Human Demonstrations cites this paper.

AdaptAgent: Adapting Multimodal Web Agents with Few-Shot Learning from Human Demonstrations Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T16:27:47.890701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:27:47.890701Z digest=sha256:01eb5666207bad376f12f8134e772a5f3c36975c81a099fdade155fd453a4f3e

Observation 3cd638b4-c102-494e-97bc-47d81fd50750 · inbound

InputSnatch: Stealing Input in LLM Services via Timing Side-Channel Attacks cites this paper.

InputSnatch: Stealing Input in LLM Services via Timing Side-Channel Attacks Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-12T11:29:29.554565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:29:29.554565Z digest=sha256:7fb355f446e2041f4d442b1d5474f8b339f2d7ac1148dfabfa1160ac5737c1a0

Observation 49783f9b-3fb7-4782-bb87-16bbfa5ff718 · inbound

The broader spectrum of in-context learning cites this paper.

The broader spectrum of in-context learning Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-11T22:10:22.937583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:10:22.937583Z digest=sha256:5797dc9c303192d8982b74cc2dae41c27f6cd75841de728b065bd2aef951a5cf

Observation 54eaeb8b-5944-4991-9dbb-49cb7d52c8a9 · inbound

Understanding the LLM-ification of CHI: Unpacking the Impact of LLMs at CHI through a Systematic Literature Review cites this paper.

Understanding the LLM-ification of CHI: Unpacking the Impact of LLMs at CHI through a Systematic Literature Review Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 144

Resolution
unresolved
no resolver link, observed 2026-08-10T17:06:55.633108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:06:55.633108Z digest=sha256:cd860dadcb7e48c2a13b6a7c6ac9d712bfc0d28f293d43a390430d96c8ba5fc9

Observation 82f21333-e037-4555-80b5-9af74df1f2ba · inbound

Normative Evaluation of Large Language Models with Everyday Moral Dilemmas cites this paper.

Normative Evaluation of Large Language Models with Everyday Moral Dilemmas Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-10T00:49:57.082783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:49:57.082783Z digest=sha256:096f3c0aaf31b73922579f2280a38b0231863f0674e8070f607a29a8fd0caa84

Observation d0441053-09e6-4cd6-ab56-93da98c43b82 · inbound

CodeSCM: Causal Analysis for Multi-Modal Code Generation cites this paper.

CodeSCM: Causal Analysis for Multi-Modal Code Generation Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-08T20:08:13.553448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:08:13.553448Z digest=sha256:a139287d276953948d7f414e6553de9666aab599dc684e0eecd5e80c2d385f50

Observation 11dfe6a2-4267-4f6a-ade1-3359f6563496 · inbound

Benchmarking Prompt Sensitivity in Large Language Models cites this paper.

Benchmarking Prompt Sensitivity in Large Language Models Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-08T16:55:45.540970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:55:45.540970Z digest=sha256:e263a19a3d56570cf37b2d85e0fa7ac95f0bbf5d86a6025dc5dfa6ba9a129ed2

Observation fff5e99a-28a3-401f-9e05-938aaa4da2a4 · inbound

Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs cites this paper.

Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-08T00:04:57.434528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:04:57.434528Z digest=sha256:613c9e2e0e3a40ee88dec5e8ce941ef4e5b30a2399ed2d5dec45165bffb96b78

Observation 9e07dffd-368e-47e8-9de9-c952a1b2c84a · inbound

Language Model Fine-Tuning on Scaled Survey Data for Predicting Distributions of Public Opinions cites this paper.

Language Model Fine-Tuning on Scaled Survey Data for Predicting Distributions of Public Opinions Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-05-23T03:02:26.773428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:e94602dc6a3831c1559d8a3540be096655da4e5440eb5bfdc304b91a83f60f7b

Observation d4d06f39-afbc-4e11-947d-80b83b9585f3 · inbound

Collaboration among Multiple Large Language Models for Medical Question Answering cites this paper.

Collaboration among Multiple Large Language Models for Medical Question Answering Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:08.057788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:08.057788Z digest=sha256:745887f34ea0405f78e8ac9a4bc03973c0708bca6a0b27b04ba6fefd54dc8353

Observation b7b88fb4-20ef-490b-a7e5-975aef909afc · inbound

Existing Large Language Model Unlearning Evaluations Are Inconclusive cites this paper.

Existing Large Language Model Unlearning Evaluations Are Inconclusive Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:56.757190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:05:56.757190Z digest=sha256:30f4334055a2b046c5778d8af83db6328f2a645a5ecdaf3363e2b73c91d1e9f9

Observation 7cbae61c-9af0-476c-8de6-d50f1e5c1aa3 · inbound

A Multimodal, Multilingual, and Multidimensional Pipeline for Fine-grained Crowdsourcing Earthquake Damage Evaluation cites this paper.

A Multimodal, Multilingual, and Multidimensional Pipeline for Fine-grained Crowdsourcing Earthquake Damage Evaluation Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:12.470417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:08:12.470417Z digest=sha256:4de067acfd5060efa6fd69cb791c8f75f4a7bf01a067cac36230d2edf1909879

Observation 32d913c0-1673-4981-82e1-4f9a59691a1f · inbound

More or Less Wrong: A Benchmark for Directional Bias in LLM Comparative Reasoning cites this paper.

More or Less Wrong: A Benchmark for Directional Bias in LLM Comparative Reasoning Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T10:59:06.932061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:59:06.932061Z digest=sha256:c69bcb5bff782ec72726a50ff3b77a41213e6d1232c9c43f06856ef9f477a294

Observation 45ab169e-069f-4672-a7ee-8567bcce31df · inbound

Understanding and Meeting Practitioner Needs When Measuring Representational Harms Caused by LLM-Based Systems cites this paper.

Understanding and Meeting Practitioner Needs When Measuring Representational Harms Caused by LLM-Based Systems Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:03.052746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:45:03.052746Z digest=sha256:e0d7a02530b20550584f14406c19a17e44a969010bf4567056e9e66ea698c40c

Observation 34196dfb-50c7-4c67-a110-faae8e7dc90c · inbound

From Prompts to Constructs: A Dual-Validity Framework for LLM Research in Psychology cites this paper.

From Prompts to Constructs: A Dual-Validity Framework for LLM Research in Psychology Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:28.044768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:28.044768Z digest=sha256:8fd63c9a6dc893b5a436d6d0254ddf9a08d4fad9a8d50a1ab72fe3c82d0c950f

Observation 4fab458f-9a4d-40f5-b211-9ed2da36fff3 · inbound

Requirements Elicitation Follow-Up Question Generation cites this paper.

Requirements Elicitation Follow-Up Question Generation Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:07.707797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:07.707797Z digest=sha256:4178bf17f601128ee16080538ba198b254d3a76f9afe255813e49c4f1ba8e983

Observation 3a2ea975-db5b-4873-ae31-2c52ef77f146 · inbound

PromptSuite: A Task-Agnostic Framework for Multi-Prompt Generation cites this paper.

PromptSuite: A Task-Agnostic Framework for Multi-Prompt Generation Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-19T03:47:02.410931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-19T03:43:28.543349Z digest=sha256:286164239ec8fa9c798bfae1418bc47005e63ec89647e013cb24786e34797c46

Observation b6dd6aea-528c-453c-bae2-a93bc18b02e6 · inbound

Understanding Human Limits in Pattern Recognition: A Computational Model of Sequential Reasoning in Rock, Paper, Scissors cites this paper.

Understanding Human Limits in Pattern Recognition: A Computational Model of Sequential Reasoning in Rock, Paper, Scissors Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T14:26:03.490201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:26:03.490201Z digest=sha256:ca7a948b489ba604f4aaa4fe687299d5275db3e7095ba4a77f3705fde6f6af3a

Observation 6230737b-69d8-4d1e-8bf6-a0bf9e508377 · inbound

Compiling Prompts, Not Crafting Them: A Reproducible Workflow for AI-Assisted Evidence Synthesis cites this paper.

Compiling Prompts, Not Crafting Them: A Reproducible Workflow for AI-Assisted Evidence Synthesis Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T17:13:52.950973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:13:52.950973Z digest=sha256:1c3626b46642394ca4c3a455d0e3cd4cd4911391ffe3b04874327b3be84db54b

Observation 950659b2-ed32-4383-a1da-4e75ab68be5f · inbound

Flaw or Artifact? Rethinking Prompt Sensitivity in Evaluating LLMs cites this paper.

Flaw or Artifact? Rethinking Prompt Sensitivity in Evaluating LLMs Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-05T12:17:20.371226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:17:20.371226Z digest=sha256:8df6345f4c316bebfadeb5cb943225c6a7f4b4619ed7ac3aaf173025fff3261e

Observation 3c0c6ae2-382e-4d24-b1ab-c7c9c75339c6 · inbound

No for Some, Yes for Others: Persona Prompts and Other Sources of False Refusal in Language Models cites this paper.

No for Some, Yes for Others: Persona Prompts and Other Sources of False Refusal in Language Models Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T21:22:14.357858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:22:14.357858Z digest=sha256:f94ebfdb44f1f8f62f6a6900d77c22377f1b2b9a2fbd3959b992eca0bda831c0

Observation 8094e4a3-c2f5-4e28-b429-c33515ecf71b · inbound

From scratch to silver: Creating trustworthy training data for patent-SDG classification using Large Language Models cites this paper.

From scratch to silver: Creating trustworthy training data for patent-SDG classification using Large Language Models Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-04T19:27:04.897757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T19:27:04.897757Z digest=sha256:3d149fb38a5cf3e12c3b81bcc96b543584fedd618c9639bc4820122761c7cf3f

Observation b551cc06-294f-4c30-a9f9-467c782b935b · inbound

Discrimination by LLMs: Cross-lingual Bias Assessment and Mitigation in Decision-Making and Summarisation cites this paper.

Discrimination by LLMs: Cross-lingual Bias Assessment and Mitigation in Decision-Making and Summarisation Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T20:19:50.468699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:19:50.468699Z digest=sha256:acbabc7126feeb68dd91e2ae9017e0c228b4ff8d23f8ad856b26a419d901697d

Observation 1048b3db-c417-4803-8f34-526566bf1e09 · inbound

Discrimination by LLMs: Cross-lingual Bias Assessment and Mitigation in Decision-Making and Summarisation cites this paper.

Discrimination by LLMs: Cross-lingual Bias Assessment and Mitigation in Decision-Making and Summarisation Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T20:19:50.473280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:19:50.473280Z digest=sha256:49d0749eb1983bb194d4adab06c197f8e5c32204b09c63b012c721f619af118f

Observation 0196a771-86c1-43c1-b180-a3060c8fcad1 · inbound

Position: AI Evaluations Should be Grounded on a Theory of Capability cites this paper.

Position: AI Evaluations Should be Grounded on a Theory of Capability Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-05-21T22:00:41.479428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-21T21:57:55.834632Z digest=sha256:4e378be02124f008cfb5f517216961e0121fe255acd0a048c3283ba7d8cfa30e

Observation f113b980-247b-40cb-9faf-2fa1e4ab22bb · inbound

Activation Steering with a Feedback Controller cites this paper.

Activation Steering with a Feedback Controller Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-21T21:50:41.383188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T21:50:05.186376Z digest=sha256:dd675337de1de1d0e7f1135e9fb73cfd5d70b6e31319efab7b6199cf596d2942

Observation a3f56d2a-92ff-4ca4-95a4-967095642230 · inbound

RegCheck: A tool for structured comparisons between study registrations and papers cites this paper.

RegCheck: A tool for structured comparisons between study registrations and papers Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T09:36:34.579749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:36:34.579749Z digest=sha256:6d726a9ac221165107c9cd3fe310b88e7650fef25ee39006bf6d02fdd8aae968

Observation eb725fc7-ab55-474c-8cc3-d4fbcb21f1dc · inbound

RegCheck: A tool for structured comparisons between study registrations and papers cites this paper.

RegCheck: A tool for structured comparisons between study registrations and papers Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T09:36:34.417517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:36:34.417517Z digest=sha256:4d47a6fc8f15a916cf2d1e5fbc1821048b566323a3b77abd8b3e1c564b792dd9

Observation a0b70b9e-7ec7-4425-9461-f809d8b0a067 · inbound

Framing Instability in LLM Ethical Stance: Auditing Negation Sensitivity in Moral Dilemmas cites this paper.

Framing Instability in LLM Ethical Stance: Auditing Negation Sensitivity in Moral Dilemmas Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T07:03:14.824744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T07:03:14.824744Z digest=sha256:55e2376cff4a23f8d51526598d6524e03abcb3fe2c23185bd19b9ec93dfda89d

Observation 59a49e8a-dbed-4b47-8ff8-9afa1e692aa5 · inbound

Visual Persuasion: What Influences Decisions of Vision-Language Models? cites this paper.

Visual Persuasion: What Influences Decisions of Vision-Language Models? Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T23:00:27.287919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T23:00:27.287919Z digest=sha256:dabd7e1f93933f376b9081258f2e89a5eba05988fd0e13406c7b66f74ee6d83a

Observation 5639a345-b6f2-4d0f-a369-79ee60fb985f · inbound

Collective AI can amplify tiny perturbations into divergent decisions cites this paper.

Collective AI can amplify tiny perturbations into divergent decisions Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T01:59:51.558892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T14:24:13.562656Z digest=sha256:fc5931485391db276719c7c89ce4ae02f2be716bc717042430f395011d1c62c8

Observation c4363e52-b8a2-456a-8a84-b2de1811b75a · inbound

Causal Evidence that Language Models use Confidence to Drive Behavior cites this paper.

Causal Evidence that Language Models use Confidence to Drive Behavior Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-21T09:44:05.678228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T09:43:05.524088Z digest=sha256:ec5559071d9247453083124d23faec5c71dff9b88ae5d4deff5564e7f79a815c

Observation f0ab09b1-d23b-4100-a869-6a77e807a213 · inbound

Dual Implications of Quark Mass Hierarchies to Flavor Structure cites this paper.

Dual Implications of Quark Mass Hierarchies to Flavor Structure Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-13T13:46:54.451163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T13:46:54.451163Z digest=sha256:7dc99b68d2b0a40a9a6a7f0bbd3413674913ef3b6cdafcb1f5305771f6951c07

Observation f5077a94-b1f6-4337-b9bb-dd86afecd2eb · inbound

Steerable but Not Decodable: Function Vectors Operate Beyond the Logit Lens cites this paper.

Steerable but Not Decodable: Function Vectors Operate Beyond the Logit Lens Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:59:51.558892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-13T20:59:02.824707Z digest=sha256:e29871cacc60eb61b0bf90acaf8af70265728ec6ca45c15096af928cc0b6a5f1

Observation 6a0da57e-c6c1-45ef-b2e3-5dade4b334c6 · inbound

The Cartesian Cut in Agentic AI cites this paper.

The Cartesian Cut in Agentic AI Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 58

Resolution
malformed identifier
arxiv_id, observed 2026-05-17T01:59:51.558892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T17:54:45.333038Z digest=sha256:fa5cad9b68d2a8ed01086c3b96c8422b489c70d88d4587ab96e33428c1c403c0

Observation 4e800469-e96b-4d01-974b-b63e89638940 · inbound

Select Smarter, Not More: Prompt-Aware Evaluation Scheduling with Submodular Guarantees cites this paper.

Select Smarter, Not More: Prompt-Aware Evaluation Scheduling with Submodular Guarantees Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:59:51.558892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T15:47:45.677717Z digest=sha256:29bf023d10450f9506879d408833cb3bf8941c5e84ece69d41dafa725d5c55a6

Observation fd3276a5-c025-4ae0-8bfe-3e05b32ad1b8 · inbound

The PICCO Framework for Large Language Model Prompting: A Taxonomy and Reference Architecture for Prompt Structure cites this paper.

The PICCO Framework for Large Language Model Prompting: A Taxonomy and Reference Architecture for Prompt Structure Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:59:51.558892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-13T20:00:19.061076Z digest=sha256:35854cedf8018535e3bbbbd0f3171907728ebdd2c1af563273c63cd291692793

Observation fd771154-d352-40ab-a386-13d231096e1b · inbound

SPAGBias: Uncovering and Tracing Structured Spatial Gender Bias in Large Language Models cites this paper.

SPAGBias: Uncovering and Tracing Structured Spatial Gender Bias in Large Language Models Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:59:51.558892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-10T11:52:49.311085Z digest=sha256:5bf6b9e0b0ec595e15a14bfd29d197a3b7b219d155887b6428f46b18a7251cb2

Observation b656be60-9453-48e1-b04d-c00fc60baf97 · inbound

Compared to What? Baselines and Metrics for Counterfactual Prompting cites this paper.

Compared to What? Baselines and Metrics for Counterfactual Prompting Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 54

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T01:59:51.558892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-09T19:02:46.991897Z digest=sha256:fe1ff2bd841a3d1109d9450ffb95969216a74c264febd40ccc5e509a961586db

Observation b3540b05-1e28-4ba6-9009-187f14869d22 · inbound

What Single-Prompt Accuracy Misses: A Multi-Variant Reliability Audit of Language Models cites this paper.

What Single-Prompt Accuracy Misses: A Multi-Variant Reliability Audit of Language Models Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:59:51.558892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-08T19:28:46.547245Z digest=sha256:18a2355580ca0ce7587a7a96dc0836eb0d1a16ec254eb5aab7f65744fd039b87

Observation 9b7c6f9c-0405-434b-a17f-816941b1d01a · inbound

Benchmarking Local Language Models for Social Robots using Edge Devices cites this paper.

Benchmarking Local Language Models for Social Robots using Edge Devices Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:59:51.558892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-08T17:42:32.154484Z digest=sha256:cda4f7b2d1065b8f75ea88be4f0f121d80f91ed20164d098074130dd937639fd

Observation 98600158-a739-4ef0-b0b4-de4212556679 · inbound

Paraphrase-Induced Output-Mode Collapse: When LLMs Break Character Under Semantically Equivalent Inputs cites this paper.

Paraphrase-Induced Output-Mode Collapse: When LLMs Break Character Under Semantically Equivalent Inputs Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:59:51.558892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-08T16:26:54.068997Z digest=sha256:3692d360d74564214a8d58a0f309787a23b01fda364638aeb16ed70f6d22cccf

Observation 83ff963b-d838-4c2d-a198-21f2622ffc4f · inbound

Paraphrase-Induced Output-Mode Collapse: When LLMs Break Character Under Semantically Equivalent Inputs cites this paper.

Paraphrase-Induced Output-Mode Collapse: When LLMs Break Character Under Semantically Equivalent Inputs Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:59:51.558892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-12T02:13:48.672711Z digest=sha256:8874a270cea048ad70215fd2cbaae16d0f4949272f6cf458b0f382632b9b1087

Observation 975fbeb2-490a-4282-bb4a-d5cd8db7c00f · inbound

Beyond Accuracy: Policy Invariance as a Reliability Test for LLM Safety Judges cites this paper.

Beyond Accuracy: Policy Invariance as a Reliability Test for LLM Safety Judges Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:59:51.558892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-08T10:23:58.198464Z digest=sha256:5d02175252938f2595caa77c7ecf1e9625ccba81725dade80c6023fe1ae65a4e

Observation b6f3d18d-dbf1-4f98-a18e-3f42002bc3b8 · inbound

Measuring Evaluation-Context Divergence in Open-Weight LLMs: A Paired-Prompt Protocol with Pilot Evidence of Alignment-Pipeline-Specific Heterogeneity cites this paper.

Measuring Evaluation-Context Divergence in Open-Weight LLMs: A Paired-Prompt Protocol with Pilot Evidence of Alignment-Pipeline-Specific Heterogeneity Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T01:59:51.558892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-08T10:23:02.697982Z digest=sha256:ca63d07a3884b04bb0bcf5d9fc5c00d8127c9e07ebeaca99c119d8dc61748e08

Observation 34ae1dde-d9bf-479c-bce8-692524314ae3 · inbound

Why Global LLM Leaderboards Are Misleading: Small Portfolios for Heterogeneous Supervised ML cites this paper.

Why Global LLM Leaderboards Are Misleading: Small Portfolios for Heterogeneous Supervised ML Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 293

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T01:59:51.558892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-08T12:05:32.881629Z digest=sha256:6f37bbc4b34c99cac79d7c94c91cfb15adf0a108f91259338b7546c02e01f4b2

Observation f333364e-a655-438f-9a68-0ccb8934eee6 · inbound

The Text Uncanny Valley: Non-Monotonic Performance Degradation in LLM Information Retrieval cites this paper.

The Text Uncanny Valley: Non-Monotonic Performance Degradation in LLM Information Retrieval Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:59:51.558892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-11T02:53:07.297356Z digest=sha256:b3393da89cce61e4754c065239de34747b1166cd04e43306fa70d41a01b97eb5

Observation c80d12c3-6a2f-40a8-a7f5-cfb844a17bd0 · inbound

CyBiasBench: Benchmarking Bias in LLM Agents for Cyber-Attack Scenarios cites this paper.

CyBiasBench: Benchmarking Bias in LLM Agents for Cyber-Attack Scenarios Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:59:51.558892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-11T01:54:46.391345Z digest=sha256:878f8a447dd47ef4c982da099dc01184aba9becb60793e7df8031c3aa907fe12

Observation b4349365-f316-4c16-aa26-d9233aba4688 · inbound

CUDABeaver: Benchmarking LLM-Based Automated CUDA Debugging cites this paper.

CUDABeaver: Benchmarking LLM-Based Automated CUDA Debugging Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:59:51.558892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-12T01:43:06.644640Z digest=sha256:da9ececaabe5d387392d738e25ca8e5fe56e2fd21ea632451b05dc50af2bfcde

Observation 526d97db-3377-46d5-baa0-7ba430a84b00 · inbound

CUDABeaver: Benchmarking LLM-Based Automated CUDA Debugging cites this paper.

CUDABeaver: Benchmarking LLM-Based Automated CUDA Debugging Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-07-01T13:35:46.961576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T22:54:20.299335Z digest=sha256:3f489bb4aa96bf399b6759cf6b561807819002d6f7226a5d4b4ed523afd20ab5

Observation 49678d8d-4a34-418b-a548-c911ef1878f8 · inbound

Beyond the Cartesian Illusion: Testing Two-Stage Multi-Modal Theory of Mind under Perceptual Bottlenecks cites this paper.

Beyond the Cartesian Illusion: Testing Two-Stage Multi-Modal Theory of Mind under Perceptual Bottlenecks Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-20T10:13:11.910514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T10:10:31.059095Z digest=sha256:5fd55b97e4f1d673cdf41cf4e90ccb6dd682512807a70bdfff20128552561e1f

Observation 3355f211-f632-46a9-b56f-75db36161e2f · inbound

Stop Drawing Scientific Claims from LLM Social Simulations Without Robustness Audits cites this paper.

Stop Drawing Scientific Claims from LLM Social Simulations Without Robustness Audits Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 63

Resolution
verified exact
local_arxiv, observed 2026-05-20T13:43:19.588773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T13:40:04.275438Z digest=sha256:d401e2a810e7d0319790ce1dd138bf98dc1a8e612f5f5d053e849ee20ac4bfb0

Observation 98ca3752-a808-47be-b31d-4478a0c7e66a · inbound

Towards Context-Invariant Safety Alignment for Large Language Models cites this paper.

Towards Context-Invariant Safety Alignment for Large Language Models Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 46

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T05:39:40.847530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-21T05:36:12.562807Z digest=sha256:b682ca9dd71017cbad0dd994193e51f584f59cee7e49263f4cb1e6f02bad57ac

Observation dd1d57e0-4ee4-4a65-90ef-28dd879c1439 · inbound

SafetyRepro: Configuration-Conditional Rank Instability on Alignment Benchmarks cites this paper.

SafetyRepro: Configuration-Conditional Rank Instability on Alignment Benchmarks Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:54:00.992662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-29T22:50:48.263600Z digest=sha256:0e195dbfee64c0f9b1e981a7dbc755d93cd4b97df9afbbb842865ac335bf47c0

Observation 2a250451-02d6-4609-8150-551b576494b5 · inbound

Paraphrase Brittleness in Production Retrieval-Augmented Commercial Recommendation: Reproducibility Below the Rerun-Stability Baseline cites this paper.

Paraphrase Brittleness in Production Retrieval-Augmented Commercial Recommendation: Reproducibility Below the Rerun-Stability Baseline Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-06-30T14:54:45.706466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T14:44:45.353335Z digest=sha256:3a8c51069d5f19a9a823f66f82fe71980fa2f987979558eb955796d41d9dc819

Observation 65addb7c-8ece-4002-9a36-c4a24c3c2aff · inbound

Chain-based Adaptive Reconfiguration Over Lattices for Hallucination Reduction cites this paper.

Chain-based Adaptive Reconfiguration Over Lattices for Hallucination Reduction Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T18:13:48.306109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:613baab46f41dfeabd6b0acdb979814b57705d691af02ecdc0fb41c28f25a46f

Observation 532c76cf-ba79-4c40-ad63-cc7413ad183f · inbound

Specialty-Specific Medical Language Model for Immune-Mediated Diseases cites this paper.

Specialty-Specific Medical Language Model for Immune-Mediated Diseases Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-12T22:27:01.885350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:27:01.885350Z digest=sha256:1249bd3447c4b476e90073111cbbd1177f97ca3abb4c01624033dda5cb2dbe7d

Observation b750e918-9a7b-430d-a5cb-c5eef6908d31 · inbound

How Consistent Are LLM Agents? Measuring Behavioral Reproducibility in Multi-Step Tool-Calling Pipelines cites this paper.

How Consistent Are LLM Agents? Measuring Behavioral Reproducibility in Multi-Step Tool-Calling Pipelines Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-04T21:40:08.966556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-07-04T21:40:04.594380Z digest=sha256:8d948232dde7f048c3d033f235961d3adabcafa538a4bd20298c8ea3790650c3

Observation 4c8908a7-5104-4e75-b152-b9504f050c72 · inbound

Mind Your Tone: Does Tone Alter LLM Performance? cites this paper.

Mind Your Tone: Does Tone Alter LLM Performance? Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T12:23:23.829759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-29T12:19:51.838450Z digest=sha256:0f4891e026068ad670cd628a0378e1dcd26a6f2dbfd3578832f18c16ebae3f4e

Observation 2a0ffda8-a8a4-46bc-abc3-8f524e2a89ca · inbound

Persona Conditioning of Brand Recommendations in Retrieval-Augmented Commercial Chat: A Prominence-Stratified Cross-Provider Audit cites this paper.

Persona Conditioning of Brand Recommendations in Retrieval-Augmented Commercial Chat: A Prominence-Stratified Cross-Provider Audit Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-06-29T07:23:12.704384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-29T07:20:32.538522Z digest=sha256:1992e5c9f4b2f03b67ad26195ba59ed04b9a88b95d527ed0b80b78f4283212fc

Observation f2bf329c-cc47-4e49-aec3-76de5a68e488 · inbound

Same Patient, Different Words, Different Diagnosis? Evaluating Semantic Stability in Clinical LLMs cites this paper.

Same Patient, Different Words, Different Diagnosis? Evaluating Semantic Stability in Clinical LLMs Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-06-29T07:23:13.270260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-29T07:14:22.444020Z digest=sha256:d6da232c7aed0e459eeec917fc609bade8f91a50d2f73573743aa5485c3d0d6e

Observation 35f46b7d-ee0d-45b8-b27a-509280c3a4ee · inbound

On the impact of retrieved content representations in RAG Pipelines cites this paper.

On the impact of retrieved content representations in RAG Pipelines Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T20:16:12.164754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T21:20:41.309842Z digest=sha256:65586c0fab15e964dfe68f3b8b8bfb06426016f5cd89604e7c719902cbe42c46

Observation 9be348f7-4089-4696-8886-d694bcc3defb · inbound

Consistency Training Can Entrench Misalignment cites this paper.

Consistency Training Can Entrench Misalignment Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 49

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T03:26:28.642293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-28T10:07:31.153337Z digest=sha256:21774d0cb6dd67649eb69811f2a2e03bb6f7d1131d26c12888aa701bf693874a

Observation 2b1fce16-aaa5-425c-bdd7-c036efbc2c5f · inbound

AIP: A Graph Representation for Learning and Governing Agent Skills cites this paper.

AIP: A Graph Representation for Learning and Governing Agent Skills Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-02T08:36:47.780631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T05:58:37.538980Z digest=sha256:077abf9598a48b44916ceba4ff91f67347745c7f55313ccc2d1ba41315553fae

Observation 15aecd79-a8fd-4cdd-8e56-ab18ad9839f6 · inbound

Stability vs. Manipulability: Evaluating Robustness Under Post-Decision Interaction in LLM Judges cites this paper.

Stability vs. Manipulability: Evaluating Robustness Under Post-Decision Interaction in LLM Judges Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 67

Resolution
verified exact
local_arxiv, observed 2026-07-02T08:36:47.667751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-28T05:58:59.870335Z digest=sha256:4f9b4a832f1ad38afa27fa41868321286cf18663197dd02679aa67e0dddbf1d7

Observation 484c52d7-3061-452d-9891-7327bd6971d1 · inbound

Self-Harness: Harnesses That Improve Themselves cites this paper.

Self-Harness: Harnesses That Improve Themselves Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-07-03T01:37:30.768446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-27T16:25:59.583722Z digest=sha256:f6c01d6f59627401d63dbd6999ab097f1149a58f9fbcfbbaa16fc8eb8684fb02

Observation 925b5beb-4f5a-4615-89d0-f13eb0d0586b · inbound

Are We Evaluating Knowledge or Phrasing? Mitigating MCQA Sensitivity with ParaEval cites this paper.

Are We Evaluating Knowledge or Phrasing? Mitigating MCQA Sensitivity with ParaEval Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-03T05:27:40.336372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-27T13:13:21.687950Z digest=sha256:84e9965ff8ce793a3bce420c07c9e926f586c6c1424d82e8c9fcfdc2e5891edb

Observation abd1163d-84b0-4ee9-bd69-342245e3dc8a · inbound

Preregistration for Experiments with AI Agents cites this paper.

Preregistration for Experiments with AI Agents Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 60

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T00:15:08.983043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-07-01T00:14:42.159357Z digest=sha256:4cad26c2fe4b4b328e5dc45ffbbdb7aa21292eca71daeb6ad71e7faf13345c1f

Observation 61da2415-daca-4f13-9f56-221c666c47ee · inbound

To Isolate or to Score? Model-Adaptive Assessment for Cost-Efficient Multi-Agent RAG cites this paper.

To Isolate or to Score? Model-Adaptive Assessment for Cost-Efficient Multi-Agent RAG Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T18:40:02.339892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-25T22:44:43.951083Z digest=sha256:1001c955c0eb39ced18e00aef9934d867fd662859509020bd9a9f07a6d11cc4f

Observation d007a58f-9e84-473f-97f8-cc5aed4a3759 · inbound

Instruction Bleed: Cross-Module Interference in Prompt-Composed Agentic Systems cites this paper.

Instruction Bleed: Cross-Module Interference in Prompt-Composed Agentic Systems Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-04T15:19:56.816194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-26T01:42:22.071423Z digest=sha256:f8d122827a20a841ebb3716c452effaeaf8090a505cba37607e13552fa6d4c58

Observation 26506fa3-9f47-42c5-a0e4-2f67a613a2a0 · inbound

A Deterministic Control Plane for LLM Coding Agents cites this paper.

A Deterministic Control Plane for LLM Coding Agents Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-07-04T14:19:54.749800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-26T03:58:28.523545Z digest=sha256:a643b7d369c2427da31df6ff4671d39bc2c5014dcb0a0eeebe53658e3efe5692

Observation 477d1fd4-b47e-44a5-b56a-38f93edfa261 · inbound

Towards Physical Intuitions for Alignment Dynamics: A Case Study With Randomness Crystallization cites this paper.

Towards Physical Intuitions for Alignment Dynamics: A Case Study With Randomness Crystallization Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T06:44:19.133463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-30T06:39:32.596823Z digest=sha256:c513265b918bd380153199748c3b3cc3755a42881a14d3ec4b628506ff61c42e

Observation 1b62b622-50c2-4937-8790-2dcd97e20740 · inbound

Falsification, Not Exposure: An Internally Preregistered Placebo-Controlled Decomposition of Self-Repair Feedback in Frozen Small Code Models cites this paper.

Falsification, Not Exposure: An Internally Preregistered Placebo-Controlled Decomposition of Self-Repair Feedback in Frozen Small Code Models Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-07-01T11:05:42.229723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-07-01T04:44:56.520156Z digest=sha256:5eb58bd8ebc727089482fd39a717311eb83c66e6e5db9304115fe86e482b4f55

Observation b1b818aa-b07c-4c93-9dc7-786f9c5ae973 · inbound

A Penny for Your Prompts: Experiments Detecting and Mitigating LLM Usage by Survey Respondents cites this paper.

A Penny for Your Prompts: Experiments Detecting and Mitigating LLM Usage by Survey Respondents Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-07-02T07:06:43.608084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-02T06:58:41.513964Z digest=sha256:ddfca16a6d5daf6a297ab9d99f07b2c27b56107266fe2cfdb0898688bfb67d43

Observation 72628e34-64c9-48b3-888c-2d2ed69436c4 · inbound

Persona Non Grata: LLM Persona-Driven Generations in MCQA are Unstable in Distinct Dimensions cites this paper.

Persona Non Grata: LLM Persona-Driven Generations in MCQA are Unstable in Distinct Dimensions Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 43

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T13:06:58.396919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-07-02T13:06:28.163476Z digest=sha256:ed4fc4f8a177e1335ee558c37e18246d0aeb774b3458d429f8a94967f5d282df

Observation 64b3a29c-8133-43ca-8f98-9fcec3499afc · inbound

The Powerless Noise: How Experimental Settings Shape the Reported Power of Noise cites this paper.

The Powerless Noise: How Experimental Settings Shape the Reported Power of Noise Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-12T01:10:52.578408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:10:52.578408Z digest=sha256:05587065207f7a18693f514697371df365136541c2ae0a682bbf40a064b55b70

Observation 6331b2e4-56c8-4ba6-82e6-43216ff0e9a5 · inbound

The Powerless Noise: How Experimental Settings Shape the Reported Power of Noise cites this paper.

The Powerless Noise: How Experimental Settings Shape the Reported Power of Noise Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-13T07:04:18.645464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T07:04:18.645464Z digest=sha256:c5e51eff8795e73c800aa00ecd637b2538cf2eec081225483d6c8918f1e5d9d7

Observation e8aa52b9-50b7-4857-ba42-2e5c7aaf4568 · inbound

Efficient Safety Alignment of Language Models via Latent Personality Traits cites this paper.

Efficient Safety Alignment of Language Models via Latent Personality Traits Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-07-10T15:27:20.131028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-07-10T15:26:23.290009Z digest=sha256:165c2e703bb6cbedae49f473681f97287bbf5b878a0c26a4ed9113a1703697c5

Observation f939d13c-bfd1-46c4-a7d4-751fde27759e · inbound

Format Sensitivity Index: Token-Controlled Prompt Wrapper Robustness and Schema Compliance in LLM Benchmarking cites this paper.

Format Sensitivity Index: Token-Controlled Prompt Wrapper Robustness and Schema Compliance in LLM Benchmarking Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-14T19:17:45.652076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T19:17:45.652076Z digest=sha256:cc924729a1c26c8dd35fefd35ca83a37175bb02f44cab5fc6cd7329a32952411

Observation 1d107bd2-1d3c-4cd8-8594-ed31dca207d5 · inbound

When Counterbalancing Hides the Bias: Access-Conditioned Position Lock in Forced-Choice LLM Evaluation cites this paper.

When Counterbalancing Hides the Bias: Access-Conditioned Position Lock in Forced-Choice LLM Evaluation Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-14T13:32:31.523845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T13:32:31.523845Z digest=sha256:8311da6515c7bea9d3976b6be66d888ba789efa2d50e053f28fea19e7d32c769

Observation 4252f286-08e7-4cb1-8455-ed83ca30c39f · inbound

Linear representations of grammaticality in neural language models cites this paper.

Linear representations of grammaticality in neural language models Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 202

Resolution
unresolved
no resolver link, observed 2026-08-01T23:59:51.912910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T23:59:51.912910Z digest=sha256:a27017b26e91e844323bc6b9dc8bd913b540cca93b636a675f9e2750d1430db4

Observation e26db494-ebb4-4594-b5bd-f2c979823a99 · inbound

Structured Output Collapses Answer Diversity Across 44 Language Models cites this paper.

Structured Output Collapses Answer Diversity Across 44 Language Models Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T15:23:45.126518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:23:45.126518Z digest=sha256:8623aa6240d4fb08016d64ba9832a3fae8ddc0f329a7bd8e85d5912131695292

Observation 8a6c7e48-a5a8-457c-95de-9d0822855397 · inbound

Prompt Design at Scale: How Format, Instruction Count, and Context Length Shape Instruction Adherence and Hallucination in Large Language Models cites this paper.

Prompt Design at Scale: How Format, Instruction Count, and Context Length Shape Instruction Adherence and Hallucination in Large Language Models Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T13:01:01.724152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:01:01.724152Z digest=sha256:024443b407b6ff1efb880cc853bb065fe5033f021483e17116a29651152c8179

Observation 0d850a5a-62e4-4273-8f1d-1f3e1059691d · inbound

A Knowledge-Injection Framework for Zero-Shot Adaptation of LLMs to Delirium Prediction cites this paper.

A Knowledge-Injection Framework for Zero-Shot Adaptation of LLMs to Delirium Prediction Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-02T13:59:21.524288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:59:21.524288Z digest=sha256:acc1bc89961963e4662b37b8cff172edda65474d5e28bc8dcbeb6776eba6abba

Observation 2d827d14-b1c1-4bd3-928a-5498c3224a70 · inbound

Reexamining zero-shot summarization: Empirical investigation of trustworthiness of LLM-summarizers cites this paper.

Reexamining zero-shot summarization: Empirical investigation of trustworthiness of LLM-summarizers Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T08:46:15.357112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:46:15.357112Z digest=sha256:1398af9410daa224c933b3678cb487105fd1a32e9bd890c8f1b1593625e7cf63

Observation 4b5a835c-c89d-4baa-bb83-3058e788e72b · inbound

Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy cites this paper.

Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T13:46:09.599956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:46:09.599956Z digest=sha256:7299191b91fcef2dc6f0131f69ecb327f54fb7d46df28692aedf81a8ea0ed138

Observation 5c0b3f9e-5023-42cf-bfe0-6d65e5718c24 · inbound

Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity cites this paper.

Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T00:48:49.223511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:48:49.223511Z digest=sha256:9e05eb374577c41f1451d9aaa77438a7b7cbb1cabf843540134b903856a88998

Observation 7b8dfe7b-3e8f-4616-9f54-dc2a0ac0623b · inbound

Routing Is Least Learnable Where It Is Most Valuable: Bounds on Representation Routing for Web Agents cites this paper.

Routing Is Least Learnable Where It Is Most Valuable: Bounds on Representation Routing for Web Agents Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 138

Resolution
unresolved
no resolver link, observed 2026-08-07T13:40:15.005783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:40:15.005783Z digest=sha256:c6176cf68364d0503182b0877911cff1a50c212dc6818fe2a753811aa486210f

Observation bb688394-056d-41d5-aaf1-007eadd68b80 · inbound

Same physical state, different collective dynamics: state encodings select synchronization outcomes in language-model agents cites this paper.

Same physical state, different collective dynamics: state encodings select synchronization outcomes in language-model agents Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T17:33:56.227086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:33:56.227086Z digest=sha256:769af35bb10e531bb17a73a5221edd3dff101642db3fe6d5a24829f42a222430

Observation bdbbb002-5899-4cc3-8b58-02836ac5f396 · inbound

Persona Conditioning as an Assessor-Sensitivity Probe for LLM-Based IR Evaluation cites this paper.

Persona Conditioning as an Assessor-Sensitivity Probe for LLM-Based IR Evaluation Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T00:52:42.047818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:52:42.047818Z digest=sha256:4629b5d6895e4e7547dca4e43a0fd8c6362622a11dbc59c7ca5f9ebd2f5cd40c

Observation 6de43ecb-a9bc-44da-99b5-bed25974c9c2 · inbound

Actions Speak Louder than Words: Measuring Cross-Lingual Policy Retention in Tool-Using Agents cites this paper.

Actions Speak Louder than Words: Measuring Cross-Lingual Policy Retention in Tool-Using Agents Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T06:04:17.071359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:04:17.071359Z digest=sha256:92647f5584679046e1accf0348525504f54d253ecbdb3cf6dc82b58d790735c8

Observation 3bb96473-236a-497f-ba68-8dd57ed9a617 · inbound

QuoteBench: How Matched Scores Can Hide Command-Path Failures cites this paper.

QuoteBench: How Matched Scores Can Hide Command-Path Failures Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:52.510712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:52.510712Z digest=sha256:efbb1e037ae155a8795d063000812fdc9ed725cd19555410e37770022346bc49