Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-17T01:59:51.326493Z
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 96 inbound Pith citation observations for arXiv:2310.11324.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-17T01:59:51.326493Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-14T04:32:52.510712Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
64 of 64 outbound references displayed
External citation measurements
0
pith, observed 2026-08-05T02:28:24.338817Z
Observation 23c5f13b-058f-4f8b-97de-a2411d423f31 · outbound
Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Tweet: Susan & I found MMLU performance jump 6-10 points in the 40s by formatting multiple choice as (A) not A in MMLU (for internal model)
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 956be0df-66bb-4914-a0db-8cd4e7b5039f · outbound
Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Falcon-40B : an open large language model with state-of-the-art performance
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 925bbcdb-d037-43c3-b065-71c87e4f4038 · outbound
Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting An empirical evaluation of thompson sampling
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation ba4e1fcf-f107-4b51-b6b3-062d01b12596 · outbound
Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Better hypothesis testing for statistical machine translation: Controlling for optimizer instability
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 88b0f953-5c8b-4220-9bcd-651eb1f61fe3 · outbound
Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting GPT 3.int8(): 8-bit matrix multiplication for transformers at scale
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 35fd58a4-464d-4701-b1f4-c35e5b1aa193 · outbound
Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Openprompt: An open-source framework for prompt-learning
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation ef7a012f-8209-4e1e-935a-00baa84995a4 · outbound
Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Measuring and improving consistency in pretrained language models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation c5b09f3e-1956-450c-8160-8942eb53edba · outbound
Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Deep reinforcement learning that matters
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 6c7f1cdf-d166-42c1-980b-8e8b26b7cfe1 · outbound
Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Editing models with task arithmetic
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation efb1adda-a288-4065-a733-9ed2f4c65e3c · outbound
Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting How can we know what language models know? Transactions of the Association for Computational Linguistics, 8: 0 423--438
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation c17940ac-8712-49eb-9a48-df739b684ee4 · outbound
Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Asymptotically efficient adaptive allocation rules
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 69a1127a-7e13-4c0a-badc-649939be276a · outbound
Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting The power of scale for parameter-efficient prompt tuning
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 1b610d35-481b-4300-8abb-3812474c5d8d · outbound
Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Rouge: A package for automatic evaluation of summaries
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 1f5438eb-dfa8-414f-86d3-4f75dddc46a1 · outbound
Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting What makes chain-of-thought prompting effective? a counterfactual study
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 9e5056de-a469-4e7b-802e-28445a228263 · outbound
Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Stereoset: Measuring stereotypical bias in pretrained language models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 67abd06c-7927-4565-bdc4-01521636eed4 · outbound
Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Learning how to ask: Querying lms with mixtures of soft prompts
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 87eabb56-a08b-463f-84ba-6b8cc280e469 · outbound
Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Self-diagnosis and self-debiasing: A proposal for reducing corpus-based bias in nlp
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 6bcac65a-229f-407b-a575-735fcb99ab1b · outbound
Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Chatgpt: Optimizing language models for dialogue
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation ce4cb433-8914-4664-9e4b-0cd8553c6402 · outbound
Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Bertscore: Evaluating text generation with bert
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 363ad71d-9efc-4221-a71f-71beafb4a411 · outbound
Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Large language models are human-level prompt engineers
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 303e6923-59d2-4188-9afa-ecef67cd933b · outbound
Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Findings of the Association for Computational Linguistics: EMNLP 2023 , pages=
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation ce95afe9-39b8-41aa-bf58-377ab12a4698 · outbound
Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Scaling Learning Algorithms Towards
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 3fdc71a9-eb1c-48dd-9eaf-2490eeefec66 · outbound
Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting and Osindero, Simon and Teh, Yee Whye , journal =
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 7cb0a65f-d2bb-4230-aa33-fabbc7b8dbdd · outbound
Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting 2016 , publisher=
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation a894853f-7809-492e-8cc1-6e3d66184bc2 · outbound
Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Transactions of the Association for Computational Linguistics , volume=
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation adddd939-ed1f-43d7-af5e-356bd620117d · outbound
Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Super- N atural I nstructions: Generalization via declarative instructions on 1600+ NLP tasks
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 6d3acaeb-17d6-4dd3-89c7-43c915e926ea · outbound
Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation e7963705-ba5a-4820-af5b-668668322802 · outbound
Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Unresolved cited work
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 0fdbc677-f88c-4be8-8d5e-7376cb06b0cb · outbound
Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Reframing Instructional Prompts to GPTk's Language
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 0b22b42e-7e46-489c-b79e-0ab25d634b23 · outbound
Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Transactions of the Association for Computational Linguistics , volume=
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 4e1ed639-7176-459f-a20c-b2f7e260e0a3 · outbound
Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting and Wallace, Eric and Singh, Sameer
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 9332d508-fd90-48aa-8f5c-5c72f609a9de · outbound
Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting RLP rompt: Optimizing Discrete Text Prompts with Reinforcement Learning
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 2d1f930c-6029-4fd1-abd4-8bb73c703fcf · outbound
Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting G r IPS : Gradient-free, Edit-based Instruction Search for Prompting Large Language Models
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 684c532d-f3c3-412c-aee9-67edfa5e69c0 · outbound
Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Toward Human Readable Prompt Tuning: Kubrick's The Shining is a good movie, and a good prompt too?
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 7802663b-0457-45e6-aaf6-30dfdac66826 · outbound
Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting The Eleventh International Conference on Learning Representations , year=
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 5ac2f298-efaf-462f-9ad2-c7d11eb7e95e · outbound
Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Automatic Prompt Optimization with "Gradient Descent" and Beam Search
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation cdd0a501-f844-4ed7-95e0-796f3a5ed4c0 · outbound
Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Text summarization branches out , pages=
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 0db92237-9bae-4670-a5f5-a1837037446d · outbound
Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting International Conference on Learning Representations , year=
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 34689bb0-5418-4fde-b316-dc369ef0ed53 · outbound
Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting The Eleventh International Conference on Learning Representations , year=
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 7acd681a-a8bc-45a2-b6ed-a7c4ae282724 · outbound
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 3597a8de-b7d7-4221-bad5-b1edc0fab879 · outbound
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 9079ed20-f698-4c86-9b07-489d1b1d3578 · outbound
Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 4b9b33c4-422a-41a7-9277-514c541c5a89 · outbound
Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Jailbroken: How Does LLM Safety Training Fail?
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation e0f1cf4c-2221-4853-9611-d9eee7638c10 · outbound
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation b62c89cf-8186-4d10-b394-f6b6d82c0e14 · outbound
Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Advances in neural information processing systems , volume=
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation ed0459fa-53df-4574-84c7-53b74c5b28bc · outbound
Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Advances in applied mathematics , volume=
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 48b41ed1-527d-43bd-bf2c-c133a2a03a5d · outbound
Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting A Prompt Pattern Catalog to Enhance Prompt Engineering with ChatGPT
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 79ef4430-c6ac-46b7-b90f-63670cb02eee · outbound
Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation fa491dbd-8986-483e-8de1-53b530d9f23e · outbound
Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Demystifying Prompts in Language Models via Perplexity Estimation
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation c09083b8-870b-4ffb-9aee-d26dcf3e37d7 · outbound
Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting doi: 10.18653/v1/2021.acl-long.295
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f0463af4-f301-44b1-8e27-58c4ebc9fd6b · outbound
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d7bd6c92-6451-4ae4-85c8-20ad0c8e3e11 · outbound
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 277fac3e-bce9-4b1f-bc44-e224a5b0c1be · outbound
Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Unresolved cited work
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 6315a095-6442-4f59-8434-5c95ad579103 · outbound
Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT) , year=
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 63a03461-8ef6-4f3b-9801-72fd56b21afb · outbound
Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics: System Demonstrations , pages=
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f0fa1666-343a-4a81-94c1-afc124fdc093 · outbound
Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing , pages=
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 06807b0d-c111-442f-9c49-bfc1f1dca216 · outbound
Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Prompt Waywardness: The Curious Case of Discretized Interpretation of Continuous Prompts
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation ac4ccb65-6b8b-41a3-83d2-722b6f0c8d2b · outbound
Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing , pages=
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 5cda8762-5e00-4b05-bfa4-73a7a1c30a49 · outbound
Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies , pages=
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 62ee2e47-b803-4397-ae12-dfa2aa8d5498 · outbound
Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Reproducibility of Benchmarked Deep Reinforcement Learning Tasks for Continuous Control
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 90da363f-d2d8-4b37-b585-cfe12678bc2a · outbound
Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting OpenAI blog , year=
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 79f94082-5c5e-4783-a0f5-42fabe4a2263 · outbound
Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Transactions of the Association for Computational Linguistics , volume=
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 1d185803-b062-4568-92b2-ee65b70b3a8a · outbound
Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Proceedings of the AAAI conference on artificial intelligence , volume=
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 017a4e09-2529-4a05-a318-df7e76e82b77 · outbound
Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting 2023 , booktitle=
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 67d1c560-d29b-4c70-b960-e933eb8b09ac · inbound
Lessons from the Trenches on Reproducible Evaluation of Language Models Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 104
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation afc502ba-3a02-4d8f-978b-05972e225180 · inbound
Leveraging LLMs for Predictive Insights in Food Policy and Behavioral Interventions Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 2013
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf5aaaec-a456-484e-9538-44747de18621 · inbound
Does Prompt Formatting Have Any Impact on LLM Performance? Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b23bf9eb-c0aa-4a52-96a4-9573e21b26b9 · inbound
BetterBench: Assessing AI Benchmarks, Uncovering Issues, and Establishing Best Practices Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a4fd0ef-9e60-46d5-a808-f148bfbb1e3f · inbound
AdaptAgent: Adapting Multimodal Web Agents with Few-Shot Learning from Human Demonstrations Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3cd638b4-c102-494e-97bc-47d81fd50750 · inbound
InputSnatch: Stealing Input in LLM Services via Timing Side-Channel Attacks Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49783f9b-3fb7-4782-bb87-16bbfa5ff718 · inbound
The broader spectrum of in-context learning Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54eaeb8b-5944-4991-9dbb-49cb7d52c8a9 · inbound
Understanding the LLM-ification of CHI: Unpacking the Impact of LLMs at CHI through a Systematic Literature Review Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 144
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82f21333-e037-4555-80b5-9af74df1f2ba · inbound
Normative Evaluation of Large Language Models with Everyday Moral Dilemmas Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0441053-09e6-4cd6-ab56-93da98c43b82 · inbound
CodeSCM: Causal Analysis for Multi-Modal Code Generation Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11dfe6a2-4267-4f6a-ade1-3359f6563496 · inbound
Benchmarking Prompt Sensitivity in Large Language Models Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fff5e99a-28a3-401f-9e05-938aaa4da2a4 · inbound
Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e07dffd-368e-47e8-9de9-c952a1b2c84a · inbound
Language Model Fine-Tuning on Scaled Survey Data for Predicting Distributions of Public Opinions Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d4d06f39-afbc-4e11-947d-80b83b9585f3 · inbound
Collaboration among Multiple Large Language Models for Medical Question Answering Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7b88fb4-20ef-490b-a7e5-975aef909afc · inbound
Existing Large Language Model Unlearning Evaluations Are Inconclusive Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7cbae61c-9af0-476c-8de6-d50f1e5c1aa3 · inbound
A Multimodal, Multilingual, and Multidimensional Pipeline for Fine-grained Crowdsourcing Earthquake Damage Evaluation Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32d913c0-1673-4981-82e1-4f9a59691a1f · inbound
More or Less Wrong: A Benchmark for Directional Bias in LLM Comparative Reasoning Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45ab169e-069f-4672-a7ee-8567bcce31df · inbound
Understanding and Meeting Practitioner Needs When Measuring Representational Harms Caused by LLM-Based Systems Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34196dfb-50c7-4c67-a110-faae8e7dc90c · inbound
From Prompts to Constructs: A Dual-Validity Framework for LLM Research in Psychology Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4fab458f-9a4d-40f5-b211-9ed2da36fff3 · inbound
Requirements Elicitation Follow-Up Question Generation Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a2ea975-db5b-4873-ae31-2c52ef77f146 · inbound
PromptSuite: A Task-Agnostic Framework for Multi-Prompt Generation Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation b6dd6aea-528c-453c-bae2-a93bc18b02e6 · inbound
Understanding Human Limits in Pattern Recognition: A Computational Model of Sequential Reasoning in Rock, Paper, Scissors Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6230737b-69d8-4d1e-8bf6-a0bf9e508377 · inbound
Compiling Prompts, Not Crafting Them: A Reproducible Workflow for AI-Assisted Evidence Synthesis Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 950659b2-ed32-4383-a1da-4e75ab68be5f · inbound
Flaw or Artifact? Rethinking Prompt Sensitivity in Evaluating LLMs Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c0c6ae2-382e-4d24-b1ab-c7c9c75339c6 · inbound
No for Some, Yes for Others: Persona Prompts and Other Sources of False Refusal in Language Models Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8094e4a3-c2f5-4e28-b429-c33515ecf71b · inbound
From scratch to silver: Creating trustworthy training data for patent-SDG classification using Large Language Models Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b551cc06-294f-4c30-a9f9-467c782b935b · inbound
Discrimination by LLMs: Cross-lingual Bias Assessment and Mitigation in Decision-Making and Summarisation Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1048b3db-c417-4803-8f34-526566bf1e09 · inbound
Discrimination by LLMs: Cross-lingual Bias Assessment and Mitigation in Decision-Making and Summarisation Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0196a771-86c1-43c1-b180-a3060c8fcad1 · inbound
Position: AI Evaluations Should be Grounded on a Theory of Capability Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f113b980-247b-40cb-9faf-2fa1e4ab22bb · inbound
Activation Steering with a Feedback Controller Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation a3f56d2a-92ff-4ca4-95a4-967095642230 · inbound
RegCheck: A tool for structured comparisons between study registrations and papers Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb725fc7-ab55-474c-8cc3-d4fbcb21f1dc · inbound
RegCheck: A tool for structured comparisons between study registrations and papers Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0b70b9e-7ec7-4425-9461-f809d8b0a067 · inbound
Framing Instability in LLM Ethical Stance: Auditing Negation Sensitivity in Moral Dilemmas Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59a49e8a-dbed-4b47-8ff8-9afa1e692aa5 · inbound
Visual Persuasion: What Influences Decisions of Vision-Language Models? Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5639a345-b6f2-4d0f-a369-79ee60fb985f · inbound
Collective AI can amplify tiny perturbations into divergent decisions Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation c4363e52-b8a2-456a-8a84-b2de1811b75a · inbound
Causal Evidence that Language Models use Confidence to Drive Behavior Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f0ab09b1-d23b-4100-a869-6a77e807a213 · inbound
Dual Implications of Quark Mass Hierarchies to Flavor Structure Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5077a94-b1f6-4337-b9bb-dd86afecd2eb · inbound
Steerable but Not Decodable: Function Vectors Operate Beyond the Logit Lens Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 6a0da57e-c6c1-45ef-b2e3-5dade4b334c6 · inbound
The Cartesian Cut in Agentic AI Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 4e800469-e96b-4d01-974b-b63e89638940 · inbound
Select Smarter, Not More: Prompt-Aware Evaluation Scheduling with Submodular Guarantees Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation fd3276a5-c025-4ae0-8bfe-3e05b32ad1b8 · inbound
The PICCO Framework for Large Language Model Prompting: A Taxonomy and Reference Architecture for Prompt Structure Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation fd771154-d352-40ab-a386-13d231096e1b · inbound
SPAGBias: Uncovering and Tracing Structured Spatial Gender Bias in Large Language Models Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation b656be60-9453-48e1-b04d-c00fc60baf97 · inbound
Compared to What? Baselines and Metrics for Counterfactual Prompting Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation b3540b05-1e28-4ba6-9009-187f14869d22 · inbound
What Single-Prompt Accuracy Misses: A Multi-Variant Reliability Audit of Language Models Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 9b7c6f9c-0405-434b-a17f-816941b1d01a · inbound
Benchmarking Local Language Models for Social Robots using Edge Devices Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 98600158-a739-4ef0-b0b4-de4212556679 · inbound
Paraphrase-Induced Output-Mode Collapse: When LLMs Break Character Under Semantically Equivalent Inputs Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 83ff963b-d838-4c2d-a198-21f2622ffc4f · inbound
Paraphrase-Induced Output-Mode Collapse: When LLMs Break Character Under Semantically Equivalent Inputs Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 975fbeb2-490a-4282-bb4a-d5cd8db7c00f · inbound
Beyond Accuracy: Policy Invariance as a Reliability Test for LLM Safety Judges Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation b6f3d18d-dbf1-4f98-a18e-3f42002bc3b8 · inbound
Measuring Evaluation-Context Divergence in Open-Weight LLMs: A Paired-Prompt Protocol with Pilot Evidence of Alignment-Pipeline-Specific Heterogeneity Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 34ae1dde-d9bf-479c-bce8-692524314ae3 · inbound
Why Global LLM Leaderboards Are Misleading: Small Portfolios for Heterogeneous Supervised ML Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 293
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f333364e-a655-438f-9a68-0ccb8934eee6 · inbound
The Text Uncanny Valley: Non-Monotonic Performance Degradation in LLM Information Retrieval Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation c80d12c3-6a2f-40a8-a7f5-cfb844a17bd0 · inbound
CyBiasBench: Benchmarking Bias in LLM Agents for Cyber-Attack Scenarios Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation b4349365-f316-4c16-aa26-d9233aba4688 · inbound
CUDABeaver: Benchmarking LLM-Based Automated CUDA Debugging Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 526d97db-3377-46d5-baa0-7ba430a84b00 · inbound
CUDABeaver: Benchmarking LLM-Based Automated CUDA Debugging Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 49678d8d-4a34-418b-a548-c911ef1878f8 · inbound
Beyond the Cartesian Illusion: Testing Two-Stage Multi-Modal Theory of Mind under Perceptual Bottlenecks Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 3355f211-f632-46a9-b56f-75db36161e2f · inbound
Stop Drawing Scientific Claims from LLM Social Simulations Without Robustness Audits Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 98ca3752-a808-47be-b31d-4478a0c7e66a · inbound
Towards Context-Invariant Safety Alignment for Large Language Models Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation dd1d57e0-4ee4-4a65-90ef-28dd879c1439 · inbound
SafetyRepro: Configuration-Conditional Rank Instability on Alignment Benchmarks Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 2a250451-02d6-4609-8150-551b576494b5 · inbound
Paraphrase Brittleness in Production Retrieval-Augmented Commercial Recommendation: Reproducibility Below the Rerun-Stability Baseline Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 65addb7c-8ece-4002-9a36-c4a24c3c2aff · inbound
Chain-based Adaptive Reconfiguration Over Lattices for Hallucination Reduction Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 532c76cf-ba79-4c40-ad63-cc7413ad183f · inbound
Specialty-Specific Medical Language Model for Immune-Mediated Diseases Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b750e918-9a7b-430d-a5cb-c5eef6908d31 · inbound
How Consistent Are LLM Agents? Measuring Behavioral Reproducibility in Multi-Step Tool-Calling Pipelines Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 4c8908a7-5104-4e75-b152-b9504f050c72 · inbound
Mind Your Tone: Does Tone Alter LLM Performance? Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 2a0ffda8-a8a4-46bc-abc3-8f524e2a89ca · inbound
Persona Conditioning of Brand Recommendations in Retrieval-Augmented Commercial Chat: A Prominence-Stratified Cross-Provider Audit Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f2bf329c-cc47-4e49-aec3-76de5a68e488 · inbound
Same Patient, Different Words, Different Diagnosis? Evaluating Semantic Stability in Clinical LLMs Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 35f46b7d-ee0d-45b8-b27a-509280c3a4ee · inbound
On the impact of retrieved content representations in RAG Pipelines Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 9be348f7-4089-4696-8886-d694bcc3defb · inbound
Consistency Training Can Entrench Misalignment Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 2b1fce16-aaa5-425c-bdd7-c036efbc2c5f · inbound
AIP: A Graph Representation for Learning and Governing Agent Skills Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 15aecd79-a8fd-4cdd-8e56-ab18ad9839f6 · inbound
Stability vs. Manipulability: Evaluating Robustness Under Post-Decision Interaction in LLM Judges Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 484c52d7-3061-452d-9891-7327bd6971d1 · inbound
Self-Harness: Harnesses That Improve Themselves Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 925b5beb-4f5a-4615-89d0-f13eb0d0586b · inbound
Are We Evaluating Knowledge or Phrasing? Mitigating MCQA Sensitivity with ParaEval Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation abd1163d-84b0-4ee9-bd69-342245e3dc8a · inbound
Preregistration for Experiments with AI Agents Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 61da2415-daca-4f13-9f56-221c666c47ee · inbound
To Isolate or to Score? Model-Adaptive Assessment for Cost-Efficient Multi-Agent RAG Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d007a58f-9e84-473f-97f8-cc5aed4a3759 · inbound
Instruction Bleed: Cross-Module Interference in Prompt-Composed Agentic Systems Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 26506fa3-9f47-42c5-a0e4-2f67a613a2a0 · inbound
A Deterministic Control Plane for LLM Coding Agents Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 477d1fd4-b47e-44a5-b56a-38f93edfa261 · inbound
Towards Physical Intuitions for Alignment Dynamics: A Case Study With Randomness Crystallization Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 1b62b622-50c2-4937-8790-2dcd97e20740 · inbound
Falsification, Not Exposure: An Internally Preregistered Placebo-Controlled Decomposition of Self-Repair Feedback in Frozen Small Code Models Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation b1b818aa-b07c-4c93-9dc7-786f9c5ae973 · inbound
A Penny for Your Prompts: Experiments Detecting and Mitigating LLM Usage by Survey Respondents Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 72628e34-64c9-48b3-888c-2d2ed69436c4 · inbound
Persona Non Grata: LLM Persona-Driven Generations in MCQA are Unstable in Distinct Dimensions Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 64b3a29c-8133-43ca-8f98-9fcec3499afc · inbound
The Powerless Noise: How Experimental Settings Shape the Reported Power of Noise Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6331b2e4-56c8-4ba6-82e6-43216ff0e9a5 · inbound
The Powerless Noise: How Experimental Settings Shape the Reported Power of Noise Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8aa52b9-50b7-4857-ba42-2e5c7aaf4568 · inbound
Efficient Safety Alignment of Language Models via Latent Personality Traits Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f939d13c-bfd1-46c4-a7d4-751fde27759e · inbound
Format Sensitivity Index: Token-Controlled Prompt Wrapper Robustness and Schema Compliance in LLM Benchmarking Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d107bd2-1d3c-4cd8-8594-ed31dca207d5 · inbound
When Counterbalancing Hides the Bias: Access-Conditioned Position Lock in Forced-Choice LLM Evaluation Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4252f286-08e7-4cb1-8455-ed83ca30c39f · inbound
Linear representations of grammaticality in neural language models Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 202
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e26db494-ebb4-4594-b5bd-f2c979823a99 · inbound
Structured Output Collapses Answer Diversity Across 44 Language Models Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a6c7e48-a5a8-457c-95de-9d0822855397 · inbound
Prompt Design at Scale: How Format, Instruction Count, and Context Length Shape Instruction Adherence and Hallucination in Large Language Models Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d850a5a-62e4-4273-8f1d-1f3e1059691d · inbound
A Knowledge-Injection Framework for Zero-Shot Adaptation of LLMs to Delirium Prediction Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d827d14-b1c1-4bd3-928a-5498c3224a70 · inbound
Reexamining zero-shot summarization: Empirical investigation of trustworthiness of LLM-summarizers Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b5a835c-c89d-4baa-bb83-3058e788e72b · inbound
Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c0b3f9e-5023-42cf-bfe0-6d65e5718c24 · inbound
Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b8dfe7b-3e8f-4616-9f54-dc2a0ac0623b · inbound
Routing Is Least Learnable Where It Is Most Valuable: Bounds on Representation Routing for Web Agents Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 138
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb688394-056d-41d5-aaf1-007eadd68b80 · inbound
Same physical state, different collective dynamics: state encodings select synchronization outcomes in language-model agents Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bdbbb002-5899-4cc3-8b58-02836ac5f396 · inbound
Persona Conditioning as an Assessor-Sensitivity Probe for LLM-Based IR Evaluation Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6de43ecb-a9bc-44da-99b5-bed25974c9c2 · inbound
Actions Speak Louder than Words: Measuring Cross-Lingual Policy Retention in Tool-Using Agents Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3bb96473-236a-497f-ba68-8dd57ed9a617 · inbound
QuoteBench: How Matched Scores Can Hide Command-Path Failures Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.