Pith. sign in

Paper Citation Record · LEDGER

Benchmarking Prompt Sensitivity in Large Language Models

As of 9 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 3 inbound Pith citation observations for arXiv:2502.06065.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.06065 v1

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T16:55:45.557734Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-20T13:40:04.275438Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T13:43:19.601845Z

Reference resolution

41 of 41 outbound references displayed

  • verified exact5
  • verified fuzzy11
  • unresolved23
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c7831059-049e-483c-b6a7-5ae62f7a32e5 · outbound

This paper cites an unresolved cited work.

Benchmarking Prompt Sensitivity in Large Language Models Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-08T16:55:46.262863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T16:55:45.426645Z digest=sha256:7c46588e48e1c1a0c6f66c5bb27c518d56de267ee886a527e2bc3eab54bcda86

Observation f0db9cb2-d921-4d21-9879-cb757f49591f · outbound

This paper cites an unresolved cited work.

Benchmarking Prompt Sensitivity in Large Language Models Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-08T16:55:46.254795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T16:55:45.430337Z digest=sha256:cbd5a003791b0093ec473983ed83e6c742cce10c2453b20dc4d9700578669bf1

Observation 5be17900-1104-40a4-9899-7b9681e98a42 · outbound

This paper cites In: Al- Onaizan,Y.,Bansal,M.,Chen,Y.N.(eds.)Proceedingsofthe2024ConferenceonEmpirical Methods in Natural Language Processing.

Benchmarking Prompt Sensitivity in Large Language Models In: Al- Onaizan,Y.,Bansal,M.,Chen,Y.N.(eds.)Proceedingsofthe2024ConferenceonEmpirical Methods in Natural Language Processing

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T16:55:45.433342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:55:45.433342Z digest=sha256:511446ca99c5ff6fc81b959152eb12dc5ea78f0fdf62782548beee8b6e978148

Observation 2eea4f62-aa24-40e7-8473-e1da6528e81f · outbound

This paper cites In: European Conference on Information Retrieval.

Benchmarking Prompt Sensitivity in Large Language Models In: European Conference on Information Retrieval

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:55:46.246742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T16:55:45.436845Z digest=sha256:64ff7034eb729d18be0b8a34694d150c6431621114971ec59e2a8827e321c770

Observation dfc17dfd-c25a-4437-8f36-400d00802611 · outbound

This paper cites In: Proceedings of the 2024 AnnualInternationalACMSIGIRConferenceonResearchandDevelopmentinInformation Retrieval in the Asia Pacific Region.

Benchmarking Prompt Sensitivity in Large Language Models In: Proceedings of the 2024 AnnualInternationalACMSIGIRConferenceonResearchandDevelopmentinInformation Retrieval in the Asia Pacific Region

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:55:46.238705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T16:55:45.440027Z digest=sha256:04e0759b5313f684ddce8254b5f415c1cfa6bed3c11b375dc108fff822507c5d

Observation afea40b4-363b-46dc-9477-8eb6156e6d6b · outbound

This paper cites In: Proceedings of the 31st ACM International Conference on Information & Knowledge Management.

Benchmarking Prompt Sensitivity in Large Language Models In: Proceedings of the 31st ACM International Conference on Information & Knowledge Management

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:55:46.230432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T16:55:45.443134Z digest=sha256:e25916c70915b7462898863586bd5be866f6258dad08b1f862d52391efea7aaf

Observation 7540a4d6-a0e8-4039-9ebf-3ebec2efc0df · outbound

This paper cites an unresolved cited work.

Benchmarking Prompt Sensitivity in Large Language Models Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-08T16:55:46.222231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T16:55:45.446388Z digest=sha256:bd56ba2937cdecfe52ef59b47b38f963cabdb304630cf0c5762c31532e244d33

Observation ba1d22d8-203c-426b-a3f1-4a7e7bc106a8 · outbound

This paper cites an unresolved cited work.

Benchmarking Prompt Sensitivity in Large Language Models Unresolved cited work

Reference 8

Resolution
verified exact
doi, observed 2026-08-08T16:55:45.695016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T16:55:45.449259Z digest=sha256:e499aad82c28e1e139a27de8b6b69637aa0c6fafc6f30d0556d5f0f55e9eccd9

Observation ca08d4bd-05fa-42ca-88c8-50f68234115a · outbound

This paper cites In: Proceedings of the 28th ACM International Conference on Information and Knowledge Management.

Benchmarking Prompt Sensitivity in Large Language Models In: Proceedings of the 28th ACM International Conference on Information and Knowledge Management

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T16:55:45.452428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:55:45.452428Z digest=sha256:8f37202c8feadeefd9a51f719f2809a41afd02a02fa8a827b617f9513d25016c

Observation 5104d4e0-1104-4669-b20e-731c025fe304 · outbound

This paper cites What's the Magic Word? A Control Theory of LLM Prompting.

Benchmarking Prompt Sensitivity in Large Language Models What's the Magic Word? A Control Theory of LLM Prompting

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T16:55:45.455564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:55:45.455564Z digest=sha256:25b44b05cc4af1434f844f95c485a7c8601ef4265d3394d9a96fabf44ed4ea70

Observation abc38505-6812-43a0-a97b-1d0c66d75fdb · outbound

This paper cites In: European Conference on Information Retrieval.

Benchmarking Prompt Sensitivity in Large Language Models In: European Conference on Information Retrieval

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:55:46.213922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T16:55:45.459104Z digest=sha256:155b85fd1bf4ec6b548d17764e7cec7e068e5b2c89a8f731e8d603e0907cb411

Observation 522b75aa-f2cc-49d0-92e0-c56cc7f0eea2 · outbound

This paper cites an unresolved cited work.

Benchmarking Prompt Sensitivity in Large Language Models Unresolved cited work

Reference 12

Resolution
verified exact
doi, observed 2026-08-08T16:55:45.674197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T16:55:45.464402Z digest=sha256:7f26edce0ebad30e6356900674224068bbbff260aa850337fdadedbce3a82875

Observation 73b7a0a9-616b-40b3-b352-2b02ab9cff01 · outbound

This paper cites In: European Conference on Information Retrieval.

Benchmarking Prompt Sensitivity in Large Language Models In: European Conference on Information Retrieval

Reference 13

Resolution
malformed identifier
raw_fallback, observed 2026-08-08T16:55:46.205708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T16:55:45.467411Z digest=sha256:1c4b3806eaaeec01d0ebdb8b25352289f372e99567eab1342ca2943dcff591f7

Observation 0737b338-7258-4e24-833d-e987613140a4 · outbound

This paper cites In: Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Informa- tion Retrieval.

Benchmarking Prompt Sensitivity in Large Language Models In: Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Informa- tion Retrieval

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T16:55:45.470420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:55:45.470420Z digest=sha256:8d6b172fda88311244a1172b8961a933e0499818af2bcc44e2df27cc9a375b7b

Observation fb873f31-db0a-42d2-b04a-62f139adcc0c · outbound

This paper cites Unveiling and Manipulating Prompt Influence in Large Language Models.

Benchmarking Prompt Sensitivity in Large Language Models Unveiling and Manipulating Prompt Influence in Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T16:55:45.473373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:55:45.473373Z digest=sha256:d243f1b138b1c1c0669839ed297c971805c66cbe8d2edc403bcaf531a67e9247

Observation 47e15cfe-6d79-4e2d-ba23-a921e1f3f21b · outbound

This paper cites Information13(2), 83 (2022).

Benchmarking Prompt Sensitivity in Large Language Models Information13(2), 83 (2022)

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:55:46.197515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T16:55:45.476636Z digest=sha256:3955dacd41be782eb0448c1d077d915a4fb06ea50eab9f982542df3cc8c5c315

Observation fb4d5384-3382-4e8a-8f76-1a495dcddcc6 · outbound

This paper cites IEEE Access11, 76581–76604 (2023).

Benchmarking Prompt Sensitivity in Large Language Models IEEE Access11, 76581–76604 (2023)

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:55:46.189086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T16:55:45.479575Z digest=sha256:805b4603056914b78329e9b555546ebcdea83233005ea38f02441f6658b17f63

Observation ba3f28d9-862f-4301-8967-b5991ef90d38 · outbound

This paper cites In: CIKM (2008).

Benchmarking Prompt Sensitivity in Large Language Models In: CIKM (2008)

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:55:46.180181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T16:55:45.482509Z digest=sha256:190adfddf4ac7cd78db693010fcb728f0167415d81f33cef91a632478a3f390f

Observation e50f2947-2f99-4298-a5e9-46bda323f118 · outbound

This paper cites In: Proceedings of the 33rd ACM International Confer- ence on Information and Knowledge Management.

Benchmarking Prompt Sensitivity in Large Language Models In: Proceedings of the 33rd ACM International Confer- ence on Information and Knowledge Management

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:55:46.171610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T16:55:45.485697Z digest=sha256:5b420fb2922391b11df88eed87c6f37b6a2a55d8cab38b356a4c3a3186157257

Observation 1108b83d-6dbb-42b0-8f06-b3c492a20636 · outbound

This paper cites Mistral 7B.

Benchmarking Prompt Sensitivity in Large Language Models Mistral 7B

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T16:55:45.488945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:55:45.488945Z digest=sha256:dbcfa2722d2206d79f5630d7b678fa1f25a9df905070a501383d587e55f96a93

Observation 26b35b12-ec95-466d-8789-17d125c14868 · outbound

This paper cites In: Barzilay, R., Kan, M.Y.

Benchmarking Prompt Sensitivity in Large Language Models In: Barzilay, R., Kan, M.Y

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T16:55:45.492434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:55:45.492434Z digest=sha256:032508a775d92924b1a6229ac6d75b0ba74754a6d5a304bd4b177a24996ea1da

Observation 5d2ef4fe-0748-4656-be65-2175709938e6 · outbound

This paper cites In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers).

Benchmarking Prompt Sensitivity in Large Language Models In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T16:55:45.495831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:55:45.495831Z digest=sha256:201e413745c717853f898b8d26228f8b9d381125ee476a17b2bf67473a9cc8c9

Observation 3b7ffde7-27fd-4fc6-b0bb-735a7538e3a7 · outbound

This paper cites an unresolved cited work.

Benchmarking Prompt Sensitivity in Large Language Models Unresolved cited work

Reference 23

Resolution
verified exact
doi, observed 2026-08-08T16:55:45.651096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T16:55:45.499110Z digest=sha256:18db0835597b500d8e5091db1b3bfbc7a574205b553bc81ddff3b07441ac4bc6

Observation b729d661-c7d7-4087-9793-392ae8c4b0ea · outbound

This paper cites Internet Reference Services Quarterly27, 203 – 210 (2023).https://doi.org/10.1080/ 10875301.2023.2227621.

Benchmarking Prompt Sensitivity in Large Language Models Internet Reference Services Quarterly27, 203 – 210 (2023).https://doi.org/10.1080/ 10875301.2023.2227621

Reference 24

Resolution
malformed identifier
raw_fallback, observed 2026-08-08T16:55:45.980823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T16:55:45.502079Z digest=sha256:0afeaf1c73e34f686a56cafb1f801af2bd2d8ee856645a9f897b9a7a73e0d142

Observation 7824525e-6804-474b-89bb-f2d8a9a39010 · outbound

This paper cites https://doi.org/10.18653/v1/2023.findings-emnlp.241, http: //dx.doi.org/10.18653/v1/2023.findings-emnlp.241.

Benchmarking Prompt Sensitivity in Large Language Models https://doi.org/10.18653/v1/2023.findings-emnlp.241, http: //dx.doi.org/10.18653/v1/2023.findings-emnlp.241

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T16:55:45.505163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:55:45.505163Z digest=sha256:3807b6d923e91a19714eacf49f2e8ff4a8fdddd4f8b0c017710260acd5f0961b

Observation e257f84d-e8bb-43cd-9f61-7491da07ad2c · outbound

This paper cites an unresolved cited work.

Benchmarking Prompt Sensitivity in Large Language Models Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T16:55:45.508370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:55:45.508370Z digest=sha256:e8493cba27df492520fd374d350d84fe39e6f0ce1b61cf62dd0c6aee2916295d

Observation d4604f86-0b0a-4cca-890d-2db2442aa080 · outbound

This paper cites Query Performance Prediction using Relevance Judgments Generated by Large Language Models.

Benchmarking Prompt Sensitivity in Large Language Models Query Performance Prediction using Relevance Judgments Generated by Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T16:55:45.511620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:55:45.511620Z digest=sha256:f78c2c2cf34e90e2f323e0251af4e5164aa193fddef3cf207b5bc2df9f5afd0a

Observation 93cfb014-0c4f-48aa-a81e-36368b23297d · outbound

This paper cites The Llama 3 Herd of Models.

Benchmarking Prompt Sensitivity in Large Language Models The Llama 3 Herd of Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T16:55:45.514952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:55:45.514952Z digest=sha256:1ccb5e4677bc3864de8fa722dbb1813c3cd235e60c6d6ac9245e07d2ab1db418

Observation 4639ae2e-21b8-4fb9-ab6f-ff2898196f44 · outbound

This paper cites Navigating Prompt Complexity for Zero-Shot Classification: A Study of Large Language Models in Computational Social Science.

Benchmarking Prompt Sensitivity in Large Language Models Navigating Prompt Complexity for Zero-Shot Classification: A Study of Large Language Models in Computational Social Science

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T16:55:45.518162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:55:45.518162Z digest=sha256:392defefe93730e1da33aeb20cdd33a4cab3b002b9b39907411181b98c09c9bb

Observation 5f4fe807-9dfd-41fe-8f26-d703dddce000 · outbound

This paper cites Testing LLMs on Code Generation with Varying Levels of Prompt Specificity.

Benchmarking Prompt Sensitivity in Large Language Models Testing LLMs on Code Generation with Varying Levels of Prompt Specificity

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T16:55:45.521371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:55:45.521371Z digest=sha256:ff048f59e1bf8f7a489a516510a0383b5837f409b8b97a7ec80992fb77ffee00

Observation 167f177b-1e6d-426a-84e8-7a3332fc9d1f · outbound

This paper cites PQPP: A Joint Benchmark for Text-to-Image Prompt and Query Performance Prediction.

Benchmarking Prompt Sensitivity in Large Language Models PQPP: A Joint Benchmark for Text-to-Image Prompt and Query Performance Prediction

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-08T16:55:45.792127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T16:55:45.524418Z digest=sha256:9b3b04d0bcca694df691fd4596ac2aa0912963ffc6495b9bfd6877ca68354328

Observation 1c0b5fbc-294d-4adf-8817-933f07008425 · outbound

This paper cites Semantic Consistency for Assuring Reliability of Large Language Models.

Benchmarking Prompt Sensitivity in Large Language Models Semantic Consistency for Assuring Reliability of Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T16:55:45.527516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:55:45.527516Z digest=sha256:a1b6f973362e831d8eddc77ebf5b00369f65cb55eb3185411064f0e41a69aade

Observation a3f0f6bb-5e0a-4821-b713-52e4edd0a935 · outbound

This paper cites In: Proceedings of the 2024 Annual International ACM SIGIR Conference on Research and Development in Information Retrieval in the Asia Pacific Region.

Benchmarking Prompt Sensitivity in Large Language Models In: Proceedings of the 2024 Annual International ACM SIGIR Conference on Research and Development in Information Retrieval in the Asia Pacific Region

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T16:55:45.531616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:55:45.531616Z digest=sha256:b45cb6e2850a8409ed1be051ab2e56e8bf093496c5d2d51289dc890c5b3a3665

Observation 38e2003d-8cca-4612-b730-81796aadcae7 · outbound

This paper cites In: Proceedings of the 32nd ACM Inter- national Conference on Information and Knowledge Management.

Benchmarking Prompt Sensitivity in Large Language Models In: Proceedings of the 32nd ACM Inter- national Conference on Information and Knowledge Management

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:55:46.162750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T16:55:45.534670Z digest=sha256:f3becda45513f9d2bd6890c64b03800077c124b8f331ec6fd87e197b0263e098

Observation f4ac3d28-1374-40aa-84f8-2b7d1076fb6c · outbound

This paper cites In: European Conference on Information Re- trieval.pp.30–39.Springer(2024).

Benchmarking Prompt Sensitivity in Large Language Models In: European Conference on Information Re- trieval.pp.30–39.Springer(2024)

Reference 35

Resolution
verified exact
doi, observed 2026-08-08T16:55:45.604745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T16:55:45.537815Z digest=sha256:5afe6f1770c11aac12914771f69d389f293833eb64aceca3f68c4de8760b3877

Observation 11dfe6a2-4267-4f6a-ade1-3359f6563496 · outbound

This paper cites Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting.

Benchmarking Prompt Sensitivity in Large Language Models Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-08T16:55:45.540970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:55:45.540970Z digest=sha256:e94adb5cff844d656059d4e07edd5cf06467f7122a3a102b259e667fa104bdbe

Observation 5502d8ec-7031-42d5-88ac-7a32dcfb7701 · outbound

This paper cites In: Proceedings of the 34th International Conference on Neural Information Processing Systems.

Benchmarking Prompt Sensitivity in Large Language Models In: Proceedings of the 34th International Conference on Neural Information Processing Systems

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:55:46.152681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T16:55:45.544449Z digest=sha256:70c1e4b5f1c04ba750c78a27d9cd2fbfcaf8c136877524dfec85deb6c2b1de29

Observation fe230fa3-32ae-4b6d-9ef6-cd5fe05a1870 · outbound

This paper cites eugeneyan.com (Aug 2024),https://eugeneyan.com/writing/llm-evaluators/.

Benchmarking Prompt Sensitivity in Large Language Models eugeneyan.com (Aug 2024),https://eugeneyan.com/writing/llm-evaluators/

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:55:46.142696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T16:55:45.547529Z digest=sha256:529a6b97af2e1a817f6bbb5d30198112809a0d6334fcc83b28da2f8ee5b19f40

Observation d5ffa9f8-6bd1-4e36-88e9-5b82e781e062 · outbound

This paper cites an unresolved cited work.

Benchmarking Prompt Sensitivity in Large Language Models Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-08T16:55:45.550593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:55:45.550593Z digest=sha256:1cfc4b326ef5f52dfe148a494db9c471f9b220dc9beb0d5e6507d6910adb6894

Observation b7267df2-c433-47f8-99f8-1efbc8a7c390 · outbound

This paper cites PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts.

Benchmarking Prompt Sensitivity in Large Language Models PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-08T16:55:45.553837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:55:45.553837Z digest=sha256:2fdc7f24b33ae4c502c9cb0b78c1c5195a93aa21bd28fc962b82a3c181e16ae1

Observation 0396fec9-6d6e-4841-9db6-cc8f4f3a90de · outbound

This paper cites ProSA: Assessing and Understanding the Prompt Sensitivity of LLMs.

Benchmarking Prompt Sensitivity in Large Language Models ProSA: Assessing and Understanding the Prompt Sensitivity of LLMs

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-08T16:55:45.557734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:55:45.557734Z digest=sha256:269ced6e5949effdc403898f0016fee93988e65020352dc375e5374f2eab02e2

Pith citing papers

Observation 7e8f3c14-1ffe-43e3-98ed-a7d2cb6e2295 · inbound

Understanding the Mechanism of Altruism in Large Language Models cites this paper.

Understanding the Mechanism of Altruism in Large Language Models Benchmarking Prompt Sensitivity in Large Language Models

Reference 192

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T13:31:02.511069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T01:36:50.329664Z digest=sha256:555877bf59036722b9ac3bfb70803d495afa20aacb6d423ebf910b3da0bc97ea

Observation ac8129f6-3efa-4cfa-9687-e94cbfea7a27 · inbound

Coordinates of Capability: A Unified MTMM-Geometric Framework for LLM Evaluation cites this paper.

Coordinates of Capability: A Unified MTMM-Geometric Framework for LLM Evaluation Benchmarking Prompt Sensitivity in Large Language Models

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T01:46:14.436745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-12T01:41:42.003483Z digest=sha256:6d4fd0cb95e15a85eeb9c1176b01f12dd8ac9b4dbed1dbd55115f9d09838213e

Observation 9eacf9cf-1ef4-420b-bfd1-49665c342223 · inbound

Stop Drawing Scientific Claims from LLM Social Simulations Without Robustness Audits cites this paper.

Stop Drawing Scientific Claims from LLM Social Simulations Without Robustness Audits Benchmarking Prompt Sensitivity in Large Language Models

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:43:19.603517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T13:40:04.275438Z digest=sha256:daae4a40ceda2884d8b720e2c8e5b68d764bc6f729c629467c6532da6c124cc1