Pith. sign in

Paper Citation Record · LEDGER

Benchmarking Prompt Sensitivity in Large Language Models

As of 9 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 3 inbound Pith citation observations for arXiv:2502.06065.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.06065 v1

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T16:55:45.557734Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-20T13:40:04.275438Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T13:43:19.601845Z

Reference resolution

41 of 41 outbound references displayed

  • verified exact5
  • verified fuzzy11
  • unresolved23
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c7831059-049e-483c-b6a7-5ae62f7a32e5 · outbound

This paper cites an unresolved cited work.

Benchmarking Prompt Sensitivity in Large Language Models Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-08T16:55:46.262863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T16:55:45.426645Z digest=sha256:2b2567995793942cdf77f47acad2e4689f0a0ec59ecc35ef04630882fce721fc

Observation f0db9cb2-d921-4d21-9879-cb757f49591f · outbound

This paper cites an unresolved cited work.

Benchmarking Prompt Sensitivity in Large Language Models Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-08T16:55:46.254795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T16:55:45.430337Z digest=sha256:932f02d4fa559642a75463840ceb8eb6ccacbe93b608dceadc07bd469b910bf4

Observation 5be17900-1104-40a4-9899-7b9681e98a42 · outbound

This paper cites In: Al- Onaizan,Y.,Bansal,M.,Chen,Y.N.(eds.)Proceedingsofthe2024ConferenceonEmpirical Methods in Natural Language Processing.

Benchmarking Prompt Sensitivity in Large Language Models In: Al- Onaizan,Y.,Bansal,M.,Chen,Y.N.(eds.)Proceedingsofthe2024ConferenceonEmpirical Methods in Natural Language Processing

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T16:55:45.433342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:55:45.433342Z digest=sha256:511446ca99c5ff6fc81b959152eb12dc5ea78f0fdf62782548beee8b6e978148

Observation 2eea4f62-aa24-40e7-8473-e1da6528e81f · outbound

This paper cites In: European Conference on Information Retrieval.

Benchmarking Prompt Sensitivity in Large Language Models In: European Conference on Information Retrieval

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:55:46.246742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T16:55:45.436845Z digest=sha256:036661643a9e68e371ed74460b12cedd87797c66e836ed9dc4f4f412ff11b585

Observation dfc17dfd-c25a-4437-8f36-400d00802611 · outbound

This paper cites In: Proceedings of the 2024 AnnualInternationalACMSIGIRConferenceonResearchandDevelopmentinInformation Retrieval in the Asia Pacific Region.

Benchmarking Prompt Sensitivity in Large Language Models In: Proceedings of the 2024 AnnualInternationalACMSIGIRConferenceonResearchandDevelopmentinInformation Retrieval in the Asia Pacific Region

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:55:46.238705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T16:55:45.440027Z digest=sha256:392d8c66f2f1f3b867dd2d129bc617dfcc57677e795fa48e7de761767d1e16fc

Observation afea40b4-363b-46dc-9477-8eb6156e6d6b · outbound

This paper cites In: Proceedings of the 31st ACM International Conference on Information & Knowledge Management.

Benchmarking Prompt Sensitivity in Large Language Models In: Proceedings of the 31st ACM International Conference on Information & Knowledge Management

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:55:46.230432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T16:55:45.443134Z digest=sha256:5a98f53d11c037b983d058f471d4d61b4c3870bc4be6bb4378ddfed5fcf2816d

Observation 7540a4d6-a0e8-4039-9ebf-3ebec2efc0df · outbound

This paper cites an unresolved cited work.

Benchmarking Prompt Sensitivity in Large Language Models Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-08T16:55:46.222231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T16:55:45.446388Z digest=sha256:a158a53d6fef2ecf84552be2c88f7349c79c87e14fa3b8e888ed35cc44e7db9c

Observation ba1d22d8-203c-426b-a3f1-4a7e7bc106a8 · outbound

This paper cites an unresolved cited work.

Benchmarking Prompt Sensitivity in Large Language Models Unresolved cited work

Reference 8

Resolution
verified exact
doi, observed 2026-08-08T16:55:45.695016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T16:55:45.449259Z digest=sha256:7711fd99e77e4688bdd97c64910c76c9e8e10e970bee2fe5d7acd72505ae74e2

Observation ca08d4bd-05fa-42ca-88c8-50f68234115a · outbound

This paper cites In: Proceedings of the 28th ACM International Conference on Information and Knowledge Management.

Benchmarking Prompt Sensitivity in Large Language Models In: Proceedings of the 28th ACM International Conference on Information and Knowledge Management

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T16:55:45.452428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:55:45.452428Z digest=sha256:8f37202c8feadeefd9a51f719f2809a41afd02a02fa8a827b617f9513d25016c

Observation 5104d4e0-1104-4669-b20e-731c025fe304 · outbound

This paper cites What's the Magic Word? A Control Theory of LLM Prompting.

Benchmarking Prompt Sensitivity in Large Language Models What's the Magic Word? A Control Theory of LLM Prompting

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T16:55:45.455564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:55:45.455564Z digest=sha256:25b44b05cc4af1434f844f95c485a7c8601ef4265d3394d9a96fabf44ed4ea70

Observation abc38505-6812-43a0-a97b-1d0c66d75fdb · outbound

This paper cites In: European Conference on Information Retrieval.

Benchmarking Prompt Sensitivity in Large Language Models In: European Conference on Information Retrieval

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:55:46.213922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T16:55:45.459104Z digest=sha256:0a589ec09bdf8e9a653a93abc96ab2e4f245e362052c0ef61033b7130de8049b

Observation 522b75aa-f2cc-49d0-92e0-c56cc7f0eea2 · outbound

This paper cites an unresolved cited work.

Benchmarking Prompt Sensitivity in Large Language Models Unresolved cited work

Reference 12

Resolution
verified exact
doi, observed 2026-08-08T16:55:45.674197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T16:55:45.464402Z digest=sha256:3d62c651e4573519254d2bf0d3c1c43a45a38e917a224231736099743e3ffd3c

Observation 73b7a0a9-616b-40b3-b352-2b02ab9cff01 · outbound

This paper cites In: European Conference on Information Retrieval.

Benchmarking Prompt Sensitivity in Large Language Models In: European Conference on Information Retrieval

Reference 13

Resolution
malformed identifier
raw_fallback, observed 2026-08-08T16:55:46.205708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T16:55:45.467411Z digest=sha256:7cfc90a1b4685b343506cfd9c98875e7a8b631f490ac1716ce3f0e109945f709

Observation 0737b338-7258-4e24-833d-e987613140a4 · outbound

This paper cites In: Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Informa- tion Retrieval.

Benchmarking Prompt Sensitivity in Large Language Models In: Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Informa- tion Retrieval

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T16:55:45.470420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:55:45.470420Z digest=sha256:8d6b172fda88311244a1172b8961a933e0499818af2bcc44e2df27cc9a375b7b

Observation fb873f31-db0a-42d2-b04a-62f139adcc0c · outbound

This paper cites Unveiling and Manipulating Prompt Influence in Large Language Models.

Benchmarking Prompt Sensitivity in Large Language Models Unveiling and Manipulating Prompt Influence in Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T16:55:45.473373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:55:45.473373Z digest=sha256:d243f1b138b1c1c0669839ed297c971805c66cbe8d2edc403bcaf531a67e9247

Observation 47e15cfe-6d79-4e2d-ba23-a921e1f3f21b · outbound

This paper cites Information13(2), 83 (2022).

Benchmarking Prompt Sensitivity in Large Language Models Information13(2), 83 (2022)

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:55:46.197515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T16:55:45.476636Z digest=sha256:15aa2b4c65236417d570511314d066245a74b0361c3487ae77e1bab490ab0ba3

Observation fb4d5384-3382-4e8a-8f76-1a495dcddcc6 · outbound

This paper cites IEEE Access11, 76581–76604 (2023).

Benchmarking Prompt Sensitivity in Large Language Models IEEE Access11, 76581–76604 (2023)

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:55:46.189086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T16:55:45.479575Z digest=sha256:ad03a1f5e89ff65052b8824a4aa79f68fff0ac6c6900bb955dffa1c28b678ec6

Observation ba3f28d9-862f-4301-8967-b5991ef90d38 · outbound

This paper cites In: CIKM (2008).

Benchmarking Prompt Sensitivity in Large Language Models In: CIKM (2008)

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:55:46.180181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T16:55:45.482509Z digest=sha256:eddc7af37ca2804162448f9705b81f1ecb029d3a16f392176713ce242b49008b

Observation e50f2947-2f99-4298-a5e9-46bda323f118 · outbound

This paper cites In: Proceedings of the 33rd ACM International Confer- ence on Information and Knowledge Management.

Benchmarking Prompt Sensitivity in Large Language Models In: Proceedings of the 33rd ACM International Confer- ence on Information and Knowledge Management

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:55:46.171610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T16:55:45.485697Z digest=sha256:4f313c3727a617fddfb5308ccabc0841f9606d97e7b15c73493301a41f7931b1

Observation 1108b83d-6dbb-42b0-8f06-b3c492a20636 · outbound

This paper cites Mistral 7B.

Benchmarking Prompt Sensitivity in Large Language Models Mistral 7B

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T16:55:45.488945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:55:45.488945Z digest=sha256:dbcfa2722d2206d79f5630d7b678fa1f25a9df905070a501383d587e55f96a93

Observation 26b35b12-ec95-466d-8789-17d125c14868 · outbound

This paper cites In: Barzilay, R., Kan, M.Y.

Benchmarking Prompt Sensitivity in Large Language Models In: Barzilay, R., Kan, M.Y

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T16:55:45.492434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:55:45.492434Z digest=sha256:032508a775d92924b1a6229ac6d75b0ba74754a6d5a304bd4b177a24996ea1da

Observation 5d2ef4fe-0748-4656-be65-2175709938e6 · outbound

This paper cites In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers).

Benchmarking Prompt Sensitivity in Large Language Models In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T16:55:45.495831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:55:45.495831Z digest=sha256:201e413745c717853f898b8d26228f8b9d381125ee476a17b2bf67473a9cc8c9

Observation 3b7ffde7-27fd-4fc6-b0bb-735a7538e3a7 · outbound

This paper cites an unresolved cited work.

Benchmarking Prompt Sensitivity in Large Language Models Unresolved cited work

Reference 23

Resolution
verified exact
doi, observed 2026-08-08T16:55:45.651096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T16:55:45.499110Z digest=sha256:25d397781ddd96edff3ac41889f296baa596a5817816bee32c5bbf4f4da3be4c

Observation b729d661-c7d7-4087-9793-392ae8c4b0ea · outbound

This paper cites Internet Reference Services Quarterly27, 203 – 210 (2023).https://doi.org/10.1080/ 10875301.2023.2227621.

Benchmarking Prompt Sensitivity in Large Language Models Internet Reference Services Quarterly27, 203 – 210 (2023).https://doi.org/10.1080/ 10875301.2023.2227621

Reference 24

Resolution
malformed identifier
raw_fallback, observed 2026-08-08T16:55:45.980823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T16:55:45.502079Z digest=sha256:6894b891431e791ce436204baea0f612ce37e42b669ca656cb27f4827b6f51ad

Observation 7824525e-6804-474b-89bb-f2d8a9a39010 · outbound

This paper cites https://doi.org/10.18653/v1/2023.findings-emnlp.241, http: //dx.doi.org/10.18653/v1/2023.findings-emnlp.241.

Benchmarking Prompt Sensitivity in Large Language Models https://doi.org/10.18653/v1/2023.findings-emnlp.241, http: //dx.doi.org/10.18653/v1/2023.findings-emnlp.241

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T16:55:45.505163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:55:45.505163Z digest=sha256:3807b6d923e91a19714eacf49f2e8ff4a8fdddd4f8b0c017710260acd5f0961b

Observation e257f84d-e8bb-43cd-9f61-7491da07ad2c · outbound

This paper cites an unresolved cited work.

Benchmarking Prompt Sensitivity in Large Language Models Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T16:55:45.508370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:55:45.508370Z digest=sha256:e8493cba27df492520fd374d350d84fe39e6f0ce1b61cf62dd0c6aee2916295d

Observation d4604f86-0b0a-4cca-890d-2db2442aa080 · outbound

This paper cites Query Performance Prediction using Relevance Judgments Generated by Large Language Models.

Benchmarking Prompt Sensitivity in Large Language Models Query Performance Prediction using Relevance Judgments Generated by Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T16:55:45.511620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:55:45.511620Z digest=sha256:f78c2c2cf34e90e2f323e0251af4e5164aa193fddef3cf207b5bc2df9f5afd0a

Observation 93cfb014-0c4f-48aa-a81e-36368b23297d · outbound

This paper cites The Llama 3 Herd of Models.

Benchmarking Prompt Sensitivity in Large Language Models The Llama 3 Herd of Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T16:55:45.514952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:55:45.514952Z digest=sha256:1ccb5e4677bc3864de8fa722dbb1813c3cd235e60c6d6ac9245e07d2ab1db418

Observation 4639ae2e-21b8-4fb9-ab6f-ff2898196f44 · outbound

This paper cites Navigating Prompt Complexity for Zero-Shot Classification: A Study of Large Language Models in Computational Social Science.

Benchmarking Prompt Sensitivity in Large Language Models Navigating Prompt Complexity for Zero-Shot Classification: A Study of Large Language Models in Computational Social Science

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T16:55:45.518162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:55:45.518162Z digest=sha256:392defefe93730e1da33aeb20cdd33a4cab3b002b9b39907411181b98c09c9bb

Observation 5f4fe807-9dfd-41fe-8f26-d703dddce000 · outbound

This paper cites Testing LLMs on Code Generation with Varying Levels of Prompt Specificity.

Benchmarking Prompt Sensitivity in Large Language Models Testing LLMs on Code Generation with Varying Levels of Prompt Specificity

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T16:55:45.521371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:55:45.521371Z digest=sha256:ff048f59e1bf8f7a489a516510a0383b5837f409b8b97a7ec80992fb77ffee00

Observation 167f177b-1e6d-426a-84e8-7a3332fc9d1f · outbound

This paper cites PQPP: A Joint Benchmark for Text-to-Image Prompt and Query Performance Prediction.

Benchmarking Prompt Sensitivity in Large Language Models PQPP: A Joint Benchmark for Text-to-Image Prompt and Query Performance Prediction

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-08T16:55:45.792127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T16:55:45.524418Z digest=sha256:4a4ff48f1bd5259386a215160c36fcd11ec4c795aa27264d3ca67c5839366fd8

Observation 1c0b5fbc-294d-4adf-8817-933f07008425 · outbound

This paper cites Semantic Consistency for Assuring Reliability of Large Language Models.

Benchmarking Prompt Sensitivity in Large Language Models Semantic Consistency for Assuring Reliability of Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T16:55:45.527516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:55:45.527516Z digest=sha256:a1b6f973362e831d8eddc77ebf5b00369f65cb55eb3185411064f0e41a69aade

Observation a3f0f6bb-5e0a-4821-b713-52e4edd0a935 · outbound

This paper cites In: Proceedings of the 2024 Annual International ACM SIGIR Conference on Research and Development in Information Retrieval in the Asia Pacific Region.

Benchmarking Prompt Sensitivity in Large Language Models In: Proceedings of the 2024 Annual International ACM SIGIR Conference on Research and Development in Information Retrieval in the Asia Pacific Region

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T16:55:45.531616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:55:45.531616Z digest=sha256:b45cb6e2850a8409ed1be051ab2e56e8bf093496c5d2d51289dc890c5b3a3665

Observation 38e2003d-8cca-4612-b730-81796aadcae7 · outbound

This paper cites In: Proceedings of the 32nd ACM Inter- national Conference on Information and Knowledge Management.

Benchmarking Prompt Sensitivity in Large Language Models In: Proceedings of the 32nd ACM Inter- national Conference on Information and Knowledge Management

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:55:46.162750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T16:55:45.534670Z digest=sha256:12ab320ab6033c7e3a4af1e375f3b5f3690e72f4352be6974a6de2982c46cd9b

Observation f4ac3d28-1374-40aa-84f8-2b7d1076fb6c · outbound

This paper cites In: European Conference on Information Re- trieval.pp.30–39.Springer(2024).

Benchmarking Prompt Sensitivity in Large Language Models In: European Conference on Information Re- trieval.pp.30–39.Springer(2024)

Reference 35

Resolution
verified exact
doi, observed 2026-08-08T16:55:45.604745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T16:55:45.537815Z digest=sha256:46eb64d65bea6f18e7490a0c29300155e20d6b0d1a811eeac14061b339c73934

Observation 11dfe6a2-4267-4f6a-ade1-3359f6563496 · outbound

This paper cites Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting.

Benchmarking Prompt Sensitivity in Large Language Models Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-08T16:55:45.540970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:55:45.540970Z digest=sha256:e94adb5cff844d656059d4e07edd5cf06467f7122a3a102b259e667fa104bdbe

Observation 5502d8ec-7031-42d5-88ac-7a32dcfb7701 · outbound

This paper cites In: Proceedings of the 34th International Conference on Neural Information Processing Systems.

Benchmarking Prompt Sensitivity in Large Language Models In: Proceedings of the 34th International Conference on Neural Information Processing Systems

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:55:46.152681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T16:55:45.544449Z digest=sha256:ba90014995a9911f86c7e34130451806e7a103a0f46824debe6e87eac9f68bbd

Observation fe230fa3-32ae-4b6d-9ef6-cd5fe05a1870 · outbound

This paper cites eugeneyan.com (Aug 2024),https://eugeneyan.com/writing/llm-evaluators/.

Benchmarking Prompt Sensitivity in Large Language Models eugeneyan.com (Aug 2024),https://eugeneyan.com/writing/llm-evaluators/

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:55:46.142696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T16:55:45.547529Z digest=sha256:dfa48ace08d91e2df1a96bd5d961cb3e62399832eef18aa23eb38bca3b5dfc0e

Observation d5ffa9f8-6bd1-4e36-88e9-5b82e781e062 · outbound

This paper cites an unresolved cited work.

Benchmarking Prompt Sensitivity in Large Language Models Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-08T16:55:45.550593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:55:45.550593Z digest=sha256:1cfc4b326ef5f52dfe148a494db9c471f9b220dc9beb0d5e6507d6910adb6894

Observation b7267df2-c433-47f8-99f8-1efbc8a7c390 · outbound

This paper cites PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts.

Benchmarking Prompt Sensitivity in Large Language Models PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-08T16:55:45.553837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:55:45.553837Z digest=sha256:2fdc7f24b33ae4c502c9cb0b78c1c5195a93aa21bd28fc962b82a3c181e16ae1

Observation 0396fec9-6d6e-4841-9db6-cc8f4f3a90de · outbound

This paper cites ProSA: Assessing and Understanding the Prompt Sensitivity of LLMs.

Benchmarking Prompt Sensitivity in Large Language Models ProSA: Assessing and Understanding the Prompt Sensitivity of LLMs

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-08T16:55:45.557734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:55:45.557734Z digest=sha256:269ced6e5949effdc403898f0016fee93988e65020352dc375e5374f2eab02e2

Pith citing papers

Observation 7e8f3c14-1ffe-43e3-98ed-a7d2cb6e2295 · inbound

Understanding the Mechanism of Altruism in Large Language Models cites this paper.

Understanding the Mechanism of Altruism in Large Language Models Benchmarking Prompt Sensitivity in Large Language Models

Reference 192

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T13:31:02.511069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T01:36:50.329664Z digest=sha256:bbf684bb311ca59ec21d36bd4e48cddc0a5722b41ce5c8faff554434160f3269

Observation ac8129f6-3efa-4cfa-9687-e94cbfea7a27 · inbound

Coordinates of Capability: A Unified MTMM-Geometric Framework for LLM Evaluation cites this paper.

Coordinates of Capability: A Unified MTMM-Geometric Framework for LLM Evaluation Benchmarking Prompt Sensitivity in Large Language Models

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T01:46:14.436745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-12T01:41:42.003483Z digest=sha256:f39f915651f385a8edc35e03ad3c5828367bc6be86cc4a89c78132de7b36fd90

Observation 9eacf9cf-1ef4-420b-bfd1-49665c342223 · inbound

Stop Drawing Scientific Claims from LLM Social Simulations Without Robustness Audits cites this paper.

Stop Drawing Scientific Claims from LLM Social Simulations Without Robustness Audits Benchmarking Prompt Sensitivity in Large Language Models

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:43:19.603517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T13:40:04.275438Z digest=sha256:bfcbcde30fa70e58c468a709ed934c9c3689bde9a6ca44cee8dd0d8bf285f737