Pith. sign in

Paper Citation Record · LEDGER

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance

As of 19 August 2026, this Paper Citation Record lists 88 of 88 outbound references and 0 inbound Pith citation observations for arXiv:2608.11694.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.11694 v1

Coverage vector

measured 88 of 88 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:39:52.407918Z

measured 88 of 88 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

88 of 88 outbound references displayed

  • verified exact0
  • verified fuzzy21
  • unresolved66
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4a0b3aa5-66dd-480f-8755-8c8ed4d79bde · outbound

This paper cites American Psychologist , volume=.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance American Psychologist , volume=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:52.062361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:39:52.062361Z digest=sha256:70bccd1fbe9e227ae74e4e879ae87e23b729c04f0d5c1ab8e8cbc3f74b8f1277

Observation 3dd46683-4846-42e5-98c0-5c1d46a7b828 · outbound

This paper cites Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , pages=.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , pages=

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:52.066230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:39:52.066230Z digest=sha256:07e2f77caf690b3119d16cf931c9f82da33b1ad965a74629e9153a9577f7b309

Observation b9642f4d-e66c-4113-b314-f253f07d72ea · outbound

This paper cites Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing , pages=.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing , pages=

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:52.073148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:39:52.073148Z digest=sha256:c5196e431dc4a9458fee79ccc3a40ed270bf33659ac79ec6171e4e6fd8d89295

Observation 109456f6-e665-4720-9b2f-d386afaa6f09 · outbound

This paper cites A Dataset for Answering Time-Sensitive Questions.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance A Dataset for Answering Time-Sensitive Questions

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:52.076592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:39:52.076592Z digest=sha256:90a97e3cb5ed2868dbaf200019095b2c381a44324000f4c6b06f2cf862157e79

Observation 24f16c8e-c5b0-4c23-aa29-5710360be67a · outbound

This paper cites Holistic Evaluation of Language Models.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance Holistic Evaluation of Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:52.080134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:39:52.080134Z digest=sha256:6509aaa02f2e27c798544b69cc19ec3a5aae9f53d390771195c63caa83ba8372

Observation 06e35ed0-ecf8-4fa0-9da9-44fa5be352f6 · outbound

This paper cites Psychometrika , volume=.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance Psychometrika , volume=

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:52.083429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:39:52.083429Z digest=sha256:d969316621ef36271743184831ea810a55bc96afafd2829ee90c28899ad5f2fe

Observation 1c3c8517-7422-43d5-9a48-0ee24479e919 · outbound

This paper cites Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics , pages=.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics , pages=

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:52.087044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:39:52.087044Z digest=sha256:e21f066c4001d3ac842424fc7832e3b21b9c03a5369315cbd524d4e80d5fcebc

Observation ed6785fc-6d2c-49ec-b036-947f4f27296f · outbound

This paper cites IEEE Access , volume=.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance IEEE Access , volume=

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:52.090158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:39:52.090158Z digest=sha256:b817af92c475b8e28f79eeb9daf93c20fc4f0a55b01df8ec6870d0351d2c2acd

Observation c93bfb25-5b3e-48ae-9732-18aef02ab399 · outbound

This paper cites Transformed-Linear Models for Time Series Extremes.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance Transformed-Linear Models for Time Series Extremes

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-08-16T00:39:52.697193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T00:39:52.093334Z digest=sha256:d4f2f3fcdbde59708baa2980483d6e5af54c0032ce0b198144acd4a3e83cc082

Observation 5f8c32a0-f5e8-4940-9a44-6dffdd8e65f1 · outbound

This paper cites ACM Transactions on Intelligent Systems and Technology , year=.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance ACM Transactions on Intelligent Systems and Technology , year=

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:52.096901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:39:52.096901Z digest=sha256:a24a944c5ff69d849666165f5d34fe926542c3d0d265fbfc95aab469b2cb1ef2

Observation fb2c94a3-3d22-4283-940d-7ec4970dcae0 · outbound

This paper cites Proceedings of ACL , pages=.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance Proceedings of ACL , pages=

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:52.100276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:39:52.100276Z digest=sha256:1fcf71856b5edf9c9fa866352affddb583a9c113433664c0058feeda5cdadd08

Observation 7f34a779-b703-44c9-9ee9-c204200b007c · outbound

This paper cites Ignore Previous Prompt: Attack Techniques For Language Models.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance Ignore Previous Prompt: Attack Techniques For Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:52.103724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:39:52.103724Z digest=sha256:9006d7e4b25222e50ad7ce85d60d82bcc831e21803ae2701170888b7cf68c2dc

Observation a368030c-3898-4c62-9d52-31ef4d66fad4 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:52.107243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:39:52.107243Z digest=sha256:2e470c53cf4bcce09d690b90ca88bd98dad84208362c44da54a8fac74a896c52

Observation 072b0d71-7e86-4954-8c51-6d0ff14ca349 · outbound

This paper cites Trusting Your Evidence: Hallucinate Less with Context-aware Decoding.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance Trusting Your Evidence: Hallucinate Less with Context-aware Decoding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:52.117507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:39:52.117507Z digest=sha256:fea551739748f364f474391d07d14078c27a3b3f5606620d66d48303d43df630

Observation f7fbf4f4-ac29-4026-be64-e95332c41bbb · outbound

This paper cites Distill , year=.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance Distill , year=

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:52.128030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:39:52.128030Z digest=sha256:009bb8d59c058644223302ba3a67c3b2b6567f97d07d9d7eb23a205b71f8dfe0

Observation 9a2fa57f-161e-4465-aaee-5312862ce30c · outbound

This paper cites ACM Computing Surveys , volume=.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance ACM Computing Surveys , volume=

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:39:53.050672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T00:39:52.131470Z digest=sha256:75748db5bfe08d17397ca8dfa4126b0ee3c9da6f058ce8fea136e3791ef351bc

Observation 672faa8e-4a09-4dd5-89ff-6203d73766f2 · outbound

This paper cites arXiv preprint arXiv:2401.xxxxx , year=.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance arXiv preprint arXiv:2401.xxxxx , year=

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:39:53.040611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T00:39:52.135446Z digest=sha256:27d480fc351a9d360c17bab79ae9a588406792cccf8271ce89ab49a7e1a6d129

Observation 8d6663f4-c8f0-4476-a826-29a807ccb08f · outbound

This paper cites Debug like a Human: A Large Language Model Debugger via Verifying Runtime Execution Step-by-step.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance Debug like a Human: A Large Language Model Debugger via Verifying Runtime Execution Step-by-step

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:52.142234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:39:52.142234Z digest=sha256:5fc52d855dcfb9f668c41d8052ae6a295699eacfc38ad72f26cd5e25ec93b6e8

Observation cac133d0-4b3c-4ea3-9cd9-07445236a2f9 · outbound

This paper cites arXiv preprint arXiv:2312.xxxxx , year=.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance arXiv preprint arXiv:2312.xxxxx , year=

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:39:53.031352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T00:39:52.146342Z digest=sha256:409e49fc6d1f7f73219190b267ddbe803a731cf27113b74222cd5a8bd043788e

Observation a473920b-5e73-46bc-99f0-d9ab77f7ac1d · outbound

This paper cites Teaching Large Language Models to Self-Debug.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance Teaching Large Language Models to Self-Debug

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:52.149556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:39:52.149556Z digest=sha256:573599340c8b7c33a168481d0a48cb72e95f4bb17c950f7a976b090d90b15f7f

Observation b01fae02-8e58-4b60-b5a3-ceea22c4625c · outbound

This paper cites Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:52.164587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:39:52.164587Z digest=sha256:6acbace8aca914642dbf5a74e43491773ac897b6c36a50a4e40bee9fd52ba29c

Observation 2d59c4a8-b4fc-40b4-b858-d269151199fd · outbound

This paper cites Proceedings of ICLR , year=.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance Proceedings of ICLR , year=

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:52.167887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:39:52.167887Z digest=sha256:fadbbb868dfdb7b90b2fd6fd0ab9b1edb4a136a66ab9c864f16769d5de9c870c

Observation c93766a7-0435-4550-b113-c690e6875793 · outbound

This paper cites Advances in Neural Information Processing Systems , pages=.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance Advances in Neural Information Processing Systems , pages=

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:39:53.018148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T00:39:52.174325Z digest=sha256:9ab1f0271887b8b5e9a9c158c0d4f4cf6c1a33259fab8ddbab652f4bd630a139

Observation 83042d05-e2c3-4d49-8890-876883ea9f6d · outbound

This paper cites Least-to-Most Prompting Enables Complex Reasoning in Large Language Models.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance Least-to-Most Prompting Enables Complex Reasoning in Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:52.177206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:39:52.177206Z digest=sha256:7f92046c53627f20241be07b9962b174dc2dd73dd14edb42fe4187b6a288e1dd

Observation fe72cc17-d14f-4721-900a-fc0fabb284f0 · outbound

This paper cites Measuring and Narrowing the Compositionality Gap in Language Models.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance Measuring and Narrowing the Compositionality Gap in Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:52.183790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:39:52.183790Z digest=sha256:4f5caf98270da9eb0ba50b742235dac01fb5b8072de2c11a0c0ceef41ddd565d

Observation 4225dfae-6b2c-4c79-bc4e-fcd2b5abc957 · outbound

This paper cites Proceedings of ICLR , year=.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance Proceedings of ICLR , year=

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:39:53.009243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T00:39:52.187177Z digest=sha256:5eeec86b3273d82743b0e0e367b30ecf559a7e58260cb88aa03c23878012aab8

Observation 820012fd-dd33-4268-abda-d175a099ce9c · outbound

This paper cites Proceedings of ICLR , year=.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance Proceedings of ICLR , year=

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:39:53.000346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T00:39:52.190121Z digest=sha256:2c7a375740e3f4c4813e62078fed4aa60aea2c10da5fb2f01df46525c6928a9c

Observation 01b1d51f-2108-4d1d-854d-f19f74a9e283 · outbound

This paper cites Large Language Models Are Human-Level Prompt Engineers.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance Large Language Models Are Human-Level Prompt Engineers

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:52.193011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:39:52.193011Z digest=sha256:c7de6e52194e0eaa159309d404a81d15aaeb1284e191c568119119b2ad33e77e

Observation 29845126-0a0a-4e85-8fa6-4a00d45bfe51 · outbound

This paper cites Query Rewriting for Retrieval-Augmented Large Language Models.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance Query Rewriting for Retrieval-Augmented Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:52.199842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:39:52.199842Z digest=sha256:d3814f76ec677b31249dba8755af077c60a5b1eaa0b14cb3ec113c18053bdbf7

Observation a6422781-e41a-43bf-bee4-6ee22d1a9e28 · outbound

This paper cites Proceedings of NAACL-HLT , pages=.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance Proceedings of NAACL-HLT , pages=

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:39:52.991815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T00:39:52.203004Z digest=sha256:75f8f0f13f1726fc5ff033eb39efd5889a0a43d8ea72c9713dd3f43853554441

Observation 11bd3b0f-0322-4c0e-b817-32c0a374db68 · outbound

This paper cites arXiv preprint arXiv:2412.xxxxx , year=.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance arXiv preprint arXiv:2412.xxxxx , year=

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:39:52.983407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T00:39:52.206705Z digest=sha256:970a671b799ca3cc3a395616ea559eb5e7bd1a1c9755110d6088c1e69a2cb406

Observation ded91785-5ec3-42e9-bc8f-5cea2ac935cd · outbound

This paper cites arXiv preprint arXiv:2412.xxxxx , year=.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance arXiv preprint arXiv:2412.xxxxx , year=

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:39:52.974864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T00:39:52.212755Z digest=sha256:325994d1bf4065563b5f87608c9caf9d40e3bf641018bbecbb0beafaaf51c75a

Observation 000d0a28-e75c-464e-896e-de176bd66c77 · outbound

This paper cites arXiv preprint arXiv:2412.xxxxx , year=.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance arXiv preprint arXiv:2412.xxxxx , year=

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:39:52.964651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T00:39:52.215494Z digest=sha256:9800db7e1083b432385bfc264ce85b44d98f9de06a1c0ceec4833d514d6484d4

Observation 3542f226-4608-4309-85d5-dc5b1ddf4658 · outbound

This paper cites arXiv preprint arXiv:2405.xxxxx , year=.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance arXiv preprint arXiv:2405.xxxxx , year=

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:39:52.956264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T00:39:52.218202Z digest=sha256:a6f274b8af030873d6b19cef1955bfa27263c8776d118b2533ffc7d152107e6d

Observation 8d6eea9f-4b1e-4593-a1ab-58dab6b12eda · outbound

This paper cites 2024 , note=.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance 2024 , note=

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:39:52.946965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T00:39:52.221408Z digest=sha256:0526b2a61c2fdc66f98c87ea5548dd4322ff9f73fd6df45169b26320abf4087b

Observation cd6ceb40-3b06-41fe-881a-9e6245e6b39e · outbound

This paper cites 2024 , note=.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance 2024 , note=

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:39:52.937812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T00:39:52.224466Z digest=sha256:612ca12f4d32c2c5854f0f39eccf0ada0e2141595ae841c055916fc2732f7e4e

Observation 57438c98-649d-487b-b6e7-142945df7d2d · outbound

This paper cites an unresolved cited work.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:39:52.928962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T00:39:52.231272Z digest=sha256:5bf236b1e6e811e6c5f8f25a7f074297ab57aa0ae86b2b22de94da734ba542fc

Observation 7d29956f-a067-4240-8eca-c66d90f06a0d · outbound

This paper cites an unresolved cited work.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:39:52.919566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T00:39:52.234450Z digest=sha256:d72689236ef69984912a8bedbf6ef6ef216c1ea232ea7edb07b0624cea7a9eb2

Observation e0aa3ae6-2b91-4268-b819-c223936dd16c · outbound

This paper cites an unresolved cited work.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:39:52.910724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T00:39:52.237614Z digest=sha256:972e7276785609f8b1d824c59f523f7b38ca262e020d427b3f049ccc6dfae399

Observation cd38b232-9403-4df0-bb7b-ba0920572adb · outbound

This paper cites an unresolved cited work.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:39:52.900886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T00:39:52.240583Z digest=sha256:7a471f422c0c73a9562f7b615c8498f7e67e7555a7e699e9cc4f1b8035319a68

Observation c6baaf8a-4275-4730-a995-b4ff038e2a11 · outbound

This paper cites an unresolved cited work.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:39:52.891221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T00:39:52.243711Z digest=sha256:87b5eeaa52b8f2e3ce8485edecc857d1fcfe12575292f8853b2096b9bc804041

Observation 4f7d71ac-6318-4ab3-b766-50d9f2c849a5 · outbound

This paper cites , title =.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance , title =

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:39:52.882272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T00:39:52.246745Z digest=sha256:cba5c077450edd71689919d7476587b54636e90496bdda5ba268824a54cfd34b

Observation b86b5a10-00a1-4759-b229-ce0607d566ea · outbound

This paper cites 2025 , eprint=.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance 2025 , eprint=

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:52.249521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:39:52.249521Z digest=sha256:29d2934155fd6d64808e99aa770eaf9776e87a028200519933d6c00a8be5256a

Observation 3b18db8e-9d6d-4c70-8581-153af402b92c · outbound

This paper cites an unresolved cited work.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:39:52.867977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T00:39:52.252429Z digest=sha256:b6021cbad91670b696bc67a5481eb6a0db835ba778cb58df6a89af40a1715563

Observation 22e85d94-484f-46fd-9483-586f0a93cb95 · outbound

This paper cites 2025 , eprint=.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance 2025 , eprint=

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:39:52.859486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T00:39:52.255531Z digest=sha256:c1a5f82c854bb6396ceeabad2837807dc08e7226b3be2754bdc6d4d8c2c0771e

Observation 2ca7759e-e40d-4ae5-8c3a-97556c259832 · outbound

This paper cites 2024 , eprint=.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance 2024 , eprint=

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:52.259041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:39:52.259041Z digest=sha256:1f4d2a26ab22b68c36155598cbe910d29518041ccb9555cd74b52fc65e2a8d7e

Observation 8a6e6c99-5829-4ce1-9e8b-ef6f7ac7fabb · outbound

This paper cites 2024 , eprint=.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance 2024 , eprint=

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:52.262447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:39:52.262447Z digest=sha256:a0807de5e8f2ad0d07caefbc3f76ea7c3a2b053ddd8f43b7a021fc8ce5a0f693

Observation f9c02850-0ce2-4ba5-96c2-1194fcceb87c · outbound

This paper cites 2024 , eprint=.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance 2024 , eprint=

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:52.265510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:39:52.265510Z digest=sha256:c9f5e062d2d39eeda67ead4b8c73e3ddfe68a573f871b0a8bc6462b95bf580fe

Observation 339028c1-69e8-4f57-8df2-8af0a7e2d536 · outbound

This paper cites 2024 , eprint=.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance 2024 , eprint=

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:52.268216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:39:52.268216Z digest=sha256:c6aafa7b1191de27a456b3025da5ede2695f81e497d01d908a3aaca1e876435a

Observation 5b055ac8-3d62-41c1-8833-ede914938f82 · outbound

This paper cites 2024 , institution=.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance 2024 , institution=

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:39:52.832136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T00:39:52.271385Z digest=sha256:850e26f80d9ce38beb3195118af733415fa3976407302eb9e1e5aa51b5883f65

Observation b17afc4a-b8b4-4f7e-aa94-6d2d817896ad · outbound

This paper cites 2024 , url=.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance 2024 , url=

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:39:52.822731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T00:39:52.275140Z digest=sha256:897950a12fd98892fb9045dafe6768242f578debcd29a8698736c3213272def7

Observation 10bcb47f-3c67-414c-9560-eaa8737641ac · outbound

This paper cites 2024 , eprint=.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance 2024 , eprint=

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:39:52.813623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T00:39:52.278063Z digest=sha256:ac79bc2f93711af52a2b9db573944dea4eaa8a63b7d2c96b9bbf7779c729068e

Observation 3cdcf9d4-da17-4dd1-816d-721de9e4eac1 · outbound

This paper cites 2024 , eprint =.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance 2024 , eprint =

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:52.280836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:39:52.280836Z digest=sha256:34d34fa59533ef44fa33206cbf08890a3579490abcb6abe8bce2ff12a86ddf4f

Observation 8bf043df-0fea-4a38-9cee-46ccd62e5ac3 · outbound

This paper cites 2021 , eprint =.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance 2021 , eprint =

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:39:52.799689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T00:39:52.287941Z digest=sha256:d6a1ceb86f6d0f4c08c725bef124ccc409222199e1a996258914bdf84ccfeb40

Observation a7c44720-1f06-4fb7-81a3-b3da376fb2e1 · outbound

This paper cites 2023 , eprint =.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance 2023 , eprint =

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:52.294509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:39:52.294509Z digest=sha256:67e6e10cfff742bb50d26ab98820dcd0d7fe3d3f28177d2e9c3716b02d539149

Observation 60db8979-57cb-42be-af4a-953372c1a009 · outbound

This paper cites Pragmatics in the Era of Large Language Models: A Survey on Datasets, Evaluation, Opportunities and Challenges.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance Pragmatics in the Era of Large Language Models: A Survey on Datasets, Evaluation, Opportunities and Challenges

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:52.300504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:39:52.300504Z digest=sha256:ca83d456aa7e0c279a06df68c78970c299886c5f194f009bb1671adf75c7b6d9

Observation b4695234-aac8-4034-beeb-46d743bf556b · outbound

This paper cites Ant \`o nia and Salam \'o , Maria.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance Ant \`o nia and Salam \'o , Maria

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:39:52.784835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T00:39:52.303872Z digest=sha256:e261ef132027a492065690fe90e78f8766bddb73c436f806a08b30bc9f0a9c77

Observation 7fb97a12-64e1-4773-bae9-942dd1e01df9 · outbound

This paper cites Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics: Industry Track , year=.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics: Industry Track , year=

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:39:52.775938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T00:39:52.313141Z digest=sha256:230d7a67da66c65293b784c2ab83b89416afc50b6dd9a7cefdc941e1b6bff2d1

Observation 8d3f84f2-1aa7-4dbd-b118-a019b77fed43 · outbound

This paper cites Phi-4 Technical Report.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance Phi-4 Technical Report

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:52.316382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:39:52.316382Z digest=sha256:d7f4c257d1d7e9dae7b45a83c501f64c790abe74412a5648434c9ad488a4772e

Observation ff4aadb3-4b63-40cc-a4f4-78639bb99f87 · outbound

This paper cites gpt-oss-120b & gpt-oss-20b Model Card.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance gpt-oss-120b & gpt-oss-20b Model Card

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:52.319984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:39:52.319984Z digest=sha256:9989f27ce6b20822909ee40e55ebcb295c142c6d23b758894a48fce08224437d

Observation 9c1837ae-7470-418e-b754-f645caeeef64 · outbound

This paper cites GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:52.323512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:39:52.323512Z digest=sha256:4a6122273b554b93fdd8b98621829459818dee7ae26070f4919af7e206ef4d29

Observation 68549af3-e057-437b-bfa1-aaa772e8e4ae · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance Training Verifiers to Solve Math Word Problems

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:52.326186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:39:52.326186Z digest=sha256:5e352820cca13fa64f102d738470027b5e7c4038471dac2549e61a02cbaff6ed

Observation ca9af2cf-4364-47ce-9036-dc4974ec3244 · outbound

This paper cites an unresolved cited work.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance Unresolved cited work

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:52.328984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:39:52.328984Z digest=sha256:7bc9fe88d6ff6f3a8a33d941d45da2b822bba47e18abb8dcd983382249b45419

Observation bed39e8d-e445-4588-995a-f50c528c54c3 · outbound

This paper cites Robustness Gym: Unifying the NLP Evaluation Landscape.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance Robustness Gym: Unifying the NLP Evaluation Landscape

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:52.331692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:39:52.331692Z digest=sha256:2adf29a5c05f9b204d334b735317e6f900d7c5339ee66ec9f54ac139fd39633b

Observation 97170ab4-e3fa-4443-a1b0-a44b2c79de52 · outbound

This paper cites an unresolved cited work.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance Unresolved cited work

Reference 84

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:39:52.765947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T00:39:52.334686Z digest=sha256:95dbe398d04008ee9bb076943651921ebdff7f25834d333f81dcd6d92585148c

Observation 5a984a37-5c10-4b45-b82d-28347514f2a5 · outbound

This paper cites The Llama 3 Herd of Models.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance The Llama 3 Herd of Models

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:52.337868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:39:52.337868Z digest=sha256:b800e126bb95dfb32b863e8ece4d8dd54cb256a01ff467c1772aaad74d7118ce

Observation 18949cbc-fe90-4c22-901a-1cd6c46ebb80 · outbound

This paper cites an unresolved cited work.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance Unresolved cited work

Reference 86

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:39:52.757017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T00:39:52.340815Z digest=sha256:18f9a41478993adf8bf558286dd63e6ce562453c1ef4be62de491dcb6e09d4f7

Observation d1f3b4a0-757d-4aca-8055-e3291e2cc6ea · outbound

This paper cites Mistral 7B.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance Mistral 7B

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:52.344190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:39:52.344190Z digest=sha256:5dbd57d34efbdc3e2bf09f22bea86226f3b9cc7358f3fb2bb771d904abf6b37b

Observation bc503bfa-b9b5-4cc7-bf20-09f759493e69 · outbound

This paper cites LLM-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative Fusion.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance LLM-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative Fusion

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:52.347392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:39:52.347392Z digest=sha256:a1059c477198a76e25a3fe7e4966420dc94f5e01c5af73fb90f0dd24867dda9d

Observation c7ff56bf-1cb0-4690-a76c-392a433eb579 · outbound

This paper cites DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:52.350464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:39:52.350464Z digest=sha256:6b62d5994dd388efc0d1b44a26996c8a74168d6320ef8a384f73812965efdb5d

Observation 3bf73bf2-5f24-446f-9077-3cee862f65d8 · outbound

This paper cites Decomposed Prompting: A Modular Approach for Solving Complex Tasks.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance Decomposed Prompting: A Modular Approach for Solving Complex Tasks

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:52.353992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:39:52.353992Z digest=sha256:24fec6cfccd2d7a3b3840811b70d207c4492d63304aed4b6fc02353463195e10

Observation 29c3c036-d30f-4a0f-9abe-30b683d33de3 · outbound

This paper cites Let's Verify Step by Step.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance Let's Verify Step by Step

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:52.356996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:39:52.356996Z digest=sha256:70287165de23b992e9da53c95ccd390a55104fb8f2c7333d98f8b41d68c2a161

Observation b01399e6-1792-4743-9e71-5c8ba7ac22bf · outbound

This paper cites Self-Refine: Iterative Refinement with Self-Feedback.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance Self-Refine: Iterative Refinement with Self-Feedback

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:52.359807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:39:52.359807Z digest=sha256:2b911832d06a018c525db6a04483bda0b7bbf4861f0355c5379b2889d6d946ed

Observation ab384b88-11a8-432f-8c17-4315a86e9ecb · outbound

This paper cites Granite Code Models: A Family of Open Foundation Models for Code Intelligence.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance Granite Code Models: A Family of Open Foundation Models for Code Intelligence

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:52.362576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:39:52.362576Z digest=sha256:1e15cfb5b0400003ffc4842a325f92c24d28bbd4c1dad15f576e1acc942ea52f

Observation 2eb428a1-3306-46a2-bb84-3358713eac60 · outbound

This paper cites an unresolved cited work.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance Unresolved cited work

Reference 94

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:39:52.747827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T00:39:52.366023Z digest=sha256:8a3b6a3f7413c3d340ce000d32aeb82a68062cf65fb7c9ade11377d5ef4d8189

Observation 71c17b2b-a50e-4d4a-9cdd-a44fb3bd3a75 · outbound

This paper cites Automatic Prompt Optimization with "Gradient Descent" and Beam Search.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance Automatic Prompt Optimization with "Gradient Descent" and Beam Search

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:52.369060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:39:52.369060Z digest=sha256:fb50fc56ac6e798672b0d81176e611745bd4c938c355e11a97d39106adc4c5dd

Observation 7ab81be8-bf9b-4928-af00-1c1a5236d669 · outbound

This paper cites an unresolved cited work.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance Unresolved cited work

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:52.372551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:39:52.372551Z digest=sha256:330b1ffbc650880693180c86b7a88adf584cea00db9f7830f8dfea91d4238c26

Observation c21fd6f6-b1c2-40a2-b2d3-69f9b15646be · outbound

This paper cites an unresolved cited work.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance Unresolved cited work

Reference 97

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:39:52.739113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T00:39:52.375646Z digest=sha256:8d5cb4a9ee83bf1c9472cda459e7edc639a95db705bd561207967af57c1a86d8

Observation 8b4eefdf-1dc6-4b40-9cf9-7a367ab500e4 · outbound

This paper cites Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:52.378889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:39:52.378889Z digest=sha256:d81947ff04fd172be653ca049d534f9cf40d322a12ed5cb1f25d351a27fab267

Observation 18b6ef97-f982-4bfd-b88a-923db5ffc3b4 · outbound

This paper cites Reflexion: Language Agents with Verbal Reinforcement Learning.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance Reflexion: Language Agents with Verbal Reinforcement Learning

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:52.382593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:39:52.382593Z digest=sha256:b843d9ebbf958da736fc3c3a291a6de0733b631ff9e7b2d6c68458d3c4e7f399

Observation f4c2d64f-7e56-40b6-b502-cf4246f02e95 · outbound

This paper cites Qwen2.5 Technical Report.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance Qwen2.5 Technical Report

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:52.386246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:39:52.386246Z digest=sha256:5290d26339b0b1bb09023e4ec0a7ee02660c9ce3ac1860fee9ba012dba3837dc

Observation 4233beed-5a8d-465d-aec7-490d776d96ae · outbound

This paper cites DebugBench: Evaluating Debugging Capability of Large Language Models.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance DebugBench: Evaluating Debugging Capability of Large Language Models

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:52.390073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:39:52.390073Z digest=sha256:02b9df2aa48ee57fe37ea33c63b8fdc9011c25d778bc71e46d92f4c32a2a889f

Observation eabaaace-4c0d-410a-a61b-39be8a7844e7 · outbound

This paper cites Universal Adversarial Triggers for Attacking and Analyzing NLP.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance Universal Adversarial Triggers for Attacking and Analyzing NLP

Reference 102

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:52.393578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:39:52.393578Z digest=sha256:50070d6ab426fcd6829a81f6ffbf30c72da0919a32f61ab62747468495d07f56

Observation b26d49e1-fb4e-4f5f-aea1-c3d37a80e3a5 · outbound

This paper cites an unresolved cited work.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance Unresolved cited work

Reference 103

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:39:52.729732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T00:39:52.396361Z digest=sha256:142f378d413ff5abbcc2f3516ec1599e29960fd39992b059f068db924bef19b2

Observation 3916d912-f60c-42a2-8593-599f8ac5ae9c · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 104

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:52.399322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:39:52.399322Z digest=sha256:4df02cba6598ac42060a2a7a4924f36ac61da7bdebf5dcecb11606d2e5b9e519

Observation ab7e5815-0df3-4a43-838e-6c7ab4076f46 · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 105

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:52.402005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:39:52.402005Z digest=sha256:c9ea34b458a56b60ebf0cb0620e72842b25a6ee15b8f80e8082ef2bc7c8aa4dc

Observation 15a16d15-6928-49d0-b0b1-7c9427d2aeb4 · outbound

This paper cites Qwen3 Technical Report.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance Qwen3 Technical Report

Reference 106

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:52.404876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:39:52.404876Z digest=sha256:dd18ba71e0a6367956323a144d263ab2c7441226c73fac23d387502d633ad03c

Observation 05d3f8d6-143b-477b-a4d8-26ab71841ca2 · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 107

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:52.407918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:39:52.407918Z digest=sha256:56ed2c582e4107644f0648ea2ace45fb62f4ed1614445a3ae9a563db1daeaa71

Pith citing papers

No inbound Pith citation observations are available.