Pith. sign in

Paper Citation Record · LEDGER

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation

As of 8 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 1 inbound Pith citation observation for arXiv:2607.14109.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.14109 v1

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T14:42:46.895939Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-31T03:19:00.468509Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

41 of 41 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved40
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a8dbedc3-b580-44a4-a9ac-7ad770bbecec · outbound

This paper cites Matharena: Evaluating llms on uncontaminated math competitions, February 2025.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Matharena: Evaluating llms on uncontaminated math competitions, February 2025

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T14:42:43.688759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:43.688759Z digest=sha256:8dcf602cea2889732a10de9647401f5223f01ac63ef99545ec8cd0acddf7a3d7

Observation 8d12022a-c777-4faa-b2d7-cf6bc25524cd · outbound

This paper cites On the Measure of Intelligence.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation On the Measure of Intelligence

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T14:42:43.736195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:43.736195Z digest=sha256:af8d6f330efa264f29a28e60c30f3a831856ecdd29b67e3d9561961b0b5f5af2

Observation 417d6c12-b74f-4a4b-a868-83dfde7673f4 · outbound

This paper cites FailureSensorIQ: A Multi-Choice QA Dataset for Understanding Sensor Relationships and Failure Modes.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation FailureSensorIQ: A Multi-Choice QA Dataset for Understanding Sensor Relationships and Failure Modes

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T14:42:43.792576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:43.792576Z digest=sha256:f6f3122b8165e130e6da124e043fae014bc4d1684879680026b1b5a63c38dad5

Observation 62de648b-6cab-427d-b068-09919e94592c · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T14:42:43.868953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:43.868953Z digest=sha256:4af1dbb364fff3823422a7962780ef2ed366c977a166e571968c980d29ca0be6

Observation c81abba5-3dd0-4b90-8299-695c12f88f6d · outbound

This paper cites The language model evaluation harness, July 2024.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation The language model evaluation harness, July 2024

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T14:42:43.942648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:43.942648Z digest=sha256:9db8938a80b6a838315963564f448afa3e67e254fd4ee516ed2683bd20c40840

Observation 3fa92299-6cb3-4ee8-9aa2-6f6e55ec04e1 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Measuring Massive Multitask Language Understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T14:42:44.008058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:44.008058Z digest=sha256:27b4981efb9f131e73c39f96bf1bff8bf68bd868dfb730bc9ff31a3bc5ed65a6

Observation 07c789f4-03fc-4bae-a2b9-ce5bbcb66dd4 · outbound

This paper cites What disease does this patient have? a large-scale open domain question answering dataset from medical exams.Applied Sciences, 11(14):6421, 2021.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation What disease does this patient have? a large-scale open domain question answering dataset from medical exams.Applied Sciences, 11(14):6421, 2021

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T14:42:44.077828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:44.077828Z digest=sha256:ed17cd6089d6024665d98ce182eca1b439795bfebddc8db8c95e0bb9cd17988a

Observation 6a17f4d5-65a3-4b16-8856-0ccc80f9cbde · outbound

This paper cites Dynabench: Rethinking benchmarking in NLP.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Dynabench: Rethinking benchmarking in NLP

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T14:42:44.125291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:44.125291Z digest=sha256:eb1c77ade9fa3aad08a32ee73b05ff155fffdd04e59852c30d6d55c700a62645

Observation 6da2fbb5-5677-4a57-80d3-639352a73d76 · outbound

This paper cites Large language models are zero-shot reasoners.Advances in Neural Information Processing Systems, 35:22199–22213, 2022.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Large language models are zero-shot reasoners.Advances in Neural Information Processing Systems, 35:22199–22213, 2022

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T14:42:44.201129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:44.201129Z digest=sha256:35e7dd41addac25d6d590586cce1bd5311cf6f36878e9843998d0c81ea7f2561

Observation b665a6cf-375d-462c-b5ac-a24d556037b0 · outbound

This paper cites Time-MQA: Time Series Multi-Task Question Answering with Context Enhancement.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Time-MQA: Time Series Multi-Task Question Answering with Context Enhancement

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T14:42:44.258314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:44.258314Z digest=sha256:669b86dfd1f80a7210e5dc765a7f80316cca663a6547a3f736a450e35e16479f

Observation 0f54f492-b7ab-455a-bc1f-b81c96cb9f13 · outbound

This paper cites Prompt repetition improves non-reasoning llms.arXiv preprint arXiv:2512.14982, 2025.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Prompt repetition improves non-reasoning llms.arXiv preprint arXiv:2512.14982, 2025

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T14:42:44.341417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:44.341417Z digest=sha256:95bde37522ca31cb2ccc3f560fd4b844cb2d0a14fc0b312cc76be426703ba343

Observation 41a00ded-0a6b-4172-bbf1-76b0eeb0c10c · outbound

This paper cites Holistic Evaluation of Language Models.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Holistic Evaluation of Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T14:42:44.399738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:44.399738Z digest=sha256:51295d80891e7abdf34c3597046c4cc653a0844f5ee2847a4b3d24f84584f00e

Observation dd81da62-65e4-44e3-9d57-0260c64f53c7 · outbound

This paper cites Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T14:42:44.484533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:44.484533Z digest=sha256:d5c2d463d96c09b999dea992cab3165be8805232fdf71337e0b4a0ec46a1e017

Observation bbd846b0-2b40-4227-8ff3-1ac792f89aa5 · outbound

This paper cites SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T14:42:44.560800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:44.560800Z digest=sha256:01e9baf4196bcb479a6a09fb60ebd6dbcc047bfa6e83df4f4e5644e2b58143ce

Observation 0c9c93e9-8fc0-4bda-a6c1-b5b050d53f65 · outbound

This paper cites Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T14:42:44.631644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:44.631644Z digest=sha256:2b52aacd93989afd5f0cbc7e3a3b4793c5264c70f7b67ce8616da541989295bf

Observation bc0dad1d-790a-46e4-9562-e30006943550 · outbound

This paper cites Can a suit of armor conduct electricity? a new dataset for open book question answering.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Can a suit of armor conduct electricity? a new dataset for open book question answering

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T14:42:44.690538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:44.690538Z digest=sha256:fb52bdcda08f6a4ecf01713cd5f27715bdfdc65dd045f33b0c3e505024d5f3c5

Observation bbe22d41-b35b-4b96-bfe4-2711edc2c830 · outbound

This paper cites State of what art? a call for multi-prompt LLM evaluation.Transactions of the Association for Computational Linguistics, 12:933–949, 2024.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation State of what art? a call for multi-prompt LLM evaluation.Transactions of the Association for Computational Linguistics, 12:933–949, 2024

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T14:42:44.771497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:44.771497Z digest=sha256:46ea152b3a28b4aa753a36d816cc1b00fed7e83b389f69c439738e4774eb88ef

Observation 2b98c736-2e7d-4f91-81f7-65991303bab1 · outbound

This paper cites s1: Simple test-time scaling.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation s1: Simple test-time scaling

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T14:42:44.851388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:44.851388Z digest=sha256:ed10d716d49b44d25d62f8d1709fc3e6da5557ff84bc6cb33debce5c7f49e6ee

Observation 803145ac-f648-4d8e-b804-397b5d9d61b3 · outbound

This paper cites Large language models sensitivity to the order of options in multiple-choice questions.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Large language models sensitivity to the order of options in multiple-choice questions

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T14:42:44.933675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:44.933675Z digest=sha256:61e3285b54ff51d718ad4555c4c0a38c5fe4aebd357be18c306cca384338a76d

Observation 7a2d1432-7497-4207-8d26-67b65e837ee8 · outbound

This paper cites Smith, and Mike Lewis.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Smith, and Mike Lewis

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T14:42:44.989333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:44.989333Z digest=sha256:a0965bc63486d662939b97548aa3c43babf9d311c0aee3fcb7a437fbf581aacb

Observation 08c60252-8d51-437c-9a79-752328250936 · outbound

This paper cites Qwen3 Technical Report.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Qwen3 Technical Report

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T14:42:45.058968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:45.058968Z digest=sha256:731dc0f936734dcee472bbb1ff6de4cc0b72ef07c8beb7590cf9c3998b200613

Observation ae5a0261-8753-4e0f-a714-525fe8aca56c · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T14:42:45.126365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:45.126365Z digest=sha256:f7962a7400b9438166697f87288d6a240c426b6119e0aa7c0d5274c4b086881f

Observation 4208621a-99dc-462f-ae7b-af01e99f96ee · outbound

This paper cites Leveraging large language models for multiple choice question answering.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Leveraging large language models for multiple choice question answering

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T14:42:45.188960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:45.188960Z digest=sha256:aae140c6b47606196ef16e9caeb6a3db68581db70caf9aad7a7a40e261193c2b

Observation d90898b8-5376-45d0-bd23-cfc1466a0610 · outbound

This paper cites Rogers, Inna Goncearenco, Giuseppe Sarli, Igor Galynker, Denis Peskoff, Marine Carpuat, Jules White, Shyamal Anadkat, Alexander Hoyle, and Philip Resnik.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Rogers, Inna Goncearenco, Giuseppe Sarli, Igor Galynker, Denis Peskoff, Marine Carpuat, Jules White, Shyamal Anadkat, Alexander Hoyle, and Philip Resnik

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T14:42:45.271611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:45.271611Z digest=sha256:8434cc665b2884f10526f45f7e61f292ce3d3a572bcd4ad9498c5792ee740b83

Observation 69da3077-c6af-40d9-8c9b-e370aff58097 · outbound

This paper cites Quantifying language models’ sensitivity to spurious features in prompt design or: How i learned to start worrying about prompt formatting.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Quantifying language models’ sensitivity to spurious features in prompt design or: How i learned to start worrying about prompt formatting

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T14:42:45.351262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:45.351262Z digest=sha256:6f92721ee3e5218529cec9d33c0e1d563d34b7c163359325707ff53c4c18136f

Observation 3c9b5ff5-58b6-40e5-9167-4b6ec5955431 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T14:42:45.427128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:45.427128Z digest=sha256:1822a674ff80ab0291f0d3d1f7352dc05973ea53afa714efa7b8c898ba47ed3f

Observation 386a2a21-f5f9-4ae0-9201-d03189116cb1 · outbound

This paper cites To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T14:42:45.567640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:45.567640Z digest=sha256:e197979e0300fdb64b7dfefba6087b20f6b370fe985e8a998d7a9ae80c0d6514

Observation 070c18b0-ffd7-407a-a68b-274a9e4b9031 · outbound

This paper cites Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T14:42:45.737278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:45.737278Z digest=sha256:d5b932f10fab545ba77addc45adca8ae1acbddf4abe2334a7f365df61a154a6a

Observation c24b6286-216d-48e8-a2b9-698aee799855 · outbound

This paper cites On the self-verification limitations of large language models on reasoning and planning tasks.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation On the self-verification limitations of large language models on reasoning and planning tasks

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T14:42:45.811194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:45.811194Z digest=sha256:fb25116f10390d4b1d897982cba1896b45d61528fff46778d3bb4f8e571683fc

Observation a086f0a8-0857-4400-b802-8181fd1b6235 · outbound

This paper cites Correctbench: A benchmark of self-correction in llms.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Correctbench: A benchmark of self-correction in llms

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T14:42:45.891639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:45.891639Z digest=sha256:c2fbca1fc8aef67a0a760ce3376942c39086e76041cb1c56d63190868d01f9be

Observation 7869f54f-2d87-4e95-bf18-6242d863dbc3 · outbound

This paper cites Plan-and-solve prompting: Improving zero-shot chain-of-thought reasoning by large language models.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Plan-and-solve prompting: Improving zero-shot chain-of-thought reasoning by large language models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T14:42:45.950957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:45.950957Z digest=sha256:348177f13d435981bb2f5b71032f670544a103a0105b84efe4714a8e209c6a51

Observation 5d944d42-32fc-4147-8eba-f70080c55be0 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T14:42:46.034794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:46.034794Z digest=sha256:b6024cfabb91d6441e22eabbb3d650f0a004ce77faa1e873304a6309998c7ed9

Observation 7ba8a91c-af99-4941-8964-d5d2feba3625 · outbound

This paper cites MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T14:42:46.087335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:46.087335Z digest=sha256:257452c27d95dae290b21b481f18207eb2d9be88a4640a4869f73d6f8137d9ef

Observation 435a630f-28ab-479a-9174-543f4c4ada6f · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Chain-of-thought prompting elicits reasoning in large language models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T14:42:46.157373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:46.157373Z digest=sha256:82ae2dcf7d5f3f0e977085bd15b26df07cb611c96e10d3ab55f98e616a8c1f92

Observation fe702455-6e0e-495a-b1e9-bc7d265db053 · outbound

This paper cites Griffiths, Yuan Cao, and Karthik Narasimhan.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Griffiths, Yuan Cao, and Karthik Narasimhan

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T14:42:46.222619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:46.222619Z digest=sha256:1380e40fe86c65b9e0c37a618ba3cbc3336b98f0aa5b14b7bf6226ff340f7bba

Observation 1a961236-da12-45fd-bd23-b8616d05477b · outbound

This paper cites Chi, and Denny Zhou.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Chi, and Denny Zhou

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T14:42:46.344230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:46.344230Z digest=sha256:45443a3a3113c93348c44702c863bff3027694eff544ed4fcf333bc027bc66a1

Observation 4ef20845-7303-4fc7-b64f-855a71505425 · outbound

This paper cites Generate rather than Retrieve: Large Language Models are Strong Context Generators.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Generate rather than Retrieve: Large Language Models are Strong Context Generators

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T14:42:46.446916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:46.446916Z digest=sha256:11bd9ee497aa2567beeec1fc94423d10f96ab928a5ce96faeb619e2e1fa00be7

Observation dfa6988e-f120-447c-9960-8645380e31a3 · outbound

This paper cites Large language models are not robust multiple choice selectors.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Large language models are not robust multiple choice selectors

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T14:42:46.560399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:46.560399Z digest=sha256:9baa359a2e3d6bd2c52bcf69f19a94f2e45e570b3856ac2a3ce548bf2617ddbd

Observation 14d137af-1c5b-4983-ba1d-0cb7b9eb896b · outbound

This paper cites Scaling physical reasoning with the physics dataset, 2025.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Scaling physical reasoning with the physics dataset, 2025

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T14:42:46.668559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:46.668559Z digest=sha256:6592d12e4f40ff0a5343fcc44aa69085bb224f891c8a2758a825f7a276d0bfca

Observation ab524a00-2979-4d1d-8413-3f01fdfb7fa5 · outbound

This paper cites PromptBench: A Unified Library for Evaluation of Large Language Models.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation PromptBench: A Unified Library for Evaluation of Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T14:42:46.785295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:46.785295Z digest=sha256:b0a1c27c3b05d1965f90c8c068580cef3b7591d32905a24964c7c2c1368d5880

Observation 0f23936f-5148-4702-9339-0d8f554e1d0b · outbound

This paper cites Avg. opt.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Avg. opt

Reference 41

Resolution
malformed identifier
no resolver link, observed 2026-08-02T14:42:46.895939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:46.895939Z digest=sha256:4ac50dad6406cbe6847d315d53b9b73c4aaf910140cad3f10f603d831d27d7f9

Pith citing papers

Observation c09d33ea-403f-4738-85b1-ac4435c38fb5 · inbound

Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B cites this paper.

Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-31T03:19:00.468509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T03:19:00.468509Z digest=sha256:076b8c2f0d06c5c6a05a2f736a02bd9dc2131035125013c2d3933dc84d6ee792