Pith. sign in

Paper Citation Record · LEDGER

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation

As of 21 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 0 inbound Pith citation observations for arXiv:2505.24263.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.24263 v1

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:35:29.070019Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

37 of 37 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved36
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 48403ae3-c436-4c4d-8b1b-b0882126a52f · outbound

This paper cites URL: " 'urlintro :=.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation URL: " 'urlintro :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:25.126963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:25.126963Z digest=sha256:34ff328041a4d4a4d98c517861c0e4513da1bb83f0e82b341d2087bbe8754917

Observation 552d78f0-3bb1-4696-8fe9-f0b5a18c4c1a · outbound

This paper cites write newline.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:25.285963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:25.285963Z digest=sha256:3d26690891a9d10498969194b2ee7d249b8acc7e485651df84e2b64851a3106a

Observation 381ca8e0-ebc8-45b7-beb2-c466d9774410 · outbound

This paper cites an unresolved cited work.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:25.410371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:25.410371Z digest=sha256:425b2d8a40c7cef2311f1ddaab4e6d87743fafcc0e0c141671d39c6999e70a54

Observation 5d579d1d-a646-4908-bd15-bd401705277f · outbound

This paper cites an unresolved cited work.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:35:30.205471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-07T12:35:25.533380Z digest=sha256:f7c13d18e12bdab9a1e5595e9e88e55fdb07b08b2c66bf4ec55fc98f8204cbf0

Observation ac6a04c4-0f96-44ff-a4d0-5d758d64bf20 · outbound

This paper cites Language Models are Few-Shot Learners.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Language Models are Few-Shot Learners

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:25.691779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:25.691779Z digest=sha256:66cf4bce2402accbe71d6d39427e361112f4f7ba8f1d076a8206e2deb4858cce

Observation 7d2be1ac-92ae-4e22-9908-3792e513e538 · outbound

This paper cites Quantifying Memorization Across Neural Language Models.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Quantifying Memorization Across Neural Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:25.837301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:25.837301Z digest=sha256:7c9e7c10d7c290edad91c338b383dbd54627b166d1722366f174e32d84b13d25

Observation 1faed743-79ba-435e-a9d9-a5ba3cb31840 · outbound

This paper cites an unresolved cited work.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:25.945153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:25.945153Z digest=sha256:3fc8aca774e14bae649931024a20e3cd48a6cfe131f99b693ef0cca6ea6c0cd4

Observation 6fb46aeb-d547-4ee8-8969-fc01656f5fca · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Training Verifiers to Solve Math Word Problems

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:26.047017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:26.047017Z digest=sha256:b4a61d6e421fa8469cb6918129446604c0cb6229f5ebf656605f1dd200c6d921

Observation d447cf9b-c72d-4059-a0f0-3344f8a8dc83 · outbound

This paper cites Does Data Contamination Detection Work (Well) for LLMs? A Survey and Evaluation on Detection Assumptions.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Does Data Contamination Detection Work (Well) for LLMs? A Survey and Evaluation on Detection Assumptions

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:26.123753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:26.123753Z digest=sha256:61599d52e6fbeade09bbf4c69dc02b2b0523f31477a778b393e9aa32fe2fbb54

Observation 9b4366e9-05f8-49cd-a94a-145685687d6d · outbound

This paper cites Are We Done with MMLU?.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Are We Done with MMLU?

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:26.208319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:26.208319Z digest=sha256:095d8aa48ce0d2797f365ca5304ea89e1c63df53c5fc102ee3558581b5a9448b

Observation 65d8c614-fee6-4973-98cc-704b3e9632b4 · outbound

This paper cites Time Travel in LLMs: Tracing Data Contamination in Large Language Models.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Time Travel in LLMs: Tracing Data Contamination in Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:26.309692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:26.309692Z digest=sha256:3e8293c98052bda8d9151c1795702e562ee65f0965e6aac7a6865d8749111253

Observation 4f148745-1121-4a19-8c0b-296d37a97674 · outbound

This paper cites The Llama 3 Herd of Models.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation The Llama 3 Herd of Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:26.414783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:26.414783Z digest=sha256:961ecc68893b7e3ba7bfe8332c6e424c7bdd0d0f7d3419e44562528da944bc17

Observation f015e85b-9744-4278-af9d-5a08ebfe5715 · outbound

This paper cites an unresolved cited work.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:26.515744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:26.515744Z digest=sha256:49128dd14b2e7dcf1c481f799772648a8c63e75287d563481bdc570b6a08e03f

Observation f872dafe-c61b-4cd0-b962-c71c5ba91f46 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Measuring Massive Multitask Language Understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:26.585502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:26.585502Z digest=sha256:25cc9d5f331b4c1887e0e2f89540b4a9dd6f13cf3450a8e158991fc94fc76843

Observation 9fa753c5-b81e-466b-9385-c969a3c49ecf · outbound

This paper cites an unresolved cited work.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:35:30.011952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-07T12:35:26.693692Z digest=sha256:b5d32e29548bb98964ede8e340ea8c522f85daabd07dda3678bc5afff5b448e7

Observation f23ba1cc-33d7-4198-8548-4725c998a439 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation LoRA: Low-Rank Adaptation of Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:26.812818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:26.812818Z digest=sha256:cb5dcbe6c6534dd1adb62d5377c788c54432a6880919069cb22cd9c4484f685c

Observation 088ddac3-4d2c-4796-97d3-8db2db9abe7b · outbound

This paper cites Membership Inference Attacks on Machine Learning: A Survey.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Membership Inference Attacks on Machine Learning: A Survey

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:26.911572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:26.911572Z digest=sha256:9cd96a6dbb994519f5257b60446ccfb348b9bb22ff32f3cd13b6b63c8c314c04

Observation 3a55477a-f0cc-4719-8e26-4ed2979c2896 · outbound

This paper cites Estimating Contamination via Perplexity: Quantifying Memorisation in Language Model Evaluation.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Estimating Contamination via Perplexity: Quantifying Memorisation in Language Model Evaluation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:26.977795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:26.977795Z digest=sha256:d07d5fcadbc093d5d2709c3bb6abc5e555c92261aa01062fed4f8154046a659e

Observation c0af4880-be2f-4f7f-b237-dd4aed103c85 · outbound

This paper cites an unresolved cited work.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:27.076291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:27.076291Z digest=sha256:1ba9134deffd536eecd397f22b9e9a9561474f36e9a050626a6d7bf0e7efa39a

Observation aca29502-3b06-4411-898b-37b2eb2c6e40 · outbound

This paper cites Lin, Jacob Hilton, and Owain Evans.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Lin, Jacob Hilton, and Owain Evans

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:35:29.821277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-07T12:35:27.190607Z digest=sha256:843182a0a53c456cbe5938a88a7b81055d6fcce0927e64ede69010b902aaea15

Observation e29cbdeb-17ed-44ca-be0f-87b2bd22759a · outbound

This paper cites DeepSeek-V3 Technical Report.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation DeepSeek-V3 Technical Report

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:27.321284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:27.321284Z digest=sha256:4efd72f45bc06566e97004ec642bdc0e0b9c66e3db6fbbb142def2d3744d973d

Observation 537db04c-fc95-4991-b246-5bc604ff91bc · outbound

This paper cites an unresolved cited work.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:27.401062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:27.401062Z digest=sha256:dbcee9c7867b8669e185d73478936d09da40d576bc04667eec9d34d52577d124

Observation b4c6546f-481d-4c0f-a8a6-1c9b078a59b3 · outbound

This paper cites Training on the Benchmark Is Not All You Need.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Training on the Benchmark Is Not All You Need

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:27.501910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:27.501910Z digest=sha256:6912593c007c54ed9d35c12e59d098068db2fbc677c6ba64a858fb241a617cf1

Observation 022bc4e4-6bca-4107-a61d-e4d2afd41e0d · outbound

This paper cites GPT-4 Technical Report.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation GPT-4 Technical Report

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:27.657313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:27.657313Z digest=sha256:6bc8c5fa8b4b6d3628efdb75c100f9330483a55513443300241e86c076304e6b

Observation 85948a63-eefd-4c53-8708-a7d5933f2fbb · outbound

This paper cites an unresolved cited work.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:27.763705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:27.763705Z digest=sha256:45ef7969a2e5d0f980e910c39bca8ee7cee55619090ceb377f3937667f4c411c

Observation 457c0c24-c0e0-4a64-bfc4-7e22fc3024c4 · outbound

This paper cites Qwen2.5 Technical Report.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Qwen2.5 Technical Report

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:27.876148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:27.876148Z digest=sha256:b1df2623fd71859ca17a94dd27344929abd862ecf202f59f7c4c7a15dc6268a4

Observation c8d45bf4-cb03-4b7e-b8af-fae0dddd9161 · outbound

This paper cites Leveraging Large Language Models for Multiple Choice Question Answering.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Leveraging Large Language Models for Multiple Choice Question Answering

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:27.989088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:27.989088Z digest=sha256:319746dd0e50b505e4675ed407d8addc42c550464e118f06884fee916665c1db

Observation a1645894-ad5c-4459-87c3-38c1f24c1685 · outbound

This paper cites Towards Data Contamination Detection for Modern Large Language Models: Limitations, Inconsistencies, and Oracle Challenges.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Towards Data Contamination Detection for Modern Large Language Models: Limitations, Inconsistencies, and Oracle Challenges

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:28.145036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:28.145036Z digest=sha256:a8c362edf32c0fbcd19e600c6a1ea0dc6e2afc56afbd9a1f194654f22ed42c5a

Observation d4a2cc5c-414e-4f72-93f1-84b947cfecfd · outbound

This paper cites SocialIQA: Commonsense Reasoning about Social Interactions.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation SocialIQA: Commonsense Reasoning about Social Interactions

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:28.259215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:28.259215Z digest=sha256:d4f1e247bd6311240bf1ee08d74052264f339719ae85fea59e2ac2d7b4ef3122

Observation 877f8036-0eb7-428d-b4b5-b2904ec2975f · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Gemini: A Family of Highly Capable Multimodal Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:28.402129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:28.402129Z digest=sha256:56cbc25aed4ea67c364a6a7f7c2b1ec75803c00b91e7aed91ce8220c46e5dbf4

Observation 17509d7d-145e-4f20-b4c0-cafb062a0914 · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Gemma: Open Models Based on Gemini Research and Technology

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:28.488255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:28.488255Z digest=sha256:0d549cc49c0c4e2e0fa6215395b65ba41b8d26b1609b17419d9fc80b1c26f125

Observation c169a4fc-480e-431e-9945-25576704a8e1 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Gemma 2: Improving Open Language Models at a Practical Size

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:28.599455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:28.599455Z digest=sha256:8009a3cc6e3b13a0bf8f8d3310ec662a57336cd0b5438d5c79a1bbca412175e4

Observation d931bec7-4dd3-40f8-b778-8aef1d04f872 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation LLaMA: Open and Efficient Foundation Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:28.705768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:28.705768Z digest=sha256:2d627c969aed0fc9b6606f5e8a43f236787772137f31322ba2bbdc95d3186f5a

Observation 7265672e-7d0f-4ee7-aab8-3d9c9f00974d · outbound

This paper cites MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:28.812524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:28.812524Z digest=sha256:b8dcd64bea598c1ae6456e1e264507530fa72e66ac289132ef4d45a760d202be

Observation cc895bca-1fb5-4f50-872f-432387d4c5bd · outbound

This paper cites HuggingFace's Transformers: State-of-the-art Natural Language Processing.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation HuggingFace's Transformers: State-of-the-art Natural Language Processing

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:28.896218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:28.896218Z digest=sha256:b53ec71cf3852f93ca58e54a71a2b3ce407b52fb0d7cc4b8fcfc1a72997d2792

Observation cbd67ca9-a106-43a1-a3b5-d6a351c50b4f · outbound

This paper cites Benchmarking Benchmark Leakage in Large Language Models.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Benchmarking Benchmark Leakage in Large Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:28.982799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:28.982799Z digest=sha256:e4e97dd3e657821ecf3aaf7c3ad5d67e72e785183d6480968a13273ea131910b

Observation dfe21523-ee1b-44fc-b0db-954c381c2c42 · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:29.070019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:29.070019Z digest=sha256:2786834c9a1a71e1ff143a5f37382c25b3eead3c120c5779ef1149a00de6392c

Pith citing papers

No inbound Pith citation observations are available.