Pith. sign in

Paper Citation Record · LEDGER

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation

As of 7 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 0 inbound Pith citation observations for arXiv:2505.24263.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.24263 v1

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:35:29.070019Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

37 of 37 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved36
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 48403ae3-c436-4c4d-8b1b-b0882126a52f · outbound

This paper cites URL: " 'urlintro :=.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation URL: " 'urlintro :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:25.126963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:25.126963Z digest=sha256:5685b12a6d67df52447ba876673d717e21310f845a653fc13b514d92ec4ce29e

Observation 552d78f0-3bb1-4696-8fe9-f0b5a18c4c1a · outbound

This paper cites write newline.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:25.285963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:25.285963Z digest=sha256:05fbf39746f4b9d49df39e9bbf41c98717dd21e3089ff448b5bcdfc5909f4a60

Observation 381ca8e0-ebc8-45b7-beb2-c466d9774410 · outbound

This paper cites an unresolved cited work.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:25.410371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:25.410371Z digest=sha256:fe576321ac3acace529c6ed19665fd439beee870df7b166d1712143d1514c921

Observation 5d579d1d-a646-4908-bd15-bd401705277f · outbound

This paper cites an unresolved cited work.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:35:30.205471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:35:25.533380Z digest=sha256:0d53d2b071069d7f16457c7f0adb4e6301645a1bcedfd859d563dc472d78c1ef

Observation ac6a04c4-0f96-44ff-a4d0-5d758d64bf20 · outbound

This paper cites Language Models are Few-Shot Learners.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Language Models are Few-Shot Learners

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:25.691779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:25.691779Z digest=sha256:09e7a5787828a6876f8cfc5df07017908e11497b2b641b72b51996a5cfb757bc

Observation 7d2be1ac-92ae-4e22-9908-3792e513e538 · outbound

This paper cites Quantifying Memorization Across Neural Language Models.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Quantifying Memorization Across Neural Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:25.837301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:25.837301Z digest=sha256:76562d41fa01a61a5bd1f8f379c1c1feba7b07bb507985294b94f1601fd5204d

Observation 1faed743-79ba-435e-a9d9-a5ba3cb31840 · outbound

This paper cites an unresolved cited work.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:25.945153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:25.945153Z digest=sha256:c4a1d70e99bcbed88c6ff154ce34545cb5a14092396088364023a6a1f4419d85

Observation 6fb46aeb-d547-4ee8-8969-fc01656f5fca · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Training Verifiers to Solve Math Word Problems

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:26.047017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:26.047017Z digest=sha256:95ea6089ad1d1d8db09250202984b1dab5b1ad4a30a86407d242f692f8044572

Observation d447cf9b-c72d-4059-a0f0-3344f8a8dc83 · outbound

This paper cites Does Data Contamination Detection Work (Well) for LLMs? A Survey and Evaluation on Detection Assumptions.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Does Data Contamination Detection Work (Well) for LLMs? A Survey and Evaluation on Detection Assumptions

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:26.123753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:26.123753Z digest=sha256:ebb9043c1b275212745c26a2a33039869d99b7f419aff6b3caa236493f6f8c7c

Observation 9b4366e9-05f8-49cd-a94a-145685687d6d · outbound

This paper cites Are We Done with MMLU?.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Are We Done with MMLU?

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:26.208319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:26.208319Z digest=sha256:c39e3261cfabf1c304b89d78ba42db0f8846d08e5342bc414fb96265834d1b3c

Observation 65d8c614-fee6-4973-98cc-704b3e9632b4 · outbound

This paper cites Time Travel in LLMs: Tracing Data Contamination in Large Language Models.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Time Travel in LLMs: Tracing Data Contamination in Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:26.309692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:26.309692Z digest=sha256:f44bf96f396daa5a862ee91bee0cc70111fb4d9bea2c65c29a2339b90d5f6dd1

Observation 4f148745-1121-4a19-8c0b-296d37a97674 · outbound

This paper cites The Llama 3 Herd of Models.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation The Llama 3 Herd of Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:26.414783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:26.414783Z digest=sha256:a7a71cb79ba2273648d802475fed52bd71ec0ed384e001ea9bfbf473c53a0219

Observation f015e85b-9744-4278-af9d-5a08ebfe5715 · outbound

This paper cites an unresolved cited work.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:26.515744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:26.515744Z digest=sha256:fa929e656c2dbba1de3f9ec30f1bec75fbfe328f4e06c274654160dc36bb245c

Observation f872dafe-c61b-4cd0-b962-c71c5ba91f46 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Measuring Massive Multitask Language Understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:26.585502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:26.585502Z digest=sha256:e2c82492ff9d500a0d996958856fc6a845070f627ca978bf090331f77c69c3e6

Observation 9fa753c5-b81e-466b-9385-c969a3c49ecf · outbound

This paper cites an unresolved cited work.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:35:30.011952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:35:26.693692Z digest=sha256:14dc326a645e3b101cc2fddc6dac7d1d85b815a2e3eeba30256dcae944dffbf0

Observation f23ba1cc-33d7-4198-8548-4725c998a439 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation LoRA: Low-Rank Adaptation of Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:26.812818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:26.812818Z digest=sha256:b3bdd637f5b16ddaacad2884f22026d04ac73f28f2c47479c1ded0a9b5510ef2

Observation 088ddac3-4d2c-4796-97d3-8db2db9abe7b · outbound

This paper cites Membership Inference Attacks on Machine Learning: A Survey.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Membership Inference Attacks on Machine Learning: A Survey

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:26.911572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:26.911572Z digest=sha256:b55c6770fc10a836718b930f1862b6d9fa0a52b04ad1b1014f7eb66fba1ddd9b

Observation 3a55477a-f0cc-4719-8e26-4ed2979c2896 · outbound

This paper cites Estimating Contamination via Perplexity: Quantifying Memorisation in Language Model Evaluation.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Estimating Contamination via Perplexity: Quantifying Memorisation in Language Model Evaluation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:26.977795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:26.977795Z digest=sha256:fd6c851d082d81675bbcfd22acfce475d3dcb1c7d991070f861ccc12ef995e28

Observation c0af4880-be2f-4f7f-b237-dd4aed103c85 · outbound

This paper cites an unresolved cited work.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:27.076291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:27.076291Z digest=sha256:b4fe260028307b46953d1266b5505107a9829af38bf16d1a6281d1594f5f115e

Observation aca29502-3b06-4411-898b-37b2eb2c6e40 · outbound

This paper cites Lin, Jacob Hilton, and Owain Evans.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Lin, Jacob Hilton, and Owain Evans

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:35:29.821277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:35:27.190607Z digest=sha256:008a6da91012c16777a02dc12019226812a2226fcaa4cbda707e90078a4dc742

Observation e29cbdeb-17ed-44ca-be0f-87b2bd22759a · outbound

This paper cites DeepSeek-V3 Technical Report.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation DeepSeek-V3 Technical Report

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:27.321284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:27.321284Z digest=sha256:30e544546c2e2f19b0553d4719800908f5ccd157d1b52195418a6014ba460c7a

Observation 537db04c-fc95-4991-b246-5bc604ff91bc · outbound

This paper cites an unresolved cited work.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:27.401062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:27.401062Z digest=sha256:d7e98497fbf048ed20bca439b58efd851d8617c2d1f4dcef405a5dbd4b569753

Observation b4c6546f-481d-4c0f-a8a6-1c9b078a59b3 · outbound

This paper cites Training on the Benchmark Is Not All You Need.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Training on the Benchmark Is Not All You Need

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:27.501910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:27.501910Z digest=sha256:6270705e8fc1e26428fa07258fb38453275468634a61db8859ffabf99c86896e

Observation 022bc4e4-6bca-4107-a61d-e4d2afd41e0d · outbound

This paper cites GPT-4 Technical Report.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation GPT-4 Technical Report

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:27.657313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:27.657313Z digest=sha256:63b01487e62848184ff6090c491203cf751ff100542cd3815301b8e255e17ec4

Observation 85948a63-eefd-4c53-8708-a7d5933f2fbb · outbound

This paper cites an unresolved cited work.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:27.763705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:27.763705Z digest=sha256:e1e843a0bb3d5d7f1f437a2cad7834d9579f1c1a73aa46adb7a663a38cb4f849

Observation 457c0c24-c0e0-4a64-bfc4-7e22fc3024c4 · outbound

This paper cites Qwen2.5 Technical Report.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Qwen2.5 Technical Report

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:27.876148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:27.876148Z digest=sha256:7ca7abeb056279c68f46ffa34e96a037253ff7411e5aabcc1db71342aa177eb0

Observation c8d45bf4-cb03-4b7e-b8af-fae0dddd9161 · outbound

This paper cites Leveraging Large Language Models for Multiple Choice Question Answering.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Leveraging Large Language Models for Multiple Choice Question Answering

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:27.989088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:27.989088Z digest=sha256:8a87114313d1cde76a6229ab96f4851bfb39b8f9c2e4ce7bdf6dcec10cba13ac

Observation a1645894-ad5c-4459-87c3-38c1f24c1685 · outbound

This paper cites Towards Data Contamination Detection for Modern Large Language Models: Limitations, Inconsistencies, and Oracle Challenges.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Towards Data Contamination Detection for Modern Large Language Models: Limitations, Inconsistencies, and Oracle Challenges

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:28.145036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:28.145036Z digest=sha256:4f0fc9d87bda9a50a2fbbe44936aae3554aacd198f727aa7480efddd1c605fad

Observation d4a2cc5c-414e-4f72-93f1-84b947cfecfd · outbound

This paper cites SocialIQA: Commonsense Reasoning about Social Interactions.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation SocialIQA: Commonsense Reasoning about Social Interactions

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:28.259215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:28.259215Z digest=sha256:5aeae3ea67cc2fec5105f90180908b9af3067125a7c048ced058e3f20319f97b

Observation 877f8036-0eb7-428d-b4b5-b2904ec2975f · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Gemini: A Family of Highly Capable Multimodal Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:28.402129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:28.402129Z digest=sha256:6c619d59210064248af7379c2581ce8d55e040d508fbca3a8fea5f93b1a2604d

Observation 17509d7d-145e-4f20-b4c0-cafb062a0914 · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Gemma: Open Models Based on Gemini Research and Technology

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:28.488255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:28.488255Z digest=sha256:dbcd9ec3cba41a758077f63364c759ad87e54a4c0dceb1ac4b7b8566ce6e3a7e

Observation c169a4fc-480e-431e-9945-25576704a8e1 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Gemma 2: Improving Open Language Models at a Practical Size

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:28.599455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:28.599455Z digest=sha256:ce821d8d143073cad41abb34e9fbae10fc7f7dce4180f584e6ff79a1f28a85df

Observation d931bec7-4dd3-40f8-b778-8aef1d04f872 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation LLaMA: Open and Efficient Foundation Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:28.705768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:28.705768Z digest=sha256:6633f7800b069e113e981124b97a61b468609603611a0d5fc8e5b0e08a13531e

Observation 7265672e-7d0f-4ee7-aab8-3d9c9f00974d · outbound

This paper cites MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:28.812524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:28.812524Z digest=sha256:aa152279eb77767be12c1b2587543fa3d29158a01f782e7d5c5be28dc3efdd1e

Observation cc895bca-1fb5-4f50-872f-432387d4c5bd · outbound

This paper cites HuggingFace's Transformers: State-of-the-art Natural Language Processing.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation HuggingFace's Transformers: State-of-the-art Natural Language Processing

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:28.896218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:28.896218Z digest=sha256:a69b1f85bae2b0e70b063b186e2f4a0572b7bc63f61ea8c05e684f91c1af4122

Observation cbd67ca9-a106-43a1-a3b5-d6a351c50b4f · outbound

This paper cites Benchmarking Benchmark Leakage in Large Language Models.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Benchmarking Benchmark Leakage in Large Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:28.982799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:28.982799Z digest=sha256:193c6fcbeff8a5c251e7b569db1273fa756809f801d60289bfa68f45faff6df0

Observation dfe21523-ee1b-44fc-b0db-954c381c2c42 · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:29.070019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:29.070019Z digest=sha256:3e3fc830c2ae54c12e8bf0ad71a8c7b75013ed0e8fa5c748c0f56fdc053bb925

Pith citing papers

No inbound Pith citation observations are available.