Pith. sign in

Paper Citation Record · LEDGER

The Non-Determinism of Small LLMs: Evidence of Low Answer Consistency in Repetition Trials of Standard Multiple-Choice Benchmarks

As of 19 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 1 inbound Pith citation observation for arXiv:2509.09705.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.09705 v1

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T05:30:12.142238Z

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-14T13:32:31.523845Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

35 of 35 outbound references displayed

  • verified exact2
  • verified fuzzy0
  • unresolved33
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 207ec1f2-74c1-4806-99a3-75d9ccbbc67c · outbound

This paper cites URL: " 'urlintro :=.

The Non-Determinism of Small LLMs: Evidence of Low Answer Consistency in Repetition Trials of Standard Multiple-Choice Benchmarks URL: " 'urlintro :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T05:30:11.720305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T05:30:11.720305Z digest=sha256:6b1b55d7eddca7cbc2736fbb2b6ad79564fdfdc652c20682656c884f18c815f3

Observation e349950d-c763-487e-82da-d14146b15091 · outbound

This paper cites write newline.

The Non-Determinism of Small LLMs: Evidence of Low Answer Consistency in Repetition Trials of Standard Multiple-Choice Benchmarks write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T05:30:11.731015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T05:30:11.731015Z digest=sha256:1dc35b04e9a8ee0b2b9c4a64a5ba813dac00a7d234de94e2f5de3c618d5714a1

Observation 7ef62e0d-e755-43c7-9b4e-7911f0e99891 · outbound

This paper cites an unresolved cited work.

The Non-Determinism of Small LLMs: Evidence of Low Answer Consistency in Repetition Trials of Standard Multiple-Choice Benchmarks Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:30:13.775939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-05T05:30:11.752300Z digest=sha256:df3509c86ab2e3fa6c50aa1e869561c035ba1c73225bbbae2eb480fd5c8a4d72

Observation 27856e2b-641e-4404-9e61-9f9bd02c79c9 · outbound

This paper cites Non-Determinism of "Deterministic" LLM Settings.

The Non-Determinism of Small LLMs: Evidence of Low Answer Consistency in Repetition Trials of Standard Multiple-Choice Benchmarks Non-Determinism of "Deterministic" LLM Settings

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T05:30:11.765974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T05:30:11.765974Z digest=sha256:5357ecf4207ae3002e3f155d4fee2ba69eb19e5b82bf70327ffa44517b866d6a

Observation 1f9eb626-f268-44a2-8aea-3cf7b858a701 · outbound

This paper cites an unresolved cited work.

The Non-Determinism of Small LLMs: Evidence of Low Answer Consistency in Repetition Trials of Standard Multiple-Choice Benchmarks Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:30:13.753752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-05T05:30:11.778515Z digest=sha256:42359cd55e5c74f02dc066f968b215e15eae9d539f2138c52bd7c3c1585a20be

Observation c0cd5982-cabc-4996-b23a-0b0358a8ed3c · outbound

This paper cites A Survey on Data Contamination for Large Language Models.

The Non-Determinism of Small LLMs: Evidence of Low Answer Consistency in Repetition Trials of Standard Multiple-Choice Benchmarks A Survey on Data Contamination for Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T05:30:11.788526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T05:30:11.788526Z digest=sha256:a718ebff732f543aecfe5512b4edef4394b53e35c651713dbf73ce3105e9f2ff

Observation fb17db78-f3c6-4598-9b8a-b653b4d6195f · outbound

This paper cites an unresolved cited work.

The Non-Determinism of Small LLMs: Evidence of Low Answer Consistency in Repetition Trials of Standard Multiple-Choice Benchmarks Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:30:13.713303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-05T05:30:11.799447Z digest=sha256:110bb5738bf071d0a45a3bbaf349df51bd25746719f0c08a18068d438ddde829

Observation aabe2cdf-f60e-4cdd-bb5f-a12c9db5e4c6 · outbound

This paper cites DeepSeek-V3 Technical Report.

The Non-Determinism of Small LLMs: Evidence of Low Answer Consistency in Repetition Trials of Standard Multiple-Choice Benchmarks DeepSeek-V3 Technical Report

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T05:30:11.816722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T05:30:11.816722Z digest=sha256:b85105d0e3bc6d43d3808a1d4f7124ea24b49039945f9c54a710462770991094

Observation 08c015af-f399-487c-8984-c97a56ef9c14 · outbound

This paper cites The Llama 3 Herd of Models.

The Non-Determinism of Small LLMs: Evidence of Low Answer Consistency in Repetition Trials of Standard Multiple-Choice Benchmarks The Llama 3 Herd of Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T05:30:11.824513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T05:30:11.824513Z digest=sha256:69df0fe173c7edb0041f944f2d84d0797538521735ecf817917dcbaf46905ee6

Observation eee64261-ae94-4e30-b07d-679bc36d1498 · outbound

This paper cites Are We Done with MMLU?.

The Non-Determinism of Small LLMs: Evidence of Low Answer Consistency in Repetition Trials of Standard Multiple-Choice Benchmarks Are We Done with MMLU?

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T05:30:11.831165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T05:30:11.831165Z digest=sha256:d6cd5928b0cd98307270b4098ff51a2bd91a4a9f43a27a4b7cb67e7e6c325445

Observation f5a04dc6-4864-4011-ad56-8d226b71bb46 · outbound

This paper cites an unresolved cited work.

The Non-Determinism of Small LLMs: Evidence of Low Answer Consistency in Repetition Trials of Standard Multiple-Choice Benchmarks Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:30:13.673780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-05T05:30:11.852164Z digest=sha256:baecd4dd0e34e757d47e9d38966d673f77c4ee0994c87e989c96450a0e19a5ec

Observation 99ad80cc-4ac6-481b-984d-5dae7d32e6e7 · outbound

This paper cites MedAlpaca -- An Open-Source Collection of Medical Conversational AI Models and Training Data.

The Non-Determinism of Small LLMs: Evidence of Low Answer Consistency in Repetition Trials of Standard Multiple-Choice Benchmarks MedAlpaca -- An Open-Source Collection of Medical Conversational AI Models and Training Data

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T05:30:11.872307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T05:30:11.872307Z digest=sha256:619a964d56e2342f0ad28bf6d4ee31223f8ec04d69e2ed6dcf41ceaf4cfdafe4

Observation db36a074-6b34-4a90-b967-71adeb2ff198 · outbound

This paper cites Mistral 7B.

The Non-Determinism of Small LLMs: Evidence of Low Answer Consistency in Repetition Trials of Standard Multiple-Choice Benchmarks Mistral 7B

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T05:30:11.881873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T05:30:11.881873Z digest=sha256:39fa32eea2e1fc6d96425ed2b0f6bec3bd3bb43f21e2978ca47a5d14bb5e7ee1

Observation 90097737-fd3f-46fd-9844-9b87c0ac3cd9 · outbound

This paper cites Mixtral of Experts.

The Non-Determinism of Small LLMs: Evidence of Low Answer Consistency in Repetition Trials of Standard Multiple-Choice Benchmarks Mixtral of Experts

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T05:30:11.889738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T05:30:11.889738Z digest=sha256:d3a92dffac2a1b011275d28b0d11210b0396d053326df938a32b406e92d0612d

Observation a4229807-4463-4e7d-8db4-dbf26670bf7b · outbound

This paper cites an unresolved cited work.

The Non-Determinism of Small LLMs: Evidence of Low Answer Consistency in Repetition Trials of Standard Multiple-Choice Benchmarks Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T05:30:11.912244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T05:30:11.912244Z digest=sha256:d12a7eac8633cbb7cd7473d97d0b691c2143e14f4066725f4aa412852821c0f5

Observation ebe18297-a12e-4106-8898-141bca6a9562 · outbound

This paper cites BioMistral: A Collection of Open-Source Pretrained Large Language Models for Medical Domains.

The Non-Determinism of Small LLMs: Evidence of Low Answer Consistency in Repetition Trials of Standard Multiple-Choice Benchmarks BioMistral: A Collection of Open-Source Pretrained Large Language Models for Medical Domains

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T05:30:11.929147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T05:30:11.929147Z digest=sha256:831f3de91ae84f86f36cad0ff423556076b4e1c7615976156610f80303c23c2b

Observation 8c171ac4-f432-4ce3-9ba7-6edd5f4bac2f · outbound

This paper cites Evaluating the Consistency of LLM Evaluators.

The Non-Determinism of Small LLMs: Evidence of Low Answer Consistency in Repetition Trials of Standard Multiple-Choice Benchmarks Evaluating the Consistency of LLM Evaluators

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-05T05:30:12.810095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-05T05:30:11.935278Z digest=sha256:65a3280af2dddc4e1a97994bd51736fe377db18aaf253286b9a0ec44120b02c9

Observation cfaf754b-7eb9-4a4a-b26c-d547803a0840 · outbound

This paper cites an unresolved cited work.

The Non-Determinism of Small LLMs: Evidence of Low Answer Consistency in Repetition Trials of Standard Multiple-Choice Benchmarks Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:30:13.608082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-05T05:30:11.948137Z digest=sha256:c9856f13fa4396a8d63b14a14c881c95c0b0016bb291f0096b92bdce8111317a

Observation d92a01ec-a93a-42f0-b1ec-822f0ae917fc · outbound

This paper cites an unresolved cited work.

The Non-Determinism of Small LLMs: Evidence of Low Answer Consistency in Repetition Trials of Standard Multiple-Choice Benchmarks Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:30:13.563492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-05T05:30:11.955337Z digest=sha256:e8dc4a6a9bf724ba513d2fc97e0457a6492d6578001045d83174583713a6fc9f

Observation 1d5af5e3-1471-418f-989f-4be0fc108edc · outbound

This paper cites GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models.

The Non-Determinism of Small LLMs: Evidence of Low Answer Consistency in Repetition Trials of Standard Multiple-Choice Benchmarks GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T05:30:11.961290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T05:30:11.961290Z digest=sha256:c3488443eeaff91e9eedf95689a0a5913d28453c83c02cb228c207f9b74a99d0

Observation 9fe68723-2e1e-42f6-8e7c-9d7111755257 · outbound

This paper cites SCORE: Systematic COnsistency and Robustness Evaluation for Large Language Models.

The Non-Determinism of Small LLMs: Evidence of Low Answer Consistency in Repetition Trials of Standard Multiple-Choice Benchmarks SCORE: Systematic COnsistency and Robustness Evaluation for Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T05:30:11.968973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T05:30:11.968973Z digest=sha256:7cd2a119595d8340fa58b1ebe34d890214942606df2c7fc7234e33ee042d20d3

Observation 32e84c90-d21c-415e-9b3a-d8d4390ea1a3 · outbound

This paper cites Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine.

The Non-Determinism of Small LLMs: Evidence of Low Answer Consistency in Repetition Trials of Standard Multiple-Choice Benchmarks Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T05:30:11.982290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T05:30:11.982290Z digest=sha256:434212008bafff463d9f9dd7bbaaecd8928d0f66ee0c7ea4f6f280dad8c7b884

Observation eb17a5ab-3b25-4dd2-821e-13418371b0cd · outbound

This paper cites an unresolved cited work.

The Non-Determinism of Small LLMs: Evidence of Low Answer Consistency in Repetition Trials of Standard Multiple-Choice Benchmarks Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:30:13.520034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-05T05:30:11.995696Z digest=sha256:3cee1a565d1cea9f657d2c890655dbee310dfdcf94931a90015468c266af23d7

Observation 64e6d07a-7fba-4cb0-b81f-c7b18967a902 · outbound

This paper cites an unresolved cited work.

The Non-Determinism of Small LLMs: Evidence of Low Answer Consistency in Repetition Trials of Standard Multiple-Choice Benchmarks Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:30:13.475849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-05T05:30:12.004059Z digest=sha256:7d08274c7d50dfcadeb2fe909e3ac0aa8f0d9a9a5cd6f8d4050184241f988fce

Observation 330975ec-2d3a-41a7-af3e-9c04f131139b · outbound

This paper cites an unresolved cited work.

The Non-Determinism of Small LLMs: Evidence of Low Answer Consistency in Repetition Trials of Standard Multiple-Choice Benchmarks Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T05:30:12.013568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T05:30:12.013568Z digest=sha256:5449069ef50525711d13df44a61058a2d1d5385bb45d8cfcd48dc3ec1444a122

Observation cb92023f-dffa-4cf7-8e2a-a2661d56cf30 · outbound

This paper cites Qwen2.5 Technical Report.

The Non-Determinism of Small LLMs: Evidence of Low Answer Consistency in Repetition Trials of Standard Multiple-Choice Benchmarks Qwen2.5 Technical Report

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T05:30:12.023639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T05:30:12.023639Z digest=sha256:416d2dc80bdf2dfa57e7c5b2db0f039718ece1500ddbc545e298bcc567b44d7a

Observation cdc53206-30bd-46d4-bf55-97148b1c1409 · outbound

This paper cites A Comprehensive Survey of Contamination Detection Methods in Large Language Models.

The Non-Determinism of Small LLMs: Evidence of Low Answer Consistency in Repetition Trials of Standard Multiple-Choice Benchmarks A Comprehensive Survey of Contamination Detection Methods in Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T05:30:12.032640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T05:30:12.032640Z digest=sha256:0bcaabb6cc6b469d7db76f9dc5427642eac179e289be20273a4f38057e003bf1

Observation df548de7-c07b-4da8-be3d-345a703d8947 · outbound

This paper cites Towards Expert-Level Medical Question Answering with Large Language Models.

The Non-Determinism of Small LLMs: Evidence of Low Answer Consistency in Repetition Trials of Standard Multiple-Choice Benchmarks Towards Expert-Level Medical Question Answering with Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T05:30:12.044204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T05:30:12.044204Z digest=sha256:5646bc42b04888979e484374fa7a40811d65d5a28e58046ba9198dbad610b038

Observation c64a07ba-1aca-4d10-bbe2-35b3646355a2 · outbound

This paper cites The Good, The Bad, and The Greedy: Evaluation of LLMs Should Not Ignore Non-Determinism.

The Non-Determinism of Small LLMs: Evidence of Low Answer Consistency in Repetition Trials of Standard Multiple-Choice Benchmarks The Good, The Bad, and The Greedy: Evaluation of LLMs Should Not Ignore Non-Determinism

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T05:30:12.058244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T05:30:12.058244Z digest=sha256:39e27f96d7399fe8d132d8747653358948824dcde6bd90106573a32ef0eb9df7

Observation 1e6d670a-6487-463e-99e5-4146d2022ec8 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

The Non-Determinism of Small LLMs: Evidence of Low Answer Consistency in Repetition Trials of Standard Multiple-Choice Benchmarks LLaMA: Open and Efficient Foundation Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T05:30:12.076195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T05:30:12.076195Z digest=sha256:7e112962b2727e3bb2981c462bb6ba280719679de395324fd45e75cb156685ac

Observation 0d3f1bbd-4f92-41df-8291-c3b58d7aa63a · outbound

This paper cites an unresolved cited work.

The Non-Determinism of Small LLMs: Evidence of Low Answer Consistency in Repetition Trials of Standard Multiple-Choice Benchmarks Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:30:13.423747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-05T05:30:12.087544Z digest=sha256:b4a364ca8746b647c827094cadc05ca73e2e2140e6ee1f081de32b88bcc5e27f

Observation 25c0a7a5-d43d-4815-abdb-7269cfbfa7e7 · outbound

This paper cites an unresolved cited work.

The Non-Determinism of Small LLMs: Evidence of Low Answer Consistency in Repetition Trials of Standard Multiple-Choice Benchmarks Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T05:30:12.108055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T05:30:12.108055Z digest=sha256:d369bb50c4ed9f1f908c513c30342a4bf684e0396b641e55553ac39cf8aa9a5a

Observation d91d2787-d1d4-4461-9d7f-e4334cdea54e · outbound

This paper cites LLMs May Perform MCQA by Selecting the Least Incorrect Option.

The Non-Determinism of Small LLMs: Evidence of Low Answer Consistency in Repetition Trials of Standard Multiple-Choice Benchmarks LLMs May Perform MCQA by Selecting the Least Incorrect Option

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T05:30:12.120281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T05:30:12.120281Z digest=sha256:ff9d70c6364b8ea65b2f200689a8fcc68305345019a23496c99cf99d5215410e

Observation 022b1ac3-f3d7-4b9e-8f2a-2a5bb6ddf8ed · outbound

This paper cites Rethinking Generative Large Language Model Evaluation for Semantic Comprehension.

The Non-Determinism of Small LLMs: Evidence of Low Answer Consistency in Repetition Trials of Standard Multiple-Choice Benchmarks Rethinking Generative Large Language Model Evaluation for Semantic Comprehension

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-08-05T05:30:12.376125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-05T05:30:12.131461Z digest=sha256:aed33a23b5ba659eb9e1d333793804d68f2b8cf99a5adb183418967d4cb9974c

Observation 04b02f54-93e5-4da2-97ae-62edbabf7138 · outbound

This paper cites Benchmark Data Contamination of Large Language Models: A Survey.

The Non-Determinism of Small LLMs: Evidence of Low Answer Consistency in Repetition Trials of Standard Multiple-Choice Benchmarks Benchmark Data Contamination of Large Language Models: A Survey

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T05:30:12.142238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T05:30:12.142238Z digest=sha256:b759ab0c1dc6d1837edc0bf21fe9e528ae4711aed1235a1b9855db6895821f7c

Pith citing papers

Observation 2eea4e8a-cf4f-4f95-b579-7ef6b518771e · inbound

When Counterbalancing Hides the Bias: Access-Conditioned Position Lock in Forced-Choice LLM Evaluation cites this paper.

When Counterbalancing Hides the Bias: Access-Conditioned Position Lock in Forced-Choice LLM Evaluation The Non-Determinism of Small LLMs: Evidence of Low Answer Consistency in Repetition Trials of Standard Multiple-Choice Benchmarks

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-14T13:32:31.523845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T13:32:31.523845Z digest=sha256:a88e249a69de9538fe1b9269a731040644e479db902d050c1ab438cb07108447