Pith. sign in

Paper Citation Record · LEDGER

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks

As of 20 August 2026, this Paper Citation Record lists 66 of 66 outbound references and 0 inbound Pith citation observations for arXiv:2608.02966.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.02966 v1

Coverage vector

measured 66 of 66 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T14:58:40.410979Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

66 of 66 outbound references displayed

  • verified exact1
  • verified fuzzy39
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 98a919b9-dec6-439a-b202-39dd39870100 · outbound

This paper cites Darrell , title =.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Darrell , title =

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T14:58:39.046108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:58:39.046108Z digest=sha256:2ff65f115ff327bef948d1125aa962cc62d479240ce7c3f9e65fe4e66bd9519c

Observation 49f83c5c-c3a7-4d2e-9b72-a61a452c5a80 · outbound

This paper cites Psychometrika , year =.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Psychometrika , year =

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:44.261192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T14:58:39.119070Z digest=sha256:1db815e741d910c41372ed341b897fff61f8e959ef721c72a9f0fae5926f806a

Observation e363868c-6342-4e44-b9f9-90da01d52558 · outbound

This paper cites , title =.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks , title =

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:44.105716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T14:58:39.207881Z digest=sha256:8a4781180820e0394eceac7920b6e4ac8a1d3955f419501d083973d4b67cae30

Observation 1cfc82f6-00a7-46a8-9d4a-6880c92a96a5 · outbound

This paper cites Measurement: Interdisciplinary Research and Perspectives , year =.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Measurement: Interdisciplinary Research and Perspectives , year =

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:44.061099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T14:58:39.212976Z digest=sha256:46cabd5de00a9e9ea58ce08c82922c8a2763c82561b1af84eb0dde3c17968c33

Observation 42f866ec-59ac-42be-b44a-cba553a23282 · outbound

This paper cites , title =.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks , title =

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:44.042634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T14:58:39.217844Z digest=sha256:2d46ed5ad3da879434a3cfac30eda151a18c7010e2a191b3656686a9ea72e940

Observation 73ab8a78-50c6-439e-9029-d8baf7bd3019 · outbound

This paper cites Findings of the Association for Computational Linguistics: EMNLP 2023 , year =.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Findings of the Association for Computational Linguistics: EMNLP 2023 , year =

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:44.024830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T14:58:39.223126Z digest=sha256:fc3f501e5022d448b0ab0c18e8fe734ae431636d39a46e263588e2aa049e0186

Observation 8f71685b-bdd6-4a45-85a3-68c9551f999a · outbound

This paper cites Findings of the Association for Computational Linguistics: NAACL 2024 , pages=.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Findings of the Association for Computational Linguistics: NAACL 2024 , pages=

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T14:58:39.227968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:58:39.227968Z digest=sha256:aea9511349c7bda6fba54a682062148f6334efb4fa811b6e9c927cc6a5ab5461

Observation 4b1556e6-2d05-4252-bf2e-4dca6929eed9 · outbound

This paper cites Journal of Educational Measurement , volume=.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Journal of Educational Measurement , volume=

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:43.949220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T14:58:39.232861Z digest=sha256:f8b592d1abaeacbd71209a5e27a62e8113946b32fda3dfcf4c9d9b11863e1d84

Observation 252a3a37-a1ed-4fc2-ada8-55cbaa1a1462 · outbound

This paper cites From Loops to Oops: Fallback Behaviors of Language Models Under Uncertainty.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks From Loops to Oops: Fallback Behaviors of Language Models Under Uncertainty

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T14:58:39.237142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:58:39.237142Z digest=sha256:b1c2a037f6a70a2274ca30a2e429c86376a407f2fbb9931061b533af651b9443

Observation bcc1d5fa-ab3e-4664-99a6-8b7354550dd0 · outbound

This paper cites and Linn, Robert L.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks and Linn, Robert L

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:43.825304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T14:58:39.243238Z digest=sha256:3932deeb124a961172719f33d219da022b9f7306f11bbd5d502dfb62b0311da8

Observation 466cfa14-f956-4623-8254-1127a8100af2 · outbound

This paper cites Applied Psychological Measurement , volume=.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Applied Psychological Measurement , volume=

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:43.806408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T14:58:39.249743Z digest=sha256:e839d3d56e003abd1331396ef9e12c2bf8239279eaafd60ffa00521ee4ca0bf0

Observation 09a27488-dcbf-43b2-b564-aad4156788fa · outbound

This paper cites Marketing science , volume=.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Marketing science , volume=

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:43.724061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T14:58:39.255201Z digest=sha256:7225fe6551438fa9f367c494bd1b9c549ea29fbe0d7492dad4259ba2164c415b

Observation e80b7b7f-c848-4957-a7e2-d9148bee9788 · outbound

This paper cites Transportation , volume=.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Transportation , volume=

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:43.659772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T14:58:39.259897Z digest=sha256:a736caec5d88ea9333e625472bd4c9201d31e049fba35d4d25579725e17c5ff7

Observation 1887a2ca-bdfd-4c54-96e9-6fa6225cbb8a · outbound

This paper cites , title =.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks , title =

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T14:58:39.265177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:58:39.265177Z digest=sha256:7cd6f1dc0e724f14aef5190103a2e695374ed262fb26f7634768b875601a5967

Observation 44b107b5-2ca9-427c-b3f0-5a49e1db1cad · outbound

This paper cites Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics , year =.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics , year =

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:43.606732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T14:58:39.269644Z digest=sha256:a32ca7baedd0bda730da4f4c7812bac99b79d6e84635709c2627b815784017e4

Observation 275e80f0-bae8-4024-802b-a7d6a3c754d9 · outbound

This paper cites Statistical Theories of Mental Test Scores , editor =.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Statistical Theories of Mental Test Scores , editor =

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T14:58:39.274850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:58:39.274850Z digest=sha256:34fc7d9413cb0cc6378f4019b98815ad2101f19529d11329059de69ca0181d12

Observation eae2c5e3-1856-432e-9edc-2209300be1a2 · outbound

This paper cites Measurement: Interdisciplinary Research and Perspectives , volume=.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Measurement: Interdisciplinary Research and Perspectives , volume=

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:43.504572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T14:58:39.279399Z digest=sha256:fccd1961f731d9f91b2695be07c3b3deee43cc7e0b0eaa27f8a32f0cfef1e6ff

Observation ba14c083-016c-4271-872a-38a1c37b0263 · outbound

This paper cites On the Unidentifiability of the Fixed-Effects.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks On the Unidentifiability of the Fixed-Effects

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:43.485264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T14:58:39.283837Z digest=sha256:97f579aec07d864128815266aea93cbd94f525f808a53d7be8961c05daaa8d81

Observation cc1e30ac-1a4f-4b26-a7c1-010f4fd28b8b · outbound

This paper cites and Lord, Frederic M.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks and Lord, Frederic M

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:43.433906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T14:58:39.342230Z digest=sha256:d2243af75accbf531f510c3b0175825433edb7a2536e20a30fa1add6ae63b2c3

Observation e629055f-9993-48d0-8623-aef14d46fb21 · outbound

This paper cites Philip and Skene, Allan M.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Philip and Skene, Allan M

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:43.308092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T14:58:39.442986Z digest=sha256:79cb70c601b8af0cff71d4dc9c5d1ac3ccff896f53994760b8ee9f74db0f1e83

Observation 16490a9c-ed94-449d-871e-c36982697ff3 · outbound

This paper cites Advances in Neural Information Processing Systems 23 , year =.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Advances in Neural Information Processing Systems 23 , year =

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:43.290454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T14:58:39.491224Z digest=sha256:9028553be3ac6bb2d7877d8f945244b5203b37e18f6c8b0aaa83046b52c451b6

Observation 31bdd3df-ac03-4c39-8953-b0efa414effd · outbound

This paper cites Advances in Neural Information Processing Systems 37, Datasets and Benchmarks Track , year =.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Advances in Neural Information Processing Systems 37, Datasets and Benchmarks Track , year =

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:43.274140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T14:58:39.496361Z digest=sha256:f56d40935491ee3e8402e1bad8de7d6bea839f6729ffb8377cf2bd9fce056cce

Observation 92b33539-a865-4026-8da0-581a48bf1772 · outbound

This paper cites , title =.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks , title =

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:43.156418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T14:58:39.501364Z digest=sha256:456ef32cd34b223aa0d976833fad115da9bc734eaf2d201a32e56f90038916b5

Observation 135fd128-a544-4b68-b6bf-2097d387b460 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T14:58:39.506630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:58:39.506630Z digest=sha256:d1ab869e8431ff1237360669711bf1778933b3fa1ed204954fca408cfe46c73d

Observation bd1913b7-16ae-4c7d-8400-bdea67ff9f2e · outbound

This paper cites Findings of the Association for Computational Linguistics: NAACL 2024 , year =.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Findings of the Association for Computational Linguistics: NAACL 2024 , year =

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:43.137872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T14:58:39.513083Z digest=sha256:4908a9a8996cece1c1d9d33a44b2722d349ee4cec1c5ad8d9b162b3f2c87042f

Observation eb5a045f-c8d8-468f-96b1-7927be04d7cc · outbound

This paper cites Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics , year =.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics , year =

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:43.045165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T14:58:39.518553Z digest=sha256:c490ca3e825de15e2dea635f26798f0e2ab67077a0165abcb51e565b8efa67ac

Observation 8449a83d-8cba-4f63-857a-b0b2aceab770 · outbound

This paper cites IEEE/ACM Transactions on Audio, Speech, and Language Processing , volume=.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks IEEE/ACM Transactions on Audio, Speech, and Language Processing , volume=

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T14:58:39.523190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:58:39.523190Z digest=sha256:395e00c3fdf790abc4f12f145042c2fd1694dc6dc4abb90c84e1b65181293759

Observation f74b1681-a0b0-4240-ba1f-1e5932b534c3 · outbound

This paper cites Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing , year =.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing , year =

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:42.931463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T14:58:39.527766Z digest=sha256:8728bed443d49560984732794bda3f77e5c87e62d2809f3b143eeff9f63d2c1e

Observation d66e6103-db5e-4788-830c-b4514b182854 · outbound

This paper cites Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics , year =.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics , year =

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:42.913179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T14:58:39.532191Z digest=sha256:9999ca7c755215be76b9ba424dd1cd62bfae7d61b790c58b05bc6aab6b2898c5

Observation 5730f55f-d330-4240-80b1-cb163c39e95c · outbound

This paper cites Applied Sciences , year =.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Applied Sciences , year =

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:42.772270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T14:58:39.537816Z digest=sha256:70622392779d474f198891a6478d8ddcb0795e5720fa7a1b780c4a26750af263

Observation 743b2775-6ef4-40ad-a896-5649c658398e · outbound

This paper cites Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing , year =.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing , year =

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:42.704715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T14:58:39.622501Z digest=sha256:7ef1fee925b93808844f0433c90ccde3bff14f953edcc0c8b36605f9b299f5fb

Observation 3f39efbf-f6f1-4052-9402-79645ac5139f · outbound

This paper cites an unresolved cited work.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T14:58:39.695919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:58:39.695919Z digest=sha256:b496dfb9a668eb07855a2c080985b8a08df5bf72a651d30b006c2b9bd0a175cf

Observation 823da1a1-170c-4777-a176-ccd8c8e3120c · outbound

This paper cites Proceedings of the 34th AAAI Conference on Artificial Intelligence , year =.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Proceedings of the 34th AAAI Conference on Artificial Intelligence , year =

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:42.676083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T14:58:39.753538Z digest=sha256:6ba8ced215c6182e09e1faebc6d00996245a0a39173fa8f624c3a68856d8c285

Observation bf456cb8-86bf-45c3-bc28-2a0a9e346376 · outbound

This paper cites International Conference on Learning Representations , year =.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks International Conference on Learning Representations , year =

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:42.659007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T14:58:39.758422Z digest=sha256:dc9790abf21597a12f7b53e4ee80686a49072639f29fa6c8ac3bf4a1a823d4ee

Observation 391415ce-7682-415f-9bf0-0c6e869888ab · outbound

This paper cites International Conference on Learning Representations , year =.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks International Conference on Learning Representations , year =

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T14:58:39.763766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:58:39.763766Z digest=sha256:889a19edc4400938d96388346e2ec83c66e1afa4c7b5cd64e2e22fae03aaf106

Observation 8b307df6-310e-49ec-a3a6-160c9715e38d · outbound

This paper cites Proceedings of the AAAI Conference on Artificial Intelligence , volume=.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Proceedings of the AAAI Conference on Artificial Intelligence , volume=

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:42.537237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T14:58:39.768489Z digest=sha256:1e97db79e1445d3faa57af35befc1844fd2f4f7911ea28833c135882934a3771

Observation fd090313-a6fe-4f36-83e3-077afbdb9f52 · outbound

This paper cites Probabilistic models for some intelligence and attainment tests.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Probabilistic models for some intelligence and attainment tests

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T14:58:39.773095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:58:39.773095Z digest=sha256:56e68880340a45a87c0ffc307467d50ac975d4d37477fd1849f8673734dfec32

Observation ac383f0e-602c-4944-a6ea-0f2e6cc0f6d1 · outbound

This paper cites 1968 , publisher=.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks 1968 , publisher=

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:42.427621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T14:58:39.778953Z digest=sha256:1d0578fb17c50ee45d467edeb0ab00a8efd32990ee0779e83fdf171c2ccc3c1b

Observation 92de8232-a970-4f2f-9670-c82f5568a34e · outbound

This paper cites Proceedings of the 12th International Conference on Educational Data Mining , year =.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Proceedings of the 12th International Conference on Educational Data Mining , year =

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:42.406286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T14:58:39.783899Z digest=sha256:f0896f419ba52b34c2a3898e5612693274ffc090fb27666a38815aa64a07935d

Observation 891d7191-25f3-40f7-85c3-bbe2997d65e4 · outbound

This paper cites an unresolved cited work.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-15T14:58:42.274400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T14:58:39.788953Z digest=sha256:835a591298b10ccd85c8902baeb40a7beac3af88799c7302235bebd9a5b05ab4

Observation 19025d95-827b-4ff1-a962-8d98ff1be4cd · outbound

This paper cites Annual Meeting of the Association for Computational Linguistics , year=.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Annual Meeting of the Association for Computational Linguistics , year=

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:42.212821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T14:58:39.794229Z digest=sha256:ee9374be3fae599e37a4ee3deae64a8168cab7df555fc8422d844a6d219b0574

Observation 7b740791-a193-4a6e-b833-e1e712d36b07 · outbound

This paper cites 2025 , eprint=.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks 2025 , eprint=

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:42.145916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T14:58:39.799355Z digest=sha256:c0219cc09008cdd5037849e77b5c6ffa345a3467781a4fdebdb69e2d454dd026

Observation 9be63788-8820-4dc1-bdb8-e98f3d805e94 · outbound

This paper cites arXiv preprint arXiv:2606.15643 , year=.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks arXiv preprint arXiv:2606.15643 , year=

Reference 43

Resolution
verified exact
raw_fallback, observed 2026-08-15T14:58:41.319337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T14:58:39.803809Z digest=sha256:4d504fd73e127189ee48f1769ab9cacf33dc070c86546de3cce3b7c96c7ccb84

Observation dcf39de4-b277-4ab4-af22-face4a94ea6a · outbound

This paper cites Adaptive Testing for LLM Evaluation: A Psychometric Alternative to Static Benchmarks.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Adaptive Testing for LLM Evaluation: A Psychometric Alternative to Static Benchmarks

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T14:58:39.837621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:58:39.837621Z digest=sha256:af0c9dbe0d4f4868584cb7d1508b96786f0f2ae8cce72cef115215d14aa3e3c3

Observation 3914437c-a570-4816-99bd-f448b11d03fe · outbound

This paper cites Proceedings of the AAAI Conference on Artificial Intelligence , volume=.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Proceedings of the AAAI Conference on Artificial Intelligence , volume=

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:42.097855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T14:58:39.942519Z digest=sha256:04c920b37898056c761e58e8445a86bdf331c9cac64877250d4349efafa9547a

Observation fc5cc222-e378-40ee-aa02-12ba49c7b4f9 · outbound

This paper cites Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages=.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages=

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:41.994121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T14:58:40.025703Z digest=sha256:63b6a80edfbf8c08a69587b0e644e2d31ffe4cec7459fff0a6c4205d586cceef

Observation 48acfe07-7d24-4984-8462-7fa5334b4551 · outbound

This paper cites Language Models (Mostly) Know What They Know.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Language Models (Mostly) Know What They Know

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T14:58:40.046423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:58:40.046423Z digest=sha256:f7195f02af65e1b1ee4601955418aed9e90c11ff4c7ef2b49c71d99a5f70ba09

Observation 2a82a322-3441-4796-b949-4df879b38398 · outbound

This paper cites Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages=.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages=

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T14:58:40.052098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:58:40.052098Z digest=sha256:5fab0fa16fffd60889c237124a311c1c1c6aa2a07de24417a1acdc1e495c716c

Observation 8965e48b-1db1-41f6-b05f-85b037253ebd · outbound

This paper cites arXiv preprint arXiv:2509.10625 , year=.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks arXiv preprint arXiv:2509.10625 , year=

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T14:58:40.056761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:58:40.056761Z digest=sha256:da3e0d24047c9eb53d01a5d2f8f39eed01d4102ec1fd5b1f0ae90431ac91482d

Observation e320e6a1-cf39-4465-9788-2e469bfe716d · outbound

This paper cites Findings of the Association for Computational Linguistics: ACL 2025 , pages=.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Findings of the Association for Computational Linguistics: ACL 2025 , pages=

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:41.865410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T14:58:40.061858Z digest=sha256:7eab27837b1adf869e94d07f123bbc2e67b5d5396cc645087511bc2e6a4b582c

Observation 9cfeaf8c-6499-47f7-a276-c9b42cfd7208 · outbound

This paper cites Auditing LLM Benchmarks with Item Response Theory.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Auditing LLM Benchmarks with Item Response Theory

Reference 51

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T14:58:40.949418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T14:58:40.067379Z digest=sha256:c53deceee5bddc0ce6aa0b9e26226a2c8b524c99a7667bbc7e3e61a333aebe46

Observation 3009f9d1-f208-4363-8315-770cc243e049 · outbound

This paper cites Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages=.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages=

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:41.846317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T14:58:40.072684Z digest=sha256:29b78e6de3eb74698399bc669a6f83b65adc7af1ab8f570cc9efebaba1c1a688

Observation c3269eb1-3b2f-4186-80ae-325b83e779cd · outbound

This paper cites Quantifying Consistency in LLM Logical Reasoning via Structural Uncertainty.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Quantifying Consistency in LLM Logical Reasoning via Structural Uncertainty

Reference 53

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T14:58:40.759762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T14:58:40.078037Z digest=sha256:160ba3981d8a574d896bac8f42763edc3a24a7b0f235dcd64f5d318bc84f68a0

Observation 86c54e22-70da-4285-8f37-224c27091867 · outbound

This paper cites Proceedings of the AAAI Conference on Artificial Intelligence , volume=.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Proceedings of the AAAI Conference on Artificial Intelligence , volume=

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:41.757465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T14:58:40.083482Z digest=sha256:a1c8d4d0545787fb08b7c61ba3fd8fe730ebed9f560eab60ae0f2cb077914090

Observation 2b4193f8-9359-4f7d-a2d4-f298689cad4d · outbound

This paper cites International Conference on Learning Representations , volume=.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks International Conference on Learning Representations , volume=

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:41.652754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T14:58:40.170145Z digest=sha256:ac72fbdc36644de2c753eef58a9dd593626f2e94c9e8618a5903319e0b72a16d

Observation c9c76461-eb64-4550-8223-1c2edc8c8d8a · outbound

This paper cites an unresolved cited work.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-15T14:58:41.573767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T14:58:40.220073Z digest=sha256:254efb34cf3cf8365169f8d76faf21df8d8fe4a6ec36a138ab55d91528ec6027

Observation 868a54fa-74fc-47cf-9642-0aa782cabb4d · outbound

This paper cites Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T14:58:40.225030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:58:40.225030Z digest=sha256:80506befda0a33866ef946f2c2b7fabd38b3098429295b12d2ceb7ecca273d94

Observation 7b8055b6-ef0c-49f8-bf80-5f5a3343f606 · outbound

This paper cites The Masked Advantage: Uncovering Local-Language Access to Cultural Knowledge in LLMs.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks The Masked Advantage: Uncovering Local-Language Access to Cultural Knowledge in LLMs

Reference 58

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T14:58:40.596362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T14:58:40.229901Z digest=sha256:a60c709d831cc906efbd0c969a0a0d9bcce5e9f9b40fdef5e96e1f19d8dfff16

Observation c31ed83a-6dba-456f-9667-c33999fe6cce · outbound

This paper cites Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing , pages =.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing , pages =

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-15T14:58:40.235979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:58:40.235979Z digest=sha256:f35fb96caf0936fe3b72d47a7f0b2d7a45934c684c0c09821c06c6a47005416a

Observation fffebd34-82ae-4557-a7e8-05c872ba85c5 · outbound

This paper cites Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) , pages =.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) , pages =

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-15T14:58:40.240467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:58:40.240467Z digest=sha256:ea66e476249711153679ea85d60d6be79ff300fcc70644110a439de54b4a4933

Observation 5c26c898-a903-4f5e-ba48-c3b34c9fe424 · outbound

This paper cites Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages =.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages =

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-15T14:58:40.244955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:58:40.244955Z digest=sha256:8fb840f52c5c2447d4d1e110ac6f4e97da84cac03efcbd7c478a018490926440

Observation 9d714cd8-41ec-4053-a260-51ae77bdf287 · outbound

This paper cites Beyond Accuracy: Behavioral Testing of.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Beyond Accuracy: Behavioral Testing of

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-15T14:58:40.252942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:58:40.252942Z digest=sha256:941167a19526cdd7d29c6dbaac383ec2f2f885bba154f3014cc1ec7e7fbeaf58

Observation fb4eb4cf-2d3e-442b-8b8b-53d72f4bdd99 · outbound

This paper cites The Hitchhiker.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks The Hitchhiker

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-15T14:58:40.320183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:58:40.320183Z digest=sha256:7d1ce2f5b8b59900f032b52647d8e24ce8522fbcc8d59e95bc4eb093043b43c7

Observation 64d0d676-a841-4f1c-a9a0-97faae517a57 · outbound

This paper cites The Spanish Journal of Psychology , volume=.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks The Spanish Journal of Psychology , volume=

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:41.487546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T14:58:40.399395Z digest=sha256:ab1e37ed26e12ec03ee94142e41091696e3f77b44f53c8eb08ff8b1a7a2ab22e

Observation 50d02924-2b61-4d7c-bbc1-40e0c6886ba2 · outbound

This paper cites tinyBenchmarks: evaluating LLMs with fewer examples.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks tinyBenchmarks: evaluating LLMs with fewer examples

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-15T14:58:40.405231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:58:40.405231Z digest=sha256:8b0ec0fd7d3eb4ddabf096da55ec9c2f2b35e03307defe3362958acd6118b77f

Observation 7c5529b9-94e5-4a3a-bb49-163c8f2f68b5 · outbound

This paper cites International Conference on Learning Representations , volume=.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks International Conference on Learning Representations , volume=

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:41.468974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T14:58:40.410979Z digest=sha256:8cb830afd4a9da8bb669a8ccaff5128f7e41d04d9456ba1c42da4c176c586e6e

Pith citing papers

No inbound Pith citation observations are available.