Pith. sign in

Paper Citation Record · LEDGER

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks

As of 20 August 2026, this Paper Citation Record lists 66 of 66 outbound references and 0 inbound Pith citation observations for arXiv:2608.02966.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.02966 v1

Coverage vector

measured 66 of 66 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T14:58:40.410979Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

66 of 66 outbound references displayed

  • verified exact1
  • verified fuzzy39
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 98a919b9-dec6-439a-b202-39dd39870100 · outbound

This paper cites Darrell , title =.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Darrell , title =

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T14:58:39.046108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:58:39.046108Z digest=sha256:2ff65f115ff327bef948d1125aa962cc62d479240ce7c3f9e65fe4e66bd9519c

Observation 49f83c5c-c3a7-4d2e-9b72-a61a452c5a80 · outbound

This paper cites Psychometrika , year =.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Psychometrika , year =

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:44.261192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T14:58:39.119070Z digest=sha256:d3894ba222ea10a62b3d3fae1bff588cb0f7fc6524c30f0c46348196984c18c4

Observation e363868c-6342-4e44-b9f9-90da01d52558 · outbound

This paper cites , title =.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks , title =

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:44.105716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T14:58:39.207881Z digest=sha256:7cdc26044b557107b55058b405bf4db27802ae0fcdbe121c52412807620d8a91

Observation 1cfc82f6-00a7-46a8-9d4a-6880c92a96a5 · outbound

This paper cites Measurement: Interdisciplinary Research and Perspectives , year =.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Measurement: Interdisciplinary Research and Perspectives , year =

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:44.061099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T14:58:39.212976Z digest=sha256:5454c83a365b78e93806265cd37011fd6076280387fdf8350c0bd92a26e09293

Observation 42f866ec-59ac-42be-b44a-cba553a23282 · outbound

This paper cites , title =.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks , title =

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:44.042634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T14:58:39.217844Z digest=sha256:ed84fee4f3571f42b87f5a66d2b7055d0ad6d452680096ba6a83996d250ad929

Observation 73ab8a78-50c6-439e-9029-d8baf7bd3019 · outbound

This paper cites Findings of the Association for Computational Linguistics: EMNLP 2023 , year =.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Findings of the Association for Computational Linguistics: EMNLP 2023 , year =

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:44.024830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T14:58:39.223126Z digest=sha256:51d8560e27cb0e871c80a48368e18dfbb462ac6fe39334f632b48413711b52a9

Observation 8f71685b-bdd6-4a45-85a3-68c9551f999a · outbound

This paper cites Findings of the Association for Computational Linguistics: NAACL 2024 , pages=.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Findings of the Association for Computational Linguistics: NAACL 2024 , pages=

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T14:58:39.227968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:58:39.227968Z digest=sha256:aea9511349c7bda6fba54a682062148f6334efb4fa811b6e9c927cc6a5ab5461

Observation 4b1556e6-2d05-4252-bf2e-4dca6929eed9 · outbound

This paper cites Journal of Educational Measurement , volume=.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Journal of Educational Measurement , volume=

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:43.949220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T14:58:39.232861Z digest=sha256:13b4d2686cdbce0a93fb1a156e98f5c1ded931d19ef2c2f4986b4d26c0a85a3a

Observation 252a3a37-a1ed-4fc2-ada8-55cbaa1a1462 · outbound

This paper cites From Loops to Oops: Fallback Behaviors of Language Models Under Uncertainty.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks From Loops to Oops: Fallback Behaviors of Language Models Under Uncertainty

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T14:58:39.237142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:58:39.237142Z digest=sha256:b1c2a037f6a70a2274ca30a2e429c86376a407f2fbb9931061b533af651b9443

Observation bcc1d5fa-ab3e-4664-99a6-8b7354550dd0 · outbound

This paper cites and Linn, Robert L.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks and Linn, Robert L

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:43.825304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T14:58:39.243238Z digest=sha256:9cc2c224b54b30477744c7a89d4accc2ea8d947b0890f2756d86e74e3ae23bfd

Observation 466cfa14-f956-4623-8254-1127a8100af2 · outbound

This paper cites Applied Psychological Measurement , volume=.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Applied Psychological Measurement , volume=

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:43.806408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T14:58:39.249743Z digest=sha256:f9c0f6fba3390e51ab7aa6bf220f6cd5c9fe382bcbacb050ff0fd9d8aa1abe26

Observation 09a27488-dcbf-43b2-b564-aad4156788fa · outbound

This paper cites Marketing science , volume=.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Marketing science , volume=

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:43.724061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T14:58:39.255201Z digest=sha256:40bc2f95e4f2a3215aa83ccaeeeb6e6de47afa8d2e251e8b2b314f9452366f57

Observation e80b7b7f-c848-4957-a7e2-d9148bee9788 · outbound

This paper cites Transportation , volume=.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Transportation , volume=

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:43.659772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T14:58:39.259897Z digest=sha256:4e13783ab9b48ba0a89e0efa861ca9681d027b8a3dbdb137c43b1404a60a20f9

Observation 1887a2ca-bdfd-4c54-96e9-6fa6225cbb8a · outbound

This paper cites , title =.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks , title =

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T14:58:39.265177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:58:39.265177Z digest=sha256:7cd6f1dc0e724f14aef5190103a2e695374ed262fb26f7634768b875601a5967

Observation 44b107b5-2ca9-427c-b3f0-5a49e1db1cad · outbound

This paper cites Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics , year =.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics , year =

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:43.606732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T14:58:39.269644Z digest=sha256:9df3bbdd722afba43be9184c28a89eb9a6e4496c2ea369ff6e22c7e16f50dd60

Observation 275e80f0-bae8-4024-802b-a7d6a3c754d9 · outbound

This paper cites Statistical Theories of Mental Test Scores , editor =.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Statistical Theories of Mental Test Scores , editor =

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T14:58:39.274850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:58:39.274850Z digest=sha256:34fc7d9413cb0cc6378f4019b98815ad2101f19529d11329059de69ca0181d12

Observation eae2c5e3-1856-432e-9edc-2209300be1a2 · outbound

This paper cites Measurement: Interdisciplinary Research and Perspectives , volume=.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Measurement: Interdisciplinary Research and Perspectives , volume=

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:43.504572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T14:58:39.279399Z digest=sha256:639f2ea9538e7e6dde4f27aa6f0e2bf588e9802feefa07adf4797d545a207ed1

Observation ba14c083-016c-4271-872a-38a1c37b0263 · outbound

This paper cites On the Unidentifiability of the Fixed-Effects.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks On the Unidentifiability of the Fixed-Effects

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:43.485264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T14:58:39.283837Z digest=sha256:dbd088e6eea451b68f9792211952ebf9732c964425917375bffcc47a119f50af

Observation cc1e30ac-1a4f-4b26-a7c1-010f4fd28b8b · outbound

This paper cites and Lord, Frederic M.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks and Lord, Frederic M

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:43.433906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T14:58:39.342230Z digest=sha256:1b1b82c8c42c3508e615dedd2c85eafb9cffb00f75ff2cb47c7b1df381745505

Observation e629055f-9993-48d0-8623-aef14d46fb21 · outbound

This paper cites Philip and Skene, Allan M.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Philip and Skene, Allan M

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:43.308092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T14:58:39.442986Z digest=sha256:2ab03ee62a60cdddb7b32e2e453fdcaeb91b4815044731dc2d01cf382b61451a

Observation 16490a9c-ed94-449d-871e-c36982697ff3 · outbound

This paper cites Advances in Neural Information Processing Systems 23 , year =.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Advances in Neural Information Processing Systems 23 , year =

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:43.290454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T14:58:39.491224Z digest=sha256:e1e65fdfe99f7782792c7f7feefbdbe4bf1b9f830f766492071ac10360c3b6c9

Observation 31bdd3df-ac03-4c39-8953-b0efa414effd · outbound

This paper cites Advances in Neural Information Processing Systems 37, Datasets and Benchmarks Track , year =.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Advances in Neural Information Processing Systems 37, Datasets and Benchmarks Track , year =

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:43.274140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T14:58:39.496361Z digest=sha256:0f3342b0b1a217f0333e486a9293f11dd47fdf396742e7f8bced98c3afb90b56

Observation 92b33539-a865-4026-8da0-581a48bf1772 · outbound

This paper cites , title =.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks , title =

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:43.156418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T14:58:39.501364Z digest=sha256:de6b7870b07e36f53bd22a96e6961a9e85d502114cb05f5b99421cb1754c3207

Observation 135fd128-a544-4b68-b6bf-2097d387b460 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T14:58:39.506630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:58:39.506630Z digest=sha256:d1ab869e8431ff1237360669711bf1778933b3fa1ed204954fca408cfe46c73d

Observation bd1913b7-16ae-4c7d-8400-bdea67ff9f2e · outbound

This paper cites Findings of the Association for Computational Linguistics: NAACL 2024 , year =.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Findings of the Association for Computational Linguistics: NAACL 2024 , year =

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:43.137872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T14:58:39.513083Z digest=sha256:965250448525c31a30e398049a9a2d9dbbaa0cc8bcceb55e65b6ba0d758a5575

Observation eb5a045f-c8d8-468f-96b1-7927be04d7cc · outbound

This paper cites Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics , year =.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics , year =

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:43.045165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T14:58:39.518553Z digest=sha256:f66b4a5988d19fff75d3d080620fe23ee5aa535f18bd0671c4a6cc6abb9afcc4

Observation 8449a83d-8cba-4f63-857a-b0b2aceab770 · outbound

This paper cites IEEE/ACM Transactions on Audio, Speech, and Language Processing , volume=.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks IEEE/ACM Transactions on Audio, Speech, and Language Processing , volume=

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T14:58:39.523190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:58:39.523190Z digest=sha256:395e00c3fdf790abc4f12f145042c2fd1694dc6dc4abb90c84e1b65181293759

Observation f74b1681-a0b0-4240-ba1f-1e5932b534c3 · outbound

This paper cites Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing , year =.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing , year =

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:42.931463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T14:58:39.527766Z digest=sha256:c6debfb1217d562783f24e1416db75bba8b800c6e756ec282f8b6f5731339cc0

Observation d66e6103-db5e-4788-830c-b4514b182854 · outbound

This paper cites Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics , year =.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics , year =

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:42.913179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T14:58:39.532191Z digest=sha256:25f898ff7624f76152491a8a574615de8509c952a12822d76066f1fb362a95d1

Observation 5730f55f-d330-4240-80b1-cb163c39e95c · outbound

This paper cites Applied Sciences , year =.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Applied Sciences , year =

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:42.772270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T14:58:39.537816Z digest=sha256:41588ff68e589f7b1947c3cf8444091a02e7a82d009fda0f0f4f7ba06d15533d

Observation 743b2775-6ef4-40ad-a896-5649c658398e · outbound

This paper cites Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing , year =.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing , year =

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:42.704715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T14:58:39.622501Z digest=sha256:8ea28322c21fa708be2db8576280d06a9058ecbb755bb4eaccc7690b5bba77be

Observation 3f39efbf-f6f1-4052-9402-79645ac5139f · outbound

This paper cites an unresolved cited work.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T14:58:39.695919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:58:39.695919Z digest=sha256:b496dfb9a668eb07855a2c080985b8a08df5bf72a651d30b006c2b9bd0a175cf

Observation 823da1a1-170c-4777-a176-ccd8c8e3120c · outbound

This paper cites Proceedings of the 34th AAAI Conference on Artificial Intelligence , year =.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Proceedings of the 34th AAAI Conference on Artificial Intelligence , year =

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:42.676083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T14:58:39.753538Z digest=sha256:1b847e1c969fb718cd3f2728ba52045a70bec2413c1305a114cafbdd7552c955

Observation bf456cb8-86bf-45c3-bc28-2a0a9e346376 · outbound

This paper cites International Conference on Learning Representations , year =.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks International Conference on Learning Representations , year =

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:42.659007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T14:58:39.758422Z digest=sha256:60524573798ca7edc9ce02bb885328498b8def8d9682516b5ae455686d6a6a92

Observation 391415ce-7682-415f-9bf0-0c6e869888ab · outbound

This paper cites International Conference on Learning Representations , year =.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks International Conference on Learning Representations , year =

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T14:58:39.763766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:58:39.763766Z digest=sha256:889a19edc4400938d96388346e2ec83c66e1afa4c7b5cd64e2e22fae03aaf106

Observation 8b307df6-310e-49ec-a3a6-160c9715e38d · outbound

This paper cites Proceedings of the AAAI Conference on Artificial Intelligence , volume=.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Proceedings of the AAAI Conference on Artificial Intelligence , volume=

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:42.537237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T14:58:39.768489Z digest=sha256:668cd21dfd5c23b79b28412dabb7b472767d477c88d7172ae04a85df56edafab

Observation fd090313-a6fe-4f36-83e3-077afbdb9f52 · outbound

This paper cites Probabilistic models for some intelligence and attainment tests.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Probabilistic models for some intelligence and attainment tests

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T14:58:39.773095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:58:39.773095Z digest=sha256:56e68880340a45a87c0ffc307467d50ac975d4d37477fd1849f8673734dfec32

Observation ac383f0e-602c-4944-a6ea-0f2e6cc0f6d1 · outbound

This paper cites 1968 , publisher=.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks 1968 , publisher=

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:42.427621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T14:58:39.778953Z digest=sha256:ce76416a00bc89a8e4d17924531f45eb80955b5fcc43b1467a7030c04db57abc

Observation 92de8232-a970-4f2f-9670-c82f5568a34e · outbound

This paper cites Proceedings of the 12th International Conference on Educational Data Mining , year =.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Proceedings of the 12th International Conference on Educational Data Mining , year =

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:42.406286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T14:58:39.783899Z digest=sha256:10c3033cedf85c921e42960ef519a3e28862c5b0f64590d32e3e92d185d5a0bd

Observation 891d7191-25f3-40f7-85c3-bbe2997d65e4 · outbound

This paper cites an unresolved cited work.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-15T14:58:42.274400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T14:58:39.788953Z digest=sha256:5e2213ee0b4b701ff4a9d11fc79c30351ecefa4e751f059642580b9f1b7df027

Observation 19025d95-827b-4ff1-a962-8d98ff1be4cd · outbound

This paper cites Annual Meeting of the Association for Computational Linguistics , year=.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Annual Meeting of the Association for Computational Linguistics , year=

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:42.212821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T14:58:39.794229Z digest=sha256:aaba261ca0987280efcab4ed30041345bf123b6c010d30295678733e8968b647

Observation 7b740791-a193-4a6e-b833-e1e712d36b07 · outbound

This paper cites 2025 , eprint=.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks 2025 , eprint=

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:42.145916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T14:58:39.799355Z digest=sha256:338bcb23e2b3df18e12c808423f74cf163152b907564b8964773cfdf2c5fe0e7

Observation 9be63788-8820-4dc1-bdb8-e98f3d805e94 · outbound

This paper cites arXiv preprint arXiv:2606.15643 , year=.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks arXiv preprint arXiv:2606.15643 , year=

Reference 43

Resolution
verified exact
raw_fallback, observed 2026-08-15T14:58:41.319337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T14:58:39.803809Z digest=sha256:1354f7ed0f10e3736e09a2fc7ae4836b3f23d6ed5cf559306fc52ae08000c676

Observation dcf39de4-b277-4ab4-af22-face4a94ea6a · outbound

This paper cites Adaptive Testing for LLM Evaluation: A Psychometric Alternative to Static Benchmarks.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Adaptive Testing for LLM Evaluation: A Psychometric Alternative to Static Benchmarks

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T14:58:39.837621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:58:39.837621Z digest=sha256:af0c9dbe0d4f4868584cb7d1508b96786f0f2ae8cce72cef115215d14aa3e3c3

Observation 3914437c-a570-4816-99bd-f448b11d03fe · outbound

This paper cites Proceedings of the AAAI Conference on Artificial Intelligence , volume=.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Proceedings of the AAAI Conference on Artificial Intelligence , volume=

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:42.097855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T14:58:39.942519Z digest=sha256:0cb6c6ca4d08de9cbb2a96900ac7e0d7cd1465d3ee39e6eb8b9fbd56d4f72096

Observation fc5cc222-e378-40ee-aa02-12ba49c7b4f9 · outbound

This paper cites Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages=.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages=

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:41.994121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T14:58:40.025703Z digest=sha256:d4fca8aa8baced2064d0edd4dea6e1a1644c5be1675b2d85593e51d2e8f46fec

Observation 48acfe07-7d24-4984-8462-7fa5334b4551 · outbound

This paper cites Language Models (Mostly) Know What They Know.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Language Models (Mostly) Know What They Know

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T14:58:40.046423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:58:40.046423Z digest=sha256:f7195f02af65e1b1ee4601955418aed9e90c11ff4c7ef2b49c71d99a5f70ba09

Observation 2a82a322-3441-4796-b949-4df879b38398 · outbound

This paper cites Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages=.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages=

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T14:58:40.052098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:58:40.052098Z digest=sha256:5fab0fa16fffd60889c237124a311c1c1c6aa2a07de24417a1acdc1e495c716c

Observation 8965e48b-1db1-41f6-b05f-85b037253ebd · outbound

This paper cites arXiv preprint arXiv:2509.10625 , year=.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks arXiv preprint arXiv:2509.10625 , year=

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T14:58:40.056761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:58:40.056761Z digest=sha256:da3e0d24047c9eb53d01a5d2f8f39eed01d4102ec1fd5b1f0ae90431ac91482d

Observation e320e6a1-cf39-4465-9788-2e469bfe716d · outbound

This paper cites Findings of the Association for Computational Linguistics: ACL 2025 , pages=.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Findings of the Association for Computational Linguistics: ACL 2025 , pages=

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:41.865410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T14:58:40.061858Z digest=sha256:a46c0333a51b9509564b1ff99c2837d67ec1ba9360e80e5e6a89817c6272980c

Observation 9cfeaf8c-6499-47f7-a276-c9b42cfd7208 · outbound

This paper cites Auditing LLM Benchmarks with Item Response Theory.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Auditing LLM Benchmarks with Item Response Theory

Reference 51

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T14:58:40.949418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T14:58:40.067379Z digest=sha256:da92dcf522471c337109e1f57c3bbe4d281c174fa4600dc24f8ec341c02c3d57

Observation 3009f9d1-f208-4363-8315-770cc243e049 · outbound

This paper cites Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages=.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages=

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:41.846317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T14:58:40.072684Z digest=sha256:0cfc6d290cea4e485ede32c3a2b346db52c9f1e198c669dda7b5c25d70b01ddc

Observation c3269eb1-3b2f-4186-80ae-325b83e779cd · outbound

This paper cites Quantifying Consistency in LLM Logical Reasoning via Structural Uncertainty.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Quantifying Consistency in LLM Logical Reasoning via Structural Uncertainty

Reference 53

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T14:58:40.759762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T14:58:40.078037Z digest=sha256:474155d6ad9cab4b1512603449407ab8abf3308f68b140416372fbf5c106bb26

Observation 86c54e22-70da-4285-8f37-224c27091867 · outbound

This paper cites Proceedings of the AAAI Conference on Artificial Intelligence , volume=.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Proceedings of the AAAI Conference on Artificial Intelligence , volume=

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:41.757465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T14:58:40.083482Z digest=sha256:8842ff41d94a768ddd549a0012837e00f6dd7b33585de7706255989b7b8b4981

Observation 2b4193f8-9359-4f7d-a2d4-f298689cad4d · outbound

This paper cites International Conference on Learning Representations , volume=.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks International Conference on Learning Representations , volume=

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:41.652754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T14:58:40.170145Z digest=sha256:cdbacaf0b93324bb7f558746f11ea24e067e66ffa7c786d85fc7b867a86965f9

Observation c9c76461-eb64-4550-8223-1c2edc8c8d8a · outbound

This paper cites an unresolved cited work.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-15T14:58:41.573767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T14:58:40.220073Z digest=sha256:05293d61cdaf31617e29059678bc274ba8a477f7a7541ae374485eb00c3a7237

Observation 868a54fa-74fc-47cf-9642-0aa782cabb4d · outbound

This paper cites Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T14:58:40.225030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:58:40.225030Z digest=sha256:80506befda0a33866ef946f2c2b7fabd38b3098429295b12d2ceb7ecca273d94

Observation 7b8055b6-ef0c-49f8-bf80-5f5a3343f606 · outbound

This paper cites The Masked Advantage: Uncovering Local-Language Access to Cultural Knowledge in LLMs.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks The Masked Advantage: Uncovering Local-Language Access to Cultural Knowledge in LLMs

Reference 58

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T14:58:40.596362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T14:58:40.229901Z digest=sha256:99811d8a606fc657f3ce1c7b46d27c9b283ccca535f623d91fd0e373839a8b60

Observation c31ed83a-6dba-456f-9667-c33999fe6cce · outbound

This paper cites Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing , pages =.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing , pages =

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-15T14:58:40.235979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:58:40.235979Z digest=sha256:f35fb96caf0936fe3b72d47a7f0b2d7a45934c684c0c09821c06c6a47005416a

Observation fffebd34-82ae-4557-a7e8-05c872ba85c5 · outbound

This paper cites Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) , pages =.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) , pages =

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-15T14:58:40.240467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:58:40.240467Z digest=sha256:ea66e476249711153679ea85d60d6be79ff300fcc70644110a439de54b4a4933

Observation 5c26c898-a903-4f5e-ba48-c3b34c9fe424 · outbound

This paper cites Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages =.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages =

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-15T14:58:40.244955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:58:40.244955Z digest=sha256:8fb840f52c5c2447d4d1e110ac6f4e97da84cac03efcbd7c478a018490926440

Observation 9d714cd8-41ec-4053-a260-51ae77bdf287 · outbound

This paper cites Beyond Accuracy: Behavioral Testing of.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks Beyond Accuracy: Behavioral Testing of

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-15T14:58:40.252942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:58:40.252942Z digest=sha256:941167a19526cdd7d29c6dbaac383ec2f2f885bba154f3014cc1ec7e7fbeaf58

Observation fb4eb4cf-2d3e-442b-8b8b-53d72f4bdd99 · outbound

This paper cites The Hitchhiker.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks The Hitchhiker

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-15T14:58:40.320183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:58:40.320183Z digest=sha256:7d1ce2f5b8b59900f032b52647d8e24ce8522fbcc8d59e95bc4eb093043b43c7

Observation 64d0d676-a841-4f1c-a9a0-97faae517a57 · outbound

This paper cites The Spanish Journal of Psychology , volume=.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks The Spanish Journal of Psychology , volume=

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:41.487546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T14:58:40.399395Z digest=sha256:e27c7542012be0fa7aea0a64b7dcd5f8e1f6d67a43c58a895343b1ff8852d8be

Observation 50d02924-2b61-4d7c-bbc1-40e0c6886ba2 · outbound

This paper cites tinyBenchmarks: evaluating LLMs with fewer examples.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks tinyBenchmarks: evaluating LLMs with fewer examples

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-15T14:58:40.405231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:58:40.405231Z digest=sha256:8b0ec0fd7d3eb4ddabf096da55ec9c2f2b35e03307defe3362958acd6118b77f

Observation 7c5529b9-94e5-4a3a-bb49-163c8f2f68b5 · outbound

This paper cites International Conference on Learning Representations , volume=.

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks International Conference on Learning Representations , volume=

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:41.468974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T14:58:40.410979Z digest=sha256:1b271b1b94b543591d0c22f08ea6de21ace372d71b53f845c001c65d1c36243b

Pith citing papers

No inbound Pith citation observations are available.