Pith. sign in

Paper Citation Record · LEDGER

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations)

As of 7 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 0 inbound Pith citation observations for arXiv:2608.06202.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.06202 v1

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:44:36.904062Z

measured 65 of 65 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

65 of 65 outbound references displayed

  • verified exact3
  • verified fuzzy16
  • unresolved44
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 26221a13-0976-4b7c-8ff5-babd1a283942 · outbound

This paper cites 2026 , note=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) 2026 , note=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:31.564761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:31.564761Z digest=sha256:13d6e3b1578c0b2551ea70590b6b1b372892345bed5e6c9c0babe841608ae7d7

Observation 622a2c7a-90e9-427b-bd74-9178d043a17c · outbound

This paper cites Findings of the.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Findings of the

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:07.990676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:44:31.623981Z digest=sha256:cb8667761792fb27bcbee96865b221fa2541dca971bdfa6c333b3b864792cf9c

Observation 462ca983-a498-4f75-b75c-16294ff53ff8 · outbound

This paper cites Proceedings of the 62nd.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Proceedings of the 62nd

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:07.837480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:44:31.700906Z digest=sha256:7313d03fdfc6e278dfe1e67b8172bcf2c8b0da8b21ab7f90d2220517d3973dd9

Observation dee58e1d-f0a9-4e4e-8c85-9406f8f71f35 · outbound

This paper cites Proceedings of the 10th SIGHAN Workshop on Chinese Language Processing (SIGHAN-10) , pages=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Proceedings of the 10th SIGHAN Workshop on Chinese Language Processing (SIGHAN-10) , pages=

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:07.677459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:44:31.813047Z digest=sha256:39bb1c927c23e293d3d02b59387f23e7143896091363f6fc104b8c44a64165d2

Observation 1d6aa5bc-2a07-47f9-8374-28c6e9e277b0 · outbound

This paper cites JustLogic: A Comprehensive Benchmark for Evaluating Deductive Reasoning in Large Language Models.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) JustLogic: A Comprehensive Benchmark for Evaluating Deductive Reasoning in Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:31.930407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:31.930407Z digest=sha256:14dc67ced110ff7d5357ab6cceebd2238559804fbab1a4cf783bc142372ea1a0

Observation 71b7280e-9268-4942-ae0b-d9bd554abc3b · outbound

This paper cites Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:07.547709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:44:32.026529Z digest=sha256:244774090ee1781d02ad891e051271453484968f21f27a7d1438a4c6f72ce229

Observation 60b402ce-04b2-4459-b5f6-e38590f0a660 · outbound

This paper cites FormalMATH: Benchmarking Formal Mathematical Reasoning of Large Language Models.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) FormalMATH: Benchmarking Formal Mathematical Reasoning of Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:32.091734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:32.091734Z digest=sha256:4d5b9ecb2debca1602585d83799320683c2acedb339a8114e1f84e9cb3020073

Observation f4f4de7f-479c-4700-a2e9-6569fef31de5 · outbound

This paper cites European Semantic Web Conference , pages=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) European Semantic Web Conference , pages=

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:07.332701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:44:32.198014Z digest=sha256:45c28173991df7ddef57121d33c164254abcc4e124bf427854fb52adc8145658

Observation 8c2e8861-ca7c-4252-972b-efc4ed7283d9 · outbound

This paper cites Transactions of the Association for Computational Linguistics , volume=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Transactions of the Association for Computational Linguistics , volume=

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:32.295139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:32.295139Z digest=sha256:7c29db8848052cf9c9ea5799bb491f025583a962c718e51e42f247154af865a1

Observation 94a50989-6931-4d0e-a9fa-d8b75f58de0a · outbound

This paper cites Proceedings of the 60th annual meeting of the association for computational linguistics (volume 1: long papers) , pages=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Proceedings of the 60th annual meeting of the association for computational linguistics (volume 1: long papers) , pages=

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:32.383467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:32.383467Z digest=sha256:91efb0b545d873ee94ffd8e65f5bd7dba5e322a5f512c1407f0840393407d470

Observation 68e4a836-098f-4d8f-b437-4c400f68a726 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Evaluating Large Language Models Trained on Code

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:32.454833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:32.454833Z digest=sha256:ae6c28bdb7befb15ff151b4c671778e94ebc71084ec847fb2eae222049952dcd

Observation 70309c71-8603-499e-99da-07f30a0a32c8 · outbound

This paper cites Proceedings of the 2021 conference of the North American chapter of the association for computational linguistics: human language technologies , pages=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Proceedings of the 2021 conference of the North American chapter of the association for computational linguistics: human language technologies , pages=

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:06.985875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:44:32.497238Z digest=sha256:825af32a1f3a14e8ddd896d7634b95f5107e778a86552d0437d8e9d374157f21

Observation cfb01546-9fc0-4fbd-97f4-7857da88a24e · outbound

This paper cites Findings of the association for computational linguistics: EMNLP 2020 , pages=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Findings of the association for computational linguistics: EMNLP 2020 , pages=

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:32.549221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:32.549221Z digest=sha256:110a9a3bfe33ccc297923dedfc456a48b233a6a0cdd90f7891e92fab5350bd70

Observation a7228e6f-6477-4bf1-98c1-836015154b6f · outbound

This paper cites Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing , pages=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing , pages=

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:06.595044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:44:32.617451Z digest=sha256:25de77925a5247b48b875e5dc12341896a2f5627c4923a5a67c96cfbee6c33e3

Observation 2d72a67a-258c-4fa6-8534-0578564f6b3c · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:06.360053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:44:32.703087Z digest=sha256:0e9cd92af05d06995957ab9272547ffe8eb2d8b60425be1fa985f8e6cb82e13c

Observation e8d60a6d-d40b-46f5-8eb6-d5ffcf91416b · outbound

This paper cites C row S -Pairs: A Challenge Dataset for Measuring Social Biases in Masked Language Models.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) C row S -Pairs: A Challenge Dataset for Measuring Social Biases in Masked Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:32.795670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:32.795670Z digest=sha256:6e6ab5b1912d370291d1208748cccda5f7d0b90d2c6c9e1bfdb7d1539f33a123

Observation b47d6d3e-4761-4dc4-a5f8-a6b66cf1012b · outbound

This paper cites Transactions on Machine Learning Research , doi =.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Transactions on Machine Learning Research , doi =

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:06.241243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:44:32.882722Z digest=sha256:85d80cf2ee6229c916626be94c38736ec34e4c47ebc72bb18198d393f1285493

Observation 8ef6a630-248e-4bea-aecf-078d979fbb64 · outbound

This paper cites ACM transactions on intelligent systems and technology , volume=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) ACM transactions on intelligent systems and technology , volume=

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:32.978703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:32.978703Z digest=sha256:01ae38fd3818971cfbd9466a5e13640f66d2ac8f79261ac20ca4f957d002c05e

Observation 7776f9c7-d798-446f-a516-ba187f2bbd87 · outbound

This paper cites Transactions on machine learning research , year=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Transactions on machine learning research , year=

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:33.102903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:33.102903Z digest=sha256:c3e65cf825df62e15dd43ba065589b98e89706f4d487287afbe3cc2b7d6788af

Observation 83be2d93-8c2c-4244-b90c-003d12d088db · outbound

This paper cites Measuring Massive Multitask Language Understanding.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Measuring Massive Multitask Language Understanding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:33.168261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:33.168261Z digest=sha256:00edea0d9157a60c0ad578ea172a3dbd22084094e3559051cc4a2414e8d97211

Observation fe0054e4-ac38-4184-9bbe-44a120d787cb · outbound

This paper cites Findings of the Association for Computational Linguistics: NAACL 2024 , pages=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Findings of the Association for Computational Linguistics: NAACL 2024 , pages=

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:33.260360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:33.260360Z digest=sha256:5a98a5d77aec9de601b33ea0ade274a643116946d97b9d09293a309cc346e696

Observation 086b3cad-130d-45e1-8776-16f61cd85221 · outbound

This paper cites P ro SA : Assessing and Understanding the Prompt Sensitivity of LLM s.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) P ro SA : Assessing and Understanding the Prompt Sensitivity of LLM s

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:33.369494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:33.369494Z digest=sha256:24b8d6d85b482f13735b216844068ab664dc58b196c2f2fe0c96c32e52923735

Observation f92216a6-268c-4b31-bd92-8fc6e0a4b4bb · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Advances in Neural Information Processing Systems , volume=

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:33.461520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:33.461520Z digest=sha256:6fce861c5a0e0eb5e9642ff79bfd38f73312c6002268b07029a90dc68c319df6

Observation 1e24bf35-9c64-4a7b-aa78-09c43b5ead4d · outbound

This paper cites Proceedings of the 1st ACM workshop on large AI systems and models with privacy and safety analysis , pages=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Proceedings of the 1st ACM workshop on large AI systems and models with privacy and safety analysis , pages=

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:05.861608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:44:33.550095Z digest=sha256:b565f95d75b637672497220e8e70a8ea53776d44ef4a819f21c1ba6cfc0f00f3

Observation fe0a5809-6866-4f94-a79c-1da3448024a0 · outbound

This paper cites Understanding User Experience in Large Language Model Interactions.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Understanding User Experience in Large Language Model Interactions

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:33.653851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:33.653851Z digest=sha256:8edfd1fbf78504d502004f6490395c7d00c91d4a5602c6632e90f7028715f28d

Observation d46bb22a-b864-4e17-9d3f-89312334bcac · outbound

This paper cites Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:33.736893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:33.736893Z digest=sha256:4c660d00f0a57e6bbf15172e3aeb3a0235ee7c251f97425f9ef9ff6fccdd42b9

Observation a8b092d3-0c98-4656-96af-edda0a1640e5 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Advances in Neural Information Processing Systems , volume=

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:05.649164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:44:33.804536Z digest=sha256:a8a5575fd8fa3637ea06f0eb2ec95cca0bfb00d92e32e26f1b71f642c22adc2d

Observation 0c553b29-49b6-4375-aa24-bdf47c7ce632 · outbound

This paper cites Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:33.856791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:33.856791Z digest=sha256:7112b81427f24f5151b55e3edf88bda9fd4d65f2be0e53b573f3920d6a2e8d31

Observation da414d8c-eecd-4044-b747-d975df546e44 · outbound

This paper cites arXiv preprint arXiv:2509.19364 , year=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) arXiv preprint arXiv:2509.19364 , year=

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:33.930656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:33.930656Z digest=sha256:9f54ccbe8a18f8879fbcc7ee55cf87e99758e7d5dd6574c9fd87a2faaa183a1c

Observation 2e3f3f2b-b54d-4663-a092-ccc87aea3b80 · outbound

This paper cites an unresolved cited work.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:45:05.404673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:44:33.991547Z digest=sha256:671705ee13cd13c4e7756cb6aa99a0cd2fbbc46042ccd584e1863bd54c450fbb

Observation 9a9935d0-f7b6-46d0-acd9-2b2a9276d937 · outbound

This paper cites 2023 , isbn =.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) 2023 , isbn =

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:34.071960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:34.071960Z digest=sha256:606c56b67543bc536bb57e87a7552b5e3e2af17ac9a255710148ba8b5a646805

Observation 6161cf9b-f2c6-4a51-bb84-19e368a3ece4 · outbound

This paper cites On the Robustness of ChatGPT: An Adversarial and Out-of-distribution Perspective.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) On the Robustness of ChatGPT: An Adversarial and Out-of-distribution Perspective

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:34.189426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:34.189426Z digest=sha256:73f62f529be013414fbf337c9cfe2a3ab78d0e8a2e6e1868153799c13d83e97b

Observation 9dccb30e-fa97-4f3c-9bda-fba0e3b7e762 · outbound

This paper cites Ask Again, Then Fail: Large Language Models' Vacillations in Judgment.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Ask Again, Then Fail: Large Language Models' Vacillations in Judgment

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:34.260166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:34.260166Z digest=sha256:84caf1bd684161cb8df62dea64138f33e46f75e2578c75b95038c2327f4ad147

Observation c0b012f5-87a5-48b6-bcab-0b05fd694512 · outbound

This paper cites S iren ' s Song in the AI Ocean: A Survey on Hallucination in Large Language Models.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) S iren ' s Song in the AI Ocean: A Survey on Hallucination in Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:34.316712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:34.316712Z digest=sha256:ae7953f69c43a332e948da13d066a84cef53756b714317b5e912e09874e20137

Observation b346a3ab-f126-4e6a-bc5c-5e514e1e25f6 · outbound

This paper cites Proceedings of the 2021 ACM conference on fairness, accountability, and transparency , pages=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Proceedings of the 2021 ACM conference on fairness, accountability, and transparency , pages=

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:34.409924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:34.409924Z digest=sha256:99918e779593da9f07c69a30bd3ecde1b6eb96d88013d70a3113333ed3ed35b7

Observation 15f074f6-1ad5-452b-af55-9103cc4219d2 · outbound

This paper cites Should ChatGPT be Biased? Challenges and Risks of Bias in Large Language Models.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Should ChatGPT be Biased? Challenges and Risks of Bias in Large Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:34.503503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:34.503503Z digest=sha256:d572f0041356d77c22abb794c3d4c8d7b3e41092350ff4229d73735d53daee62

Observation 1313c07f-0d7c-4d4e-9d24-d81a7fb31217 · outbound

This paper cites LLM Spirals of Delusion: A Benchmarking Audit Study of AI Chatbot Interfaces.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) LLM Spirals of Delusion: A Benchmarking Audit Study of AI Chatbot Interfaces

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:34.565332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:34.565332Z digest=sha256:4bade638443cb893a5d4c1a2164c1cef5149965af37c10ba1c58baf808e303ee

Observation a458f07a-0ffa-4f31-865f-231be2c48040 · outbound

This paper cites Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society , number=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society , number=

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:05.151838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:44:34.623847Z digest=sha256:001b8400ce56eeb9976a027ddde56de51573b227a52b41042c452428d62d0a13

Observation 96133c27-3520-4a5a-9715-6591749ebe40 · outbound

This paper cites Practices for.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Practices for

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:04.934779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:44:34.700820Z digest=sha256:daa3b78150c84b625937e54fce5461254296e4f4b685962d10bc8af37ab3c59b

Observation 856bbd95-784c-4347-abaf-daea68706d34 · outbound

This paper cites an unresolved cited work.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Unresolved cited work

Reference 40

Resolution
parse uncertain
raw_fallback, observed 2026-08-07T12:45:04.655486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:44:34.816885Z digest=sha256:58bfb073c8b427939eba79ffff08d219a4a9fd30e163305a49379f7bef81951d

Observation f43044d6-d349-4364-bbd8-fb8fb3bedbbb · outbound

This paper cites Aligning AI With Shared Human Values.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Aligning AI With Shared Human Values

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:34.935769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:34.935769Z digest=sha256:03c55dbcea0cc5559e804f697fa69be92d5e176d4f0d2224e334ea3e9c4ce2a0

Observation d834dcae-40d0-4cbc-95ab-648cf68eb808 · outbound

This paper cites Toward an Evaluation Science for Generative AI Systems.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Toward an Evaluation Science for Generative AI Systems

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:35.042357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:35.042357Z digest=sha256:624078a7b23b5bb6940aac64b53c5dffb550ed2e5be0db60a07760f295fcfaec

Observation daa23e92-20a0-4295-a933-fcdb476df63d · outbound

This paper cites Science , volume=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Science , volume=

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:35.125805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:35.125805Z digest=sha256:8eac80cb7ce1e0771cb0a99d5ee544c031e5f2db7469f292d373a72ed3d42e37

Observation 0ca37549-c4c5-45e5-970c-15f40d1808ec · outbound

This paper cites an unresolved cited work.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Unresolved cited work

Reference 44

Resolution
verified exact
doi, observed 2026-08-07T12:45:02.931158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:44:35.232606Z digest=sha256:51271bcb5b1d34b33ae006380df773d3185743a446c87727d5ab3bb5ce9ce32d

Observation 7bcf96d0-68e2-4d3c-b056-ed6c588043bb · outbound

This paper cites The Vulnerability of Language Model Benchmarks: Do They Accurately Reflect True LLM Performance?.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) The Vulnerability of Language Model Benchmarks: Do They Accurately Reflect True LLM Performance?

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:35.307850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:35.307850Z digest=sha256:e0719974aebeddc3dd479592f33e3b8cd93042834c0d8dd71d911ca361eedeee

Observation 05b8b964-fc11-4886-a522-ddbc7f20ed3f · outbound

This paper cites Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:35.393890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:35.393890Z digest=sha256:6f89e813714235ff91e1cc7108bc17dfad5aab68858bf4fc1fe96ae75d2d6606

Observation 09ebe1db-820d-48ac-9d25-7ef4f3d17a2b · outbound

This paper cites AI and the Everything in the Whole Wide World Benchmark.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) AI and the Everything in the Whole Wide World Benchmark

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:35.451657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:35.451657Z digest=sha256:206c1391a618aeeca22276944fe5d785004aa4e92fd698cd8f88df5fda8eb94c

Observation c3f7131e-5e38-46df-9f2a-7ea4c19805f0 · outbound

This paper cites Proceedings of the 2025.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Proceedings of the 2025

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:35.498690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:35.498690Z digest=sha256:8ce38cfc45fa57e5f4818347a2924559f5cd3fa2dc06218ee03b80f89d7670d5

Observation 2808b874-9e30-4988-997d-b30941a981eb · outbound

This paper cites Rethinking.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Rethinking

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:35.593629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:35.593629Z digest=sha256:60344ebb909084a6c9e9e115fe86b90441bf7b84344108d85d9df449e3a6a104

Observation 90fdb2a7-d6a1-4bdf-a364-b147057f1c14 · outbound

This paper cites Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:35.707630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:35.707630Z digest=sha256:adb425aac86640ac0638ef4d8527111062d85c085ff0662ff736d8bd7b61259a

Observation 2ff3cc6b-80e4-4297-a894-ab5196b95f4d · outbound

This paper cites Outsider.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Outsider

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:35.776696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:35.776696Z digest=sha256:80e98fc4f4eaf9c6d27c48294301531fec02cdc2fbc5d1467e71ca061c904f1b

Observation 7a18bd0c-da49-4965-981a-f5c1d621b604 · outbound

This paper cites Sociotechnical Safety Evaluation of Generative AI Systems.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Sociotechnical Safety Evaluation of Generative AI Systems

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:35.836468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:35.836468Z digest=sha256:f7e9790ba13705e647ac4593a3bfe9e3291eb7e766f2b2a442b26e87c30377fe

Observation 3825143f-fd23-4b60-8b37-122987b4325c · outbound

This paper cites Lessons from the Trenches on Reproducible Evaluation of Language Models.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Lessons from the Trenches on Reproducible Evaluation of Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:35.911671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:35.911671Z digest=sha256:81ab8a0071db6c147099d04551f23be3d2027fe24382f23ede43e9f603f10cea

Observation 095f31e1-3024-4640-83ae-201405d6b4e2 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Advances in Neural Information Processing Systems , volume=

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:04.439156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:44:35.996295Z digest=sha256:a754393644ff4eb0419297ae3e797f3cd00dabadc3be21a29225e49d0d508bee

Observation 19b54a70-5013-4aab-b741-c30f055a0581 · outbound

This paper cites Proceedings of the 29th international conference on computational linguistics , pages=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Proceedings of the 29th international conference on computational linguistics , pages=

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:04.212267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:44:36.106520Z digest=sha256:d3d8c419dd86129fed053ead1c77056b499226e8360c697378011ce97154f560

Observation b7ea2215-28e0-437e-82e8-bb84696ebfb7 · outbound

This paper cites International Conference on Learning Representations , volume=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) International Conference on Learning Representations , volume=

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:36.193313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:36.193313Z digest=sha256:0815c372d13788f94beafc75ec2ef65cfec75a30475e550eec97dd9f1eaeedfd

Observation 4e9c0f34-3323-44d4-89bd-3ce922ffa349 · outbound

This paper cites Findings of the Association for Computational Linguistics: ACL 2024 , pages=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Findings of the Association for Computational Linguistics: ACL 2024 , pages=

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:36.278667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:36.278667Z digest=sha256:6ca10a7a933251813745714267b349d0ced4dd52ff16d005951fc2d4d585e404

Observation d7b4a3c1-3323-44c2-8021-4b673020cabd · outbound

This paper cites BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:36.345910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:36.345910Z digest=sha256:c02aa81d2a3fb268cb988c71f68fea4292c6f35b088dd44cf0f4d662e1a4066f

Observation 91789132-3150-4b20-8918-30ab698702b7 · outbound

This paper cites and Darrell, Trevor and Norouzi, Narges and Gonzalez, Joseph E.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) and Darrell, Trevor and Norouzi, Narges and Gonzalez, Joseph E

Reference 59

Resolution
verified exact
doi, observed 2026-08-07T12:45:02.704440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:44:36.435665Z digest=sha256:4c7ca5cebaafa77e29b9127cb040d3f48a56b014f82790039c03416cce992876

Observation 198004bf-5cd5-4244-8dbd-4afd50b28c74 · outbound

This paper cites an unresolved cited work.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Unresolved cited work

Reference 60

Resolution
verified exact
doi, observed 2026-08-07T12:44:37.153884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:44:36.563039Z digest=sha256:ea93367f9c3617949c7b60a49d56289ef0a7327c1cb6ae3c61edda9bb0e64515

Observation 26b99a95-b552-483f-91de-7b8bd9d9cc69 · outbound

This paper cites Sacred or.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Sacred or

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:03.955485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:44:36.672670Z digest=sha256:84ef0dd42e50af5cda6e6e27a01bd1732a432f9e7f52743292e826c71a15924d

Observation 5a6ad17b-8612-40c6-a6b1-a94cab07c62f · outbound

This paper cites AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:36.752949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:36.752949Z digest=sha256:963b48feda7900332de3f69e709f21e546bc25f836dd331733bcf04d4b81228b

Observation b1e89b97-aa78-4cdf-87ea-310d3df49387 · outbound

This paper cites and Metaxa, Dana.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) and Metaxa, Dana

Reference 63

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T12:45:03.305534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:44:36.796636Z digest=sha256:3c5f34089075ae9c472ee8cd6adb283a68a7500c00fc36bbb4837bf5ce19ad9e

Observation af29f3f7-4e3a-4bc3-9c8a-b23f72efdb87 · outbound

This paper cites Proceedings of the 2020 conference on fairness, accountability, and transparency , pages=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Proceedings of the 2020 conference on fairness, accountability, and transparency , pages=

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:36.850378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:36.850378Z digest=sha256:cc8c3b8cbfe9ab0df468e0bd82466fe2d8393a1b9f3332b53cbcdbe9a925054b

Observation 4b748991-47bb-4794-885b-a35805ba49ce · outbound

This paper cites In-House Evaluation Is Not Enough: Towards Robust Third-Party Flaw Disclosure for General-Purpose AI.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) In-House Evaluation Is Not Enough: Towards Robust Third-Party Flaw Disclosure for General-Purpose AI

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:36.904062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:36.904062Z digest=sha256:64cbecd8b38b843afbd8eb9542ec53c3c936670302a8e260719c7e20c3176639

Pith citing papers

No inbound Pith citation observations are available.