Pith. sign in

Paper Citation Record · LEDGER

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations)

As of 8 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 0 inbound Pith citation observations for arXiv:2608.06202.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.06202 v1

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:44:36.904062Z

measured 65 of 65 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

65 of 65 outbound references displayed

  • verified exact3
  • verified fuzzy16
  • unresolved44
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 26221a13-0976-4b7c-8ff5-babd1a283942 · outbound

This paper cites 2026 , note=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) 2026 , note=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:31.564761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:31.564761Z digest=sha256:934c5ab67cafa39cb557c8b63bdaed845db539c33c951aef52d88cf92f7dab35

Observation 622a2c7a-90e9-427b-bd74-9178d043a17c · outbound

This paper cites Findings of the.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Findings of the

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:07.990676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:44:31.623981Z digest=sha256:7b4573a0294a2571121f378a3e8741566aa71f62bb063b250f7ec0d91e48843e

Observation 462ca983-a498-4f75-b75c-16294ff53ff8 · outbound

This paper cites Proceedings of the 62nd.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Proceedings of the 62nd

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:07.837480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:44:31.700906Z digest=sha256:e8dc46567cbbe1a7754767d6ab30e79c40a6f79985eb96040d9151605954b82b

Observation dee58e1d-f0a9-4e4e-8c85-9406f8f71f35 · outbound

This paper cites Proceedings of the 10th SIGHAN Workshop on Chinese Language Processing (SIGHAN-10) , pages=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Proceedings of the 10th SIGHAN Workshop on Chinese Language Processing (SIGHAN-10) , pages=

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:07.677459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:44:31.813047Z digest=sha256:d3785ac5b060ba8fc7d1b7e5b9301594f4f42c922ff13a11449abb1a7446511c

Observation 1d6aa5bc-2a07-47f9-8374-28c6e9e277b0 · outbound

This paper cites JustLogic: A Comprehensive Benchmark for Evaluating Deductive Reasoning in Large Language Models.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) JustLogic: A Comprehensive Benchmark for Evaluating Deductive Reasoning in Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:31.930407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:31.930407Z digest=sha256:9736232b913a8a542c88a1edb6907879751515a293d93842cc5dc3262d1a41c1

Observation 71b7280e-9268-4942-ae0b-d9bd554abc3b · outbound

This paper cites Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:07.547709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:44:32.026529Z digest=sha256:dd4e9ed2cfe5b7cf105bd92000bd569d8bb2bf28812387cbdc61da06e32979df

Observation 60b402ce-04b2-4459-b5f6-e38590f0a660 · outbound

This paper cites FormalMATH: Benchmarking Formal Mathematical Reasoning of Large Language Models.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) FormalMATH: Benchmarking Formal Mathematical Reasoning of Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:32.091734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:32.091734Z digest=sha256:6ebdc3304395ba2ecfb9fa0c30be44c2a3cb33d24d87493e7cbbece43c34954a

Observation f4f4de7f-479c-4700-a2e9-6569fef31de5 · outbound

This paper cites European Semantic Web Conference , pages=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) European Semantic Web Conference , pages=

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:07.332701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:44:32.198014Z digest=sha256:edabff42c863f053a92510d1030f9429de8a74621810e50d12434a5be88cd9fe

Observation 8c2e8861-ca7c-4252-972b-efc4ed7283d9 · outbound

This paper cites Transactions of the Association for Computational Linguistics , volume=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Transactions of the Association for Computational Linguistics , volume=

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:32.295139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:32.295139Z digest=sha256:6db2f473dc0836f3acffc89bece41e2341a2c85478744b23af53df1be19f5c51

Observation 94a50989-6931-4d0e-a9fa-d8b75f58de0a · outbound

This paper cites Proceedings of the 60th annual meeting of the association for computational linguistics (volume 1: long papers) , pages=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Proceedings of the 60th annual meeting of the association for computational linguistics (volume 1: long papers) , pages=

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:32.383467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:32.383467Z digest=sha256:9bb038339af94d16734e5f11a82b2b8bda1db9bfb3eecc2f893cb63148319f46

Observation 68e4a836-098f-4d8f-b437-4c400f68a726 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Evaluating Large Language Models Trained on Code

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:32.454833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:32.454833Z digest=sha256:141deca9d7b138507cf478d2591756ffaa74a78ca19a61cb574d91cceac16e91

Observation 70309c71-8603-499e-99da-07f30a0a32c8 · outbound

This paper cites Proceedings of the 2021 conference of the North American chapter of the association for computational linguistics: human language technologies , pages=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Proceedings of the 2021 conference of the North American chapter of the association for computational linguistics: human language technologies , pages=

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:06.985875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:44:32.497238Z digest=sha256:946c23766f86859002921858f9fd3efa279eaeb55493c44f32fa73ac6e6ad90a

Observation cfb01546-9fc0-4fbd-97f4-7857da88a24e · outbound

This paper cites Findings of the association for computational linguistics: EMNLP 2020 , pages=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Findings of the association for computational linguistics: EMNLP 2020 , pages=

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:32.549221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:32.549221Z digest=sha256:c11e349ff1424ae8c16cc9604cae26b9c15d30c71bb6415009aea0c54184634c

Observation a7228e6f-6477-4bf1-98c1-836015154b6f · outbound

This paper cites Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing , pages=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing , pages=

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:06.595044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:44:32.617451Z digest=sha256:587f4de230a7a995932268f1714afa19c6f030d4207443b5314c3bd246fc3abf

Observation 2d72a67a-258c-4fa6-8534-0578564f6b3c · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:06.360053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:44:32.703087Z digest=sha256:2e042bc4c9025f3c517e153cd12009b1b44d3159a5a677969993ea2b5f7e673f

Observation e8d60a6d-d40b-46f5-8eb6-d5ffcf91416b · outbound

This paper cites C row S -Pairs: A Challenge Dataset for Measuring Social Biases in Masked Language Models.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) C row S -Pairs: A Challenge Dataset for Measuring Social Biases in Masked Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:32.795670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:32.795670Z digest=sha256:d98a7303330f119785796f4a2e51245c2e091d4fd05754a78d6704934b51666c

Observation b47d6d3e-4761-4dc4-a5f8-a6b66cf1012b · outbound

This paper cites Transactions on Machine Learning Research , doi =.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Transactions on Machine Learning Research , doi =

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:06.241243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:44:32.882722Z digest=sha256:59f9785ccea7a9a0fe8bfa7d69ccb83bf1f7c64f722f1e5698d24ac2e16aa1b8

Observation 8ef6a630-248e-4bea-aecf-078d979fbb64 · outbound

This paper cites ACM transactions on intelligent systems and technology , volume=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) ACM transactions on intelligent systems and technology , volume=

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:32.978703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:32.978703Z digest=sha256:ae2ab0285e679e0071c9ab82cfec67b35f196958946e5f8a15b727784977e7ad

Observation 7776f9c7-d798-446f-a516-ba187f2bbd87 · outbound

This paper cites Transactions on machine learning research , year=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Transactions on machine learning research , year=

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:33.102903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:33.102903Z digest=sha256:8ac7493d5299ed1fd40f6f20320540279ebcdcf4b46661f6bc633e680f2583d4

Observation 83be2d93-8c2c-4244-b90c-003d12d088db · outbound

This paper cites Measuring Massive Multitask Language Understanding.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Measuring Massive Multitask Language Understanding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:33.168261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:33.168261Z digest=sha256:c3f476ea697b9078473c43fe06ce042b8ef1e26cc19a1d163551dad5ed3bee25

Observation fe0054e4-ac38-4184-9bbe-44a120d787cb · outbound

This paper cites Findings of the Association for Computational Linguistics: NAACL 2024 , pages=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Findings of the Association for Computational Linguistics: NAACL 2024 , pages=

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:33.260360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:33.260360Z digest=sha256:77946229d0927cfc978c9d8c12761349ce007d012f89bf78d14e12ace80bd7e6

Observation 086b3cad-130d-45e1-8776-16f61cd85221 · outbound

This paper cites P ro SA : Assessing and Understanding the Prompt Sensitivity of LLM s.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) P ro SA : Assessing and Understanding the Prompt Sensitivity of LLM s

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:33.369494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:33.369494Z digest=sha256:51656c4bba148387a8e05d0c4b8876da02d2265749f560fbf624128e28314da5

Observation f92216a6-268c-4b31-bd92-8fc6e0a4b4bb · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Advances in Neural Information Processing Systems , volume=

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:33.461520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:33.461520Z digest=sha256:ee3b22d1194d233fae97c6b10cfcb126684b1cbd80ecc270f38fea7c0fde67ac

Observation 1e24bf35-9c64-4a7b-aa78-09c43b5ead4d · outbound

This paper cites Proceedings of the 1st ACM workshop on large AI systems and models with privacy and safety analysis , pages=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Proceedings of the 1st ACM workshop on large AI systems and models with privacy and safety analysis , pages=

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:05.861608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:44:33.550095Z digest=sha256:ecb112534d26c4f2c60671e431c332f276ad9b211b95e99f1fbcdf8d9114ea21

Observation fe0a5809-6866-4f94-a79c-1da3448024a0 · outbound

This paper cites Understanding User Experience in Large Language Model Interactions.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Understanding User Experience in Large Language Model Interactions

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:33.653851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:33.653851Z digest=sha256:46f1a8b418d77e12a75997ca5ba78b9802e0f2c9675cbfc168362aeb14faa96a

Observation d46bb22a-b864-4e17-9d3f-89312334bcac · outbound

This paper cites Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:33.736893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:33.736893Z digest=sha256:b956911a73cd505516197af91345980481081cc52cf0b24229d89155405c251d

Observation a8b092d3-0c98-4656-96af-edda0a1640e5 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Advances in Neural Information Processing Systems , volume=

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:05.649164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:44:33.804536Z digest=sha256:e1217ab4c000ddab514280afd6158ffaf04de18b0f9ad4047e4a8b9dc3684e60

Observation 0c553b29-49b6-4375-aa24-bdf47c7ce632 · outbound

This paper cites Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:33.856791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:33.856791Z digest=sha256:983b7eccad4558fcdc2f93ff8dead2e3fd9643f766b37ee9807ea464d9e1d88b

Observation da414d8c-eecd-4044-b747-d975df546e44 · outbound

This paper cites arXiv preprint arXiv:2509.19364 , year=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) arXiv preprint arXiv:2509.19364 , year=

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:33.930656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:33.930656Z digest=sha256:620e66e27072dcf0653e5fc5f23438683a519545e5590c1d1c4a9d98df0c76c3

Observation 2e3f3f2b-b54d-4663-a092-ccc87aea3b80 · outbound

This paper cites an unresolved cited work.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:45:05.404673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:44:33.991547Z digest=sha256:ecb3e190b4617f4fc6c50b387847257bbf99d0cd3dc265f749fd5de8b324a5a7

Observation 9a9935d0-f7b6-46d0-acd9-2b2a9276d937 · outbound

This paper cites 2023 , isbn =.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) 2023 , isbn =

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:34.071960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:34.071960Z digest=sha256:3884861aeb099d1e4620a927ef55732de93a42bff79b86fcc0de793d100d6304

Observation 6161cf9b-f2c6-4a51-bb84-19e368a3ece4 · outbound

This paper cites On the Robustness of ChatGPT: An Adversarial and Out-of-distribution Perspective.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) On the Robustness of ChatGPT: An Adversarial and Out-of-distribution Perspective

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:34.189426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:34.189426Z digest=sha256:69f0b47cf9a0b8ab9ee8bed26878a38256939bd3144e9cc95d4968b01136d6dc

Observation 9dccb30e-fa97-4f3c-9bda-fba0e3b7e762 · outbound

This paper cites Ask Again, Then Fail: Large Language Models' Vacillations in Judgment.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Ask Again, Then Fail: Large Language Models' Vacillations in Judgment

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:34.260166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:34.260166Z digest=sha256:4c4188920a189b7200ec45bcd5a48b4c6f06300585e3cde6f1125af8caa7f558

Observation c0b012f5-87a5-48b6-bcab-0b05fd694512 · outbound

This paper cites S iren ' s Song in the AI Ocean: A Survey on Hallucination in Large Language Models.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) S iren ' s Song in the AI Ocean: A Survey on Hallucination in Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:34.316712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:34.316712Z digest=sha256:2361b58bf982e2e87c1f84da599167c47ad4758b1cb015a613f5e180a7086277

Observation b346a3ab-f126-4e6a-bc5c-5e514e1e25f6 · outbound

This paper cites Proceedings of the 2021 ACM conference on fairness, accountability, and transparency , pages=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Proceedings of the 2021 ACM conference on fairness, accountability, and transparency , pages=

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:34.409924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:34.409924Z digest=sha256:2740061c563ae88c43e265dc6bd210f939889ffe2d339aeae2a835f56ee4605e

Observation 15f074f6-1ad5-452b-af55-9103cc4219d2 · outbound

This paper cites Should ChatGPT be Biased? Challenges and Risks of Bias in Large Language Models.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Should ChatGPT be Biased? Challenges and Risks of Bias in Large Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:34.503503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:34.503503Z digest=sha256:f254922a400a0d672b9c9a69abe54f80e313b87af949d0bf821201fe32863350

Observation 1313c07f-0d7c-4d4e-9d24-d81a7fb31217 · outbound

This paper cites LLM Spirals of Delusion: A Benchmarking Audit Study of AI Chatbot Interfaces.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) LLM Spirals of Delusion: A Benchmarking Audit Study of AI Chatbot Interfaces

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:34.565332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:34.565332Z digest=sha256:2d25294491951a92b9c2a8c95404f594b881119b2a3c6b02133707fbf5893b37

Observation a458f07a-0ffa-4f31-865f-231be2c48040 · outbound

This paper cites Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society , number=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society , number=

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:05.151838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:44:34.623847Z digest=sha256:2e158b775b7d25035068b70e39c11e3a5c6c80ffa3f4bc7b5b8c8952251220e7

Observation 96133c27-3520-4a5a-9715-6591749ebe40 · outbound

This paper cites Practices for.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Practices for

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:04.934779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:44:34.700820Z digest=sha256:af7317a78e4d53b89c32093b2e848bac4c0705738b5e3a120716223b82661f7c

Observation 856bbd95-784c-4347-abaf-daea68706d34 · outbound

This paper cites an unresolved cited work.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Unresolved cited work

Reference 40

Resolution
parse uncertain
raw_fallback, observed 2026-08-07T12:45:04.655486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:44:34.816885Z digest=sha256:ddf0c5bb1d7661961ccf6b3e723f2ef3bea8d555933b5502bcf8a53945b3b545

Observation f43044d6-d349-4364-bbd8-fb8fb3bedbbb · outbound

This paper cites Aligning AI With Shared Human Values.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Aligning AI With Shared Human Values

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:34.935769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:34.935769Z digest=sha256:356b252416d87fae57eb1e5086235cf8cfa863ca31912a848e25ffce09dc57a5

Observation d834dcae-40d0-4cbc-95ab-648cf68eb808 · outbound

This paper cites Toward an Evaluation Science for Generative AI Systems.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Toward an Evaluation Science for Generative AI Systems

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:35.042357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:35.042357Z digest=sha256:0605626c30501be707876fd3c03b0a765e109d4d0628aeefa04721e8adf7dea3

Observation daa23e92-20a0-4295-a933-fcdb476df63d · outbound

This paper cites Science , volume=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Science , volume=

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:35.125805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:35.125805Z digest=sha256:2bcc24eb124cb61840e92ff64372da4de70155fb2b192c82b94ffb98dea984ba

Observation 0ca37549-c4c5-45e5-970c-15f40d1808ec · outbound

This paper cites an unresolved cited work.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Unresolved cited work

Reference 44

Resolution
verified exact
doi, observed 2026-08-07T12:45:02.931158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:44:35.232606Z digest=sha256:437cc63414a003367ce6b1513c477ea734141e9d1792b760d089d083b28ce362

Observation 7bcf96d0-68e2-4d3c-b056-ed6c588043bb · outbound

This paper cites The Vulnerability of Language Model Benchmarks: Do They Accurately Reflect True LLM Performance?.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) The Vulnerability of Language Model Benchmarks: Do They Accurately Reflect True LLM Performance?

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:35.307850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:35.307850Z digest=sha256:3dbc6e990ed70628bc7f4196286546c59aee2eb58742c720801f515289a6f543

Observation 05b8b964-fc11-4886-a522-ddbc7f20ed3f · outbound

This paper cites Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:35.393890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:35.393890Z digest=sha256:802c9e42a5a0676162572cb6c4bf179f133a8cb1ef259710b485f569e0f3fc02

Observation 09ebe1db-820d-48ac-9d25-7ef4f3d17a2b · outbound

This paper cites AI and the Everything in the Whole Wide World Benchmark.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) AI and the Everything in the Whole Wide World Benchmark

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:35.451657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:35.451657Z digest=sha256:25b7428d2ea3d988655bbb1ffcd17299a23a61cd2350d81d3cb08ae65f90e8d7

Observation c3f7131e-5e38-46df-9f2a-7ea4c19805f0 · outbound

This paper cites Proceedings of the 2025.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Proceedings of the 2025

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:35.498690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:35.498690Z digest=sha256:e7bf41775f4a0ab9392a2f9e13c9f72e5669e772521551548a3a68726c308ca3

Observation 2808b874-9e30-4988-997d-b30941a981eb · outbound

This paper cites Rethinking.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Rethinking

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:35.593629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:35.593629Z digest=sha256:b680f59be719a0d967c407d2fb95e6dbfc4c93fcf406b3b47884312c7f94acdd

Observation 90fdb2a7-d6a1-4bdf-a364-b147057f1c14 · outbound

This paper cites Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:35.707630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:35.707630Z digest=sha256:1438aaf493ef0f034dfb8bf01b205eca1246e3ccfac6468bdda6f6d94633c360

Observation 2ff3cc6b-80e4-4297-a894-ab5196b95f4d · outbound

This paper cites Outsider.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Outsider

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:35.776696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:35.776696Z digest=sha256:aea8cb14024dd639b7dce63f2ad82d921531aa0640715ded1441fa933df3661a

Observation 7a18bd0c-da49-4965-981a-f5c1d621b604 · outbound

This paper cites Sociotechnical Safety Evaluation of Generative AI Systems.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Sociotechnical Safety Evaluation of Generative AI Systems

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:35.836468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:35.836468Z digest=sha256:d6a6690883a29de997f6a3cb007dc9690ab3b31f7d430cb252444ae69de9098e

Observation 3825143f-fd23-4b60-8b37-122987b4325c · outbound

This paper cites Lessons from the Trenches on Reproducible Evaluation of Language Models.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Lessons from the Trenches on Reproducible Evaluation of Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:35.911671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:35.911671Z digest=sha256:153c92c7ad427974d00ca8713d663651745c29ec7a65880f6f2bf10078860e92

Observation 095f31e1-3024-4640-83ae-201405d6b4e2 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Advances in Neural Information Processing Systems , volume=

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:04.439156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:44:35.996295Z digest=sha256:e30ed6ad0af76b2e0cba3a1c7c31213cb8d7b5bfb63756c932c73dc3b6a4e123

Observation 19b54a70-5013-4aab-b741-c30f055a0581 · outbound

This paper cites Proceedings of the 29th international conference on computational linguistics , pages=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Proceedings of the 29th international conference on computational linguistics , pages=

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:04.212267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:44:36.106520Z digest=sha256:c2d581b30ed94f6afc273f16912752a90a4399815927e703e2b5018e1114411f

Observation b7ea2215-28e0-437e-82e8-bb84696ebfb7 · outbound

This paper cites International Conference on Learning Representations , volume=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) International Conference on Learning Representations , volume=

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:36.193313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:36.193313Z digest=sha256:a186af29dbef31d557382d6820e9b7dd8524abc274368cfa3473ee3ff0da532e

Observation 4e9c0f34-3323-44d4-89bd-3ce922ffa349 · outbound

This paper cites Findings of the Association for Computational Linguistics: ACL 2024 , pages=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Findings of the Association for Computational Linguistics: ACL 2024 , pages=

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:36.278667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:36.278667Z digest=sha256:870c3c3e3647f82901d51e87de87a7ed5930881e3c7e883889ec8430d3a7d849

Observation d7b4a3c1-3323-44c2-8021-4b673020cabd · outbound

This paper cites BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:36.345910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:36.345910Z digest=sha256:a336d80f6e3d2106fc497fda4b7af2a2c2618b6cf746a1680a7f7d20b6f58f57

Observation 91789132-3150-4b20-8918-30ab698702b7 · outbound

This paper cites and Darrell, Trevor and Norouzi, Narges and Gonzalez, Joseph E.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) and Darrell, Trevor and Norouzi, Narges and Gonzalez, Joseph E

Reference 59

Resolution
verified exact
doi, observed 2026-08-07T12:45:02.704440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:44:36.435665Z digest=sha256:be68d95db8ab46014f34de60903d6aa5664458c9bb1c7b9e425de7ff7d16eba2

Observation 198004bf-5cd5-4244-8dbd-4afd50b28c74 · outbound

This paper cites an unresolved cited work.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Unresolved cited work

Reference 60

Resolution
verified exact
doi, observed 2026-08-07T12:44:37.153884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:44:36.563039Z digest=sha256:7420a075ef404c9c256619bb65eddd3aae8e2731716062a5c374848788a2e5d7

Observation 26b99a95-b552-483f-91de-7b8bd9d9cc69 · outbound

This paper cites Sacred or.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Sacred or

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:03.955485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:44:36.672670Z digest=sha256:a423817235ad1e1da2807b57b642d47a53a975697df751461c5e4d4c6daef067

Observation 5a6ad17b-8612-40c6-a6b1-a94cab07c62f · outbound

This paper cites AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:36.752949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:36.752949Z digest=sha256:c4d18547e3abb6dbb3b3dfcf15e2d512c3092b5315a5e8f5b9cf36ba4483cbf9

Observation b1e89b97-aa78-4cdf-87ea-310d3df49387 · outbound

This paper cites and Metaxa, Dana.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) and Metaxa, Dana

Reference 63

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T12:45:03.305534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:44:36.796636Z digest=sha256:fc14053c00fb5305be8700aed825fd1c06933f214980730eced5aa93ee073163

Observation af29f3f7-4e3a-4bc3-9c8a-b23f72efdb87 · outbound

This paper cites Proceedings of the 2020 conference on fairness, accountability, and transparency , pages=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Proceedings of the 2020 conference on fairness, accountability, and transparency , pages=

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:36.850378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:36.850378Z digest=sha256:a5c95602125f7e7763fc181ae3a13570bbc863ee8b43d9c5e4bc2a7b78737b35

Observation 4b748991-47bb-4794-885b-a35805ba49ce · outbound

This paper cites In-House Evaluation Is Not Enough: Towards Robust Third-Party Flaw Disclosure for General-Purpose AI.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) In-House Evaluation Is Not Enough: Towards Robust Third-Party Flaw Disclosure for General-Purpose AI

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:36.904062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:36.904062Z digest=sha256:5ef19ffa6f41c431db76b4606caaa86e1034a8096148773d4dca6685dea00ef0

Pith citing papers

No inbound Pith citation observations are available.