Pith. sign in

Paper Citation Record · LEDGER

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks

As of 10 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 0 inbound Pith citation observations for arXiv:2608.03340.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.03340 v1

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T20:45:25.334279Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

60 of 60 outbound references displayed

  • verified exact2
  • verified fuzzy32
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation aa26d19c-07a0-4ccb-8212-206ecd3ff29c · outbound

This paper cites an unresolved cited work.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-05T20:45:33.624847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T20:45:20.449417Z digest=sha256:c63a7a0e563e32028e8ec6b177f8b09dbf953a35898399b427f1a35734624ab9

Observation afb8098b-667d-4ded-baee-5e717cb87a17 · outbound

This paper cites an unresolved cited work.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-05T20:45:33.337602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T20:45:20.550110Z digest=sha256:c612154793918e4ec2efc68e614d857196a28d95d116518dc00b660e38b0fab8

Observation 2a135fc8-f5ec-4625-b6f3-7c38ef41eae2 · outbound

This paper cites 2026 , month = apr, url =.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks 2026 , month = apr, url =

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:45:33.203841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T20:45:20.658346Z digest=sha256:d6e407c539dc6645b14ca04cfdcac673c078fc67d45510f5586699d3318f7f06

Observation 2e939414-952f-4a61-a5f7-ae748decb158 · outbound

This paper cites When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T20:45:20.763940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:45:20.763940Z digest=sha256:4d48c65f9a9dd3d438542b82cd8f72c353dc5167f420e6d0fd1f4187d374147e

Observation de45cbaf-98e6-4ad1-b2ca-d258a1fb85c6 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks Advances in Neural Information Processing Systems , volume=

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T20:45:20.845031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:45:20.845031Z digest=sha256:3a21b8a978dd49cd1962e2920d9774bdf772180e7e7074d0a3274aa83674163f

Observation 5282a702-3aa6-4dd0-b274-c032aa99ae9c · outbound

This paper cites 2026 , month = apr, url =.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks 2026 , month = apr, url =

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:45:33.008976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T20:45:20.922632Z digest=sha256:ddb29c829c33136e455896d5e89a64d4a0266619b1f0c799c0cced65980675dc

Observation 43ec1675-4301-409e-8cb1-e5d1bfc2bc60 · outbound

This paper cites Or Not? , author=.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks Or Not? , author=

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:45:32.790256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T20:45:20.995809Z digest=sha256:3b5443e4b25074885852d1afde642fa8854564b5f097c81a7a2cd6bcf996d6b5

Observation 01fef925-c661-4b57-90f2-304a50d79323 · outbound

This paper cites How Reasonable are Common-Sense Reasoning Tasks: A Case-Study on the W inograd Schema Challenge and SWAG.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks How Reasonable are Common-Sense Reasoning Tasks: A Case-Study on the W inograd Schema Challenge and SWAG

Reference 8

Resolution
verified exact
doi, observed 2026-08-05T20:45:25.518534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T20:45:21.108941Z digest=sha256:6c637c15ab40a8d21516f20c434304a4394c5a963c5c672930a79a5e3b9cf3b5

Observation 127e310a-6821-4188-8f0c-d2facd005688 · outbound

This paper cites Benchmark\^.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks Benchmark\^

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:45:32.604562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T20:45:21.186853Z digest=sha256:32ed3f034281cd2e73e26b40ce562c39d8fe9c2ab64d3b0328e3aa5037ea5889

Observation 5397c555-4833-4e7e-ac7a-dcebfe4fb56a · outbound

This paper cites Findings of the Association for Computational Linguistics: ACL 2025 , pages=.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks Findings of the Association for Computational Linguistics: ACL 2025 , pages=

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:45:32.424635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T20:45:21.263902Z digest=sha256:dad33c706d526d893537d15aaeaf9bacbe7031d83c71dd20a6942d29c94b8dff

Observation db2dd29a-4624-4216-882a-828003b9b3a0 · outbound

This paper cites arXiv preprint arXiv:2602.10657 , year=.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks arXiv preprint arXiv:2602.10657 , year=

Reference 11

Resolution
verified exact
raw_fallback, observed 2026-08-05T20:45:26.143824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T20:45:21.316558Z digest=sha256:9692fe7b126d4efbd1ddebf5a09d217b5d91436ddfe7d15e14f9c81695a38862

Observation 0160079d-44a8-449a-b2be-c9a8af6fa2d3 · outbound

This paper cites BenchMarker: An Education-Inspired Toolkit for Highlighting Flaws in Multiple-Choice Benchmarks.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks BenchMarker: An Education-Inspired Toolkit for Highlighting Flaws in Multiple-Choice Benchmarks

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T20:45:25.924443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T20:45:21.383348Z digest=sha256:1761d57fabad58abcd38b27d77c80913b8bbb6b2d48a6f7324ebc834fcdb3bdd

Observation 56379f78-0bb1-4456-978f-6501e01a31ab · outbound

This paper cites Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:45:32.241564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T20:45:21.444789Z digest=sha256:e3765ea55b00c848b37c8546676fbaec0d43e541dfb3a978d656e07266631f52

Observation ca020e86-6664-42d7-9928-97f3e0eab768 · outbound

This paper cites My answer is C.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks My answer is C

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:45:32.058435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T20:45:21.505260Z digest=sha256:de291df7ec0b9f447ae214cef09fd3cd3efbf82ecbdb8698e4450dd331d0920d

Observation 733952c6-b42d-4a14-80ea-6fc49a159c37 · outbound

This paper cites Findings of the Association for Computational Linguistics: EMNLP 2024 , pages=.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks Findings of the Association for Computational Linguistics: EMNLP 2024 , pages=

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:45:31.905048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T20:45:21.603442Z digest=sha256:9b19716ac9a9f124f99e1f3da21e535e5504123514e1efe539152414c8c77fa9

Observation 9e96d133-b9c0-4a7d-b1d1-6637c9cbba91 · outbound

This paper cites PNAS nexus , volume=.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks PNAS nexus , volume=

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:45:31.763326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T20:45:21.676758Z digest=sha256:80470729eac70415ef32dc2ff9fdaa33f762d96717534e696d64fa8587cf1ce7

Observation 88e24ca8-4eed-4a48-a8ae-7ffbf82131b4 · outbound

This paper cites Scientific Reports , volume=.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks Scientific Reports , volume=

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:45:31.537238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T20:45:21.724056Z digest=sha256:5344b175d7a08544a504c5366064cc6d9a1fe03b6fba3b5d36fc3c137f008d22

Observation 229af391-81c8-4443-b82d-06b4295f31fc · outbound

This paper cites Open-World Evaluations for Measuring Frontier AI Capabilities.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks Open-World Evaluations for Measuring Frontier AI Capabilities

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T20:45:21.806249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:45:21.806249Z digest=sha256:be81feb22c953cc1941340d2d91aa233ec303d361c62c09d88eac5c21eaf483b

Observation a9ae06f8-1a30-4c7d-a166-8bd9fa0cd825 · outbound

This paper cites arXiv preprint arXiv:2502.14359 , year=.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks arXiv preprint arXiv:2502.14359 , year=

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T20:45:21.881351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:45:21.881351Z digest=sha256:75c74939103040096e6849a57e8b6568b5a6d29287ec960b9ae2bc081bbc6254

Observation ab190fd9-26ea-4b1e-9164-de9af70698f2 · outbound

This paper cites Proceedings of the Teddington Conference on the Mechanization of Thought Processes , pages =.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks Proceedings of the Teddington Conference on the Mechanization of Thought Processes , pages =

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:45:31.406129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T20:45:21.988606Z digest=sha256:6f7fdc95bfe4923fbb316d1deef9c31381cbceef0b9bf3a250f5e864d9adc6ae

Observation 06981e3a-8a59-40fd-8f30-9fd20198ea41 · outbound

This paper cites Communications of the ACM , volume=.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks Communications of the ACM , volume=

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T20:45:22.050562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:45:22.050562Z digest=sha256:567907633262ac004805174c06349ef536b3328300fcdffb03ab4e58d4c5bafb

Observation db2dff29-6f6d-4a1d-87b7-d3f9368acb1a · outbound

This paper cites , author=.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks , author=

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T20:45:22.129241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:45:22.129241Z digest=sha256:dc882980d94c45ca2f7b17ac1bf6ca03864de830024e6da8b6aa7c87727680b2

Observation 52719735-10b8-45f1-868b-0fcd1b6670fe · outbound

This paper cites Communications of the ACM , volume=.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks Communications of the ACM , volume=

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T20:45:22.203610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:45:22.203610Z digest=sha256:3d9b13f9c85af298e2d5da97f8f7d347b912e9e92753ddb880d2bc7a47fde095

Observation 4ec8f9f7-b2bc-4807-8087-fa7e1402fe32 · outbound

This paper cites Proceedings of the 57th annual meeting of the association for computational linguistics , pages=.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks Proceedings of the 57th annual meeting of the association for computational linguistics , pages=

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T20:45:22.268373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:45:22.268373Z digest=sha256:d47da83d2e6a295994cf34f03fd2da5b6232b11f9b8422216589bed71ef08d16

Observation 7eac5cff-975f-4b49-ba7c-e35f7aca30af · outbound

This paper cites an unresolved cited work.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T20:45:22.341993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:45:22.341993Z digest=sha256:4ffbbc27700a69d2d9b074bc427fd7ba722cc5b02bfd9d1133f15adfc68cf246

Observation 2294bfab-459e-4110-ac5c-71ea8238e437 · outbound

This paper cites an unresolved cited work.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T20:45:22.417617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:45:22.417617Z digest=sha256:2c7285be7ea606a2b3e8f4a723ba878e57f5cfa3d93e95ecc066c687b03a8aef

Observation 0421babf-0f9c-4a70-9d31-da33b8c21e3d · outbound

This paper cites an unresolved cited work.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-05T20:45:31.127482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T20:45:22.512162Z digest=sha256:2a5bd07a59ca3734c01275916b0202b04684348eb201b60470dd8e021d8e0ffb

Observation f75f59ed-422a-4043-ad8a-b7b3b8e06dee · outbound

This paper cites Proceedings of the National Academy of Sciences , volume=.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks Proceedings of the National Academy of Sciences , volume=

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T20:45:22.595110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:45:22.595110Z digest=sha256:fad97dd19ef3c8dc9dc3030b7cbb64d1e5d7eeae45ae28de39550f4bc6551ec8

Observation 48271d95-d7f6-4b66-b52c-923a698f89b8 · outbound

This paper cites Nature human behaviour , volume=.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks Nature human behaviour , volume=

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:45:30.891981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T20:45:22.693317Z digest=sha256:9a4154832406fda32e9f1888fe6a2acef210c4026c1fbd5fb48f483eedc990bd

Observation 58550614-5a60-412d-9396-d7a0fe37dd61 · outbound

This paper cites Proceedings of the Second Workshop on Insights from Negative Results in NLP , pages=.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks Proceedings of the Second Workshop on Insights from Negative Results in NLP , pages=

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:45:30.682265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T20:45:22.773396Z digest=sha256:6b13b212fc47f7dcaedf27c1455466e831528e3ee43b7a21b6f318a260ff66ce

Observation 5002aff7-5756-4581-85ed-c1810b55e424 · outbound

This paper cites an unresolved cited work.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-05T20:45:30.347897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T20:45:22.867977Z digest=sha256:2d99aa8bf44034e1018b47d79880c26b653c13947c2cca2dcdc76e63f72b19e9

Observation c9e7186c-6e95-4e8f-8c91-66001519c0ee · outbound

This paper cites Speech acts , pages=.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks Speech acts , pages=

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T20:45:22.910903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:45:22.910903Z digest=sha256:3d23784ad10835893f34e74e31556cacdb75a31b28f9a927bb9384f09b1ce9c1

Observation de5ab4dc-791a-4f89-8024-4038d17feb4c · outbound

This paper cites Linguistics and philosophy , volume=.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks Linguistics and philosophy , volume=

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:45:30.011257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T20:45:23.019090Z digest=sha256:c93d649b7937d1816211c8cb62dd18e89d9e09826b8ec7bd312b3c66058bd917

Observation 1bf4a69f-2297-4daf-b649-28a88359af36 · outbound

This paper cites The Stanford Encyclopedia of Philosophy , editor=.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks The Stanford Encyclopedia of Philosophy , editor=

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:45:29.675791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T20:45:23.114665Z digest=sha256:8eecdddd81a4a792bed56e1c0391c7a92b80927a82ccb24d0ceb19c61fb26612

Observation 240b3dd8-878e-4115-b473-bc08330fae02 · outbound

This paper cites Proceedings of the 21st International Conference on Natural Language Processing (ICON) , pages=.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks Proceedings of the 21st International Conference on Natural Language Processing (ICON) , pages=

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:45:29.424806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T20:45:23.181949Z digest=sha256:a99516ec77bc5f4ac64db7e89e508f43018b14652da06b86cf7043238ec70970

Observation 91ed73d1-4662-4d14-98b0-fad4713952e3 · outbound

This paper cites an unresolved cited work.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-05T20:45:29.151850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T20:45:23.249941Z digest=sha256:389e01dfc6e7199d777cc0f19d09ef501a9855e2087c36d670484d05eb281403

Observation f61ccf09-b85f-4386-8058-b9b243eefe7b · outbound

This paper cites Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:45:28.970581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T20:45:23.336492Z digest=sha256:7e3659968cb37f054c6e34dbf48f09e1b52afe4dce0aaf6087e1c935e8193c63

Observation fd920a19-0991-495e-a77d-fd2c7c3dbaf1 · outbound

This paper cites an unresolved cited work.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-05T20:45:28.818855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T20:45:23.447424Z digest=sha256:b877a2d3940e3b871925ef9e27a9d86312350f6b13f73111dd2a2453615debaf

Observation 0078b89e-1746-4d7e-bf5f-3c29e74fab64 · outbound

This paper cites Findings of the Association for Computational Linguistics: EMNLP 2021 , pages=.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks Findings of the Association for Computational Linguistics: EMNLP 2021 , pages=

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:45:28.675032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T20:45:23.555239Z digest=sha256:87a05b29d224895fb0927c2195b4452b5489a936067b78a3f8d211d1df4e8e4f

Observation 6873148c-fe7f-467b-a6f8-8f5722991a1b · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks Advances in Neural Information Processing Systems , volume=

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:45:28.497198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T20:45:23.618974Z digest=sha256:8d46fd071c2b413a18c0e78c3c9658bf08002fbac88d8a63cb0a58409b05fcd1

Observation 12c4f787-7c6d-4bbe-84f2-2285d4aa93d6 · outbound

This paper cites arXiv preprint arXiv:2602.17594 , year=.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks arXiv preprint arXiv:2602.17594 , year=

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T20:45:23.717026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:45:23.717026Z digest=sha256:29bcbb8c89643d3a4b18749751cca8f0d82a947a3d774038cd1f06170fde28a5

Observation bc39cb14-d273-4ff2-9027-d4f9655b9c73 · outbound

This paper cites International conference on machine learning , pages=.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks International conference on machine learning , pages=

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T20:45:23.795013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:45:23.795013Z digest=sha256:445ba45a77dc265f8352361c5a7d55c54624ac5f0f7a68564b1875f08e1ea579

Observation 84b1cb74-6818-459b-be54-daa536b16bac · outbound

This paper cites an unresolved cited work.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T20:45:23.882802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:45:23.882802Z digest=sha256:e3006e49d6f958a6ee360237d23fdb56b0948d283d0da712bf18d34528689cf0

Observation 2c96cfc5-15b8-4a7f-b6c3-a30548af8fcb · outbound

This paper cites IEEE Access , year=.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks IEEE Access , year=

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:45:28.310393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T20:45:23.930324Z digest=sha256:fee10dc67755b7fcf9863866c3a631de2ece4a7e6cf21147011e7db4c9d24272

Observation 854b1981-6016-41b7-9e49-1ee231d6d212 · outbound

This paper cites Proceedings of the 29th Conference on Computational Natural Language Learning , pages=.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks Proceedings of the 29th Conference on Computational Natural Language Learning , pages=

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:45:28.160734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T20:45:24.008348Z digest=sha256:c67c57ddf6e7f0db6a4c94e448fe96951677a7b1ba6110f6aa4ca48200f22b79

Observation a38a459a-0d2f-4b6e-be60-1c74acd89b48 · outbound

This paper cites What the HellaSwag? On the Validity of Common-Sense Reasoning Benchmarks.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks What the HellaSwag? On the Validity of Common-Sense Reasoning Benchmarks

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-05T20:45:24.096011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:45:24.096011Z digest=sha256:7d1e00dd6e1eb5f1ab7ed01b3b31fdd058f7967e8578dc0b2e08db3f19f1716d

Observation b1a956fd-bbb9-4e7b-9c02-8668d5417fcf · outbound

This paper cites Findings of the Association for Computational Linguistics: EACL 2026 , pages=.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks Findings of the Association for Computational Linguistics: EACL 2026 , pages=

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:45:27.948372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T20:45:24.184144Z digest=sha256:08464b7f8a49c02e9eb49c3d8729f247ab4eca5ab0d756720356751fa452a2a9

Observation 311eb501-d071-4230-9228-4c7e3f35a4eb · outbound

This paper cites RUPBench: Benchmarking Reasoning Under Perturbations for Robustness Evaluation in Large Language Models.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks RUPBench: Benchmarking Reasoning Under Perturbations for Robustness Evaluation in Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T20:45:24.247556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:45:24.247556Z digest=sha256:fb645fce36b776329d855bee7184a15e93026eeb1bf2a138ce4a03fb26792ed8

Observation 09932863-acc0-41ed-b01c-36e43225482e · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-05T20:45:24.328597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:45:24.328597Z digest=sha256:b037175aaa26a9449ad5c8df782060932cb65cdeebc4dd79309171eb30ef8d64

Observation 1497d779-b546-44d7-b2dd-ab0c09bd559d · outbound

This paper cites Transactions of the Association for Computational Linguistics , volume=.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks Transactions of the Association for Computational Linguistics , volume=

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-05T20:45:24.389948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:45:24.389948Z digest=sha256:1b1a24f7c804c8770898d505facb032c93b46c651041493b31bb4c846d16ea25

Observation 4f64fa96-73a0-4e9a-a269-273326984b0c · outbound

This paper cites Innovative Journal of Applied Science , pages=.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks Innovative Journal of Applied Science , pages=

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:45:27.808382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T20:45:24.512569Z digest=sha256:d53c3ff38a57c9e070cd12ac6cb2cba0a4874d0f8328fc80e7a9e3de37f726c1

Observation dc36a8b4-d02a-4906-9f36-a9568a5d8666 · outbound

This paper cites Expert Systems , volume =.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks Expert Systems , volume =

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:45:27.626372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T20:45:24.612540Z digest=sha256:f90bedade01a7ea3a6f41a4d0ad9d38dcb1f578bd838aa7c7ff3d1b08568208b

Observation cfb00901-91b5-4d91-89b7-987a2f2881da · outbound

This paper cites 2026 , eprint =.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks 2026 , eprint =

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:45:27.427352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T20:45:24.704523Z digest=sha256:727bf3df1adb8b2b079f95f3b9568d735664485eae7fa5e079fad42157c75f3c

Observation 08b8dbf1-42f4-4ea3-bb39-3c57b724d0f0 · outbound

This paper cites Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics , year =.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics , year =

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:45:27.245639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T20:45:24.801821Z digest=sha256:51ddd991cf816fe618bf44df9b19fb5c14453dfcbe00c548f23a3d4b83ceb0a0

Observation 2839b1e6-23ef-42fb-bc6d-273e6745d3a4 · outbound

This paper cites Journal of the Royal Statistical Society: Series B (Methodological) , volume =.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks Journal of the Royal Statistical Society: Series B (Methodological) , volume =

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:45:27.045407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T20:45:24.926045Z digest=sha256:8ed6b8572fd560964104c5a6c8c99ee6ae3855f1ffaba42fd4ebb217e483a3df

Observation d02853ed-6643-44ce-84a4-2b90765cb88d · outbound

This paper cites AI Magazine , volume=.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks AI Magazine , volume=

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:45:26.893930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T20:45:24.998130Z digest=sha256:ab9ede2f4666f6228f59f7cfe03eada88ae778dc36cf4b236b189402181f9a66

Observation 4e4206c7-0ba1-4179-a766-04f0138960dc · outbound

This paper cites Transactions of the Association for Computational Linguistics , volume=.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks Transactions of the Association for Computational Linguistics , volume=

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:45:26.734269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T20:45:25.082085Z digest=sha256:dd316677150d712659c925e4e110002768fd40cd2bcd31f300b4f3ac48412f1f

Observation e6832bf7-f5df-4d04-b6e8-73c470f1627f · outbound

This paper cites Proceedings of the 9th Widening NLP Workshop , pages=.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks Proceedings of the 9th Widening NLP Workshop , pages=

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:45:26.588331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T20:45:25.171035Z digest=sha256:4a2a0e15b6722a4b5c5bf19c0f160c2d3cfb555a9befc98beb2cef80838975bf

Observation b9f74602-3860-44b0-9da1-e520bc2815a5 · outbound

This paper cites Transactions of the Association for Computational Linguistics , volume=.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks Transactions of the Association for Computational Linguistics , volume=

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:45:26.430611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T20:45:25.259248Z digest=sha256:20d2a3aa10df4185815f21b2b70457042171c31ec797e5a714c7c310f76fae1e

Observation 93664a2b-5556-4d70-8de5-bb9cd8492617 · outbound

This paper cites Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , pages=.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , pages=

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:45:26.280577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T20:45:25.334279Z digest=sha256:5057582d3a37f1fb885950a424cac535da415fb8d3114057faaaf050e42cb8d3

Pith citing papers

No inbound Pith citation observations are available.