Pith. sign in

Paper Citation Record · LEDGER

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench

As of 19 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 4 inbound Pith citation observations for arXiv:2507.21476.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.21476 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T12:48:09.853658Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:37:54.823963Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-13T23:08:25.106938Z

Reference resolution

27 of 27 outbound references displayed

  • verified exact0
  • verified fuzzy5
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch4

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a953727c-fb1c-4391-b553-3239a5d0b3bf · outbound

This paper cites Phi-4-reasoning Technical Report.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench Phi-4-reasoning Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T12:48:07.539696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:48:07.539696Z digest=sha256:4782da2ab4c3585fc6a24d1fa96649ca102871a76fea32ee98c5404695103a9c

Observation 12470ab8-9ba2-4b65-8aff-82ea03bf0790 · outbound

This paper cites https: //blog.google/technology/google-deepmind/ gemini-model-thinking-updates-march-2025/.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench https: //blog.google/technology/google-deepmind/ gemini-model-thinking-updates-march-2025/

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:48:12.204493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:48:07.793387Z digest=sha256:adc8cf7db5d2436c928a761cea5ea42299472491c3c0454ae442666da5b685c4

Observation 683f9eb8-8304-4b22-9cc8-ac4ddc7001b2 · outbound

This paper cites In Working Notes of CLEF 2025 - Conference and Labs of the Evaluation F orum.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench In Working Notes of CLEF 2025 - Conference and Labs of the Evaluation F orum

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:48:12.130059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:48:07.887911Z digest=sha256:95672fe696d88e1c233649cd4fed092614cd97f46994332d15cde0aeb31d1e28

Observation 62bbde9c-ac16-4bc1-92d8-1c737c8671eb · outbound

This paper cites BIG-Bench Extra Hard.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench BIG-Bench Extra Hard

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T12:48:07.971609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:48:07.971609Z digest=sha256:a9a4226b65f143afb59e1aa75da564dd57ffaafc75a63a4e657b5559f395d39a

Observation 80ac73db-40d3-45a8-86ed-31cbec686418 · outbound

This paper cites When 'YES' Meets 'BUT': Can Large Models Comprehend Contradictory Humor Through Comparative Reasoning?.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench When 'YES' Meets 'BUT': Can Large Models Comprehend Contradictory Humor Through Comparative Reasoning?

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T12:48:11.099671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:48:08.109209Z digest=sha256:13a371890a26ea3bacdcd5b3a86d221c244adecbe677a5ba69cd8cfd20c21fe6

Observation f28ad304-3e3b-4fcb-a08e-165d092eff2d · outbound

This paper cites Let's Verify Step by Step.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench Let's Verify Step by Step

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T12:48:08.158360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:48:08.158360Z digest=sha256:7a7d5989ed2f40d491426198d03aca34fde2d7421a9ef86a2481a956cfe75be6

Observation 7cc49c60-75f7-49d8-b82f-f6234cad64d4 · outbound

This paper cites In Proceed- ings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 11069–11081.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench In Proceed- ings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 11069–11081

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:48:11.956878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:48:08.199060Z digest=sha256:6b529ab356d70fdfdaa5188b9116a8fd6f3d5c7d85de9bd23f1495135cee444d

Observation ad9be2f5-65c2-4393-bae4-9f82ba291152 · outbound

This paper cites WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T12:48:08.248926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:48:08.248926Z digest=sha256:7262d389caf23be2a9071f49f756bc9962d8ce3377bd3716bc13e2e3b78584bc

Observation b084aec8-6fee-4087-a879-fb1e266d68de · outbound

This paper cites Inverse Scaling: When Bigger Isn't Better.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench Inverse Scaling: When Bigger Isn't Better

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T12:48:08.300856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:48:08.300856Z digest=sha256:f09ccd984d60755043640aad3561aedd8122607ce3cff2f64c5c4abe6192e950

Observation 75acb1a9-c225-43a2-9ff5-3e318ecf04b2 · outbound

This paper cites CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T12:48:08.423730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:48:08.423730Z digest=sha256:a30ba053cf3f91b81a564f05d82640cabb75b186bcd25f3dd917087db0ccb8bf

Observation 1c712a67-ee2b-4547-b5fb-db7e402e4fd9 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T12:48:08.472560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:48:08.472560Z digest=sha256:eed5ea11d580ecbc3d9c493fc3b09557523f1834eb4edbabc6ff945ca399bbaf

Observation 06d8d6d1-29d0-44c1-8461-11d8fd80da06 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T12:48:08.489572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:48:08.489572Z digest=sha256:d567f4ad1551d15bb3efbff60c58642e25a3dc35c484cfcd07ceb319ba197439

Observation c7ae82a2-ee3b-4401-854d-6be08571409b · outbound

This paper cites Phishing Awareness via Game-Based Learning.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench Phishing Awareness via Game-Based Learning

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T12:48:10.816118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:48:08.645463Z digest=sha256:b6b95b5962e600792ae32130320e7d5deb25752121f8cb300476c3826878064d

Observation f61fd422-528b-483d-81a9-585fc134065a · outbound

This paper cites Between Underthinking and Overthinking: An Empirical Study of Reasoning Length and correctness in LLMs.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench Between Underthinking and Overthinking: An Empirical Study of Reasoning Length and correctness in LLMs

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T12:48:08.790009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:48:08.790009Z digest=sha256:7469fd01d31cb87b0798af064223ee8d3905b4fbb792c470244b0f1f94694022

Observation a26467da-b9b2-4995-ac92-11b90ac777cb · outbound

This paper cites Challenging the Boundaries of Reasoning: An Olympiad-Level Math Benchmark for Large Language Models.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench Challenging the Boundaries of Reasoning: An Olympiad-Level Math Benchmark for Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T12:48:08.957774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:48:08.957774Z digest=sha256:6928bbd5ff1281710702ad4cbc65e19ef77f2d816055196eadb810a2c458fbb2

Observation 96465f1a-7000-4334-ae1f-ad044ac3a944 · outbound

This paper cites In Proceedings of the 2022 Confer- ence on Empirical Methods in Natural Language Processing, pages 2866–2879.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench In Proceedings of the 2022 Confer- ence on Empirical Methods in Natural Language Processing, pages 2866–2879

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:48:11.719843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:48:09.095642Z digest=sha256:e80496334dfd77faf48f105e15a89c19a2e81cd9653d2da46fe18ad5b0f3117e

Observation a0c34740-672c-4533-a901-2407262fa64c · outbound

This paper cites Signatures of room-temperature superconductivity emerging in two-dimensional domains within the new Bi/Pb-based ceramic cuprate superconductors at ambient pressure.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench Signatures of room-temperature superconductivity emerging in two-dimensional domains within the new Bi/Pb-based ceramic cuprate superconductors at ambient pressure

Reference 22

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T12:48:10.523027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:48:09.185449Z digest=sha256:f772317c622200a7bb19a3bdca4e5ab2c50f54072928e24c980885114dd0dd0f

Observation b61d379f-95df-47da-bca1-d3318f3f7d37 · outbound

This paper cites In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Tech- nologies, pages 4213–4228.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Tech- nologies, pages 4213–4228

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:48:11.432296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:48:09.311008Z digest=sha256:72e6086c55ceb20bd27cc57dec22c747551539d81de71db51ae5c431057caf2f

Observation 656c4b5c-aa05-4ebe-8311-f2e126843037 · outbound

This paper cites Uniformly rotating vortices for the lake equation.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench Uniformly rotating vortices for the lake equation

Reference 24

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T12:48:10.276133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:48:09.407105Z digest=sha256:16b1250535ed992798e720ddd0f382f85618a50cfdd0bfd337937fb2be17092d

Observation 9c75b147-50f5-4294-a4eb-4c281786a722 · outbound

This paper cites arXiv preprint arXiv:2502.18080.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench arXiv preprint arXiv:2502.18080

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T12:48:09.550871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:48:09.550871Z digest=sha256:a63b6197d6915491194a3f9bc0fcf1c08e26d658512825c2ae94f5c04b40205b

Observation b202ac87-475a-49a2-a998-8a4f9e111b8b · outbound

This paper cites Humor in AI: Massive Scale Crowd-Sourced Preferences and Benchmarks for Cartoon Captioning.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench Humor in AI: Massive Scale Crowd-Sourced Preferences and Benchmarks for Cartoon Captioning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T12:48:09.697591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:48:09.697591Z digest=sha256:c1c18f67c4fab91b3b5ca9d75dbfa9cc80a1938b1aaf08baf1dbd8882e21c22f

Observation 2720ae89-6683-4c2a-a731-9dbeb5c2d960 · outbound

This paper cites Bridging the Creativity Understanding Gap: Small-Scale Human Alignment Enables Expert-Level Humor Ranking in LLMs.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench Bridging the Creativity Understanding Gap: Small-Scale Human Alignment Enables Expert-Level Humor Ranking in LLMs

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T12:48:09.853658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:48:09.853658Z digest=sha256:82d44b55a829798f9baa18096eff8ca0d2082ea3c5962f35fe8bb111d8f8aa84

Observation 66882ed1-a3c6-4913-8e54-06cd911d2f7f · outbound

This paper cites AmbigQA: Answering Ambiguous Open-domain Questions.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench AmbigQA: Answering Ambiguous Open-domain Questions

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T12:48:08.379481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:48:08.379481Z digest=sha256:e75baf47b857610c8162cb8f5f765bdcac51392d7d34839284586628af630b30

Observation 5a21e266-f982-45af-b54a-c35634c4fe36 · outbound

This paper cites Solving Quantitative Reasoning Problems with Language Models.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench Solving Quantitative Reasoning Problems with Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T12:48:08.040390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:48:08.040390Z digest=sha256:5d5d61b8d49ec49df07964efc618250fbd85e27f83bdbf9f1771107b00e54cea

Observation 1c50d8ef-ea44-4d1b-a529-5caa53466861 · outbound

This paper cites ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T12:48:07.929176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:48:07.929176Z digest=sha256:08041827997f976230e9f852d87fd4a1829723d30948d3ba360a25582978aeda

Observation 64403be8-ce55-4562-a051-38c6ef8c9e4b · outbound

This paper cites Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T12:48:07.648384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:48:07.648384Z digest=sha256:7da74344ad48dba9b07183459497f4dd2003c5e78c9c9fed8992fec084ef32ef

Observation 3118d44f-b9c8-4fe4-8ee5-0a339bc9a225 · outbound

This paper cites ARC Prize 2024: Technical Report.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench ARC Prize 2024: Technical Report

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T12:48:07.697172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:48:07.697172Z digest=sha256:7bd3c4fccd632e74b43a7aa093ef01c985d8813ddd6f920b10c925899d676100

Pith citing papers

Observation 7da47090-9728-4baa-8a07-054361c8b7ce · inbound

Relationship-Centered Care: Relatedness and Responsible Design for Human Connections in Mental-Health Care cites this paper.

Relationship-Centered Care: Relatedness and Responsible Design for Human Connections in Mental-Health Care Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-14T20:22:12.729190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T20:22:12.729190Z digest=sha256:53ae780da3b2804448b1c7bdd0f8b65a34d32275620e72a4c188f1703772b083

Observation 0dc073df-7454-4b5b-beb0-9f15dbc037eb · inbound

HumorRank: A Tournament-Based Leaderboard for Evaluating Humor Generation in Large Language Models cites this paper.

HumorRank: A Tournament-Based Leaderboard for Evaluating Humor Generation in Large Language Models Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:08:25.110234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-13T23:07:38.352198Z digest=sha256:2d0b019c3197e3bde2efcfc5a4a0a4206e00037d2812d387d3545c36c6fa2f8e

Observation 6af44ac6-42c1-4c0b-87bb-15bc9f946ce6 · inbound

HumorRank: A Tournament-Based Leaderboard for Evaluating Humor Generation in Large Language Models cites this paper.

HumorRank: A Tournament-Based Leaderboard for Evaluating Humor Generation in Large Language Models Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T17:05:10.690456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T17:05:10.690456Z digest=sha256:296c630b98dae74c2efd513fd88d26219a011e9958dde2403bd19eefa14b2f7f

Observation 3d3eebc3-281f-4c1e-9d41-a48f9b77f061 · inbound

Computational Humor with Multimodal LLMs: Methods, Datasets, Evaluation, and Challenges cites this paper.

Computational Humor with Multimodal LLMs: Methods, Datasets, Evaluation, and Challenges Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T15:37:54.823963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:37:54.823963Z digest=sha256:a41756e5a882a505b7087357405622b848f66e7b06db1f91ce9f1fb97d2eaa39