Pith. sign in

Paper Citation Record · LEDGER

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench

As of 7 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 3 inbound Pith citation observations for arXiv:2507.21476.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.21476 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T12:48:09.853658Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T17:05:10.690456Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-13T23:08:25.106938Z

Reference resolution

27 of 27 outbound references displayed

  • verified exact0
  • verified fuzzy5
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch4

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a953727c-fb1c-4391-b553-3239a5d0b3bf · outbound

This paper cites Phi-4-reasoning Technical Report.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench Phi-4-reasoning Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T12:48:07.539696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:48:07.539696Z digest=sha256:c4d68c4b079c86d7fb2785d1cfb07468c3bdf5ad264a72630e2c44f874849d3b

Observation 12470ab8-9ba2-4b65-8aff-82ea03bf0790 · outbound

This paper cites https: //blog.google/technology/google-deepmind/ gemini-model-thinking-updates-march-2025/.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench https: //blog.google/technology/google-deepmind/ gemini-model-thinking-updates-march-2025/

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:48:12.204493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:48:07.793387Z digest=sha256:634e7af06eb82b59643c8ff099268bc016504ff7d65265146d5fc36068810475

Observation 683f9eb8-8304-4b22-9cc8-ac4ddc7001b2 · outbound

This paper cites In Working Notes of CLEF 2025 - Conference and Labs of the Evaluation F orum.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench In Working Notes of CLEF 2025 - Conference and Labs of the Evaluation F orum

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:48:12.130059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:48:07.887911Z digest=sha256:c65af65cf4288ebfdd61d1c88110717c7b7d19039b020a8bf0025f0548ce288e

Observation 62bbde9c-ac16-4bc1-92d8-1c737c8671eb · outbound

This paper cites BIG-Bench Extra Hard.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench BIG-Bench Extra Hard

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T12:48:07.971609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:48:07.971609Z digest=sha256:adaa39bc56c4064fc6c243046f19e4fca1c18a138b604db32d33bcb56b8b1420

Observation 80ac73db-40d3-45a8-86ed-31cbec686418 · outbound

This paper cites When 'YES' Meets 'BUT': Can Large Models Comprehend Contradictory Humor Through Comparative Reasoning?.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench When 'YES' Meets 'BUT': Can Large Models Comprehend Contradictory Humor Through Comparative Reasoning?

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T12:48:11.099671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:48:08.109209Z digest=sha256:cdfe5428baf14ae0dde7ac05581b5886feec29f6454502afcc2ea8f6fcdc9b7b

Observation f28ad304-3e3b-4fcb-a08e-165d092eff2d · outbound

This paper cites Let's Verify Step by Step.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench Let's Verify Step by Step

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T12:48:08.158360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:48:08.158360Z digest=sha256:b99e000e9de3233b8419fc89771db1bbf3d6e54fc487ac625b937eb9ea4cff72

Observation 7cc49c60-75f7-49d8-b82f-f6234cad64d4 · outbound

This paper cites In Proceed- ings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 11069–11081.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench In Proceed- ings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 11069–11081

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:48:11.956878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:48:08.199060Z digest=sha256:9ac0f6928fd2776e662dc3e1f7a8fd0b9f6f77cd0819f5a194623e8885fc8f8c

Observation ad9be2f5-65c2-4393-bae4-9f82ba291152 · outbound

This paper cites WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T12:48:08.248926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:48:08.248926Z digest=sha256:3edf81e905245b3e18957c6e03ac24e579cf863a69a5ef16bd1af013298956f6

Observation b084aec8-6fee-4087-a879-fb1e266d68de · outbound

This paper cites Inverse Scaling: When Bigger Isn't Better.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench Inverse Scaling: When Bigger Isn't Better

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T12:48:08.300856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:48:08.300856Z digest=sha256:4750684b1d92d12eb97eac97ffe5d7d530259419e693a1bfeca5cf45e787acf8

Observation 75acb1a9-c225-43a2-9ff5-3e318ecf04b2 · outbound

This paper cites CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T12:48:08.423730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:48:08.423730Z digest=sha256:e435cf659d981f338aca2e0da589c1a5abd5ef0f73537b6623f9c907a8a85181

Observation 1c712a67-ee2b-4547-b5fb-db7e402e4fd9 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T12:48:08.472560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:48:08.472560Z digest=sha256:9621eab38490b5c6f52bca9b9b2fcbf53e4cea1c8ca05b7870f5ee60e24fcd8b

Observation 06d8d6d1-29d0-44c1-8461-11d8fd80da06 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T12:48:08.489572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:48:08.489572Z digest=sha256:63f8fde3b9f3d1f0e57833ca78d1024a95c3c1d2b88b2ec60c79727503adb2b7

Observation c7ae82a2-ee3b-4401-854d-6be08571409b · outbound

This paper cites Phishing Awareness via Game-Based Learning.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench Phishing Awareness via Game-Based Learning

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T12:48:10.816118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:48:08.645463Z digest=sha256:ab502075bcb01cc5bfbc1e7484acdaf98b73d5fa003b3ad0eef33e243c985fae

Observation f61fd422-528b-483d-81a9-585fc134065a · outbound

This paper cites Between Underthinking and Overthinking: An Empirical Study of Reasoning Length and correctness in LLMs.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench Between Underthinking and Overthinking: An Empirical Study of Reasoning Length and correctness in LLMs

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T12:48:08.790009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:48:08.790009Z digest=sha256:af0e80959f785258d2fd83c042066527e9a0ae15721f14cd168fed8deacd5d07

Observation a26467da-b9b2-4995-ac92-11b90ac777cb · outbound

This paper cites Challenging the Boundaries of Reasoning: An Olympiad-Level Math Benchmark for Large Language Models.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench Challenging the Boundaries of Reasoning: An Olympiad-Level Math Benchmark for Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T12:48:08.957774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:48:08.957774Z digest=sha256:1bbe400d469be7432613b0d22bceceeacc81acbd77b162eecb19f94c5beb9458

Observation 96465f1a-7000-4334-ae1f-ad044ac3a944 · outbound

This paper cites In Proceedings of the 2022 Confer- ence on Empirical Methods in Natural Language Processing, pages 2866–2879.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench In Proceedings of the 2022 Confer- ence on Empirical Methods in Natural Language Processing, pages 2866–2879

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:48:11.719843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:48:09.095642Z digest=sha256:4bcba1cf7b422f121ec9441d3adb7eaf1eb1a7d0f8c4edf9b611a85756c03d8e

Observation a0c34740-672c-4533-a901-2407262fa64c · outbound

This paper cites Signatures of room-temperature superconductivity emerging in two-dimensional domains within the new Bi/Pb-based ceramic cuprate superconductors at ambient pressure.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench Signatures of room-temperature superconductivity emerging in two-dimensional domains within the new Bi/Pb-based ceramic cuprate superconductors at ambient pressure

Reference 22

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T12:48:10.523027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:48:09.185449Z digest=sha256:348b2e387dd3306f0282f3331f15d323bf77c82b397a709ceea8c7c3268ebdcf

Observation b61d379f-95df-47da-bca1-d3318f3f7d37 · outbound

This paper cites In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Tech- nologies, pages 4213–4228.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Tech- nologies, pages 4213–4228

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:48:11.432296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:48:09.311008Z digest=sha256:095e21b76b3529590146d7e3132d8df95a9ac434051f90e9686d0c34b8d8c2d6

Observation 656c4b5c-aa05-4ebe-8311-f2e126843037 · outbound

This paper cites Uniformly rotating vortices for the lake equation.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench Uniformly rotating vortices for the lake equation

Reference 24

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T12:48:10.276133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:48:09.407105Z digest=sha256:effacfc9225eab0b5a5e53be70f409709ec190a043c917c22f9e14ccdaaa1407

Observation 9c75b147-50f5-4294-a4eb-4c281786a722 · outbound

This paper cites arXiv preprint arXiv:2502.18080.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench arXiv preprint arXiv:2502.18080

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T12:48:09.550871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:48:09.550871Z digest=sha256:3d5e2132d792c97a4b534c374731b0fe7842a07a5ac34e7e744281d75e73e443

Observation b202ac87-475a-49a2-a998-8a4f9e111b8b · outbound

This paper cites Humor in AI: Massive Scale Crowd-Sourced Preferences and Benchmarks for Cartoon Captioning.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench Humor in AI: Massive Scale Crowd-Sourced Preferences and Benchmarks for Cartoon Captioning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T12:48:09.697591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:48:09.697591Z digest=sha256:98ca85ef03eff65b8fbc02f74808a1c95853127680cd5885ba4ab5115487b43c

Observation 2720ae89-6683-4c2a-a731-9dbeb5c2d960 · outbound

This paper cites Bridging the Creativity Understanding Gap: Small-Scale Human Alignment Enables Expert-Level Humor Ranking in LLMs.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench Bridging the Creativity Understanding Gap: Small-Scale Human Alignment Enables Expert-Level Humor Ranking in LLMs

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T12:48:09.853658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:48:09.853658Z digest=sha256:1ee339229e690bc0d7ea0138b912e6e1f35a131670ad53ce23ecabbd1a64bd95

Observation 66882ed1-a3c6-4913-8e54-06cd911d2f7f · outbound

This paper cites AmbigQA: Answering Ambiguous Open-domain Questions.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench AmbigQA: Answering Ambiguous Open-domain Questions

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T12:48:08.379481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:48:08.379481Z digest=sha256:aed6fc98a2a48466806e09903fb3ef6297fdbedc74488b6ba7fb00f7d7168c00

Observation 5a21e266-f982-45af-b54a-c35634c4fe36 · outbound

This paper cites Solving Quantitative Reasoning Problems with Language Models.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench Solving Quantitative Reasoning Problems with Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T12:48:08.040390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:48:08.040390Z digest=sha256:e7bd6a18572697c834514285ed8e262afc30a71799bbf6f838830454642309dc

Observation 1c50d8ef-ea44-4d1b-a529-5caa53466861 · outbound

This paper cites ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T12:48:07.929176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:48:07.929176Z digest=sha256:39b699da8e93b1da96b3b407220fc8d07f62d1ddf826ef3eaf0a8ce4f15166ba

Observation 64403be8-ce55-4562-a051-38c6ef8c9e4b · outbound

This paper cites Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T12:48:07.648384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:48:07.648384Z digest=sha256:dd65f9f41d5989721b43b1d602320adcd23857bec7e35d8fe98368f76d89afb4

Observation 3118d44f-b9c8-4fe4-8ee5-0a339bc9a225 · outbound

This paper cites ARC Prize 2024: Technical Report.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench ARC Prize 2024: Technical Report

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T12:48:07.697172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:48:07.697172Z digest=sha256:345464821b1ab401b2fd70c79d698760ebe41ec5d23746157c8da9c64f7d7398

Pith citing papers

Observation 7da47090-9728-4baa-8a07-054361c8b7ce · inbound

Relationship-Centered Care: Relatedness and Responsible Design for Human Connections in Mental-Health Care cites this paper.

Relationship-Centered Care: Relatedness and Responsible Design for Human Connections in Mental-Health Care Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-14T20:22:12.729190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T20:22:12.729190Z digest=sha256:73d729173473a13404518155d8953ff5051ef4213e0c709157bb3530cc839550

Observation 0dc073df-7454-4b5b-beb0-9f15dbc037eb · inbound

HumorRank: A Tournament-Based Leaderboard for Evaluating Humor Generation in Large Language Models cites this paper.

HumorRank: A Tournament-Based Leaderboard for Evaluating Humor Generation in Large Language Models Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:08:25.110234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-13T23:07:38.352198Z digest=sha256:c8872c13744e69e17d8a062856a16ce986b13c5741c2fc71475ff57d0606c60a

Observation 6af44ac6-42c1-4c0b-87bb-15bc9f946ce6 · inbound

HumorRank: A Tournament-Based Leaderboard for Evaluating Humor Generation in Large Language Models cites this paper.

HumorRank: A Tournament-Based Leaderboard for Evaluating Humor Generation in Large Language Models Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T17:05:10.690456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T17:05:10.690456Z digest=sha256:5b9bebec4ad9bc9ee7bb9a509952221fcfb8ed342b332be733701ad54bf2b0b4