Pith. sign in

Paper Citation Record · LEDGER

Humanity's Last Exam

As of 1 August 2026, this Paper Citation Record lists 100 of 300 outbound references and 100 inbound Pith citation observations for arXiv:2501.14249.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.14249 v10

Coverage vector

measured 100 of 300 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T18:40:50.139345Z

measured 200 of 200 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-01T06:32:01.292127+00:00

measured 100 of 185 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-15T15:03:20.637041Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T09:37:00.704616Z

Reference resolution

100 of 300 outbound references displayed

  • verified exact44
  • verified fuzzy16
  • unresolved23
  • parse uncertain17
  • malformed identifier0
  • metadata mismatch0

External citation measurements

8
pith, observed 2026-07-10T09:37:00.704616Z

Outbound references

Observation 0674c55d-a792-44bc-a066-f4f42c468349 · outbound

This paper cites A BERT Baseline for the Natural Questions.

Humanity's Last Exam A BERT Baseline for the Natural Questions

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-10T18:40:50.419005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:24da97093b855a35c739435358077fa045b9b3c8a535ada3fbb736feed2efc08

Observation 27dd35ee-06db-478e-9197-f4b86b5f2a17 · outbound

This paper cites AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents.

Humanity's Last Exam AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-14T01:35:51.514012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:d441edda3539d914df1bf2e1bfa284ef6670ab8feda946c514801446daba5dab

Observation 57b80c17-719f-4f8a-83ac-5665aa42fb40 · outbound

This paper cites The claude 3 model family: Opus, sonnet, haiku.

Humanity's Last Exam The claude 3 model family: Opus, sonnet, haiku

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:33:07.045304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:394472a488f9638ab8a633d1043c25859335e4489e9fb45c7beffb3b9e9a6e61

Observation c2f01758-effb-4793-ac7c-c3ec81d7b1d5 · outbound

This paper cites Model card addendum: Claude 3.5 haiku and upgraded claude 3.5 son- net.

Humanity's Last Exam Model card addendum: Claude 3.5 haiku and upgraded claude 3.5 son- net

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:33:07.071488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:850edc33962e4b4757209c351fbb48eddbf09161593ebb2ff8fd52d84e4eae77

Observation 4ccaf35b-d213-41e4-b046-a8fe382ced6f · outbound

This paper cites Responsible scaling policy updates.

Humanity's Last Exam Responsible scaling policy updates

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:41:07.148523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:72efc2e76b699c8e0684cdcbc74321397ea375940d47fbbf43cf4d31bf8dbe4f

Observation 90da4a4c-43cb-44d3-8020-bf418ed27da1 · outbound

This paper cites HealthBench: Evaluating Large Language Models Towards Improved Human Health.

Humanity's Last Exam HealthBench: Evaluating Large Language Models Towards Improved Human Health

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:18:21.444858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:68d3978e9913ee4028e6be70ca7340af4409c4a66c5bbc95f77520438770063d

Observation 3886ecc9-1fa3-4ed3-a3a2-17f45ef1e50f · outbound

This paper cites Austin, A.

Humanity's Last Exam Austin, A

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:41:07.155325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:0e8b7292edde9f6451427cdd152f38153d8c20ed9e59c3ed24e2c927de09b41b

Observation 4820e0a9-d957-4348-afc6-fa61d76d7bcf · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Humanity's Last Exam Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-10T18:40:50.335048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:5d0dc6e41937ee9c336c7a07a0160891b1b01223f4f307648123030f01964abd

Observation 4af1527c-2c4c-4b65-b9d4-9fd5279242a8 · outbound

This paper cites MS MARCO: A Human Generated MAchine Reading COmprehension Dataset.

Humanity's Last Exam MS MARCO: A Human Generated MAchine Reading COmprehension Dataset

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:54:21.897164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:5a04fa22c8d4c7d4a25ecde93e57dcbe38e834e0f2ee9626ad26cbee27b4e88d

Observation f303afb9-da19-4460-9313-d870f3e455bd · outbound

This paper cites Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models.

Humanity's Last Exam Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-10T18:40:50.368040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:47c031af60e67c348106f7cc25e475cd3bf6fa039330df1a5947ed00162cf441

Observation 2c8a8636-7558-4f7c-b51b-62ff7e988de6 · outbound

This paper cites MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering.

Humanity's Last Exam MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-10T18:40:50.382147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:80d5ec34f254c60db99d56928b407ab0accbff44bf550b913e5a738da7ec1c7e

Observation 0d5be82c-f123-468f-8f98-fa312324b292 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Humanity's Last Exam Evaluating Large Language Models Trained on Code

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-10T18:40:50.403484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T08:08:23.404839+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:4b7988430e247dbee2bc97137883c74483544d978a46fdd24c48518cd8171228

Observation d8d0bcd6-df78-4e00-95f3-09e4107a34c8 · outbound

This paper cites ARC Prize 2024: Technical Report.

Humanity's Last Exam ARC Prize 2024: Technical Report

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-10T18:40:50.406934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:b51682827f1a10ba8f381ec3ea9cc1cdfcb626f6c54e2bb9a6328e01e453598c

Observation 87080e3d-765c-475e-936f-4c4b66546c95 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Humanity's Last Exam Training Verifiers to Solve Math Word Problems

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-10T18:40:50.410534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:5eb759dc46b02a8ebd0863b73892f7cda6a77bb7a310877222caa8396197ce00

Observation 83ae0e2b-70ba-43c1-bb42-936cc43057af · outbound

This paper cites Deepseek-v3 technical report.

Humanity's Last Exam Deepseek-v3 technical report

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:41:07.063241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:0b6bd01639d301d8315323f064150553303be0cf4f050c43ba7848c65ddf20a9

Observation d8ca28f3-9b8a-46a8-a159-20dc3426d302 · outbound

This paper cites an unresolved cited work.

Humanity's Last Exam Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-05-16T16:41:07.046097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:60ab347fd4b1880bb9c03081beb24191193f87741f6152500d3ca4a7a7576dba

Observation 92cbd9cc-d252-4f4d-b3c5-dbf31c438bf0 · outbound

This paper cites The Llama 3 Herd of Models.

Humanity's Last Exam The Llama 3 Herd of Models

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-10T18:40:50.276608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:14c66a217eab6f89f3ade889de3718517381df955bee65210bf47a32c641be43

Observation 6e914fe3-e237-4b2a-9699-84239dbf0294 · outbound

This paper cites Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models.

Humanity's Last Exam Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-15T09:09:15.216863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:1d913ef584b886435c14e91769bfb7a650bb7a27ef81eeab77f4394f74e57f19

Observation 72a19ef2-e197-4494-ac82-e0e76871d46d · outbound

This paper cites FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI.

Humanity's Last Exam FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:44:01.791449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:f5d8b0db0ecc6b22ae2414cce21b95c87406a67757a276f9ea24eb65ff529339

Observation 79639eed-ebbe-4204-89d1-c7e00761b6df · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

Humanity's Last Exam OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:38:21.105975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:8af1947f6c0c5d17037caf32a4a8fc383ccd20bdcde02e41819253e8233bc274

Observation 1347b7c1-7c61-480c-a52d-d73e92e1782b · outbound

This paper cites Measuring Coding Challenge Competence With APPS.

Humanity's Last Exam Measuring Coding Challenge Competence With APPS

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:11:36.763016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:1033ec60b60ff0c55e969e6350f95248072818177c62fec2ad65354aec24712c

Observation f8d191e4-a193-4e65-b9bd-e12827d9ed1a · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Humanity's Last Exam Measuring Massive Multitask Language Understanding

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-10T18:40:50.330913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:16abf4efc3de0e9d0f254d0d34e572d1045128914c2889be5c0020c75145b334

Observation d9e7ea17-ef7b-4f03-9440-809463a40b70 · outbound

This paper cites Hendrycks, C.

Humanity's Last Exam Hendrycks, C

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:33:07.067401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:9c1e2b19ad5b9731713bcd7d6cac24fb2ce17a98b20446928ac44abc31c69c9c

Observation e9afb5ec-2d3f-45e4-a87a-855acb944474 · outbound

This paper cites PixMix: Dreamlike Pictures Comprehensively Improve Safety Measures.

Humanity's Last Exam PixMix: Dreamlike Pictures Comprehensively Improve Safety Measures

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-10T18:40:50.357840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:92fd49422f681a90528fff7867193cebbcf8b52e1cd5ae2e25b3f342122135bb

Observation 440d44c2-14e0-496c-8f9a-7cea5b6dcb63 · outbound

This paper cites Hosseini, A.

Humanity's Last Exam Hosseini, A

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:33:07.075575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:bc2cb262637a97e9d0b1e01bb26cdcb620ee5cda4f78502326d8768e3495a8ef

Observation 6dcaa95d-167d-409f-b60e-8f373bfa74e2 · outbound

This paper cites Not All LLM Reasoners Are Created Equal.

Humanity's Last Exam Not All LLM Reasoners Are Created Equal

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-10T18:40:50.364289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:2b1d7de7c1a494aaadd50ad128f0df0cb24d1190a9399f497785c5a5a43accad

Observation 2aed4d3f-5445-4a3e-ba59-b4571a683f11 · outbound

This paper cites Jacovi, A.

Humanity's Last Exam Jacovi, A

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:41:07.128242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:dfb6ce409b24c2f717154a6d8d8e8f0b142b2b0b106b20fe74adbca5dac00bf6

Observation e55c4d1a-215b-4cfa-9403-99d879450489 · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

Humanity's Last Exam SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-10T18:40:50.371563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:e099f2c2dc9a317ded2df04569789a9ded42c532b5689301f4788688e2e0add1

Observation 39a34301-2c01-4c36-8d7d-52060d5d47d8 · outbound

This paper cites Dynabench: Rethinking Benchmarking in NLP.

Humanity's Last Exam Dynabench: Rethinking Benchmarking in NLP

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-10T18:40:50.375124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:02111b496d8d9a294f42100af2945c432b59e96912b8b528bf58ae3d0f37fcad

Observation 1f5deb70-d016-4da3-993c-b3740f09e8d6 · outbound

This paper cites Refusal-Trained LLMs Are Easily Jailbroken As Browser Agents.

Humanity's Last Exam Refusal-Trained LLMs Are Easily Jailbroken As Browser Agents

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-10T18:40:50.378768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:57234e29d387ec14ec6fbfb069e76913992b9848616df482caa2416b7e060bc7

Observation 58c43289-107a-469a-854c-8b5c5a5ddc88 · outbound

This paper cites an unresolved cited work.

Humanity's Last Exam Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-05-16T16:41:07.134763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:7acad387896c3ffefd30613b092fcdd695c64f980854e16ac155da2d233b5ba6

Observation 0f5c6172-edb0-46f3-858b-4a2aaeea5411 · outbound

This paper cites LAB-Bench: Measuring Capabilities of Language Models for Biology Research.

Humanity's Last Exam LAB-Bench: Measuring Capabilities of Language Models for Biology Research

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:58:18.270625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:bb23c16b350a8ff7c0baa921ba7742b19d48bed145d02e6ac04b44913e024eab

Observation 1c3a6b99-f76a-4bda-85c9-11dda50e1b5a · outbound

This paper cites The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning.

Humanity's Last Exam The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:55:50.014661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:82a71c1e21cd84fdd4ba51dde529f33bd913c42ce553e423b1530955a10a2dc2

Observation 4607f30d-4d4b-45f4-8578-92e51303e4cb · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

Humanity's Last Exam MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:30:15.750849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:a052d3b1322335e8d9a8aba052d72b53379a13e3f06f07c06e0521754496aa85

Observation 2f38d407-5bf1-4fe6-ad94-b89fce08349a · outbound

This paper cites Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence.

Humanity's Last Exam Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-10T18:40:50.396135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:95ac5705a6215a9671c9e184f29bbb2cdbe59e5aff875498b1be0278a82edf6f

Observation c8b692e0-1fa8-4884-af36-e51b1e4e1f2a · outbound

This paper cites Adversarial NLI: A New Benchmark for Natural Language Understanding.

Humanity's Last Exam Adversarial NLI: A New Benchmark for Natural Language Understanding

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-10T18:40:50.399670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:67537979fa5e0bf5935167342e5170d6cc27e567c82f2c055b3bbc860e4d49df

Observation 0ed849cd-c769-4fbf-b734-9339adf4a1f2 · outbound

This paper cites Openai o1 system card.

Humanity's Last Exam Openai o1 system card

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:41:07.169492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:6a47d91dd7512f3f314c98fe9b99011d62818942831311008294a49a952f1ba6

Observation 6440a1fa-822a-41db-9817-b1c3db2063ba · outbound

This paper cites Openai and los alamos national laboratory announce bio- science research partnership.

Humanity's Last Exam Openai and los alamos national laboratory announce bio- science research partnership

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:33:07.049969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:66bcd78ad578179f1deb28def4fe1107074297480dfbc6bef2b7fdc5076834ab

Observation 1c45bb77-14a1-41fc-9566-85874a56befb · outbound

This paper cites Introducing swe-bench verified.

Humanity's Last Exam Introducing swe-bench verified

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:33:07.054377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:fab889b67a73e74409d96844965e0a421deba6293e8ca474cabd8ce28694c985

Observation 1411e5e9-95db-4ac6-9915-17da206b6e4e · outbound

This paper cites GPT-4 Technical Report.

Humanity's Last Exam GPT-4 Technical Report

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-10T18:40:50.413217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:63af7987e367f9534c697a4aac329331ffb73dcc9788d6a1c97cf40708f95316

Observation 883d4f43-ce39-4518-8d35-8861439d16cd · outbound

This paper cites an unresolved cited work.

Humanity's Last Exam Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-05-16T16:41:07.079699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:85aa34bdbbfd1c95efe32261c4805e3313d45b93929064cc4336692ddedf9f0f

Observation 919639a0-3045-4e79-bedf-518059ad8d24 · outbound

This paper cites How predictable is language model benchmark performance?.

Humanity's Last Exam How predictable is language model benchmark performance?

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-10T18:40:50.245598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:18b71a6823fe67eb263c02ab8c96241b94542e0229f0d6a9e34a070140a14191

Observation f6c7dc91-efe5-455d-8128-3832e33e9a07 · outbound

This paper cites Discovering Language Model Behaviors with Model-Written Evaluations.

Humanity's Last Exam Discovering Language Model Behaviors with Model-Written Evaluations

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:20:04.692181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:490948c898666a4707cafe7cabdb8af7e46cadfa58459872cbdc064d0ce841b8

Observation 2dc370b6-c932-4027-b9fe-d0e494d97546 · outbound

This paper cites Phuong, M.

Humanity's Last Exam Phuong, M

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:41:07.142840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:63883a61793d05930aae1d12fec2e3da69480846a10a28aede072c1fa2f74acb

Observation 2b4d9014-65cb-4474-9f90-daba5bb4bdf7 · outbound

This paper cites SQuAD: 100,000+ Questions for Machine Comprehension of Text.

Humanity's Last Exam SQuAD: 100,000+ Questions for Machine Comprehension of Text

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:13:01.811276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:2d9e5cc9b161d79fb7026a133e29af7dbfc0f11b5bf8ef0b35e8b5aae00d644a

Observation ace79c53-f700-4255-9305-baf66308ab7c · outbound

This paper cites Know What You Don't Know: Unanswerable Questions for SQuAD.

Humanity's Last Exam Know What You Don't Know: Unanswerable Questions for SQuAD

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-10T18:40:50.266938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:7b5e94cf4945556cfe27fba5f0a6d7d2a4eb28f0c369c6810cd1392603fc85fb

Observation 8cbf44d9-0a14-4efa-a824-a56615833a89 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

Humanity's Last Exam GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:00:34.729655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:cf9746108e339b9ea1393d7e3a5f7ad3e3988bbfce7b58e4ce7a4e662f50174a

Observation a10528b8-6a51-45db-8ac9-59197fac8ec9 · outbound

This paper cites Singhal, S.

Humanity's Last Exam Singhal, S

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:41:07.032519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:8cb85cea51de329fb94cdedf413c8c0d8608df06d31667346875baaf9f810e49

Observation 1e78985d-f07d-4eeb-936d-5bd115cf80ff · outbound

This paper cites Skarlinski, J.

Humanity's Last Exam Skarlinski, J

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:41:07.096669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:f1a64d0fe5d3a4a73f195d7d67ccb7de11c90976cba21db67f7379783f1d3fce

Observation 2f050906-8ac6-4008-94d3-a4650ba2ced7 · outbound

This paper cites an unresolved cited work.

Humanity's Last Exam Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-05-16T16:41:07.106613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:070425f1206d26e59fa49d2e6888ca46a7491f3578b02d9e0d6c0e7b7d4213dc

Observation 7ceabc31-aa91-46dd-97e1-43cef4c1a185 · outbound

This paper cites Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models.

Humanity's Last Exam Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:26:25.883762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:b4536920d664a4438919688755de4efbafc40e9a6635982aa42ed74ed4855eb3

Observation 1655239c-4031-40f4-a026-1ef5bfe7dc84 · outbound

This paper cites MMLU-Pro+: Evaluating Higher-Order Reasoning and Shortcut Learning in LLMs.

Humanity's Last Exam MMLU-Pro+: Evaluating Higher-Order Reasoning and Shortcut Learning in LLMs

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-10T18:40:50.292773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:8b404d8afb35944853d39b1347f53fa139571cde631dab5797942f3d34eb78af

Observation b724b403-03ac-4247-9f7e-e0a31c941ccc · outbound

This paper cites Team et al.

Humanity's Last Exam Team et al

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:41:07.024468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:51c5f003df8da9f4f06ae83842bb63e933fafee815bfbdf023bae9df284ce215

Observation 100b26cf-83cd-4197-870f-95c5f9c1c371 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Humanity's Last Exam Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-05-10T18:40:50.302563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:42fed8bf3771f3043594ba587d4e983c6972cfd275946b1dfe8e60b2a20fa3bc

Observation 78453e3f-994b-4be4-b14a-32a9eb819004 · outbound

This paper cites PutnamBench: Evaluating Neural Theorem-Provers on the Putnam Mathematical Competition.

Humanity's Last Exam PutnamBench: Evaluating Neural Theorem-Provers on the Putnam Mathematical Competition

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-10T18:40:50.306620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:ff2ea4f1eb88537daf447aebdcc63874ac4e2c0ab79386c37ec0092f927c44a3

Observation 2e4f46f9-7ba7-4cd0-b546-cdcc40773e21 · outbound

This paper cites an unresolved cited work.

Humanity's Last Exam Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-05-16T16:41:07.122991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:d9a9eaeca6d5f5649c354ee7a030901816e532f9cab9cd2e521167b2990d3ecb

Observation beb2fdbc-5f20-49aa-8fbe-015686fdb728 · outbound

This paper cites SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems.

Humanity's Last Exam SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:34:11.049162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:9e4e7bff47b374da91476bd7870d287853126019cb4138e2bedb7d1e6eedbb58

Observation 06000166-b00c-4927-90f0-36eb17c4098f · outbound

This paper cites MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark.

Humanity's Last Exam MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:51:14.588885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:7aafeb00b6a65a80b35071230c5f5615bdfc473382c4121d8fac2b0d98285173

Observation ea073248-6907-44e2-b96d-353a4f384ba4 · outbound

This paper cites Measuring short-form factuality in large language models.

Humanity's Last Exam Measuring short-form factuality in large language models

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:45:50.345693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:16a538e4aac14e72b82e0f26cee1e938f568e34189efc4af1920f8b903d39ebe

Observation d9fa4ac5-63e7-45de-986b-c9ba862334fb · outbound

This paper cites RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts.

Humanity's Last Exam RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-10T18:40:50.326504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:f9e5970e1cdb7e6356f5a1b987b2e8b47c72562870fbdf6c2f721a8afef61830

Observation bf768341-2706-4c8f-8444-c2761bda2e09 · outbound

This paper cites Grok-2 beta release.

Humanity's Last Exam Grok-2 beta release

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:41:07.044606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:cbec8471d31391299beb49619121d21a30745b84f578d610fd1070f22b393914

Observation 3005621b-41d3-4fdb-8966-ea498356f3ba · outbound

This paper cites an unresolved cited work.

Humanity's Last Exam Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-05-16T16:41:07.139183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:42befd7ec7c0b437cdcb2aa4104b04bd3bc01b7c1454447a20a5be96cbfc4d5e

Observation 92015453-8c83-49d8-9571-929546769cf9 · outbound

This paper cites HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering.

Humanity's Last Exam HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:05:12.673916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:fc822cc7fb3886147dd42400672eecc4c564a09823f14c34314b22ff8a523add

Observation 40150a9d-5215-424a-8158-f50db7e7c045 · outbound

This paper cites $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains.

Humanity's Last Exam $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:19:00.987380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:ba98272487dd8d06539c325a132436cc7c0c91ca5cdb1e3af9453b82b02de90a

Observation be75ed09-53c6-4815-a5d7-86f23bc03ac8 · outbound

This paper cites Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models.

Humanity's Last Exam Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-10T18:40:50.350304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:0e0d802a691630ab6a420fa714488614d2ef2ada393ff834f4c4dcd64e3fb8ef

Observation d6a8eaea-9549-4d9b-a2f0-5d5f2227275f · outbound

This paper cites AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models.

Humanity's Last Exam AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:03:59.570579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:86d7ed92175555cbf55ce73b8a6d6b95b62ea418439aa8875198591a22f0021c

Observation f9e68d2c-f3a3-4b32-952e-fe8202b07ea9 · outbound

This paper cites an unresolved cited work.

Humanity's Last Exam Unresolved cited work

Reference 67

Resolution
parse uncertain
raw_fallback, observed 2026-05-16T16:41:07.022055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:5fe1f9082ba7ad2660df60b7873e70bdf555e99d2263c6a0364d8a9ab4ce03db

Observation 115fe9a7-59cc-4847-ba18-a2aa615b6dec · outbound

This paper cites an unresolved cited work.

Humanity's Last Exam Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-05-16T16:41:07.130414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:66871b9c55592ce3de811c24256ec1957c6be7fd1e4ef33ba4182d4b59e51082

Observation 6aef037c-2dba-4df3-97a5-5b4830dd2665 · outbound

This paper cites an unresolved cited work.

Humanity's Last Exam Unresolved cited work

Reference 69

Resolution
unresolved
raw_fallback, observed 2026-05-16T16:41:07.113432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:5e9c43b2ce069c0645db33fedb561c43f4273262bd9435516548779f86de9fc5

Observation 2af478b3-01e0-4aaa-8a60-5ebfe48137eb · outbound

This paper cites an unresolved cited work.

Humanity's Last Exam Unresolved cited work

Reference 70

Resolution
parse uncertain
raw_fallback, observed 2026-05-16T16:41:07.055726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:8a91a85f111b39db0a49f6bdb38c599e92f2a7d26413afa23610e4bf48067402

Observation 834aafa8-ba35-4c2f-b1d2-4a4ba3319c90 · outbound

This paper cites an unresolved cited work.

Humanity's Last Exam Unresolved cited work

Reference 71

Resolution
parse uncertain
raw_fallback, observed 2026-05-16T16:41:07.028491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:6f04c82516c87f2d735e8aa64bd94db62470138e3ba62311175df3c6be4cf37d

Observation dd9b55ed-9998-498c-97a5-fb71f6d44acb · outbound

This paper cites an unresolved cited work.

Humanity's Last Exam Unresolved cited work

Reference 72

Resolution
parse uncertain
raw_fallback, observed 2026-05-16T16:41:07.108493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:f66e18d4aba20184fdc676c501ff35faa10da216ca88b1cff68812c198827d25

Observation ab2ab1e6-be25-4d32-87ee-dccf527b6e6e · outbound

This paper cites an unresolved cited work.

Humanity's Last Exam Unresolved cited work

Reference 73

Resolution
unresolved
raw_fallback, observed 2026-05-16T16:41:07.132352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:26e23b11f0642fc5fa4c1457baecc588ae518adcccd6df17ff0e01a791c33608

Observation 8d4fffab-453f-4c3f-829d-d219399d42fa · outbound

This paper cites an unresolved cited work.

Humanity's Last Exam Unresolved cited work

Reference 74

Resolution
parse uncertain
raw_fallback, observed 2026-05-16T16:41:07.095371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:c92a23bf9403325470f16d5d737b1953f2084d3863fc4435b4efa1ff4f95d6d4

Observation a2261d8d-fc99-4845-8b19-9bea067d4bb7 · outbound

This paper cites an unresolved cited work.

Humanity's Last Exam Unresolved cited work

Reference 75

Resolution
parse uncertain
raw_fallback, observed 2026-05-16T16:41:07.111614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:eba26dd87fddab377e5a86b48ca52435172bd75e2d8b05caad625a4606de6203

Observation c88780f8-9662-4b0c-b2f7-b921a2f29254 · outbound

This paper cites an unresolved cited work.

Humanity's Last Exam Unresolved cited work

Reference 76

Resolution
unresolved
raw_fallback, observed 2026-05-16T16:41:07.042300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:0f3370fd827a2256a684d95fd39fba2d9e2727f48b2b3facd943c6426897d5f1

Observation 1e93a7a6-3cc9-4183-a2ba-28cf482cf21e · outbound

This paper cites an unresolved cited work.

Humanity's Last Exam Unresolved cited work

Reference 77

Resolution
parse uncertain
raw_fallback, observed 2026-05-16T16:41:07.082815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:951601a24e16039acf4859b0f1c08687dc96f6defe3e4330b23a641feb3199cf

Observation b4dbd5f2-452f-4626-828f-17493f851d24 · outbound

This paper cites an unresolved cited work.

Humanity's Last Exam Unresolved cited work

Reference 78

Resolution
unresolved
raw_fallback, observed 2026-05-16T16:41:07.026500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:d0885280be5132f18c938c8aa9f6a503c859453bf896e0c84abdcc4913d58cbe

Observation bbbfc201-a98c-48d1-aa5f-e51059b4b3db · outbound

This paper cites an unresolved cited work.

Humanity's Last Exam Unresolved cited work

Reference 79

Resolution
parse uncertain
raw_fallback, observed 2026-05-16T16:41:07.072384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:c92ba66019b42abbd5a07a20ba6c38c3e0dc1f4b9378c19b2328fbe01cc789d6

Observation b48e72fc-f145-46ce-96f2-4c8ebd11de22 · outbound

This paper cites an unresolved cited work.

Humanity's Last Exam Unresolved cited work

Reference 80

Resolution
unresolved
raw_fallback, observed 2026-05-16T16:41:07.094040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:7289325e68bdc97ec6d7c9c541137cef40c754349b5a1a29437afc3bf7d92d88

Observation 0097681b-2932-42f7-995d-2a6616449c0b · outbound

This paper cites an unresolved cited work.

Humanity's Last Exam Unresolved cited work

Reference 81

Resolution
parse uncertain
raw_fallback, observed 2026-05-16T16:41:07.047875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:92d006cb4483e75677f831a69464d947545d13472b2f0629d545839fecbd1115

Observation 3a4768da-b57f-47fd-a87e-f48968b4ac67 · outbound

This paper cites an unresolved cited work.

Humanity's Last Exam Unresolved cited work

Reference 82

Resolution
parse uncertain
raw_fallback, observed 2026-05-16T16:41:07.068924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:65032ac83da2bbd0175e7927f33600829db7f33528cdc21ae472035faf5a382e

Observation a7cfb8fe-c67f-4928-b8e2-af431192b3e2 · outbound

This paper cites an unresolved cited work.

Humanity's Last Exam Unresolved cited work

Reference 83

Resolution
parse uncertain
raw_fallback, observed 2026-05-16T16:41:07.091887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:d9e21b33ab96208897f5fe6fc73b043e22dedbd95ab5b33110df9ef71f5d2a8e

Observation 21f33310-3f1d-4e2c-8e15-d8ea4c91ae14 · outbound

This paper cites an unresolved cited work.

Humanity's Last Exam Unresolved cited work

Reference 84

Resolution
unresolved
raw_fallback, observed 2026-05-16T16:41:07.153189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:e0fff48aaab44e210961ce6f280ea4a294cb3bf311a09d1424947f15dd03524b

Observation fcb466d3-ddf7-4d30-9dd0-297d55545e40 · outbound

This paper cites an unresolved cited work.

Humanity's Last Exam Unresolved cited work

Reference 85

Resolution
parse uncertain
raw_fallback, observed 2026-05-16T16:41:07.115307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:81e6dd4d03c1345fa644cc9670edf03a1d6712e33ecb77fb9acd5bfc6e290068

Observation 394f059a-487f-4f9d-8715-df14cc106321 · outbound

This paper cites an unresolved cited work.

Humanity's Last Exam Unresolved cited work

Reference 86

Resolution
parse uncertain
raw_fallback, observed 2026-05-16T16:41:07.056123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:242799f413b44ae5610d88dc4c6503796fb2a4d7c6ac278a2b3e3356a5740fb6

Observation 8cf3a8e3-eb9c-4c15-860d-08292da12896 · outbound

This paper cites an unresolved cited work.

Humanity's Last Exam Unresolved cited work

Reference 87

Resolution
unresolved
raw_fallback, observed 2026-05-16T16:41:07.040393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:63bf132bf56c36eddd02e7ecab43c15d4570f98033aa6ce5e817d1259a368025

Observation ae511d0e-8e38-4ad8-9c92-e73ed437e444 · outbound

This paper cites an unresolved cited work.

Humanity's Last Exam Unresolved cited work

Reference 88

Resolution
unresolved
raw_fallback, observed 2026-05-10T18:40:50.421293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:83cfb4f61a5fa3870ae1ed7eddd2ceb315dfd0748a3da1af53c096e257b27883

Observation 170224d3-f45b-4841-98a1-186e993c2f9a · outbound

This paper cites an unresolved cited work.

Humanity's Last Exam Unresolved cited work

Reference 89

Resolution
unresolved
raw_fallback, observed 2026-05-10T18:40:50.423503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:2b4809a8ec1a7f9cc377b11fe8baa8e441391a5c17231898ca3451c47e5a4c74

Observation 680ba952-a6e3-4d67-9eb3-718225dd72fe · outbound

This paper cites an unresolved cited work.

Humanity's Last Exam Unresolved cited work

Reference 90

Resolution
unresolved
raw_fallback, observed 2026-05-10T18:40:50.425594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:0592c77c2aa1b5caa55d779b3d1aa35b7ccea052b39f1527c8d08471d97a62bb

Observation a1324879-39d5-4d91-97d8-aa0e1e480e6b · outbound

This paper cites an unresolved cited work.

Humanity's Last Exam Unresolved cited work

Reference 91

Resolution
unresolved
raw_fallback, observed 2026-05-10T18:40:50.428055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:9cd57e4579290d035c45820ec503d1612695f9c438760a79734c274ea2b11521

Observation a8a39b86-2224-4d6a-b129-b1511f51d834 · outbound

This paper cites an unresolved cited work.

Humanity's Last Exam Unresolved cited work

Reference 92

Resolution
parse uncertain
raw_fallback, observed 2026-05-10T18:40:50.430346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:36030009b4b13bf8efc4edc7e94d67d7092d4dbced0276ed8c885ac3e11a324c

Observation b086cac7-3d11-4fdf-9f13-2754e58e3e11 · outbound

This paper cites an unresolved cited work.

Humanity's Last Exam Unresolved cited work

Reference 93

Resolution
unresolved
raw_fallback, observed 2026-05-10T18:40:50.432780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:74c6e8fd08056da3feb864512ca95b86f77fb51a1ca716cef547f5ada69a409f

Observation b190e4a7-dcf2-4edf-8b21-0e22de8f6eea · outbound

This paper cites an unresolved cited work.

Humanity's Last Exam Unresolved cited work

Reference 94

Resolution
unresolved
raw_fallback, observed 2026-05-10T18:40:50.434777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:8ac4d1c2976cc2f8fc2dd4490c3b71d0e16f77c541cbac19f72a7cb6bcc86175

Observation 11a7faac-286b-4157-98a0-e2c8fcea1fb3 · outbound

This paper cites an unresolved cited work.

Humanity's Last Exam Unresolved cited work

Reference 95

Resolution
unresolved
raw_fallback, observed 2026-05-10T18:40:50.436821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:513a27291a61e3c3f4115217fe35ad0a92808f72aa8f14bb503dde67d50bf8c2

Observation 01972e42-2b52-4802-b974-2482f5f47abc · outbound

This paper cites an unresolved cited work.

Humanity's Last Exam Unresolved cited work

Reference 96

Resolution
unresolved
raw_fallback, observed 2026-05-10T18:40:50.439029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:fd910b1495264472362db2706c2f4d1fc2cd9231c95ba29fb03ca18e57a63123

Observation a157bdc3-9147-4bed-8c29-9d3934faad5a · outbound

This paper cites an unresolved cited work.

Humanity's Last Exam Unresolved cited work

Reference 97

Resolution
parse uncertain
raw_fallback, observed 2026-05-10T18:40:50.440912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:37da85b048ad8b185ec1d274aa3af86bdf92808f5479447f278686fdfddaad13

Observation f2820dd9-59e5-4cfe-9dc5-5b007e13dc0d · outbound

This paper cites an unresolved cited work.

Humanity's Last Exam Unresolved cited work

Reference 98

Resolution
parse uncertain
raw_fallback, observed 2026-05-10T18:40:50.442841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:c770a293b2c726073b9405035f7462e32a557d1edfec40119abe74312560983c

Observation c469d83d-50f6-44dc-b9aa-e89215019bbc · outbound

This paper cites an unresolved cited work.

Humanity's Last Exam Unresolved cited work

Reference 99

Resolution
parse uncertain
raw_fallback, observed 2026-05-10T18:40:50.444579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:ccaee2b12b9c65341ba4d5a7407ead2d0ce42f32588eb14ed0b3c87d923b06f7

Observation f6a39c93-a05c-4be0-82de-5429e7b21b92 · outbound

This paper cites an unresolved cited work.

Humanity's Last Exam Unresolved cited work

Reference 100

Resolution
unresolved
raw_fallback, observed 2026-05-10T18:40:50.446540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:bb7bb8fa11ecab34beefe95189a08bd138c7f945efecb97a711eb5deb2c7eb0f

Pith citing papers

Observation c5c67edb-4092-48d5-af5f-889782e8cbed · inbound

R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization cites this paper.

R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization Humanity's Last Exam

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-16T00:19:20.579670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-16T00:19:20.462455Z digest=sha256:c872964dcdfa2d7f55022bc92631d7a3f4275aa0829a97b5ac41c355e45623a2

Observation 0b24b582-3f80-4e6c-9ad0-ce2ad16b1162 · inbound

From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review cites this paper.

From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review Humanity's Last Exam

Reference 68

Resolution
verified exact
local_arxiv, observed 2026-05-15T02:57:38.323991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-15T02:57:37.873567Z digest=sha256:55e39e40ea95627e581de588888c6ea52b6c42788fe1340dfdb80e826bd9c5a1

Observation 2f88b2bd-086c-4bb1-adac-c01e782a700c · inbound

WebThinker: Empowering Large Reasoning Models with Deep Research Capability cites this paper.

WebThinker: Empowering Large Reasoning Models with Deep Research Capability Humanity's Last Exam

Reference 37

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T19:14:25.395886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-16T19:14:25.283645Z digest=sha256:380e9b4cb26fa006fa6d986d87861ba802daa64aa4cb46957167ab61df5cf305

Observation 2780a976-73ef-415f-a0c0-18e83e4696a3 · inbound

LLMs Get Lost In Multi-Turn Conversation cites this paper.

LLMs Get Lost In Multi-Turn Conversation Humanity's Last Exam

Reference 64

Resolution
verified exact
local_arxiv, observed 2026-05-14T01:11:09.306334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-14T00:57:10.262350Z digest=sha256:70d007b78aa3c17a7ad297305d4f5bc6f8713d9c9406e362ddccbc6fd256f7ef

Observation acb0b9af-d810-4ad6-9010-ae68311b7a2c · inbound

DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents cites this paper.

DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents Humanity's Last Exam

Reference 21

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T08:07:39.493185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-16T08:07:39.384613Z digest=sha256:9bdc56575fbf9a557b8cd736d4fac3a7da1385f033237ded2189ac4b7fe43032

Observation 1c9fa2bd-f427-4dbb-a00c-f6156be8024b · inbound

MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention cites this paper.

MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention Humanity's Last Exam

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-12T09:28:16.460030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-12T09:28:16.189617Z digest=sha256:9ad5cb9c8f3764e4510086a1d7f7064ef0df87635bac4d9a8b031c6630890084

Observation 235c6809-5865-40dd-a270-382bd15e9b4d · inbound

Do LLMs Overthink Basic Math Reasoning? Benchmarking the Accuracy-Efficiency Tradeoff in Language Models cites this paper.

Do LLMs Overthink Basic Math Reasoning? Benchmarking the Accuracy-Efficiency Tradeoff in Language Models Humanity's Last Exam

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-19T06:27:07.304384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-19T06:25:12.799097Z digest=sha256:cc58bdf4237bb33e269932e1705db7ff3565a27ff8ca17d0580e4891d0cf771e

Observation e250b5d9-6323-4cc7-bfae-2bf1c5055381 · inbound

Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities cites this paper.

Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities Humanity's Last Exam

Reference 63

Resolution
verified exact
local_arxiv, observed 2026-05-19T05:52:07.788016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-19T05:48:02.828938Z digest=sha256:6b427247f158e2b32bd394910f2a2f6ee153153f0a634d8b7a89b0aafab9de37

Observation cade9202-7510-4f7a-860d-f0dd0926abf3 · inbound

Kimi K2: Open Agentic Intelligence cites this paper.

Kimi K2: Open Agentic Intelligence Humanity's Last Exam

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-10T18:40:50.715344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T17:49:27.926646Z digest=sha256:61b330341795871d0a36a8ac3c6aaeb7cbd09e1db2bf362a9b1f551cc395f651

Observation 357233fc-e9f1-40f4-a029-ff02ee2e2435 · inbound

WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent cites this paper.

WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent Humanity's Last Exam

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T18:56:23.980780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-15T18:56:23.817544Z digest=sha256:5b389955def1c34ee4d0a670e26f322c9e94b4e5efac48380d879442f14612dc

Observation eb428d9d-dd67-4c32-a37f-cc4792bdd90d · inbound

GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models cites this paper.

GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models Humanity's Last Exam

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-11T17:50:08.487338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T17:50:08.399160Z digest=sha256:0a7edc0993a99aeef55d889dbc8cb843a8449d07bcdf580ad2b7f40f07591913

Observation af23c177-43d3-434b-a80c-06cc5d34421c · inbound

gpt-oss-120b & gpt-oss-20b Model Card cites this paper.

gpt-oss-120b & gpt-oss-20b Model Card Humanity's Last Exam

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-10T18:40:50.715344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T12:22:54.633089Z digest=sha256:3a8d2934f49d3c93bace9adde615b33903af4fdee377b1ab96393266526dfbe7

Observation 6b57a97e-a594-484a-b4a4-d559c80f503a · inbound

Why Johnny Can't Use Agents: Industry Aspirations vs. User Realities with AI Agents cites this paper.

Why Johnny Can't Use Agents: Industry Aspirations vs. User Realities with AI Agents Humanity's Last Exam

Reference 64

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T16:56:38.089406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-18T16:55:47.922639Z digest=sha256:e3da17d45b4ddfd530e322a6554a132948fcd194162390c791d3a0a3bb8ebf4c

Observation a83b9adb-8aba-484e-82dc-f993a61451a9 · inbound

Probing the Critical Point (CritPt) of AI Reasoning: a Frontier Physics Research Benchmark cites this paper.

Probing the Critical Point (CritPt) of AI Reasoning: a Frontier Physics Research Benchmark Humanity's Last Exam

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-05-18T11:52:35.410734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-18T11:52:10.205796Z digest=sha256:fecd9086d4c33a4f8ce42c5dfec5d2785c748ecd8f7b01cb53ad863adb84a46d

Observation 6569d432-1366-4b81-93d8-84a3215b0c54 · inbound

Scaling Latent Reasoning via Looped Language Models cites this paper.

Scaling Latent Reasoning via Looped Language Models Humanity's Last Exam

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-05-15T07:43:11.681092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-15T07:43:11.620446Z digest=sha256:aa5c588ae630456eecd636d4fd15deebfd638cd1eb71208401eb628063a0133a

Observation d750ec29-25e5-4091-89f5-59367f5ebebf · inbound

MiroThinker: Pushing the Performance Boundaries of Open-Source Research Agents via Model, Context, and Interactive Scaling cites this paper.

MiroThinker: Pushing the Performance Boundaries of Open-Source Research Agents via Model, Context, and Interactive Scaling Humanity's Last Exam

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-17T21:45:17.911735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-17T21:44:18.744201Z digest=sha256:531d44465fef7d813ae1449a048c516e3850f774a1c950428f7a26276949d893

Observation ef2627d0-f6a3-4118-b287-6c71d3802a02 · inbound

DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models cites this paper.

DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models Humanity's Last Exam

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-10T18:40:50.715344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T13:05:26.667750Z digest=sha256:b1749ab9d1bbf293de3212ba5e6facff016fb4b3a0918e36f60360640e955844

Observation 6a0b5123-7090-47a9-97df-330e23b7fb72 · inbound

DVPO: Distributional Value Modeling-based Policy Optimization for LLM Post-Training cites this paper.

DVPO: Distributional Value Modeling-based Policy Optimization for LLM Post-Training Humanity's Last Exam

Reference 19

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T01:48:51.041625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-17T01:46:21.744857Z digest=sha256:ffb3927ce978d79f5c480f1394bcc45e1377f57a1dd6827e8e3f96fabb5f78bd

Observation 86e57c20-5c9b-4cd3-b69f-1291e91d39c0 · inbound

Asynchronous Reasoning: Training-Free Interactive Thinking LLMs cites this paper.

Asynchronous Reasoning: Training-Free Interactive Thinking LLMs Humanity's Last Exam

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T22:58:38.448193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-16T22:54:39.663737Z digest=sha256:11f732f8542f680525ef78d899eb80a62f547207fcbc54b168dbc7f9ea02acb9

Observation caa43568-44e2-46c8-a1c4-f85b36185768 · inbound

Evaluating Large Language Models in Scientific Discovery cites this paper.

Evaluating Large Language Models in Scientific Discovery Humanity's Last Exam

Reference 44

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T21:48:34.124182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:a2e698a0d98df13550a866aa27f6bba490c89da13b2b4d052336fac9aed1d807

Observation 396a44e8-1fbe-40bc-aa5f-f2aea7f6fb39 · inbound

MemEvolve: Meta-Evolution of Agent Memory Systems cites this paper.

MemEvolve: Meta-Evolution of Agent Memory Systems Humanity's Last Exam

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-15T22:18:15.256151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-15T22:18:15.149975Z digest=sha256:640b21f432779bf5bf369a71be38499e9bb628e50dcffe166b98f84a23a358ba

Observation 400c1496-485a-4561-94f6-350b98b17048 · inbound

MiMo-V2-Flash Technical Report cites this paper.

MiMo-V2-Flash Technical Report Humanity's Last Exam

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-12T11:33:32.866723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-12T11:33:32.568261Z digest=sha256:8e3d1c61c4b428b2f909a679b3b254a0e62ba18da2cbcb301f972728ed0ed22a

Observation 3f2b8049-4646-4474-83a5-d03314efe5c6 · inbound

Agentic Search in the Wild: Intents and Trajectory Dynamics from 14M+ Real Search Requests cites this paper.

Agentic Search in the Wild: Intents and Trajectory Dynamics from 14M+ Real Search Requests Humanity's Last Exam

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-16T11:20:53.136392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-16T11:20:20.651794Z digest=sha256:d885e2f834368d0aa952c0c2c01ce9eec876801fa316ba47eee7e53f9308108f

Observation d7332dd0-10eb-4179-969a-cf055cd9af95 · inbound

Do Not Waste Your Rollouts: Recycling Search Experience for Efficient Test-Time Scaling cites this paper.

Do Not Waste Your Rollouts: Recycling Search Experience for Efficient Test-Time Scaling Humanity's Last Exam

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-16T09:57:43.003924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-16T09:53:14.045466Z digest=sha256:01823b79d69893da68d0dc32acb1fca6463bdd91dca12a4513668753a6308308

Observation 5943b713-9af5-4807-891f-368815df5bb8 · inbound

Kimi K2.5: Visual Agentic Intelligence cites this paper.

Kimi K2.5: Visual Agentic Intelligence Humanity's Last Exam

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-10T18:40:50.715344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T16:09:05.225767Z digest=sha256:bf4ad22a923d340603b41dacc312b4d900c55aba9611a8e501ccb6a05df26bda

Observation 469ac4cb-d377-4e6b-ade6-4e3dc6dc74f0 · inbound

Do MLLMs Really Understand Space? A Mathematical Reasoning Evaluation cites this paper.

Do MLLMs Really Understand Space? A Mathematical Reasoning Evaluation Humanity's Last Exam

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-16T03:40:33.102576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-16T03:39:28.364183Z digest=sha256:4fb804f2b2dd750a92bfe81b04021bf55c9fa6231da1c367a0dd2e6b3a9d8e66

Observation dc8d06e3-b9a4-4fa2-8689-dd5f23de5987 · inbound

MedXIAOHE: A Comprehensive Recipe for Building Medical MLLMs cites this paper.

MedXIAOHE: A Comprehensive Recipe for Building Medical MLLMs Humanity's Last Exam

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-05-15T22:56:50.495675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-15T22:52:30.992054Z digest=sha256:d0881d6b55da8b809a76c36241cd431435880790ba20a8b6c6c8025f1c7d2437

Observation 8a56c459-cc88-4583-b9f6-47b17caac1b7 · inbound

GLM-5: from Vibe Coding to Agentic Engineering cites this paper.

GLM-5: from Vibe Coding to Agentic Engineering Humanity's Last Exam

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-11T05:46:41.170275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T05:46:40.836161Z digest=sha256:31fc9c1c1e54036250b7cae316b8930ecef6de85c078c3448219be0135f3c027

Observation e3d50908-f29f-4b89-8587-7e65f570c4f0 · inbound

From Human-Level AI Tales to AI Leveling Human Scales cites this paper.

From Human-Level AI Tales to AI Leveling Human Scales Humanity's Last Exam

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-15T20:10:18.091199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-15T20:08:29.604466Z digest=sha256:7be12c12b6c55fccb66cbe8c3c13c30c622cd952fadaeb838869aaca4a2bca9a

Observation 20a09fb7-cd91-443d-bab2-0c86289d43ad · inbound

Multistage Stochastic Programming for Rare Event Risk Mitigation in Power Systems Management cites this paper.

Multistage Stochastic Programming for Rare Event Risk Mitigation in Power Systems Management Humanity's Last Exam

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-15T15:03:20.637041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T15:03:20.637041Z digest=sha256:d4b13b33118729a0e6588867a071985093de3f3697d41f0a7f70b12621774a9f

Observation bc226aca-a8bd-45fd-8edf-07d89701d35c · inbound

Seed1.8 Model Card: Towards Generalized Real-World Agency cites this paper.

Seed1.8 Model Card: Towards Generalized Real-World Agency Humanity's Last Exam

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-05-15T07:45:14.348947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-15T07:44:02.827006Z digest=sha256:1af455c5477008996db69a42896abacbfe3a399b06fca6d46cb23ac1be9fddb1

Observation 94e3a864-539a-4580-8588-57f901d818db · inbound

The limits of bio-molecular modeling with large language models : a cross-scale evaluation cites this paper.

The limits of bio-molecular modeling with large language models : a cross-scale evaluation Humanity's Last Exam

Reference 25

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T20:13:13.871091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T20:09:30.460861Z digest=sha256:cde710b0a2e73c5f7a3590db99b423f9716161f739d430ed1298e326303e1aef

Observation 5acdd8f7-f516-4d95-8ba2-6425adb32b92 · inbound

Representational Collapse in Multi-Agent LLM Committees: Measurement and Diversity-Aware Consensus cites this paper.

Representational Collapse in Multi-Agent LLM Committees: Measurement and Diversity-Aware Consensus Humanity's Last Exam

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-13T18:03:04.238800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-07-11T11:50:26.030339Z digest=sha256:8ef8305eff86b8201ae9e53f773330f47289fac725628bbc0aa351315726f65f

Observation 3470c507-64f5-4e7b-91bd-ab11fbc8ea69 · inbound

GeoBrowse: A Geolocation Benchmark for Agentic Tool Use with Expert-Annotated Reasoning Traces cites this paper.

GeoBrowse: A Geolocation Benchmark for Agentic Tool Use with Expert-Annotated Reasoning Traces Humanity's Last Exam

Reference 39

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T17:33:02.351776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T17:31:08.575993Z digest=sha256:810d81e0d617393d4b8cca97a8a44a99f0ea4d9af2306f10c3e6d380a7825708

Observation 3e171097-db80-4e03-b37a-992df2e9163e · inbound

WebExpert: domain-aware web agents with critic-guided expert experience for high-precision search cites this paper.

WebExpert: domain-aware web agents with critic-guided expert experience for high-precision search Humanity's Last Exam

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:07:34.182793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-16T08:03:53.920339Z digest=sha256:9d284a22a8087fd97cc8e0bfbc26f9ca7d4a893851d4dcdac0e2314738802fec

Observation d786e55b-2582-459c-96ee-0842414b32a8 · inbound

Select-then-Solve: Paradigm Routing as Inference-Time Optimization for LLM Agents cites this paper.

Select-then-Solve: Paradigm Routing as Inference-Time Optimization for LLM Agents Humanity's Last Exam

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-11T00:25:50.918265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:32:37.703624Z digest=sha256:d2bc5341eb031e32d4251c9439f35346f6928c11f698f3a7c7407b1b90f1f1e8

Observation 68d96985-c399-419a-895e-96905da22bee · inbound

Towards Knowledgeable Deep Research: Framework and Benchmark cites this paper.

Towards Knowledgeable Deep Research: Framework and Benchmark Humanity's Last Exam

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T18:40:50.715344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:17:35.879705Z digest=sha256:4d630cb7fa9a8473097995464934a26201546f3c67a14012a05e0c81aa418100

Observation e6f9120b-8a70-437a-b054-dc36b2730da5 · inbound

PokeGym: A Visually-Driven Long-Horizon Benchmark for Vision-Language Models cites this paper.

PokeGym: A Visually-Driven Long-Horizon Benchmark for Vision-Language Models Humanity's Last Exam

Reference 6

Resolution
verified exact
doi, observed 2026-05-10T18:40:50.715344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T17:59:48.877783Z digest=sha256:bb8fb9750d5b69b224e8270dd10b6e7e6e952bc07b1bae39b94c926aa3b01081

Observation 37b3d874-5f83-45a8-b578-4440eba879b2 · inbound

Medical Reasoning with Large Language Models: A Survey and MR-Bench cites this paper.

Medical Reasoning with Large Language Models: A Survey and MR-Bench Humanity's Last Exam

Reference 108

Resolution
verified exact
local_arxiv, observed 2026-05-15T10:25:26.765112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-15T10:21:39.892271Z digest=sha256:d8267dca2d7336424988ccecf63e5f773cd8613c888341f14ffa10d847117838

Observation 6bf76044-d969-43db-bcbc-a4e11ea9710c · inbound

LABBench2: An Improved Benchmark for AI Systems Performing Biology Research cites this paper.

LABBench2: An Improved Benchmark for AI Systems Performing Biology Research Humanity's Last Exam

Reference 24

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T07:17:30.275383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-16T07:16:57.796927Z digest=sha256:d018aff80a9a39dbf6e1e13372d137a860e6d3964f408b5c79eb4a14e6cd6fd2

Observation b7fc2157-8cf1-4b95-9ec2-5600a1bc64de · inbound

SPEED-Bench: A Unified and Diverse Benchmark for Speculative Decoding cites this paper.

SPEED-Bench: A Unified and Diverse Benchmark for Speculative Decoding Humanity's Last Exam

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-16T03:17:12.502379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-16T03:13:37.202362Z digest=sha256:43806457b2e5d6498d514b3d888f69098f05708638335272c38fa0a0a0bc7cfe

Observation b4735671-518d-4406-855a-475182ed5ea6 · inbound

COMPOSITE-Stem cites this paper.

COMPOSITE-Stem Humanity's Last Exam

Reference 1

Resolution
metadata mismatch
doi, observed 2026-05-10T18:40:50.715344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-10T17:29:15.163053Z digest=sha256:218b8bc03f43680db1cfb3dea4972c3512eab1eec4a818403be95db86493112a

Observation 69eac277-e294-462c-bbae-ed0d45198427 · inbound

Frontier-Eng: Benchmarking Self-Evolving Agents on Real-World Engineering Tasks with Generative Optimization cites this paper.

Frontier-Eng: Benchmarking Self-Evolving Agents on Real-World Engineering Tasks with Generative Optimization Humanity's Last Exam

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T09:05:57.687364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T16:17:32.290531Z digest=sha256:03984a5eee9d51687797ea98d7d04eac7cc184ea1865ceedf2c0d6d7812ef3a5

Observation fccc1c0d-ae7f-4241-a74a-271ed97009a2 · inbound

PolyBench: Benchmarking LLM Forecasting and Trading Capabilities on Live Prediction Market Data cites this paper.

PolyBench: Benchmarking LLM Forecasting and Trading Capabilities on Live Prediction Market Data Humanity's Last Exam

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-13T19:18:09.294629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T19:17:22.608745Z digest=sha256:ee98df7ce4e0bfb857870566060d77e70fcff420e9fd6d8a1da350fe9ee02dd8

Observation 54cd7de7-1aa8-425c-b352-d2576954b524 · inbound

PRL-Bench: A Comprehensive Benchmark Evaluating LLMs' Capabilities in Frontier Physics Research cites this paper.

PRL-Bench: A Comprehensive Benchmark Evaluating LLMs' Capabilities in Frontier Physics Research Humanity's Last Exam

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-10T18:40:50.715344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T11:10:21.639856Z digest=sha256:d42084c5508ddc9e07bf69ef7c5f46897f4e79976faccf8471b4445ac1770658

Observation af119831-ac34-4ee4-ac46-1fc76a0aedf8 · inbound

Stargazer: A Scalable Model-Fitting Benchmark Environment for AI Agents under Astrophysical Constraints cites this paper.

Stargazer: A Scalable Model-Fitting Benchmark Environment for AI Agents under Astrophysical Constraints Humanity's Last Exam

Reference 2

Resolution
metadata mismatch
doi, observed 2026-05-10T18:40:50.715344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T09:11:56.768863Z digest=sha256:fce7ccd6240e08bc1dde1a215ef59941f2f90a54ea68a59e24016fe1be77cff6

Observation 2333c35a-667d-4fbc-ab7b-5c47f1e0544a · inbound

Stargazer: A Scalable Model-Fitting Benchmark Environment for AI Agents under Astrophysical Constraints cites this paper.

Stargazer: A Scalable Model-Fitting Benchmark Environment for AI Agents under Astrophysical Constraints Humanity's Last Exam

Reference 2

Resolution
metadata mismatch
doi, observed 2026-05-12T03:06:18.232999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-12T03:05:18.174474Z digest=sha256:43d155355109528b8227a37c4bc0e7ce8cbf3842be48c799dbccd7da803fce4d

Observation 01f1d2e7-604e-4169-b3a8-658a09b742e0 · inbound

neuralCAD-Edit: An Expert Benchmark for Multimodal-Instructed 3D CAD Model Editing cites this paper.

neuralCAD-Edit: An Expert Benchmark for Multimodal-Instructed 3D CAD Model Editing Humanity's Last Exam

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T18:40:50.715344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T09:10:57.869089Z digest=sha256:aa14fd3f6862ab71b0450182bb97ff6ff86b421026807609b60ea4056bf7d8fe

Observation 3be499db-7d7f-43e6-b15a-63a20a079131 · inbound

Federation over Text: Insight Sharing for Multi-Agent Reasoning cites this paper.

Federation over Text: Insight Sharing for Multi-Agent Reasoning Humanity's Last Exam

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-12T19:16:10.884254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T19:16:10.884254Z digest=sha256:8802b990a9520c0809cf51acf2d71d12dce3d83c7abcdd441e26b3cf6a23f77d

Observation 11ac7ecf-9467-4806-ba69-bc25174eae46 · inbound

EvoMaster: A Foundational Evolving Agent Framework for Agentic Science at Scale cites this paper.

EvoMaster: A Foundational Evolving Agent Framework for Agentic Science at Scale Humanity's Last Exam

Reference 12

Resolution
metadata mismatch
doi, observed 2026-05-10T18:40:50.715344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T05:59:01.010437Z digest=sha256:436d911cc3d7d26db8a93f2357a9d64ddc977e21100f9cf19fc7bab05b3a008c

Observation 0b4e4c8b-a014-443d-9a8e-a91490b8e86f · inbound

EvoMaster: A Foundational Evolving Agent Framework for Agentic Science at Scale cites this paper.

EvoMaster: A Foundational Evolving Agent Framework for Agentic Science at Scale Humanity's Last Exam

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-07-05T17:51:14.675306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:acf2690415da2a636a1fb6ed7ad03c54993a46a92e85e0f0b8894a060842868a

Observation f8be7b5c-3a17-4164-9b10-d179952b2019 · inbound

LiteResearcher: A Scalable Agentic RL Training Framework for Deep Research Agent cites this paper.

LiteResearcher: A Scalable Agentic RL Training Framework for Deep Research Agent Humanity's Last Exam

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-07-05T15:11:10.831345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-07-05T15:03:50.420072Z digest=sha256:e22a846455bc74a845aa14b9a65cf316beda035095852349253bc818b92643e5

Observation b7ed9a19-990e-4afa-bdf9-52efa2b4c056 · inbound

Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence cites this paper.

Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence Humanity's Last Exam

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-10T18:40:50.715344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T05:24:00.503836Z digest=sha256:a7c70e4da7183091796f4322812ff3dc3e0f51e7ae58e16a42be11a3aecc7c99

Observation 3af814b7-8cae-49f3-acba-a0627785ddcb · inbound

Wan-Image: Pushing the Boundaries of Generative Visual Intelligence cites this paper.

Wan-Image: Pushing the Boundaries of Generative Visual Intelligence Humanity's Last Exam

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-11T13:11:04.010005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T02:16:03.854650Z digest=sha256:d5c53c7770de359f3a02917ed9f976ae47f015f4172a49596dafff0029043927

Observation b1ba6cf3-7336-473e-9bc0-beaa13b7c57a · inbound

Super Apriel: One Checkpoint, Many Speeds cites this paper.

Super Apriel: One Checkpoint, Many Speeds Humanity's Last Exam

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-11T13:01:24.829456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-10T02:27:11.553553Z digest=sha256:0dcba33ee0144e6cc1ec0fa2c7999f4fa2087c1ef049538d3ced2e017337bf18

Observation 7cc4153a-ba0c-4dd7-adba-8a8f36478836 · inbound

ActuBench: A Multi-Agent LLM Pipeline for Generation and Evaluation of Actuarial Reasoning Tasks cites this paper.

ActuBench: A Multi-Agent LLM Pipeline for Generation and Evaluation of Actuarial Reasoning Tasks Humanity's Last Exam

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-11T13:46:05.225277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T00:35:24.397273Z digest=sha256:54a9edc44c20097652f75567d4496e32607d67dfd5b9ba88e8afdd9f6f21ce6c

Observation 5effcfa2-e5cf-488f-b863-ebec7be31306 · inbound

pAI/MSc: ML Theory Research with Humans on the Loop cites this paper.

pAI/MSc: ML Theory Research with Humans on the Loop Humanity's Last Exam

Reference 90

Resolution
verified exact
local_arxiv, observed 2026-05-11T13:51:03.142069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T00:00:02.883095Z digest=sha256:9bb4b1f7b88bf76f7abfd82db5f46067e03c5456825a30efd7b9c7c7043acbfb

Observation aa081844-6444-40e1-be0d-1614b1f5ab47 · inbound

Supplement Generation Training for Enhancing Agentic Task Performance cites this paper.

Supplement Generation Training for Enhancing Agentic Task Performance Humanity's Last Exam

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T18:40:50.715344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-10T00:58:27.655909Z digest=sha256:3e224e8f949ccc49ed2353541d4f8f00289198c47ccd286f8fa12f2db25a8dea

Observation a42f4c73-560f-44a9-a57b-35b187b674c7 · inbound

Large Language Models Decide Early and Explain Later cites this paper.

Large Language Models Decide Early and Explain Later Humanity's Last Exam

Reference 2

Resolution
metadata mismatch
doi, observed 2026-05-10T18:40:50.715344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-08T12:06:33.200365Z digest=sha256:ae69fd091051b9435408b4434754bdd3e4d50b176190405d66c420b86a6537e3

Observation 6b67bf1a-39d1-400f-a726-0712f74f82ff · inbound

Superminds Test: Actively Evaluating Collective Intelligence of Agent Society via Probing Agents cites this paper.

Superminds Test: Actively Evaluating Collective Intelligence of Agent Society via Probing Agents Humanity's Last Exam

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-11T19:26:08.755099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T12:00:07.345611Z digest=sha256:771d110a3015951fbe1456e460eb0851375b13f874e0a21d08576c5dd797c373

Observation a5a17ceb-94bf-42a0-9e24-cc92fad2ea69 · inbound

ProEval: Proactive Failure Discovery and Efficient Performance Estimation for Generative AI Evaluation cites this paper.

ProEval: Proactive Failure Discovery and Efficient Performance Estimation for Generative AI Evaluation Humanity's Last Exam

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-05-11T20:41:10.396791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:21:55.648930Z digest=sha256:5efc2d83a2790185a5bfc00c4273a1a29e08a3cc291153cfe2191b93de698b9d

Observation f8b53be1-b769-4bfa-bb38-c228e6c9029c · inbound

SciResearcher: Scaling Deep Research Agents for Frontier Scientific Reasoning cites this paper.

SciResearcher: Scaling Deep Research Agents for Frontier Scientific Reasoning Humanity's Last Exam

Reference 30

Resolution
metadata mismatch
doi, observed 2026-05-10T18:40:50.715344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-09T14:18:14.048230Z digest=sha256:64af12f4b144d8ba3578daecad83d50c4b3000c13d14c355412d5bf121ecf31b

Observation 2b57c698-7252-4632-9eba-b01b64b6f2fd · inbound

Cripping AI: Reimagining AI Through Lived Disability Experiences cites this paper.

Cripping AI: Reimagining AI Through Lived Disability Experiences Humanity's Last Exam

Reference 194

Resolution
verified exact
arxiv_id, observed 2026-05-10T18:40:50.715344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-08T19:00:14.226153Z digest=sha256:ab4450ad937ce1934c1e5b343ce5eae1a540d5cd743d9e9e9368df07fd236532

Observation 63681f3f-248a-41ff-b2eb-4d99ded906fc · inbound

Measuring AI Reasoning: A Guide for Researchers cites this paper.

Measuring AI Reasoning: A Guide for Researchers Humanity's Last Exam

Reference 136

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T18:40:50.715344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T18:53:18.586923Z digest=sha256:c1971f30e3d134d780cbb90195d81c1024b70ff50758327f2560f42c6828300d

Observation 0b5d8157-47eb-4945-859b-aadad2e54d3a · inbound

AcademiClaw: When Students Set Challenges for AI Agents cites this paper.

AcademiClaw: When Students Set Challenges for AI Agents Humanity's Last Exam

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-10T18:40:50.715344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T19:24:29.696454Z digest=sha256:783e462a9bdcea82a808233c7cce10ae1fc6eff7128738617ea153dfd871e989

Observation 7434b7e6-b1e7-4817-ac88-3af880204d94 · inbound

Reward Hacking Benchmark: Measuring Exploits in LLM Agents with Tool Use cites this paper.

Reward Hacking Benchmark: Measuring Exploits in LLM Agents with Tool Use Humanity's Last Exam

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-11T10:06:03.570514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-10T15:38:41.821264Z digest=sha256:d86d7ab3f1d85d3a3658d90bf7aa6b7282214dec8883949c45afa10b4d43412a

Observation a9e1de5f-f0ee-40ae-aa27-5b79aade00f5 · inbound

Toward Human-AI Complementarity Across Diverse Tasks cites this paper.

Toward Human-AI Complementarity Across Diverse Tasks Humanity's Last Exam

Reference 42

Resolution
verified exact
doi, observed 2026-05-10T18:40:50.715344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-10T15:57:52.262941Z digest=sha256:3539dbc05be83750a72a5e8148927846364d32ce2fc2971e2ab4d67560b50dfe

Observation 8f30cfd5-972a-491c-acfd-5c3507a7704e · inbound

Learning Agent Routing From Early Experience cites this paper.

Learning Agent Routing From Early Experience Humanity's Last Exam

Reference 52

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T04:35:57.568032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-11T01:15:07.381414Z digest=sha256:6a5c63445fc979def0798a33df1399830c4bec5ea5ce244b1702fb681940cbde

Observation e3f8eb24-3a03-40d3-9211-d39e8913fc96 · inbound

Memory-Efficient Looped Transformer: Decoupling Compute from Memory in Looped Language Models cites this paper.

Memory-Efficient Looped Transformer: Decoupling Compute from Memory in Looped Language Models Humanity's Last Exam

Reference 42

Resolution
metadata mismatch
doi, observed 2026-05-11T02:45:54.087333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T02:44:36.697450Z digest=sha256:f0814448292754e39e6f1de3d1e8000faadbca0ad4263ece43a3dc79eb8af8a7

Observation 0d29d7a9-bfd0-4fb9-9fcc-ec7a4e91f52c · inbound

Memory-Efficient Looped Transformer: Decoupling Compute from Memory in Looped Language Models cites this paper.

Memory-Efficient Looped Transformer: Decoupling Compute from Memory in Looped Language Models Humanity's Last Exam

Reference 42

Resolution
metadata mismatch
doi, observed 2026-05-20T23:03:49.766931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-20T23:01:17.957634Z digest=sha256:db1f12fb882af198af5b7f3247bca92543593bc45dab3443b2081e7f1248e489

Observation d3a59025-5c41-4773-80ee-1e1547439112 · inbound

Benchmarking EngGPT2-16B-A3B against Comparable Italian and International Open-source LLMs cites this paper.

Benchmarking EngGPT2-16B-A3B against Comparable Italian and International Open-source LLMs Humanity's Last Exam

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-05-11T03:40:53.432499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T03:40:04.692279Z digest=sha256:1ee089f64ac00941fe79a67d15479fad17c8c4846c9581a6e794b17b367627cc

Observation c79a6f68-3c9d-4ccd-adb0-1c0bc22e6918 · inbound

Benchmarking EngGPT2-16B-A3B against Comparable Italian and International Open-source LLMs cites this paper.

Benchmarking EngGPT2-16B-A3B against Comparable Italian and International Open-source LLMs Humanity's Last Exam

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-05-21T08:19:52.992780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-21T08:14:55.858466Z digest=sha256:61aac593af6bba0a319b09055dbd1558543718d3f40cb10f80b3b84111649a12

Observation b36eb6e8-871c-4944-be90-b689ee514ba5 · inbound

A Semantic-Sampling Framework for Evaluating Calibration in Open-Ended Question Answering cites this paper.

A Semantic-Sampling Framework for Evaluating Calibration in Open-Ended Question Answering Humanity's Last Exam

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:26:24.288692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-12T01:10:26.349351Z digest=sha256:d2053622903a30c9001fe38aad62584b15a3df4424cbfa35875427f72c7cf6c0

Observation e9d641e8-605a-49ec-813a-0a9f8f03cf34 · inbound

DiagnosticIQ: A Benchmark for LLM-Based Industrial Maintenance Action Recommendation from Symbolic Rules cites this paper.

DiagnosticIQ: A Benchmark for LLM-Based Industrial Maintenance Action Recommendation from Symbolic Rules Humanity's Last Exam

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:31:25.958831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-12T01:00:13.017290Z digest=sha256:f5c9cb44f4cbcd3a8a78e41bc20b84ab5ef5b933ca4fdb7f92c69549aff82645

Observation 6f579abd-e020-4a37-8fe1-b8e10ecfaac6 · inbound

EvoMAS: Learning Execution-Time Workflows for Multi-Agent Systems cites this paper.

EvoMAS: Learning Execution-Time Workflows for Multi-Agent Systems Humanity's Last Exam

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:36:33.120348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-12T02:29:55.683565Z digest=sha256:8231ab66c7f1ac2d02e7683339c156e0f9bf2bdf2238e5d00fbbb2a8ca0317d4

Observation 3ba90dca-fd89-452d-838f-6f1301c4465e · inbound

Soohak: A Mathematician-Curated Benchmark for Evaluating Research-level Math Capabilities of LLMs cites this paper.

Soohak: A Mathematician-Curated Benchmark for Evaluating Research-level Math Capabilities of LLMs Humanity's Last Exam

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:46:44.593173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-12T01:53:12.708706Z digest=sha256:ebc3e01b570004e638dfe9f718b071b003b04002043cdf44bc62fa53f6158437

Observation 3e7aab8a-d5e1-4247-bed9-e661859d4a38 · inbound

Soohak: A Mathematician-Curated Benchmark for Evaluating Research-level Math Capabilities of LLMs cites this paper.

Soohak: A Mathematician-Curated Benchmark for Evaluating Research-level Math Capabilities of LLMs Humanity's Last Exam

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-20T22:24:07.747968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-20T22:24:04.625823Z digest=sha256:55fea9d4dc1ef64e296049784c6d3b9578d430d198b139aeccc9bd2e3d55ba1f

Observation 72c12794-7a39-4bb2-a5e0-8bc4fbe92281 · inbound

LLM-Guided Monte Carlo Tree Search over Knowledge Graphs: Composing Mechanistic Explanations for Drug-Disease Pairs cites this paper.

LLM-Guided Monte Carlo Tree Search over Knowledge Graphs: Composing Mechanistic Explanations for Drug-Disease Pairs Humanity's Last Exam

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T03:06:18.250276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-12T03:04:57.158157Z digest=sha256:73cc53bb982f2473066ed76a8e76c356e2a7957ac5e04b08a46755bc4cd9eb11

Observation 11f613a6-aa79-4863-91bb-98a293df85e5 · inbound

Personalized Deep Research: A User-Centric Framework, Dataset, and Hybrid Evaluation for Knowledge Discovery cites this paper.

Personalized Deep Research: A User-Centric Framework, Dataset, and Hybrid Evaluation for Knowledge Discovery Humanity's Last Exam

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-12T05:46:31.365269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-12T04:55:22.596920Z digest=sha256:273f9108722cef5760072c3629be4a06bcf5623cee2a0a720f7db4f1a1c26dec

Observation 6db75031-179d-4e79-b952-ee2f4fc64bf0 · inbound

MaD Physics: Evaluating information seeking under constraints in physical environments cites this paper.

MaD Physics: Evaluating information seeking under constraints in physical environments Humanity's Last Exam

Reference 5

Resolution
metadata mismatch
doi, observed 2026-05-12T03:56:20.996631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-12T03:52:34.339865Z digest=sha256:a2619d3615aa271c6440d122ad42f9559ccc26fb6d36d0e19ccba921f845b799

Observation 77b78ed1-faa0-4140-afe4-2810514019e7 · inbound

Towards On-Policy Data Evolution for Visual-Native Multimodal Deep Search Agents cites this paper.

Towards On-Policy Data Evolution for Visual-Native Multimodal Deep Search Agents Humanity's Last Exam

Reference 3

Resolution
malformed identifier
local_arxiv, observed 2026-05-12T06:26:26.384258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-12T04:15:40.042348Z digest=sha256:a9ad1445db9b75ad1324c4f11038588dc52ba43453eac3673f7cdc052dac88f8

Observation 49549d33-548b-4285-bf2a-90fe7c53a390 · inbound

Towards On-Policy Data Evolution for Visual-Native Multimodal Deep Search Agents cites this paper.

Towards On-Policy Data Evolution for Visual-Native Multimodal Deep Search Agents Humanity's Last Exam

Reference 3

Resolution
malformed identifier
local_arxiv, observed 2026-07-01T13:55:46.395938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-06-30T22:30:28.649803Z digest=sha256:89e69ba11031e9f072cf29ab8bdeacbd0fc7f0d9c05963d94cb06686e1ac2ded

Observation de75815d-0b53-44b3-97cd-1eeabc6679eb · inbound

The Generalized Turing Test: A Foundation for Comparing Intelligence cites this paper.

The Generalized Turing Test: A Foundation for Comparing Intelligence Humanity's Last Exam

Reference 4

Resolution
metadata mismatch
doi, observed 2026-05-12T03:26:18.176183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-12T03:25:48.145137Z digest=sha256:98af6387305716e638f033917a68fb1358d32bdd370a29c1a5a54871696b1b09

Observation 2f24f210-6d5b-41f8-a315-bf72e6d78cc5 · inbound

AssayBench: An Assay-Level Virtual Cell Benchmark for LLMs and Agents cites this paper.

AssayBench: An Assay-Level Virtual Cell Benchmark for LLMs and Agents Humanity's Last Exam

Reference 56

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T05:56:24.177352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-12T04:48:54.032360Z digest=sha256:2d0cf020def255901b90616836ebedfd35c6820bda2de5c47207b6b17e612f44

Observation 802fe269-f6f9-4b27-a61c-6c047c0beab2 · inbound

Instructions Shape Production of Language, not Processing cites this paper.

Instructions Shape Production of Language, not Processing Humanity's Last Exam

Reference 188

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T03:12:09.408578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-13T03:09:02.902912Z digest=sha256:3f1fecbcca0b7fdbe38c50343db3b9c809a9d332c5aa18c6970f0ed0193168bf

Observation a9255eb0-954e-4891-91c1-d4a756e89fb9 · inbound

Instructions Shape Production of Language, not Processing cites this paper.

Instructions Shape Production of Language, not Processing Humanity's Last Exam

Reference 188

Resolution
metadata mismatch
local_arxiv, observed 2026-05-14T21:02:58.080470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-14T21:02:02.135970Z digest=sha256:6c53b7e69f1e072f46c3c9464f99c1b4c43b86ec52603515bb73dbf12c115b9d

Observation b668be92-2506-4068-b4b7-1e5e4edfa892 · inbound

Measuring Five-Nines Reliability: Sample-Efficient LLM Evaluation in Saturated Benchmarks cites this paper.

Measuring Five-Nines Reliability: Sample-Efficient LLM Evaluation in Saturated Benchmarks Humanity's Last Exam

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-13T03:07:08.632489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-13T03:06:41.504993Z digest=sha256:866dc06aee43925a81956b42587ae72772b635102e06a53bf055b7d2bdc078c4

Observation 8cbd53d9-a058-405e-97a9-a7c7c768d7ad · inbound

Formal Conjectures: An Open and Evolving Benchmark for Verified Discovery in Mathematics cites this paper.

Formal Conjectures: An Open and Evolving Benchmark for Verified Discovery in Mathematics Humanity's Last Exam

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-05-14T20:22:54.143167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:2d6aaecb0fb13ea99b87052c78180657c0acec31298d586bb7dd7fc9aea7b897

Observation 01856572-01d8-457a-bd13-3b047a257cc3 · inbound

TRIAGE: Evaluating Prospective Metacognitive Control in LLMs under Resource Constraints cites this paper.

TRIAGE: Evaluating Prospective Metacognitive Control in LLMs under Resource Constraints Humanity's Last Exam

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-05-14T19:29:23.875168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-14T19:28:25.525889Z digest=sha256:9ff5940ba2423d79ea0ac163e2f9bc144e66fd35d0771c8693be381a0e4b13d2

Observation d829a676-ab20-4782-9ce7-cd8f6e39b53d · inbound

Unsteady Metrics and Benchmarking Cultures of AI Model Builders cites this paper.

Unsteady Metrics and Benchmarking Cultures of AI Model Builders Humanity's Last Exam

Reference 49

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T04:55:01.570219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-15T04:54:26.888562Z digest=sha256:4a2cb69d6f2edb471be9fe54538835837a9d1e2babd354f7b50b53eba24b7132

Observation 4297c688-9dfc-4cce-9e77-1e041df101e2 · inbound

Boosting Reinforcement Learning with Verifiable Rewards via Randomly Selected Few-Shot Guidance cites this paper.

Boosting Reinforcement Learning with Verifiable Rewards via Randomly Selected Few-Shot Guidance Humanity's Last Exam

Reference 67

Resolution
verified exact
local_arxiv, observed 2026-05-15T03:19:43.088196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-15T03:18:26.590871Z digest=sha256:e3feca7182b74b6432c99b06a99a7b860cbf6121eef6168c1faaa2c0181fb7ef

Observation 11863d6f-f31c-4f39-aa29-3a28ef9139af · inbound

OpenDeepThink: Parallel Reasoning via Bradley-Terry Aggregation cites this paper.

OpenDeepThink: Parallel Reasoning via Bradley-Terry Aggregation Humanity's Last Exam

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-15T03:08:58.426763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-15T03:06:54.507594Z digest=sha256:6a18fca0068904717b769eb4b1480fcbc5005de86cd9a61fc6f52f0704e8f083

Observation c5f9caa1-08a5-45a7-ac7d-c4384980924f · inbound

OpenDeepThink: Parallel Reasoning via Bradley-Terry Aggregation cites this paper.

OpenDeepThink: Parallel Reasoning via Bradley-Terry Aggregation Humanity's Last Exam

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-20T20:53:43.646581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-20T20:51:47.393589Z digest=sha256:08dbef9e60afa37b0982dcd8afbf14acfb702ecdb6f4a34eb96bb502cd7fd2dd

Observation 7528d9cb-2937-41a1-b004-bf40bf8a8085 · inbound

Argus: Evidence Assembly for Scalable Deep Research Agents cites this paper.

Argus: Evidence Assembly for Scalable Deep Research Agents Humanity's Last Exam

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-20T18:43:38.705246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-20T18:39:24.925753Z digest=sha256:e66ba7ef77e41086740fb7b92ee61dd9012057db22065c0fd3e59686c1870160

Observation 38c3bfc5-04db-426f-af04-75c37f12efca · inbound

Argus: Evidence Assembly for Scalable Deep Research Agents cites this paper.

Argus: Evidence Assembly for Scalable Deep Research Agents Humanity's Last Exam

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:44:02.967776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-21T07:42:24.148398Z digest=sha256:d22e614f9284a0c718506a33fe80c039fa16a13cd9d746bc33551c25557a5d78

Observation f452c0fb-200c-4991-8f43-99bb544f3378 · inbound

Customizing an LLM for Enterprise Software Engineering cites this paper.

Customizing an LLM for Enterprise Software Engineering Humanity's Last Exam

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-20T16:23:35.630470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-20T16:18:44.870072Z digest=sha256:1db793033b1e33f661fdc514c358c547fdfabdc46204a7d0ab586ae1fbfdf5f0

Observation 232e5aeb-d6af-48c3-8cd5-039af4234293 · inbound

Customizing an LLM for Enterprise Software Engineering cites this paper.

Customizing an LLM for Enterprise Software Engineering Humanity's Last Exam

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-21T08:34:05.918381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-21T08:31:00.191393Z digest=sha256:82e9fc718c1aaa7dd5c76036ebcf1ad02488fd482b94b6a0c452279f926fd7b8

Observation a589650e-7a7d-4ede-9639-84d4ae5c1266 · inbound

Evaluating Cognitive Age Alignment in Interactive AI Agents cites this paper.

Evaluating Cognitive Age Alignment in Interactive AI Agents Humanity's Last Exam

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-20T10:38:12.569905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-20T10:35:57.643385Z digest=sha256:6fbbd7be93c3a79f7e1b06231485ad55896922057ff9c86ffb7f760012c2d8d5

Observation ed025b07-a591-4210-84c6-6285a05749f0 · inbound

Forecasting Downstream Performance of LLMs With Proxy Metrics cites this paper.

Forecasting Downstream Performance of LLMs With Proxy Metrics Humanity's Last Exam

Reference 75

Resolution
metadata mismatch
doi, observed 2026-05-20T10:33:12.048459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-20T10:30:20.575552Z digest=sha256:0a905409ab83568eb1cdcf060e2e25f5d4325580f57c188451fca780dc72ba51

Observation 1003b0c1-c8da-475b-a6cc-76b97a292cfd · inbound

OpenCompass: A Universal Evaluation Platform for Large Language Models cites this paper.

OpenCompass: A Universal Evaluation Platform for Large Language Models Humanity's Last Exam

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-20T06:18:05.134943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-20T06:17:28.043388Z digest=sha256:4a232b3cc035b068f73bf35e9d7a04e80a7742986d8f203d1bbe5fceef0ab0f8