Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-10T18:40:50.139345Z
Paper Citation Record · LEDGER
As of 1 August 2026, this Paper Citation Record lists 100 of 300 outbound references and 100 inbound Pith citation observations for arXiv:2501.14249.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-10T18:40:50.139345Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-01T06:32:01.292127+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-15T15:03:20.637041Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-10T09:37:00.704616Z
100 of 300 outbound references displayed
External citation measurements
8
pith, observed 2026-07-10T09:37:00.704616Z
Observation 0674c55d-a792-44bc-a066-f4f42c468349 · outbound
Humanity's Last Exam A BERT Baseline for the Natural Questions
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 27dd35ee-06db-478e-9197-f4b86b5f2a17 · outbound
Humanity's Last Exam AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 57b80c17-719f-4f8a-83ac-5665aa42fb40 · outbound
Humanity's Last Exam The claude 3 model family: Opus, sonnet, haiku
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation c2f01758-effb-4793-ac7c-c3ec81d7b1d5 · outbound
Humanity's Last Exam Model card addendum: Claude 3.5 haiku and upgraded claude 3.5 son- net
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 4ccaf35b-d213-41e4-b046-a8fe382ced6f · outbound
Humanity's Last Exam Responsible scaling policy updates
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 90da4a4c-43cb-44d3-8020-bf418ed27da1 · outbound
Humanity's Last Exam HealthBench: Evaluating Large Language Models Towards Improved Human Health
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 3886ecc9-1fa3-4ed3-a3a2-17f45ef1e50f · outbound
Humanity's Last Exam Austin, A
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 4820e0a9-d957-4348-afc6-fa61d76d7bcf · outbound
Humanity's Last Exam Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 4af1527c-2c4c-4b65-b9d4-9fd5279242a8 · outbound
Humanity's Last Exam MS MARCO: A Human Generated MAchine Reading COmprehension Dataset
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation f303afb9-da19-4460-9313-d870f3e455bd · outbound
Humanity's Last Exam Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 2c8a8636-7558-4f7c-b51b-62ff7e988de6 · outbound
Humanity's Last Exam MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 0d5be82c-f123-468f-8f98-fa312324b292 · outbound
Humanity's Last Exam Evaluating Large Language Models Trained on Code
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T08:08:23.404839+00:00.
Observation d8d0bcd6-df78-4e00-95f3-09e4107a34c8 · outbound
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 87080e3d-765c-475e-936f-4c4b66546c95 · outbound
Humanity's Last Exam Training Verifiers to Solve Math Word Problems
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 83ae0e2b-70ba-43c1-bb42-936cc43057af · outbound
Humanity's Last Exam Deepseek-v3 technical report
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation d8ca28f3-9b8a-46a8-a159-20dc3426d302 · outbound
Humanity's Last Exam Unresolved cited work
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 92cbd9cc-d252-4f4d-b3c5-dbf31c438bf0 · outbound
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 6e914fe3-e237-4b2a-9699-84239dbf0294 · outbound
Humanity's Last Exam Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 72a19ef2-e197-4494-ac82-e0e76871d46d · outbound
Humanity's Last Exam FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 79639eed-ebbe-4204-89d1-c7e00761b6df · outbound
Humanity's Last Exam OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 1347b7c1-7c61-480c-a52d-d73e92e1782b · outbound
Humanity's Last Exam Measuring Coding Challenge Competence With APPS
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation f8d191e4-a193-4e65-b9bd-e12827d9ed1a · outbound
Humanity's Last Exam Measuring Massive Multitask Language Understanding
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation d9e7ea17-ef7b-4f03-9440-809463a40b70 · outbound
Humanity's Last Exam Hendrycks, C
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation e9afb5ec-2d3f-45e4-a87a-855acb944474 · outbound
Humanity's Last Exam PixMix: Dreamlike Pictures Comprehensively Improve Safety Measures
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 440d44c2-14e0-496c-8f9a-7cea5b6dcb63 · outbound
Humanity's Last Exam Hosseini, A
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 6dcaa95d-167d-409f-b60e-8f373bfa74e2 · outbound
Humanity's Last Exam Not All LLM Reasoners Are Created Equal
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 2aed4d3f-5445-4a3e-ba59-b4571a683f11 · outbound
Humanity's Last Exam Jacovi, A
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation e55c4d1a-215b-4cfa-9403-99d879450489 · outbound
Humanity's Last Exam SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 39a34301-2c01-4c36-8d7d-52060d5d47d8 · outbound
Humanity's Last Exam Dynabench: Rethinking Benchmarking in NLP
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 1f5deb70-d016-4da3-993c-b3740f09e8d6 · outbound
Humanity's Last Exam Refusal-Trained LLMs Are Easily Jailbroken As Browser Agents
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 58c43289-107a-469a-854c-8b5c5a5ddc88 · outbound
Humanity's Last Exam Unresolved cited work
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 0f5c6172-edb0-46f3-858b-4a2aaeea5411 · outbound
Humanity's Last Exam LAB-Bench: Measuring Capabilities of Language Models for Biology Research
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 1c3a6b99-f76a-4bda-85c9-11dda50e1b5a · outbound
Humanity's Last Exam The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 4607f30d-4d4b-45f4-8578-92e51303e4cb · outbound
Humanity's Last Exam MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 2f38d407-5bf1-4fe6-ad94-b89fce08349a · outbound
Humanity's Last Exam Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation c8b692e0-1fa8-4884-af36-e51b1e4e1f2a · outbound
Humanity's Last Exam Adversarial NLI: A New Benchmark for Natural Language Understanding
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 0ed849cd-c769-4fbf-b734-9339adf4a1f2 · outbound
Humanity's Last Exam Openai o1 system card
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 6440a1fa-822a-41db-9817-b1c3db2063ba · outbound
Humanity's Last Exam Openai and los alamos national laboratory announce bio- science research partnership
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 1c45bb77-14a1-41fc-9566-85874a56befb · outbound
Humanity's Last Exam Introducing swe-bench verified
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 1411e5e9-95db-4ac6-9915-17da206b6e4e · outbound
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 883d4f43-ce39-4518-8d35-8861439d16cd · outbound
Humanity's Last Exam Unresolved cited work
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 919639a0-3045-4e79-bedf-518059ad8d24 · outbound
Humanity's Last Exam How predictable is language model benchmark performance?
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation f6c7dc91-efe5-455d-8128-3832e33e9a07 · outbound
Humanity's Last Exam Discovering Language Model Behaviors with Model-Written Evaluations
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 2dc370b6-c932-4027-b9fe-d0e494d97546 · outbound
Humanity's Last Exam Phuong, M
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 2b4d9014-65cb-4474-9f90-daba5bb4bdf7 · outbound
Humanity's Last Exam SQuAD: 100,000+ Questions for Machine Comprehension of Text
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation ace79c53-f700-4255-9305-baf66308ab7c · outbound
Humanity's Last Exam Know What You Don't Know: Unanswerable Questions for SQuAD
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 8cbf44d9-0a14-4efa-a824-a56615833a89 · outbound
Humanity's Last Exam GPQA: A Graduate-Level Google-Proof Q&A Benchmark
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation a10528b8-6a51-45db-8ac9-59197fac8ec9 · outbound
Humanity's Last Exam Singhal, S
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 1e78985d-f07d-4eeb-936d-5bd115cf80ff · outbound
Humanity's Last Exam Skarlinski, J
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 2f050906-8ac6-4008-94d3-a4650ba2ced7 · outbound
Humanity's Last Exam Unresolved cited work
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 7ceabc31-aa91-46dd-97e1-43cef4c1a185 · outbound
Humanity's Last Exam Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 1655239c-4031-40f4-a026-1ef5bfe7dc84 · outbound
Humanity's Last Exam MMLU-Pro+: Evaluating Higher-Order Reasoning and Shortcut Learning in LLMs
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation b724b403-03ac-4247-9f7e-e0a31c941ccc · outbound
Humanity's Last Exam Team et al
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 100b26cf-83cd-4197-870f-95c5f9c1c371 · outbound
Humanity's Last Exam Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 78453e3f-994b-4be4-b14a-32a9eb819004 · outbound
Humanity's Last Exam PutnamBench: Evaluating Neural Theorem-Provers on the Putnam Mathematical Competition
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 2e4f46f9-7ba7-4cd0-b546-cdcc40773e21 · outbound
Humanity's Last Exam Unresolved cited work
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation beb2fdbc-5f20-49aa-8fbe-015686fdb728 · outbound
Humanity's Last Exam SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 06000166-b00c-4927-90f0-36eb17c4098f · outbound
Humanity's Last Exam MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation ea073248-6907-44e2-b96d-353a4f384ba4 · outbound
Humanity's Last Exam Measuring short-form factuality in large language models
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation d9fa4ac5-63e7-45de-986b-c9ba862334fb · outbound
Humanity's Last Exam RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation bf768341-2706-4c8f-8444-c2761bda2e09 · outbound
Humanity's Last Exam Grok-2 beta release
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 3005621b-41d3-4fdb-8966-ea498356f3ba · outbound
Humanity's Last Exam Unresolved cited work
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 92015453-8c83-49d8-9571-929546769cf9 · outbound
Humanity's Last Exam HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 40150a9d-5215-424a-8158-f50db7e7c045 · outbound
Humanity's Last Exam $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation be75ed09-53c6-4815-a5d7-86f23bc03ac8 · outbound
Humanity's Last Exam Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation d6a8eaea-9549-4d9b-a2f0-5d5f2227275f · outbound
Humanity's Last Exam AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation f9e68d2c-f3a3-4b32-952e-fe8202b07ea9 · outbound
Humanity's Last Exam Unresolved cited work
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 115fe9a7-59cc-4847-ba18-a2aa615b6dec · outbound
Humanity's Last Exam Unresolved cited work
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 6aef037c-2dba-4df3-97a5-5b4830dd2665 · outbound
Humanity's Last Exam Unresolved cited work
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 2af478b3-01e0-4aaa-8a60-5ebfe48137eb · outbound
Humanity's Last Exam Unresolved cited work
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 834aafa8-ba35-4c2f-b1d2-4a4ba3319c90 · outbound
Humanity's Last Exam Unresolved cited work
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation dd9b55ed-9998-498c-97a5-fb71f6d44acb · outbound
Humanity's Last Exam Unresolved cited work
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation ab2ab1e6-be25-4d32-87ee-dccf527b6e6e · outbound
Humanity's Last Exam Unresolved cited work
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 8d4fffab-453f-4c3f-829d-d219399d42fa · outbound
Humanity's Last Exam Unresolved cited work
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation a2261d8d-fc99-4845-8b19-9bea067d4bb7 · outbound
Humanity's Last Exam Unresolved cited work
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation c88780f8-9662-4b0c-b2f7-b921a2f29254 · outbound
Humanity's Last Exam Unresolved cited work
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 1e93a7a6-3cc9-4183-a2ba-28cf482cf21e · outbound
Humanity's Last Exam Unresolved cited work
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation b4dbd5f2-452f-4626-828f-17493f851d24 · outbound
Humanity's Last Exam Unresolved cited work
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation bbbfc201-a98c-48d1-aa5f-e51059b4b3db · outbound
Humanity's Last Exam Unresolved cited work
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation b48e72fc-f145-46ce-96f2-4c8ebd11de22 · outbound
Humanity's Last Exam Unresolved cited work
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 0097681b-2932-42f7-995d-2a6616449c0b · outbound
Humanity's Last Exam Unresolved cited work
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 3a4768da-b57f-47fd-a87e-f48968b4ac67 · outbound
Humanity's Last Exam Unresolved cited work
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation a7cfb8fe-c67f-4928-b8e2-af431192b3e2 · outbound
Humanity's Last Exam Unresolved cited work
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 21f33310-3f1d-4e2c-8e15-d8ea4c91ae14 · outbound
Humanity's Last Exam Unresolved cited work
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation fcb466d3-ddf7-4d30-9dd0-297d55545e40 · outbound
Humanity's Last Exam Unresolved cited work
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 394f059a-487f-4f9d-8715-df14cc106321 · outbound
Humanity's Last Exam Unresolved cited work
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 8cf3a8e3-eb9c-4c15-860d-08292da12896 · outbound
Humanity's Last Exam Unresolved cited work
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation ae511d0e-8e38-4ad8-9c92-e73ed437e444 · outbound
Humanity's Last Exam Unresolved cited work
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 170224d3-f45b-4841-98a1-186e993c2f9a · outbound
Humanity's Last Exam Unresolved cited work
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 680ba952-a6e3-4d67-9eb3-718225dd72fe · outbound
Humanity's Last Exam Unresolved cited work
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation a1324879-39d5-4d91-97d8-aa0e1e480e6b · outbound
Humanity's Last Exam Unresolved cited work
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation a8a39b86-2224-4d6a-b129-b1511f51d834 · outbound
Humanity's Last Exam Unresolved cited work
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation b086cac7-3d11-4fdf-9f13-2754e58e3e11 · outbound
Humanity's Last Exam Unresolved cited work
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation b190e4a7-dcf2-4edf-8b21-0e22de8f6eea · outbound
Humanity's Last Exam Unresolved cited work
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 11a7faac-286b-4157-98a0-e2c8fcea1fb3 · outbound
Humanity's Last Exam Unresolved cited work
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 01972e42-2b52-4802-b974-2482f5f47abc · outbound
Humanity's Last Exam Unresolved cited work
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation a157bdc3-9147-4bed-8c29-9d3934faad5a · outbound
Humanity's Last Exam Unresolved cited work
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation f2820dd9-59e5-4cfe-9dc5-5b007e13dc0d · outbound
Humanity's Last Exam Unresolved cited work
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation c469d83d-50f6-44dc-b9aa-e89215019bbc · outbound
Humanity's Last Exam Unresolved cited work
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation f6a39c93-a05c-4be0-82de-5429e7b21b92 · outbound
Humanity's Last Exam Unresolved cited work
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation c5c67edb-4092-48d5-af5f-889782e8cbed · inbound
R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization Humanity's Last Exam
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 0b24b582-3f80-4e6c-9ad0-ce2ad16b1162 · inbound
From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review Humanity's Last Exam
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 2f88b2bd-086c-4bb1-adac-c01e782a700c · inbound
WebThinker: Empowering Large Reasoning Models with Deep Research Capability Humanity's Last Exam
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 2780a976-73ef-415f-a0c0-18e83e4696a3 · inbound
LLMs Get Lost In Multi-Turn Conversation Humanity's Last Exam
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation acb0b9af-d810-4ad6-9010-ae68311b7a2c · inbound
DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents Humanity's Last Exam
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 1c9fa2bd-f427-4dbb-a00c-f6156be8024b · inbound
MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention Humanity's Last Exam
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 235c6809-5865-40dd-a270-382bd15e9b4d · inbound
Do LLMs Overthink Basic Math Reasoning? Benchmarking the Accuracy-Efficiency Tradeoff in Language Models Humanity's Last Exam
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation e250b5d9-6323-4cc7-bfae-2bf1c5055381 · inbound
Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities Humanity's Last Exam
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation cade9202-7510-4f7a-860d-f0dd0926abf3 · inbound
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 357233fc-e9f1-40f4-a029-ff02ee2e2435 · inbound
WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent Humanity's Last Exam
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation eb428d9d-dd67-4c32-a37f-cc4792bdd90d · inbound
GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models Humanity's Last Exam
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation af23c177-43d3-434b-a80c-06cc5d34421c · inbound
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 6b57a97e-a594-484a-b4a4-d559c80f503a · inbound
Why Johnny Can't Use Agents: Industry Aspirations vs. User Realities with AI Agents Humanity's Last Exam
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation a83b9adb-8aba-484e-82dc-f993a61451a9 · inbound
Probing the Critical Point (CritPt) of AI Reasoning: a Frontier Physics Research Benchmark Humanity's Last Exam
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 6569d432-1366-4b81-93d8-84a3215b0c54 · inbound
Scaling Latent Reasoning via Looped Language Models Humanity's Last Exam
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation d750ec29-25e5-4091-89f5-59367f5ebebf · inbound
MiroThinker: Pushing the Performance Boundaries of Open-Source Research Agents via Model, Context, and Interactive Scaling Humanity's Last Exam
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation ef2627d0-f6a3-4118-b287-6c71d3802a02 · inbound
DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models Humanity's Last Exam
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 6a0b5123-7090-47a9-97df-330e23b7fb72 · inbound
DVPO: Distributional Value Modeling-based Policy Optimization for LLM Post-Training Humanity's Last Exam
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 86e57c20-5c9b-4cd3-b69f-1291e91d39c0 · inbound
Asynchronous Reasoning: Training-Free Interactive Thinking LLMs Humanity's Last Exam
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation caa43568-44e2-46c8-a1c4-f85b36185768 · inbound
Evaluating Large Language Models in Scientific Discovery Humanity's Last Exam
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 396a44e8-1fbe-40bc-aa5f-f2aea7f6fb39 · inbound
MemEvolve: Meta-Evolution of Agent Memory Systems Humanity's Last Exam
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 400c1496-485a-4561-94f6-350b98b17048 · inbound
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 3f2b8049-4646-4474-83a5-d03314efe5c6 · inbound
Agentic Search in the Wild: Intents and Trajectory Dynamics from 14M+ Real Search Requests Humanity's Last Exam
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation d7332dd0-10eb-4179-969a-cf055cd9af95 · inbound
Do Not Waste Your Rollouts: Recycling Search Experience for Efficient Test-Time Scaling Humanity's Last Exam
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 5943b713-9af5-4807-891f-368815df5bb8 · inbound
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 469ac4cb-d377-4e6b-ade6-4e3dc6dc74f0 · inbound
Do MLLMs Really Understand Space? A Mathematical Reasoning Evaluation Humanity's Last Exam
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation dc8d06e3-b9a4-4fa2-8689-dd5f23de5987 · inbound
MedXIAOHE: A Comprehensive Recipe for Building Medical MLLMs Humanity's Last Exam
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 8a56c459-cc88-4583-b9f6-47b17caac1b7 · inbound
GLM-5: from Vibe Coding to Agentic Engineering Humanity's Last Exam
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation e3d50908-f29f-4b89-8587-7e65f570c4f0 · inbound
From Human-Level AI Tales to AI Leveling Human Scales Humanity's Last Exam
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 20a09fb7-cd91-443d-bab2-0c86289d43ad · inbound
Multistage Stochastic Programming for Rare Event Risk Mitigation in Power Systems Management Humanity's Last Exam
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc226aca-a8bd-45fd-8edf-07d89701d35c · inbound
Seed1.8 Model Card: Towards Generalized Real-World Agency Humanity's Last Exam
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 94e3a864-539a-4580-8588-57f901d818db · inbound
The limits of bio-molecular modeling with large language models : a cross-scale evaluation Humanity's Last Exam
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 5acdd8f7-f516-4d95-8ba2-6425adb32b92 · inbound
Representational Collapse in Multi-Agent LLM Committees: Measurement and Diversity-Aware Consensus Humanity's Last Exam
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 3470c507-64f5-4e7b-91bd-ab11fbc8ea69 · inbound
GeoBrowse: A Geolocation Benchmark for Agentic Tool Use with Expert-Annotated Reasoning Traces Humanity's Last Exam
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 3e171097-db80-4e03-b37a-992df2e9163e · inbound
WebExpert: domain-aware web agents with critic-guided expert experience for high-precision search Humanity's Last Exam
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation d786e55b-2582-459c-96ee-0842414b32a8 · inbound
Select-then-Solve: Paradigm Routing as Inference-Time Optimization for LLM Agents Humanity's Last Exam
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 68d96985-c399-419a-895e-96905da22bee · inbound
Towards Knowledgeable Deep Research: Framework and Benchmark Humanity's Last Exam
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation e6f9120b-8a70-437a-b054-dc36b2730da5 · inbound
PokeGym: A Visually-Driven Long-Horizon Benchmark for Vision-Language Models Humanity's Last Exam
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 37b3d874-5f83-45a8-b578-4440eba879b2 · inbound
Medical Reasoning with Large Language Models: A Survey and MR-Bench Humanity's Last Exam
Reference 108
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 6bf76044-d969-43db-bcbc-a4e11ea9710c · inbound
LABBench2: An Improved Benchmark for AI Systems Performing Biology Research Humanity's Last Exam
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation b7fc2157-8cf1-4b95-9ec2-5600a1bc64de · inbound
SPEED-Bench: A Unified and Diverse Benchmark for Speculative Decoding Humanity's Last Exam
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation b4735671-518d-4406-855a-475182ed5ea6 · inbound
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 69eac277-e294-462c-bbae-ed0d45198427 · inbound
Frontier-Eng: Benchmarking Self-Evolving Agents on Real-World Engineering Tasks with Generative Optimization Humanity's Last Exam
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation fccc1c0d-ae7f-4241-a74a-271ed97009a2 · inbound
PolyBench: Benchmarking LLM Forecasting and Trading Capabilities on Live Prediction Market Data Humanity's Last Exam
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 54cd7de7-1aa8-425c-b352-d2576954b524 · inbound
PRL-Bench: A Comprehensive Benchmark Evaluating LLMs' Capabilities in Frontier Physics Research Humanity's Last Exam
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation af119831-ac34-4ee4-ac46-1fc76a0aedf8 · inbound
Stargazer: A Scalable Model-Fitting Benchmark Environment for AI Agents under Astrophysical Constraints Humanity's Last Exam
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 2333c35a-667d-4fbc-ab7b-5c47f1e0544a · inbound
Stargazer: A Scalable Model-Fitting Benchmark Environment for AI Agents under Astrophysical Constraints Humanity's Last Exam
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 01f1d2e7-604e-4169-b3a8-658a09b742e0 · inbound
neuralCAD-Edit: An Expert Benchmark for Multimodal-Instructed 3D CAD Model Editing Humanity's Last Exam
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 3be499db-7d7f-43e6-b15a-63a20a079131 · inbound
Federation over Text: Insight Sharing for Multi-Agent Reasoning Humanity's Last Exam
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11ac7ecf-9467-4806-ba69-bc25174eae46 · inbound
EvoMaster: A Foundational Evolving Agent Framework for Agentic Science at Scale Humanity's Last Exam
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 0b4e4c8b-a014-443d-9a8e-a91490b8e86f · inbound
EvoMaster: A Foundational Evolving Agent Framework for Agentic Science at Scale Humanity's Last Exam
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation f8be7b5c-3a17-4164-9b10-d179952b2019 · inbound
LiteResearcher: A Scalable Agentic RL Training Framework for Deep Research Agent Humanity's Last Exam
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation b7ed9a19-990e-4afa-bdf9-52efa2b4c056 · inbound
Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence Humanity's Last Exam
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 3af814b7-8cae-49f3-acba-a0627785ddcb · inbound
Wan-Image: Pushing the Boundaries of Generative Visual Intelligence Humanity's Last Exam
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation b1ba6cf3-7336-473e-9bc0-beaa13b7c57a · inbound
Super Apriel: One Checkpoint, Many Speeds Humanity's Last Exam
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 7cc4153a-ba0c-4dd7-adba-8a8f36478836 · inbound
ActuBench: A Multi-Agent LLM Pipeline for Generation and Evaluation of Actuarial Reasoning Tasks Humanity's Last Exam
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 5effcfa2-e5cf-488f-b863-ebec7be31306 · inbound
pAI/MSc: ML Theory Research with Humans on the Loop Humanity's Last Exam
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation aa081844-6444-40e1-be0d-1614b1f5ab47 · inbound
Supplement Generation Training for Enhancing Agentic Task Performance Humanity's Last Exam
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation a42f4c73-560f-44a9-a57b-35b187b674c7 · inbound
Large Language Models Decide Early and Explain Later Humanity's Last Exam
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 6b67bf1a-39d1-400f-a726-0712f74f82ff · inbound
Superminds Test: Actively Evaluating Collective Intelligence of Agent Society via Probing Agents Humanity's Last Exam
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation a5a17ceb-94bf-42a0-9e24-cc92fad2ea69 · inbound
ProEval: Proactive Failure Discovery and Efficient Performance Estimation for Generative AI Evaluation Humanity's Last Exam
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation f8b53be1-b769-4bfa-bb38-c228e6c9029c · inbound
SciResearcher: Scaling Deep Research Agents for Frontier Scientific Reasoning Humanity's Last Exam
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 2b57c698-7252-4632-9eba-b01b64b6f2fd · inbound
Cripping AI: Reimagining AI Through Lived Disability Experiences Humanity's Last Exam
Reference 194
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 63681f3f-248a-41ff-b2eb-4d99ded906fc · inbound
Measuring AI Reasoning: A Guide for Researchers Humanity's Last Exam
Reference 136
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 0b5d8157-47eb-4945-859b-aadad2e54d3a · inbound
AcademiClaw: When Students Set Challenges for AI Agents Humanity's Last Exam
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 7434b7e6-b1e7-4817-ac88-3af880204d94 · inbound
Reward Hacking Benchmark: Measuring Exploits in LLM Agents with Tool Use Humanity's Last Exam
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation a9e1de5f-f0ee-40ae-aa27-5b79aade00f5 · inbound
Toward Human-AI Complementarity Across Diverse Tasks Humanity's Last Exam
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 8f30cfd5-972a-491c-acfd-5c3507a7704e · inbound
Learning Agent Routing From Early Experience Humanity's Last Exam
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation e3f8eb24-3a03-40d3-9211-d39e8913fc96 · inbound
Memory-Efficient Looped Transformer: Decoupling Compute from Memory in Looped Language Models Humanity's Last Exam
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 0d29d7a9-bfd0-4fb9-9fcc-ec7a4e91f52c · inbound
Memory-Efficient Looped Transformer: Decoupling Compute from Memory in Looped Language Models Humanity's Last Exam
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation d3a59025-5c41-4773-80ee-1e1547439112 · inbound
Benchmarking EngGPT2-16B-A3B against Comparable Italian and International Open-source LLMs Humanity's Last Exam
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation c79a6f68-3c9d-4ccd-adb0-1c0bc22e6918 · inbound
Benchmarking EngGPT2-16B-A3B against Comparable Italian and International Open-source LLMs Humanity's Last Exam
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation b36eb6e8-871c-4944-be90-b689ee514ba5 · inbound
A Semantic-Sampling Framework for Evaluating Calibration in Open-Ended Question Answering Humanity's Last Exam
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation e9d641e8-605a-49ec-813a-0a9f8f03cf34 · inbound
DiagnosticIQ: A Benchmark for LLM-Based Industrial Maintenance Action Recommendation from Symbolic Rules Humanity's Last Exam
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 6f579abd-e020-4a37-8fe1-b8e10ecfaac6 · inbound
EvoMAS: Learning Execution-Time Workflows for Multi-Agent Systems Humanity's Last Exam
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 3ba90dca-fd89-452d-838f-6f1301c4465e · inbound
Soohak: A Mathematician-Curated Benchmark for Evaluating Research-level Math Capabilities of LLMs Humanity's Last Exam
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 3e7aab8a-d5e1-4247-bed9-e661859d4a38 · inbound
Soohak: A Mathematician-Curated Benchmark for Evaluating Research-level Math Capabilities of LLMs Humanity's Last Exam
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 72c12794-7a39-4bb2-a5e0-8bc4fbe92281 · inbound
LLM-Guided Monte Carlo Tree Search over Knowledge Graphs: Composing Mechanistic Explanations for Drug-Disease Pairs Humanity's Last Exam
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 11f613a6-aa79-4863-91bb-98a293df85e5 · inbound
Personalized Deep Research: A User-Centric Framework, Dataset, and Hybrid Evaluation for Knowledge Discovery Humanity's Last Exam
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 6db75031-179d-4e79-b952-ee2f4fc64bf0 · inbound
MaD Physics: Evaluating information seeking under constraints in physical environments Humanity's Last Exam
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 77b78ed1-faa0-4140-afe4-2810514019e7 · inbound
Towards On-Policy Data Evolution for Visual-Native Multimodal Deep Search Agents Humanity's Last Exam
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 49549d33-548b-4285-bf2a-90fe7c53a390 · inbound
Towards On-Policy Data Evolution for Visual-Native Multimodal Deep Search Agents Humanity's Last Exam
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation de75815d-0b53-44b3-97cd-1eeabc6679eb · inbound
The Generalized Turing Test: A Foundation for Comparing Intelligence Humanity's Last Exam
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 2f24f210-6d5b-41f8-a315-bf72e6d78cc5 · inbound
AssayBench: An Assay-Level Virtual Cell Benchmark for LLMs and Agents Humanity's Last Exam
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 802fe269-f6f9-4b27-a61c-6c047c0beab2 · inbound
Instructions Shape Production of Language, not Processing Humanity's Last Exam
Reference 188
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation a9255eb0-954e-4891-91c1-d4a756e89fb9 · inbound
Instructions Shape Production of Language, not Processing Humanity's Last Exam
Reference 188
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation b668be92-2506-4068-b4b7-1e5e4edfa892 · inbound
Measuring Five-Nines Reliability: Sample-Efficient LLM Evaluation in Saturated Benchmarks Humanity's Last Exam
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 8cbd53d9-a058-405e-97a9-a7c7c768d7ad · inbound
Formal Conjectures: An Open and Evolving Benchmark for Verified Discovery in Mathematics Humanity's Last Exam
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 01856572-01d8-457a-bd13-3b047a257cc3 · inbound
TRIAGE: Evaluating Prospective Metacognitive Control in LLMs under Resource Constraints Humanity's Last Exam
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation d829a676-ab20-4782-9ce7-cd8f6e39b53d · inbound
Unsteady Metrics and Benchmarking Cultures of AI Model Builders Humanity's Last Exam
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 4297c688-9dfc-4cce-9e77-1e041df101e2 · inbound
Boosting Reinforcement Learning with Verifiable Rewards via Randomly Selected Few-Shot Guidance Humanity's Last Exam
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 11863d6f-f31c-4f39-aa29-3a28ef9139af · inbound
OpenDeepThink: Parallel Reasoning via Bradley-Terry Aggregation Humanity's Last Exam
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation c5f9caa1-08a5-45a7-ac7d-c4384980924f · inbound
OpenDeepThink: Parallel Reasoning via Bradley-Terry Aggregation Humanity's Last Exam
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 7528d9cb-2937-41a1-b004-bf40bf8a8085 · inbound
Argus: Evidence Assembly for Scalable Deep Research Agents Humanity's Last Exam
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 38c3bfc5-04db-426f-af04-75c37f12efca · inbound
Argus: Evidence Assembly for Scalable Deep Research Agents Humanity's Last Exam
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation f452c0fb-200c-4991-8f43-99bb544f3378 · inbound
Customizing an LLM for Enterprise Software Engineering Humanity's Last Exam
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 232e5aeb-d6af-48c3-8cd5-039af4234293 · inbound
Customizing an LLM for Enterprise Software Engineering Humanity's Last Exam
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation a589650e-7a7d-4ede-9639-84d4ae5c1266 · inbound
Evaluating Cognitive Age Alignment in Interactive AI Agents Humanity's Last Exam
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation ed025b07-a591-4210-84c6-6285a05749f0 · inbound
Forecasting Downstream Performance of LLMs With Proxy Metrics Humanity's Last Exam
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 1003b0c1-c8da-475b-a6cc-76b97a292cfd · inbound
OpenCompass: A Universal Evaluation Platform for Large Language Models Humanity's Last Exam
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.