Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-27T01:28:59.297634Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 1 inbound Pith citation observation for arXiv:2606.17574.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-27T01:28:59.297634Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-01T13:53:46.937166Z
A source-named dated measurement, never combined with another source.
Source: cited_works
59 of 59 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 9d0c4706-c772-476e-9fd2-7021a3ae3beb · outbound
DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack Introducing helix 02: Full-body autonomy.https://www.figure.ai/news/helix-02, 2026
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a4d40f1-8eaf-4dba-8d94-5aae873f49ca · outbound
DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack Measuring massive multitask language understanding
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48078aec-2fea-4d24-aff6-87aebdebd11f · outbound
DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack Training Verifiers to Solve Math Word Problems
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c84b1117-b8b6-49c7-8ee3-83dc731d9542 · outbound
DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack Evaluating Large Language Models Trained on Code
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fe194f51-4a96-4743-bdf1-2b51635f6ba6 · outbound
DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62ea7505-cd8e-45bb-a08b-a7a49c4c33f9 · outbound
DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack GAIA: A benchmark for general ai assistants
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3cb0d74-26c9-41e2-be37-23f4c6ebe0f3 · outbound
DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack OSWorld: Benchmarking multimodal agents for open-ended tasks in real computer environments
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c39ab67-c49a-4b2d-93c0-5a9e8e4f8e5d · outbound
DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 000c0c81-3f88-4ea4-843e-be737591c6bc · outbound
DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c8fdb92-4fcb-4e3d-84ea-6387fc01c00a · outbound
DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack CALVIN: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks.IEEE Robotics and Automation Letters, 2022
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f023eaf6-12c2-44f8-af86-73a8fc803482 · outbound
DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack LIBERO: Benchmarking knowledge transfer for lifelong robot learning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ffb73923-4378-425c-b7fa-3740155cf60d · outbound
DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 803db83d-5b69-4758-9fb2-a62484252f7e · outbound
DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack Unresolved cited work
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e197a75-b186-4970-92c4-6063ff075bd2 · outbound
DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack Open X-Embodiment: Robotic Learning Datasets and RT-X Models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2842f995-f892-42d8-a907-f08737ddf8d1 · outbound
DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack Evaluating Real-World Robot Manipulation Policies in Simulation
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0c305888-fd76-44de-a7fa-b8843614417f · outbound
DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack HumanoidBench: Simulated Humanoid Benchmark for Whole-Body Locomotion and Manipulation
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3cbf843e-fd60-4ace-9808-6ce7c2b1d27f · outbound
DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack RoboHive: A unified framework for robot learning
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1fc877c-7310-4b1b-ba20-2cfd3cdde422 · outbound
DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack Orbit: A unified simulation framework for interactive robot learning environments.IEEE Robotics and Automation Letters, 2023
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc8c09d1-53bb-4cfa-b45e-63f95be25cd9 · outbound
DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack Michaelov, Hanwool Albert Lee, Janna, Leonid Sinev, Khalid, Kiersten Stokes, Zden ˇek Kasner, and KonradSzafer
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c36b662-255b-401e-84b3-a806ef671913 · outbound
DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack OpenCompass: A universal evaluation platform for foundation models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9222d0e-af71-4f2e-a0cd-697abb9fe1f0 · outbound
DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack Holistic Evaluation of Language Models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 14906927-c213-4dfe-a9f6-9bdad2c30e1f · outbound
DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d84ed359-c4f1-4ecf-b189-b9f9f51441fe · outbound
DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 33ddd818-f224-4ec7-82b0-6b68f15064bd · outbound
DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack Inspect AI: Framework for large language model evaluations, 2024
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 19739858-a943-4b72-a280-d472425b5bc0 · outbound
DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack DeepSeek-V4: Technical report
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb147307-4954-4cb6-90e1-cad360ba02e4 · outbound
DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack Measuring short-form factuality in large language models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7ca212be-2a01-4a30-893e-a4e4eee31410 · outbound
DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack MMLU-Pro: A more robust and challenging multi-task language understanding benchmark
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b268ac5f-8580-4dba-978a-667f7cf58e88 · outbound
DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack Unresolved cited work
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1384d45-0785-4ce6-828d-26a9db2da0f9 · outbound
DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 06ce4879-bd0f-49c3-9665-58c2006139cb · outbound
DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack Unresolved cited work
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ea4d919-05d7-4791-ae17-fea528ed8d31 · outbound
DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack C-Eval: A multi-level multi-discipline chinese evaluation suite for foundation models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e40060e2-602d-4874-b0f8-1c3eab0efcc8 · outbound
DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack Humanity's Last Exam
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b5333460-ca37-4550-bf93-a25d36ddb22b · outbound
DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack MathArena: Evaluating LLMs on uncontaminated math competitions.Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks, 2025
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34b96082-ade2-47b2-83a2-89d1eda7741e · outbound
DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack LiveCodeBench: Holistic and contamination free eval- uation of large language models for code
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation adc0dac5-123a-431a-9e1c-be3542abaa94 · outbound
DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack LiveBench: A challenging, contamination-limited LLM benchmark
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2da44b45-b6d3-4e67-8e9e-7cc4af0ef013 · outbound
DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack Measuring mathematical problem solving with the MATH dataset
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00533017-745f-46bd-a3f9-134c08684039 · outbound
DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack Let’s verify step by step
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b97a20f-30f4-4dee-92ad-a318ec69de01 · outbound
DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack Instruction-Following Evaluation for Large Language Models
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0fdf88e8-565c-4fc6-8074-0e48289f69ff · outbound
DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack Patil, Huanzhi Mao, Fanjia Yan, Charlie Cheng-Jie Ji, Vishnu Suresh, Ion Stoica, and Joseph E
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9dda837a-7c18-48f1-9ae9-a24761cdf027 · outbound
DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack MMMU: A massive multi-discipline multimodal understanding and reasoning benchmark for expert AGI
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c6ac217-6049-465c-9e9c-f844fb638155 · outbound
DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack MMMU-Pro: A more robust multi-discipline multimodal understanding benchmark
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77dee267-f976-4c4c-82e2-b7d670d75d5a · outbound
DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack MathVista: Evaluating mathematical reasoning of foundation models in visual contexts
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa191aea-507c-4f38-9585-1e15a12dc8c2 · outbound
DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack DynaMath: A dynamic visual benchmark for evaluating mathematical reasoning robustness of vision language models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f744dd39-d822-4954-8706-62829f968395 · outbound
DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack Vision language models are blind: Failing to translate detailed visual features into words
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a028ef52-8ac0-4a21-978a-d948f3d36d48 · outbound
DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack MMBench: Is your multi-modal model an all-around player? InEuropean Conference on Computer Vision, 2024
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce73596b-7664-4865-8a9c-87b449fd91d7 · outbound
DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack Are we on the right way for evaluating large vision-language models? InAdvances in Neural Information Processing Systems, 2024
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e326c79d-3a34-4c16-b58a-46a89197fb72 · outbound
DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack RealWorldQA.https://huggingface.co/datasets/xai-org/RealworldQA, 2024
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb5b73ff-e5c9-4798-b510-ab292685f955 · outbound
DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack SimpleVQA: Multimodal factuality evaluation for multimodal large language models
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01ce1a4f-893e-42d7-89e5-2674561c10c6 · outbound
DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3837a3ae-13d8-4197-8b4f-b6f25aa63e6b · outbound
DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack OCRBench: On the hidden mystery of OCR in large multimodal models.Science China Information Sciences, 2024
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b22ae11c-b951-4a68-b116-772869b8094b · outbound
DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack Teaching CLIP to count to ten
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f38a86a-af7b-468d-a387-eda6450f3ce1 · outbound
DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack Berg, and Tamara L
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f232f156-15f7-4dec-aafb-c003e7d3e8cc · outbound
DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack Gemini Robotics: Bringing AI into the Physical World
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 641c6212-c570-4e80-a20f-e66f1911b946 · outbound
DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack Video-MME: The first-ever comprehensive evaluation benchmark of multi-modal LLMs in video analysis
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36e951a3-edad-45d2-a32d-6f16d9758464 · outbound
DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack LibriSpeech: An ASR corpus based on public domain audio books
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10da4ae3-4e52-4445-927a-f27bd83a8724 · outbound
DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack WenetSpeech: A 10000+ hours multi-domain mandarin corpus for speech recognition
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fba57328-79e0-4014-b496-74d7a04f73e5 · outbound
DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 724b29f9-f116-4780-9c31-e3cbe9a5e483 · outbound
DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack Qwen3-Omni Technical Report
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4e90127e-528e-49bb-8855-90707b08cc2f · outbound
DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack Beyond the nav-graph: Vision and language navigation in continuous environments
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf181efd-5be5-4ada-8ec4-6ade15169a2e · inbound
Athena-Brain Technical Report: An Efficient Robot Brain for General Intelligence and Embodied Interaction DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.