Pith. sign in

Paper Citation Record · LEDGER

DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack

As of 7 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 1 inbound Pith citation observation for arXiv:2606.17574.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.17574 v1

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-27T01:28:59.297634Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T13:53:46.937166Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

59 of 59 outbound references displayed

  • verified exact16
  • verified fuzzy0
  • unresolved40
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9d0c4706-c772-476e-9fd2-7021a3ae3beb · outbound

This paper cites Introducing helix 02: Full-body autonomy.https://www.figure.ai/news/helix-02, 2026.

DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack Introducing helix 02: Full-body autonomy.https://www.figure.ai/news/helix-02, 2026

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-27T01:28:59.297634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:28:59.297634Z digest=sha256:5aabe5e8d0b449b1ec5d53fe93f58651a3d3c001a9e1b17abb6dfac5c8c1f543

Observation 3a4d40f1-8eaf-4dba-8d94-5aae873f49ca · outbound

This paper cites Measuring massive multitask language understanding.

DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack Measuring massive multitask language understanding

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-27T01:28:59.297634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:28:59.297634Z digest=sha256:68258da5488dce58cebcf2767d90651c4725ce1463ca214c165463b5fc657b4a

Observation 48078aec-2fea-4d24-aff6-87aebdebd11f · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack Training Verifiers to Solve Math Word Problems

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:18:56.519889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T01:28:59.297634Z digest=sha256:4a7fc93a983905ca4e7abf2922fa8bafb0435566c19b6f8b5d831af6ccc184bb

Observation c84b1117-b8b6-49c7-8ee3-83dc731d9542 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack Evaluating Large Language Models Trained on Code

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:18:56.535162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T01:28:59.297634Z digest=sha256:61bf4ae7891b50d96f30941ab4abbf2df3a65fbb86f25a2ee0c9b57d7c0664c3

Observation fe194f51-4a96-4743-bdf1-2b51635f6ba6 · outbound

This paper cites Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan.

DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-27T01:28:59.297634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:28:59.297634Z digest=sha256:01093de4d9f3346739c582c887a4661a0d183199d27b6375e30b1a31ff64777b

Observation 62ea7505-cd8e-45bb-a08b-a7a49c4c33f9 · outbound

This paper cites GAIA: A benchmark for general ai assistants.

DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack GAIA: A benchmark for general ai assistants

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-27T01:28:59.297634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:28:59.297634Z digest=sha256:6bc85a1cf786fbdd065a473379a106cf7f3bedd292b8300c1930d3e003b507e7

Observation b3cb0d74-26c9-41e2-be37-23f4c6ebe0f3 · outbound

This paper cites OSWorld: Benchmarking multimodal agents for open-ended tasks in real computer environments.

DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack OSWorld: Benchmarking multimodal agents for open-ended tasks in real computer environments

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-27T01:28:59.297634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:28:59.297634Z digest=sha256:9f7f9c6e01ce685891ac02eb35e69cd4de738591be9842bd9e103235d2f01874

Observation 7c39ab67-c49a-4b2d-93c0-5a9e8e4f8e5d · outbound

This paper cites $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains.

DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:18:56.541207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T01:28:59.297634Z digest=sha256:95c038d085d80d5ba4308d61fc7a742cc8a8f7341ab869da1aa767d8689b92ea

Observation 000c0c81-3f88-4ea4-843e-be737591c6bc · outbound

This paper cites Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig.

DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-27T01:28:59.297634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:28:59.297634Z digest=sha256:feed2ef11d9b131d650f89287fda3f0ef224b9d4fb3e01bf79c7edfddc974e60

Observation 2c8fdb92-4fcb-4e3d-84ea-6387fc01c00a · outbound

This paper cites CALVIN: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks.IEEE Robotics and Automation Letters, 2022.

DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack CALVIN: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks.IEEE Robotics and Automation Letters, 2022

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-27T01:28:59.297634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:28:59.297634Z digest=sha256:d3ac8bb3830f1764788e676a3b7a33a5415494bd95e1d5f404429a2f7c299079

Observation f023eaf6-12c2-44f8-af86-73a8fc803482 · outbound

This paper cites LIBERO: Benchmarking knowledge transfer for lifelong robot learning.

DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack LIBERO: Benchmarking knowledge transfer for lifelong robot learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-27T01:28:59.297634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:28:59.297634Z digest=sha256:1b72d2c207b152b8f2491cceeee0fbf7962b9364b79dab22ca170785dd2b19c9

Observation ffb73923-4378-425c-b7fa-3740155cf60d · outbound

This paper cites Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning.

DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-27T01:28:59.297634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:28:59.297634Z digest=sha256:9ef6338d75d111553b8e659a08df3650581c487d02e31a618c5d48fc3ec51d5b

Observation 803db83d-5b69-4758-9fb2-a62484252f7e · outbound

This paper cites an unresolved cited work.

DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-27T01:28:59.297634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:28:59.297634Z digest=sha256:14d2d97b9b447645fef4f390334bd26bf090babd3fc4f77672909bd9727a31a9

Observation 9e197a75-b186-4970-92c4-6063ff075bd2 · outbound

This paper cites Open X-Embodiment: Robotic Learning Datasets and RT-X Models.

DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack Open X-Embodiment: Robotic Learning Datasets and RT-X Models

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:18:56.572379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T01:28:59.297634Z digest=sha256:a61f008a4b5f842ee60dcf779f214a4d4d45b628e6ca60c91ee67b1fca151d70

Observation 2842f995-f892-42d8-a907-f08737ddf8d1 · outbound

This paper cites Evaluating Real-World Robot Manipulation Policies in Simulation.

DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack Evaluating Real-World Robot Manipulation Policies in Simulation

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T20:18:56.513442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T01:28:59.297634Z digest=sha256:c355e9846225ee3bf55d9e27d1dfad58c1e20db3dd29609e4ae191241a8dbffa

Observation 0c305888-fd76-44de-a7fa-b8843614417f · outbound

This paper cites HumanoidBench: Simulated Humanoid Benchmark for Whole-Body Locomotion and Manipulation.

DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack HumanoidBench: Simulated Humanoid Benchmark for Whole-Body Locomotion and Manipulation

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:18:56.580913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T01:28:59.297634Z digest=sha256:1daa2864cfb6248a4ac996ed7531cfda09f1e0e4af6b25ada5e7d83ba22ac34f

Observation 3cbf843e-fd60-4ace-9808-6ce7c2b1d27f · outbound

This paper cites RoboHive: A unified framework for robot learning.

DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack RoboHive: A unified framework for robot learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-27T01:28:59.297634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:28:59.297634Z digest=sha256:4ea2a847bc679a6aa8fcabbb81ce44883bab4f806836ff671bb2902cd2f41bb8

Observation f1fc877c-7310-4b1b-ba20-2cfd3cdde422 · outbound

This paper cites Orbit: A unified simulation framework for interactive robot learning environments.IEEE Robotics and Automation Letters, 2023.

DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack Orbit: A unified simulation framework for interactive robot learning environments.IEEE Robotics and Automation Letters, 2023

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-27T01:28:59.297634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:28:59.297634Z digest=sha256:0308d4348f12302d55fba2939770c56b3d37343c93d2e56f83e7cb83dfa604e2

Observation fc8c09d1-53bb-4cfa-b45e-63f95be25cd9 · outbound

This paper cites Michaelov, Hanwool Albert Lee, Janna, Leonid Sinev, Khalid, Kiersten Stokes, Zden ˇek Kasner, and KonradSzafer.

DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack Michaelov, Hanwool Albert Lee, Janna, Leonid Sinev, Khalid, Kiersten Stokes, Zden ˇek Kasner, and KonradSzafer

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-27T01:28:59.297634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:28:59.297634Z digest=sha256:cd036143321394fe47a86e1562aa4020a6dcae7ea0f804435b9014e2144f2d38

Observation 3c36b662-255b-401e-84b3-a806ef671913 · outbound

This paper cites OpenCompass: A universal evaluation platform for foundation models.

DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack OpenCompass: A universal evaluation platform for foundation models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-27T01:28:59.297634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:28:59.297634Z digest=sha256:7946cd7d7b395c05e37a1cc7b98c45ea6eafde92e970829caef2ec074bd5f34f

Observation e9222d0e-af71-4f2e-a0cd-697abb9fe1f0 · outbound

This paper cites Holistic Evaluation of Language Models.

DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack Holistic Evaluation of Language Models

Reference 21

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T20:18:56.587390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T01:28:59.297634Z digest=sha256:5b8ee38d7918eac2e1f5701f59329142503fc284b513bf9146d276daeb9a4b4a

Observation 14906927-c213-4dfe-a9f6-9bdad2c30e1f · outbound

This paper cites VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models.

DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-07T02:19:36.652705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T01:28:59.297634Z digest=sha256:6f748c4beef6a7aed127a992ebf5efb34c53f7bc3bc2c838e740243d699f2993

Observation d84ed359-c4f1-4ecf-b189-b9f9f51441fe · outbound

This paper cites LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models.

DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:18:56.562509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T01:28:59.297634Z digest=sha256:4dc030ca39407cf6f0c5d4d716e3ba80716053d7285ba0fb944d3237d173b56c

Observation 33ddd818-f224-4ec7-82b0-6b68f15064bd · outbound

This paper cites Inspect AI: Framework for large language model evaluations, 2024.

DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack Inspect AI: Framework for large language model evaluations, 2024

Reference 24

Resolution
verified exact
doi, observed 2026-06-27T01:30:20.408273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T01:28:59.297634Z digest=sha256:0d66a9bba17f684928a3ed9863eb830a944b030e464669f7de0f50f43bd9c07c

Observation 19739858-a943-4b72-a280-d472425b5bc0 · outbound

This paper cites DeepSeek-V4: Technical report.

DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack DeepSeek-V4: Technical report

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-27T01:28:59.297634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:28:59.297634Z digest=sha256:bd44077917f42d4c032e4468be30aa24aea53697b33f4d14e9ac8db44c887360

Observation eb147307-4954-4cb6-90e1-cad360ba02e4 · outbound

This paper cites Measuring short-form factuality in large language models.

DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack Measuring short-form factuality in large language models

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:18:56.565461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T01:28:59.297634Z digest=sha256:a3b68edfc0c3e65bb2e3c31dcd243bee3021d063fbfa2ade1f84153bc3a8c819

Observation 7ca212be-2a01-4a30-893e-a4e4eee31410 · outbound

This paper cites MMLU-Pro: A more robust and challenging multi-task language understanding benchmark.

DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack MMLU-Pro: A more robust and challenging multi-task language understanding benchmark

Reference 27

Resolution
unresolved
no resolver link, observed 2026-06-27T01:28:59.297634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:28:59.297634Z digest=sha256:83da1d04640487f1a7b8c92e5b68409f815f73ff2ba9a801e5067939d585ef21

Observation b268ac5f-8580-4dba-978a-667f7cf58e88 · outbound

This paper cites an unresolved cited work.

DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-27T01:28:59.297634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:28:59.297634Z digest=sha256:fdf0b84733eed6f7b1b3abb0d27c2d841e11f7044eba17e728cf4cbc1703299c

Observation d1384d45-0785-4ce6-828d-26a9db2da0f9 · outbound

This paper cites SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines.

DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:18:56.568783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T01:28:59.297634Z digest=sha256:74f91592cbf7ec2d88d2c153688f0ce4c27fded99fb244a32dbc133f3fcf64a9

Observation 06ce4879-bd0f-49c3-9665-58c2006139cb · outbound

This paper cites an unresolved cited work.

DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-06-27T01:28:59.297634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:28:59.297634Z digest=sha256:adf7cf2a5379c5297612608d548254d6ce4578b07e4bcdacca5a4911bddd8811

Observation 8ea4d919-05d7-4791-ae17-fea528ed8d31 · outbound

This paper cites C-Eval: A multi-level multi-discipline chinese evaluation suite for foundation models.

DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack C-Eval: A multi-level multi-discipline chinese evaluation suite for foundation models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-06-27T01:28:59.297634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:28:59.297634Z digest=sha256:e817353c1b24b754cc78bc201240d20cea29ec309d5b661260a3948ef65ca9c1

Observation e40060e2-602d-4874-b0f8-1c3eab0efcc8 · outbound

This paper cites Humanity's Last Exam.

DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack Humanity's Last Exam

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:18:56.582548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T01:28:59.297634Z digest=sha256:2af472f269c1016e25a93e2b00c6d60f78e76ad7aaac36622e918f2763ba0f35

Observation b5333460-ca37-4550-bf93-a25d36ddb22b · outbound

This paper cites MathArena: Evaluating LLMs on uncontaminated math competitions.Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks, 2025.

DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack MathArena: Evaluating LLMs on uncontaminated math competitions.Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks, 2025

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-27T01:28:59.297634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:28:59.297634Z digest=sha256:04cfbac635a4f85b4e100fd230a2d7c07a9a7865f06d7814784692d2c5b0e066

Observation 34b96082-ade2-47b2-83a2-89d1eda7741e · outbound

This paper cites LiveCodeBench: Holistic and contamination free eval- uation of large language models for code.

DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack LiveCodeBench: Holistic and contamination free eval- uation of large language models for code

Reference 34

Resolution
unresolved
no resolver link, observed 2026-06-27T01:28:59.297634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:28:59.297634Z digest=sha256:b040c115d689322c526630e520f87984609562d6e96157c8f25e6096d9599bae

Observation adc0dac5-123a-431a-9e1c-be3542abaa94 · outbound

This paper cites LiveBench: A challenging, contamination-limited LLM benchmark.

DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack LiveBench: A challenging, contamination-limited LLM benchmark

Reference 35

Resolution
unresolved
no resolver link, observed 2026-06-27T01:28:59.297634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:28:59.297634Z digest=sha256:cdbfb60ccc6f5f6273b6c9c789e2de83348abb2a38be426284da7655f7b7ccb9

Observation 2da44b45-b6d3-4e67-8e9e-7cc4af0ef013 · outbound

This paper cites Measuring mathematical problem solving with the MATH dataset.

DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack Measuring mathematical problem solving with the MATH dataset

Reference 36

Resolution
unresolved
no resolver link, observed 2026-06-27T01:28:59.297634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:28:59.297634Z digest=sha256:bb90f7cad9c93f03fbfd818b1d69e6631cd46a7a748de0c8a7bdbc142d2d48de

Observation 00533017-745f-46bd-a3f9-134c08684039 · outbound

This paper cites Let’s verify step by step.

DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack Let’s verify step by step

Reference 37

Resolution
unresolved
no resolver link, observed 2026-06-27T01:28:59.297634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:28:59.297634Z digest=sha256:d8009869fe2ea80d7cabd647eae25772017cbd7f120db804d419c9dc3e7d73d4

Observation 6b97a20f-30f4-4dee-92ad-a318ec69de01 · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack Instruction-Following Evaluation for Large Language Models

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:18:56.577524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T01:28:59.297634Z digest=sha256:58899e104c18a6df79d67f268f8dd5735e726d85d15db3b10e0c3d43f2013d83

Observation 0fdf88e8-565c-4fc6-8074-0e48289f69ff · outbound

This paper cites Patil, Huanzhi Mao, Fanjia Yan, Charlie Cheng-Jie Ji, Vishnu Suresh, Ion Stoica, and Joseph E.

DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack Patil, Huanzhi Mao, Fanjia Yan, Charlie Cheng-Jie Ji, Vishnu Suresh, Ion Stoica, and Joseph E

Reference 39

Resolution
unresolved
no resolver link, observed 2026-06-27T01:28:59.297634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:28:59.297634Z digest=sha256:2c52dd38a94160cede91650a3c96b642931248f40cd04722222632abece12415

Observation 9dda837a-7c18-48f1-9ae9-a24761cdf027 · outbound

This paper cites MMMU: A massive multi-discipline multimodal understanding and reasoning benchmark for expert AGI.

DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack MMMU: A massive multi-discipline multimodal understanding and reasoning benchmark for expert AGI

Reference 40

Resolution
unresolved
no resolver link, observed 2026-06-27T01:28:59.297634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:28:59.297634Z digest=sha256:fa8bafe2ecb8020f947f9eb991262d93d41875a2ab6f0bfa2303aeede4a2ab08

Observation 5c6ac217-6049-465c-9e9c-f844fb638155 · outbound

This paper cites MMMU-Pro: A more robust multi-discipline multimodal understanding benchmark.

DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack MMMU-Pro: A more robust multi-discipline multimodal understanding benchmark

Reference 41

Resolution
unresolved
no resolver link, observed 2026-06-27T01:28:59.297634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:28:59.297634Z digest=sha256:8925d60211612ab075915e0e2fa86b8d75e618f557c2be97d90e9b4daf8ebd9f

Observation 77dee267-f976-4c4c-82e2-b7d670d75d5a · outbound

This paper cites MathVista: Evaluating mathematical reasoning of foundation models in visual contexts.

DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack MathVista: Evaluating mathematical reasoning of foundation models in visual contexts

Reference 42

Resolution
unresolved
no resolver link, observed 2026-06-27T01:28:59.297634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:28:59.297634Z digest=sha256:36ae2159b9920350e7bbf83a44e7e6b070c1255dc5690780be0e18b446646970

Observation fa191aea-507c-4f38-9585-1e15a12dc8c2 · outbound

This paper cites DynaMath: A dynamic visual benchmark for evaluating mathematical reasoning robustness of vision language models.

DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack DynaMath: A dynamic visual benchmark for evaluating mathematical reasoning robustness of vision language models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-06-27T01:28:59.297634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:28:59.297634Z digest=sha256:31f11d854a97f537305e99a45a45c5f479fe40a7493d6d9d264c4a2f8f156d78

Observation f744dd39-d822-4954-8706-62829f968395 · outbound

This paper cites Vision language models are blind: Failing to translate detailed visual features into words.

DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack Vision language models are blind: Failing to translate detailed visual features into words

Reference 44

Resolution
unresolved
no resolver link, observed 2026-06-27T01:28:59.297634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:28:59.297634Z digest=sha256:babcc1bcd042e890324b83e2aaaa1e09c4f0b966ed45e490e41a687795dfc167

Observation a028ef52-8ac0-4a21-978a-d948f3d36d48 · outbound

This paper cites MMBench: Is your multi-modal model an all-around player? InEuropean Conference on Computer Vision, 2024.

DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack MMBench: Is your multi-modal model an all-around player? InEuropean Conference on Computer Vision, 2024

Reference 45

Resolution
unresolved
no resolver link, observed 2026-06-27T01:28:59.297634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:28:59.297634Z digest=sha256:cf3409599986a350c3aafd00a7c1c9044dfab058a1198493e9ad1dca131530a6

Observation ce73596b-7664-4865-8a9c-87b449fd91d7 · outbound

This paper cites Are we on the right way for evaluating large vision-language models? InAdvances in Neural Information Processing Systems, 2024.

DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack Are we on the right way for evaluating large vision-language models? InAdvances in Neural Information Processing Systems, 2024

Reference 46

Resolution
unresolved
no resolver link, observed 2026-06-27T01:28:59.297634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:28:59.297634Z digest=sha256:d0492610718f2e59b72b04d340b833c1cd60419eb41c95b7e8f9a1244be29bf2

Observation e326c79d-3a34-4c16-b58a-46a89197fb72 · outbound

This paper cites RealWorldQA.https://huggingface.co/datasets/xai-org/RealworldQA, 2024.

DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack RealWorldQA.https://huggingface.co/datasets/xai-org/RealworldQA, 2024

Reference 47

Resolution
unresolved
no resolver link, observed 2026-06-27T01:28:59.297634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:28:59.297634Z digest=sha256:2ef12a5b728de850af5c70ed054e7c4f8b061707954b5f186837103e732deece

Observation eb5b73ff-e5c9-4798-b510-ab292685f955 · outbound

This paper cites SimpleVQA: Multimodal factuality evaluation for multimodal large language models.

DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack SimpleVQA: Multimodal factuality evaluation for multimodal large language models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-06-27T01:28:59.297634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:28:59.297634Z digest=sha256:4d34cb03ba5c7675953bc78851858635aba5a6c3a3b98c10c28de0bbfadeeb59

Observation 01ce1a4f-893e-42d7-89e5-2674561c10c6 · outbound

This paper cites CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs.

DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:18:56.584395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T01:28:59.297634Z digest=sha256:835e167bfb614f66b13fabf1a0679572b9807c2472e6382c83f3862d4d56f0b6

Observation 3837a3ae-13d8-4197-8b4f-b6f25aa63e6b · outbound

This paper cites OCRBench: On the hidden mystery of OCR in large multimodal models.Science China Information Sciences, 2024.

DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack OCRBench: On the hidden mystery of OCR in large multimodal models.Science China Information Sciences, 2024

Reference 50

Resolution
unresolved
no resolver link, observed 2026-06-27T01:28:59.297634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:28:59.297634Z digest=sha256:49918fecdb5529c64f915fd0beb4bd2b117a65d857b55f2235d15efaf0a36a8f

Observation b22ae11c-b951-4a68-b116-772869b8094b · outbound

This paper cites Teaching CLIP to count to ten.

DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack Teaching CLIP to count to ten

Reference 51

Resolution
unresolved
no resolver link, observed 2026-06-27T01:28:59.297634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:28:59.297634Z digest=sha256:860cbf080eee1bd7f50aaab8d3de93477400cafa80c540e3dc3347e61df590b4

Observation 6f38a86a-af7b-468d-a387-eda6450f3ce1 · outbound

This paper cites Berg, and Tamara L.

DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack Berg, and Tamara L

Reference 52

Resolution
unresolved
no resolver link, observed 2026-06-27T01:28:59.297634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:28:59.297634Z digest=sha256:3ea536eb753d050778b997260c64b339bee231b3d06a13b90068203f7f08a127

Observation f232f156-15f7-4dec-aafb-c003e7d3e8cc · outbound

This paper cites Gemini Robotics: Bringing AI into the Physical World.

DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack Gemini Robotics: Bringing AI into the Physical World

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:18:56.554037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T01:28:59.297634Z digest=sha256:651c7beb3d950983c3e4d57d13defdcc00d9b573d49c8bd417ae6d8c29c9ec4d

Observation 641c6212-c570-4e80-a20f-e66f1911b946 · outbound

This paper cites Video-MME: The first-ever comprehensive evaluation benchmark of multi-modal LLMs in video analysis.

DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack Video-MME: The first-ever comprehensive evaluation benchmark of multi-modal LLMs in video analysis

Reference 54

Resolution
unresolved
no resolver link, observed 2026-06-27T01:28:59.297634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:28:59.297634Z digest=sha256:dd7228273668325f8c02970f4656a639547c070e19afbae05a8528b6313041fe

Observation 36e951a3-edad-45d2-a32d-6f16d9758464 · outbound

This paper cites LibriSpeech: An ASR corpus based on public domain audio books.

DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack LibriSpeech: An ASR corpus based on public domain audio books

Reference 55

Resolution
unresolved
no resolver link, observed 2026-06-27T01:28:59.297634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:28:59.297634Z digest=sha256:865b1937368e6273071aa03adc9f21dd5d11d2f81479532e8199b6e202c40657

Observation 10da4ae3-4e52-4445-927a-f27bd83a8724 · outbound

This paper cites WenetSpeech: A 10000+ hours multi-domain mandarin corpus for speech recognition.

DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack WenetSpeech: A 10000+ hours multi-domain mandarin corpus for speech recognition

Reference 56

Resolution
unresolved
no resolver link, observed 2026-06-27T01:28:59.297634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:28:59.297634Z digest=sha256:046ec518f17e71c4afcd18d37433173fd6baa783bde2855676c3e006b4bacfd9

Observation fba57328-79e0-4014-b496-74d7a04f73e5 · outbound

This paper cites WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs.

DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:18:56.555124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T01:28:59.297634Z digest=sha256:24ced8107195f71f876618a20c3446751bb467c7255553a3f03e873f7d6de236

Observation 724b29f9-f116-4780-9c31-e3cbe9a5e483 · outbound

This paper cites Qwen3-Omni Technical Report.

DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack Qwen3-Omni Technical Report

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:18:56.558413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T01:28:59.297634Z digest=sha256:b67650fbc093411d3e7f1ecc38b34146b78cb51f725f310af1f1b45a408b1147

Observation 4e90127e-528e-49bb-8855-90707b08cc2f · outbound

This paper cites Beyond the nav-graph: Vision and language navigation in continuous environments.

DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack Beyond the nav-graph: Vision and language navigation in continuous environments

Reference 59

Resolution
malformed identifier
no resolver link, observed 2026-06-27T01:28:59.297634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:28:59.297634Z digest=sha256:9681fef8b6f9d43d41b23872244cbefe86666bda650fccec93830e3e69198ed8

Pith citing papers

Observation cf181efd-5be5-4ada-8ec4-6ade15169a2e · inbound

Athena-Brain Technical Report: An Efficient Robot Brain for General Intelligence and Embodied Interaction cites this paper.

Athena-Brain Technical Report: An Efficient Robot Brain for General Intelligence and Embodied Interaction DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T13:53:46.937166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:53:46.937166Z digest=sha256:cefeae24cc680ce5cd52bb4cac9207c2fd25a12c7ce686bcafb3d64595eda6fd