Pith. sign in

Paper Citation Record · LEDGER

EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

As of 6 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 57 inbound Pith citation observations for arXiv:2502.09560.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.09560 v3

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-17T00:24:45.343430Z

measured 79 of 79 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 57 of 57 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T00:13:24.699600Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

22 of 22 outbound references displayed

  • verified exact1
  • verified fuzzy18
  • unresolved2
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

1
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 077f13c7-f4f6-4281-9a93-74407eed553e · outbound

This paper cites Put washed lettuce in the refrigerator.

EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents Put washed lettuce in the refrigerator

Reference 1

Resolution
verified exact
doi, observed 2026-05-17T00:24:45.369129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:24:45.343430Z digest=sha256:ddcca891534128ade29ff44ce288e1443e9228720816efec6cda7ec2ceabb5cf

Observation 83542d9a-a771-4696-8036-4bc678362fac · outbound

This paper cites an unresolved cited work.

EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-05-17T00:24:45.372187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:24:45.343430Z digest=sha256:a3fca5270fd7606fd968f8bf7cd4124fa0d4d77943f0caaf35d8248d8d1937c9

Observation 5e3f3d06-7814-40bf-8c03-e2845c01765d · outbound

This paper cites Avoid performing actions that do not meet the defined validity criteria.

EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents Avoid performing actions that do not meet the defined validity criteria

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:24:45.375431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:24:45.343430Z digest=sha256:00819df9cdca1e25c2925ef75ae499bbe30c736f1789d466662d3ed0c05f8ad8

Observation 35028448-e93f-4a45-8109-acf312d1bafe · outbound

This paper cites You can explore these instances if you do not find the desired object in the current receptacle.

EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents You can explore these instances if you do not find the desired object in the current receptacle

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:24:45.378460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:24:45.343430Z digest=sha256:2b1ba2fb012f4c3a53140bbc4c3e9fab0ca80b8213f69da1f4e10ffa478a08c5

Observation 797cd098-8497-4ada-862e-56bd9744abb0 · outbound

This paper cites If the last action is invalid, reflect on the reason, such as not adhering to action rules or missing preliminary actions, and adjust your plan accordingly.

EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents If the last action is invalid, reflect on the reason, such as not adhering to action rules or missing preliminary actions, and adjust your plan accordingly

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:24:45.381475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:24:45.343430Z digest=sha256:d2f604682ac8faedbc087ff8963104abdedcf01a05478c9a71e9b0fdea201960

Observation 71e67403-b852-4a13-97c4-542ec1991b58 · outbound

This paper cites Each plan should include no more than 20 actions.

EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents Each plan should include no more than 20 actions

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:24:45.384181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:24:45.343430Z digest=sha256:abd88baddef366785d5d5ad2bbaa7e634c78dede9d1d12fcf1db9e36cfea0b55

Observation 977209b6-ef16-472f-8278-f2a71daa2f90 · outbound

This paper cites an unresolved cited work.

EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-05-17T00:24:45.386905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:24:45.343430Z digest=sha256:6403d0199ebca4b4f59fb7c1d8a54a366534289e65b9890e5fc99545226da29d

Observation dee88191-3283-4676-a5c7-88ac72b6fe26 · outbound

This paper cites Avoid performing actions that do not meet the defined validity criteria.

EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents Avoid performing actions that do not meet the defined validity criteria

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:24:45.389411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:24:45.343430Z digest=sha256:b7b1efd9eeb5bd05ddbed74a3253b2ab22435a745e869d994a8ac4c93d25efdc

Observation 36546a12-1133-4870-ac32-3cb138ea2feb · outbound

This paper cites Try to modify the action sequence because previous actions do not lead to success.

EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents Try to modify the action sequence because previous actions do not lead to success

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:24:45.392036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:24:45.343430Z digest=sha256:2ec71a77d12e92948b7f5607da407fb1f5a86d420c69e15fe9d04776ee993972

Observation 299b9357-228e-485f-b20b-91522f8445d7 · outbound

This paper cites You can explore these instances if you do not find the desired object in the current receptacle.

EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents You can explore these instances if you do not find the desired object in the current receptacle

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:24:45.394687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:24:45.343430Z digest=sha256:5765adce15b9c2c9aea8dd61f989d835b8a977077997f55a62be087243f92fd0

Observation 403ea813-25c6-4af7-9e4d-837180734fce · outbound

This paper cites If the last action is invalid, reflect on the reason, such as not adhering to action rules or missing preliminary actions, and adjust your plan accordingly.

EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents If the last action is invalid, reflect on the reason, such as not adhering to action rules or missing preliminary actions, and adjust your plan accordingly

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:24:45.397482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:24:45.343430Z digest=sha256:a6dc9c37b403b45a73faf5302d98145246e7b35bbcc0984168b5f41944b2ee3f

Observation 72382bf5-739c-46d4-a07b-a06d9db221e9 · outbound

This paper cites try to be as close as possible.

EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents try to be as close as possible

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:24:45.399850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:24:45.343430Z digest=sha256:6e663219a1cbb029616c9d3aec219bd540c11808a34235d4a8a82f0232480f31

Observation 7350b6a4-db4a-4193-9f47-bf12a4df4cdd · outbound

This paper cites on the front left side, a few steps from the current standing point).

EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents on the front left side, a few steps from the current standing point)

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:24:45.402302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:24:45.343430Z digest=sha256:712a1869da4e6ad654e54f4e20d195a7618d52052e8f7046a4d7ec46152af745

Observation b5693947-0715-44f1-8fc4-047add080fa7 · outbound

This paper cites When planning for movement, reason based on target object’s location and obstacles around you.

EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents When planning for movement, reason based on target object’s location and obstacles around you

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:24:45.404755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:24:45.343430Z digest=sha256:b24dfb633101a01654261456db966ae428b6b07792efe82282e907e59943f91b

Observation aa04e17b-1ded-4bf1-9cd1-2f1961248a69 · outbound

This paper cites In other words, do not overly focus on correcting invalid actions when direct movement toward the target object can still bring you closer.

EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents In other words, do not overly focus on correcting invalid actions when direct movement toward the target object can still bring you closer

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:24:45.407479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:24:45.343430Z digest=sha256:526b347c2d5c404b903f85d5e9b0c4c335ef50e23e56e81282fb06252b9db524

Observation 72944b41-f8aa-43b6-9ec0-a8cb8d3cadc9 · outbound

This paper cites If so, plan nothing but ONE ROTATION at a step until that object appears in your view.

EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents If so, plan nothing but ONE ROTATION at a step until that object appears in your view

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:24:45.410831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:24:45.343430Z digest=sha256:e9c057828474c8c75cd774bfae3d845fc62579fffa76bb789fe720b70089bc99

Observation 40c206ad-1f5e-46a5-9c25-0ff9c6852f30 · outbound

This paper cites red", "maroon.

EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents red", "maroon

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:24:45.413310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:24:45.343430Z digest=sha256:720c3c140a1555606f70a6773eafc4dd3d595f465c5d156e53ef77ea7493dd33

Observation 8097f701-ec5a-42d1-9348-8ca586228f76 · outbound

This paper cites There are two copper-colored pots visible on the stovetop.

EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents There are two copper-colored pots visible on the stovetop

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:24:45.415740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:24:45.343430Z digest=sha256:15a36bd617d53e5b6c8d61750a02a3314a66015e32ad9377a048a7ade9fbe830

Observation a08ff14e-16e8-4dae-b047-36d8e03a1dc4 · outbound

This paper cites pick up the spoon3.

EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents pick up the spoon3

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:24:45.418183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:24:45.343430Z digest=sha256:2164d91bb7309c351241c02ff31dedeb8ea2f87ff99c23b4afcdcfbdd9b85904

Observation acb841d4-5da2-4479-b6ed-3428036d8655 · outbound

This paper cites pick up the sponge9.

EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents pick up the sponge9

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:24:45.420522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:24:45.343430Z digest=sha256:aa67b9ab863118087b9875c3abccebba154e69d4781e4bc433572c28d629e481

Observation 9d54dd95-e210-44a7-81ba-c1d0bcfb4efb · outbound

This paper cites action_id.

EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents action_id

Reference 27

Resolution
malformed identifier
raw_fallback, observed 2026-05-17T00:24:45.423186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:24:45.343430Z digest=sha256:2a8cb2abe73999f0416c1441f89a648f9b0927cf61275afec997136c0e4a67ea

Observation 7cc72fbe-ec6d-4038-8692-4a40cfdd318f · outbound

This paper cites Rotate to the left by 90 degrees18.

EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents Rotate to the left by 90 degrees18

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:24:45.425569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:24:45.343430Z digest=sha256:25413671bbe70daa0d6a9cb47341ab5fe1e14f6b112a6566800c4481d9a864df

Pith citing papers

Observation cd0faa7a-c349-44bf-baa1-866bfa5a84ca · inbound

From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review cites this paper.

From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Reference 124

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:24:45.426358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T02:57:37.873567Z digest=sha256:fb3f70aab67932e344eb5017775878d9a12f2f9cde1d11051579a433ea169a79

Observation 1c01cf7b-0cb0-43b0-a351-1d7b11dc2069 · inbound

VS-Bench: Evaluating VLMs for Strategic Abilities in Multi-Agent Environments cites this paper.

VS-Bench: Evaluating VLMs for Strategic Abilities in Multi-Agent Environments EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Reference 83

Resolution
verified exact
local_arxiv, observed 2026-05-19T11:57:16.257323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:57:08.314088Z digest=sha256:4009362986e66c70c26ae83d70deabce7b7a211fde24e6418da4e5d877346f97

Observation 2fbb1055-bac0-4de1-afab-f3b463cd6f0b · inbound

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence cites this paper.

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Reference 113

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T00:24:45.426358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-14T22:23:14.621091Z digest=sha256:bade9b00b9889c5f732990cdd123a32b9c6b58045c26df0c81f11df2fe3d6d5c

Observation 2ce6b19d-f7b0-40cc-bf15-783a3356c17d · inbound

RoboGPT-R1: Enhancing Robot Task Planning with Reinforcement Learning cites this paper.

RoboGPT-R1: Enhancing Robot Task Planning with Reinforcement Learning EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-04T09:33:41.839773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:33:41.839773Z digest=sha256:86a0bbfe167c819118916f43ba2b230653530f46c843865be8aa123958bff858

Observation 0c3c7d1d-2359-447e-acd6-b3946edd022e · inbound

World Simulation with Video Foundation Models for Physical AI cites this paper.

World Simulation with Video Foundation Models for Physical AI EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Reference 90

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:24:45.426358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:a56f165ec723cac5c18f4b28bbfd118d08d4afc2ecc07b328422073e9dc02479

Observation 001bf018-0459-48d5-a0b6-fc31a33ace18 · inbound

SWITCH: Benchmarking Modeling and Handling of Tangible Interfaces in Long-horizon Embodied Scenarios cites this paper.

SWITCH: Benchmarking Modeling and Handling of Tangible Interfaces in Long-horizon Embodied Scenarios EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T21:14:11.784486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:14:11.784486Z digest=sha256:917df47ba53811ccc463f726586767a710a80b0ccdc00fd53cccf1c783f28d59

Observation e46753b1-8384-4fb1-8788-a17ba1a5348e · inbound

Beyond Description: Cognitively Benchmarking Fine-Grained Action for Embodied Agents cites this paper.

Beyond Description: Cognitively Benchmarking Fine-Grained Action for Embodied Agents EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T20:42:56.921728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:42:56.921728Z digest=sha256:7c0fac1183259077b8aed073dcef86f876bbcfc2661920634da333b18797b71e

Observation 6c686cc6-2972-49c6-a179-b04d943a1882 · inbound

Vision-Language Memory for Spatial Reasoning cites this paper.

Vision-Language Memory for Spatial Reasoning EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-03T20:15:35.989249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:15:35.989249Z digest=sha256:dd6cf6210cbaa0bcae484d75726d65653f10c9056d21eaf81be99aef31616545

Observation 9c4ac8e5-6d71-49fb-a73a-d9d4a7249104 · inbound

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents cites this paper.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:11:29.855452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T03:09:31.161760Z digest=sha256:49b7610c9a626eb443ccc4815c9e008e4a9b95efafa5474b15189d27f5214abc

Observation f552b55e-3a86-4cb2-9843-7e8bfaf3efeb · inbound

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents cites this paper.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:48.033515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:48.033515Z digest=sha256:105bc11683c029bd31509aad438705a1f1eeed300998ad2f9ffcfdf77eeba7ca

Observation 9718d69a-9d11-4238-a967-d9ec96a96563 · inbound

Unified Embodied VLM Reasoning with Robotic Action via Autoregressive Discretized Pre-training cites this paper.

Unified Embodied VLM Reasoning with Robotic Action via Autoregressive Discretized Pre-training EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-03T13:29:51.752016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:29:51.752016Z digest=sha256:b806fec6fe93a18afe3b805507a8ef951ea5497782a9bb0f496d5ffc0945c458

Observation c503a699-ab9c-40d3-8e94-86609930ed8f · inbound

TCAP: Tri-Component Attention Profiling for Unsupervised Backdoor Detection in MLLM Fine-Tuning cites this paper.

TCAP: Tri-Component Attention Profiling for Unsupervised Backdoor Detection in MLLM Fine-Tuning EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-05-25T07:35:28.658157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-25T07:33:15.865358Z digest=sha256:939e6918b8827efe48b053e313187d5aa4cfc06048394546c4c3c96240bc0532

Observation 96c42348-e4f9-4d2a-9dd5-373bf61df92a · inbound

PLanAR: Planning-Language-Grounded Agentic Reasoning for Robot Manipulation cites this paper.

PLanAR: Planning-Language-Grounded Agentic Reasoning for Robot Manipulation EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T05:38:50.023173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:38:50.023173Z digest=sha256:a0612a68a14acfb9b7cca4bf1a230b475d9942cc7f0fa4ebb1bf5bf2ff1f20a7

Observation 134909e8-416c-4e2d-bab2-370b3069cce1 · inbound

ST-BiBench: Benchmarking Multi-Stream Multimodal Coordination in Bimanual Embodied Tasks for MLLMs cites this paper.

ST-BiBench: Benchmarking Multi-Stream Multimodal Coordination in Bimanual Embodied Tasks for MLLMs EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Reference 121

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:24:45.426358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T06:00:02.043029Z digest=sha256:8aaad45701e700115f2edd6b88b3cdb5d69247a21d019659e3288fbc857ebe34

Observation 0275c72d-a767-486b-becb-64c8b9991453 · inbound

LLM-WikiRace Benchmark: How Far Can LLMs Plan over Real-World Knowledge Graphs? cites this paper.

LLM-WikiRace Benchmark: How Far Can LLMs Plan over Real-World Knowledge Graphs? EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T22:26:39.032152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T22:26:39.032152Z digest=sha256:e4ee8e1a98ec255c9085771a7cb2977bfba5a88079b1046395c2818750deb502

Observation b1c5c87f-0953-4960-91d0-503c13e898c5 · inbound

MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs cites this paper.

MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Reference 91

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:24:45.426358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:12:46.385646Z digest=sha256:562f502b00870fa7be23b610853cc35fd5c35da194ab1232e21e764d06d44377

Observation cb1a712b-eeda-4722-a10e-094923d1432e · inbound

MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs cites this paper.

MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Reference 91

Resolution
verified exact
local_arxiv, observed 2026-05-22T10:31:25.382916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T10:30:06.829915Z digest=sha256:fea897c30c730842ce440b2d94e42590b429824b4568b08a9653868f4447c645

Observation e8579d2d-49ba-4eb7-b5a9-329f5edd818a · inbound

MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs cites this paper.

MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-02T22:00:12.003747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:00:12.003747Z digest=sha256:c1fea51caa372684cfdbd3640b394caf08923c42b2aa8aa21112135b06ac46e9

Observation b79e4a07-6e16-4773-b3f8-617128355912 · inbound

Evaluation as Evolution: Transforming Adversarial Diffusion into Closed-Loop Curricula for Autonomous Vehicles cites this paper.

Evaluation as Evolution: Transforming Adversarial Diffusion into Closed-Loop Curricula for Autonomous Vehicles EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:24:45.426358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:55:30.816224Z digest=sha256:c20dc6abe1e0f588ec3a733008571b6bd17ae2298043f2489cfdb4abdd77b111

Observation 941cf336-ccfb-4be9-8e3c-6ccd52b2dcf8 · inbound

Online Reasoning Video Object Segmentation cites this paper.

Online Reasoning Video Object Segmentation EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T00:24:45.426358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:29:03.843441Z digest=sha256:79ff33979730d89085d7afdcce99d83133a9d4123a9f627dfa0a73c31dd41d86

Observation 0c4c9b45-3d9f-4660-81e4-5f8065c3b317 · inbound

ESCAPE: Episodic Spatial Memory and Adaptive Execution Policy for Long-Horizon Mobile Manipulation cites this paper.

ESCAPE: Episodic Spatial Memory and Adaptive Execution Policy for Long-Horizon Mobile Manipulation EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T00:24:45.426358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T13:09:32.132811Z digest=sha256:74667c1b5c6b23bdc1562132605c4da391a5817e8e4d371cfa1236d73e14bcc6

Observation e4e02e69-4502-4cd0-a939-dcc4063492cc · inbound

MirrorBench: Evaluating Self-centric Intelligence in MLLMs by Introducing a Mirror cites this paper.

MirrorBench: Evaluating Self-centric Intelligence in MLLMs by Introducing a Mirror EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:24:45.426358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T10:53:48.374637Z digest=sha256:c413d4ee24822bf0d9bd4ce02d456da80abda3aba622e6c3124c5c764af47051

Observation 9f71d0a8-205d-4d73-b998-9c77e60e71ee · inbound

BrainMem: Brain-Inspired Evolving Memory for Embodied Agent Task Planning cites this paper.

BrainMem: Brain-Inspired Evolving Memory for Embodied Agent Task Planning EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T00:24:45.426358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T11:52:58.630590Z digest=sha256:0dd9f1db79bdf30ec3cc4fe6465dea4060fd00e95b5a36bd9b6c50e4ce8f8b68

Observation a29bcb38-f12a-4216-bd2c-f3aeee16f432 · inbound

Chain Of Interaction Benchmark (COIN): When Reasoning meets Embodied Interaction cites this paper.

Chain Of Interaction Benchmark (COIN): When Reasoning meets Embodied Interaction EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:24:45.426358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:50:09.031009Z digest=sha256:a7b6e3e39e24d2a5d356f120aa9455697e8684be5ebce760ed44b990a65501ad

Observation eac2dfab-431f-4838-84c8-48b4b5a2c81f · inbound

GaLa: Hypergraph-Guided Visual Language Models for Procedural Planning cites this paper.

GaLa: Hypergraph-Guided Visual Language Models for Procedural Planning EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T00:24:45.426358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-10T06:26:06.208036Z digest=sha256:3a05296567d94dec13e5edb36f204e50bfec0083a50b89b9fabc1d5f23e1089d

Observation 55884cfe-fbf7-41af-a01c-d9a6181c63f4 · inbound

Environmental Understanding Vision-Language Model for Embodied Agent cites this paper.

Environmental Understanding Vision-Language Model for Embodied Agent EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:24:45.426358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T03:36:52.685696Z digest=sha256:c7c06e7a29e12ae8e5e3245af6dff571521ca0722ab8fefc19ada0ba00074127

Observation 3389aba7-1e28-4e1d-aac3-d84a8eb2fd1f · inbound

Safactory: A Scalable Agentic Infrastructure for Training Trustworthy Autonomous Intelligence cites this paper.

Safactory: A Scalable Agentic Infrastructure for Training Trustworthy Autonomous Intelligence EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Reference 92

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:24:45.426358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T10:12:58.421050Z digest=sha256:acf1c15b46460715f974fc259bc309cd303a1355297c9988b4722972950e152d

Observation 2437dcb2-8887-46af-a3ac-3b02cdb6aefd · inbound

Safactory: A Scalable Agentic Infrastructure for Training Trustworthy Autonomous Intelligence cites this paper.

Safactory: A Scalable Agentic Infrastructure for Training Trustworthy Autonomous Intelligence EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Reference 92

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:24:45.426358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T00:56:48.838028Z digest=sha256:2b4786eace7dce235aeeddd9fc3e96ace72beacc969472317f77718d6a6998ff

Observation 5e155ef8-cdd4-4b86-8fb9-022f5a54d294 · inbound

MemCompiler: Compile, Don't Inject -- State-Conditioned Memory for Embodied Agents cites this paper.

MemCompiler: Compile, Don't Inject -- State-Conditioned Memory for Embodied Agents EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:24:45.426358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T02:36:53.775557Z digest=sha256:03bdbcaa29d21d6c72b341e57ba7a2c090b2ec6d5e2e92cd6d8454866046c5fe

Observation b37b9033-653c-4d26-b790-b11e552c08be · inbound

MemCompiler: Compile, Don't Inject -- State-Conditioned Memory for Embodied Agents cites this paper.

MemCompiler: Compile, Don't Inject -- State-Conditioned Memory for Embodied Agents EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:24:45.426358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T05:57:05.259111Z digest=sha256:1902948b6865d2022c1b7e4041620aa6f66208890d8bc69efa6a5e734f37d185

Observation 0246b88a-f056-425e-9c3f-0ac5c4bef107 · inbound

Done, But Not Sure: Disentangling World Completion from Self-Termination in Embodied Agents cites this paper.

Done, But Not Sure: Disentangling World Completion from Self-Termination in Embodied Agents EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:24:45.426358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:23:37.488121Z digest=sha256:d9c016f8489c730e7d82176de0c8fe1db25a6fb307c1c6444d8994442614a7dd

Observation 3bfa3d12-6d43-4469-bee3-b4bd9c966058 · inbound

Done, But Not Sure: Disentangling World Completion from Self-Termination in Embodied Agents cites this paper.

Done, But Not Sure: Disentangling World Completion from Self-Termination in Embodied Agents EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:24:45.426358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:03:30.995808Z digest=sha256:7707c995f616733c67d665ea75ddccc63daa49caf32b7c703109b74f05aaf9bd

Observation 6ae12fb6-355e-4744-98df-f7c72b39cc31 · inbound

Done, But Not Sure: Disentangling World Completion from Self-Termination in Embodied Agents cites this paper.

Done, But Not Sure: Disentangling World Completion from Self-Termination in Embodied Agents EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:24:45.426358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T06:25:15.510083Z digest=sha256:ccc94320a68065192bd00d7109ac7eeeae0d4e8ca9ee44ad77f603bf74afcaf6

Observation 33fd9655-e68a-4a65-b101-b72f71d2bb8b · inbound

SceneFunRI: Reasoning the Invisible for Task-Driven Functional Object Localization cites this paper.

SceneFunRI: Reasoning the Invisible for Task-Driven Functional Object Localization EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:24:45.426358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T05:03:54.408677Z digest=sha256:9fb7f8bb27ed26e34cb9805c450bd9527ccb587f336a786afbb3afbc67668d39

Observation 3bfad6e8-56bb-4bb2-8fa6-a580031e4416 · inbound

CosFly-Track: A Large-Scale Multi-Modal Dataset for UAV Visual Tracking via Multi-Constraint Trajectory Optimization cites this paper.

CosFly-Track: A Large-Scale Multi-Modal Dataset for UAV Visual Tracking via Multi-Constraint Trajectory Optimization EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-20T11:23:14.211508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T11:18:31.696442Z digest=sha256:39f6b6643d3aefcdc5f31ccab14ff02217dfd0692c76ad7e97516d9d72c08765

Observation 0617a6c1-888f-49ab-9de0-57633cbab29f · inbound

CosFly-Track: A Large-Scale Multi-Modal Dataset for UAV Visual Tracking via Multi-Constraint Trajectory Optimization cites this paper.

CosFly-Track: A Large-Scale Multi-Modal Dataset for UAV Visual Tracking via Multi-Constraint Trajectory Optimization EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-21T09:04:04.596726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T09:00:46.267463Z digest=sha256:cc11b7da44589ede6fe96922a8c116bc263b62b83db7913cf90b937b53206160

Observation ea6908c6-e8a1-42ca-b176-5751481be93c · inbound

AtlasVA: Self-Evolving Visual Skill Memory for Teacher-Free VLM Agents cites this paper.

AtlasVA: Self-Evolving Visual Skill Memory for Teacher-Free VLM Agents EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Reference 51

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T11:28:14.479443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T11:24:48.558423Z digest=sha256:6505ce053bcac57cbb75b0ac70da9619fddefe781ea67e31107634953ef2e4d7

Observation 2bed4e9a-b408-4a18-bcde-da92fcdf242c · inbound

DexHoldem: Playing Texas Hold'em with Dexterous Embodied System cites this paper.

DexHoldem: Playing Texas Hold'em with Dexterous Embodied System EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-05-20T09:43:11.119191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T09:38:43.801252Z digest=sha256:8ca75e19df9df862f512430b0803c9d5c7fb328af05b1ef0b1337fb76b0c6c35

Observation 572d38ee-4d2e-4dab-97d9-c762949e15f2 · inbound

WorldMemArena: Evaluating Multimodal Agent Memory Through Action-World Interaction cites this paper.

WorldMemArena: Evaluating Multimodal Agent Memory Through Action-World Interaction EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T08:53:16.490784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T08:33:38.011236Z digest=sha256:7529be9ad6c5d0beae19d75836371574468291603f42a08adf5e34b4dc1ef6b0

Observation e4c219b1-1b7e-43f8-8449-93a2d4a6f311 · inbound

OpenWebRL: Demystifying Online Multi-turn Reinforcement Learning for Visual Web Agents cites this paper.

OpenWebRL: Demystifying Online Multi-turn Reinforcement Learning for Visual Web Agents EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:06:16.401189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T15:46:50.587684Z digest=sha256:1094aef21614c7c541182e5702cf83b04ddfcae9eae6b724630b918bcbb716d4

Observation 45dda8e8-3acf-491b-b450-a7a1226d1c4c · inbound

Sample-Efficient Post-Training for LEGO Spatial-Physics Reasoning cites this paper.

Sample-Efficient Post-Training for LEGO Spatial-Physics Reasoning EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Reference 44

Resolution
metadata mismatch
local_arxiv, observed 2026-06-28T23:42:49.557377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T23:38:38.127345Z digest=sha256:822de9fda10d7c866415c0c5829a1c6ee000a76edaddbfe3845e340245433959

Observation 5541e0bd-d6fb-49d9-9bbf-de7e114fb2f1 · inbound

ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? cites this paper.

ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-07-02T20:27:22.654943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T20:23:18.667677Z digest=sha256:317923f0f9311efb358d45ee7a42bb260468f3710148d19fca35eabd233b47a1

Observation e4bdfe65-5a21-4455-9f69-007194d073a2 · inbound

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks cites this paper.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Reference 82

Resolution
verified exact
local_arxiv, observed 2026-07-03T01:27:30.613186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:2879c884a868bdf5358a5ca1a1760b76c429002ef222b85c8617accb178926b1

Observation 5fba39a7-eb8e-43a2-9743-d7a4f974a161 · inbound

Embodied-BenchClaw: An Autonomous Multi-Agent System for Embodied Spatial Intelligence Benchmark Construction cites this paper.

Embodied-BenchClaw: An Autonomous Multi-Agent System for Embodied Spatial Intelligence Benchmark Construction EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-03T10:48:02.874775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T09:48:49.786021Z digest=sha256:c579990cf111a8dd47103beb65b1932e6e2911fff0201a7c51e12a09c5bb3805

Observation 77b3e495-0d87-411f-a1c9-1adb18b4b53b · inbound

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning cites this paper.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Reference 103

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T10:48:03.120025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:c009d2c57c9ea1f6b43b52433e559198bfb93419b9296ec60779171d49181144

Observation d6be2457-669e-4684-8a08-77c73f53f3e0 · inbound

Intelligent Automation for Embodied Benchmark Construction: Pipelines, Embodiments, Simulators, and Trends cites this paper.

Intelligent Automation for Embodied Benchmark Construction: Pipelines, Embodiments, Simulators, and Trends EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Reference 43

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T11:58:05.940585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-27T09:17:50.473747Z digest=sha256:e131841b006a4c552d7d574b32e49266c3206dde00a79073d4c692504f63f025

Observation e5c0c01f-3b83-4292-8ff4-68cc6d9a4205 · inbound

GroundControl: Anticipating Navigation Failures in Vision-Language Agents via Trajectory-Consistent Uncertainty Estimates cites this paper.

GroundControl: Anticipating Navigation Failures in Vision-Language Agents via Trajectory-Consistent Uncertainty Estimates EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-07-04T04:09:34.272698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T17:13:47.070153Z digest=sha256:fdb303a7a40f344267179d8b293a97975c09654681f7695b038c5f3f27da7db7

Observation f773a281-2f9c-4f7e-a320-495ce06fbbb2 · inbound

MultiUAV-Plat: An LLM-Oriented Platform, Benchmark and Framework for Multi-UAV Collaborative Task Planning cites this paper.

MultiUAV-Plat: An LLM-Oriented Platform, Benchmark and Framework for Multi-UAV Collaborative Task Planning EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Reference 33

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T06:15:26.578752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T06:09:22.431356Z digest=sha256:6960eebbcb3facd637766cc5d3fc54630e00485671a5c36e75b16665052cf993

Observation dbb92018-0b8a-4cb0-8eed-9f322a2f62f9 · inbound

Multi-scale Mixture of World Models for Embodied Agents in Evolving Environments cites this paper.

Multi-scale Mixture of World Models for Embodied Agents in Evolving Environments EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Reference 19

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T13:06:58.502976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-02T13:01:58.866764Z digest=sha256:f928a3b944469eb7d955942c52c30e1d88a3a555fad95e89120fc5c6da3b3b9b

Observation ed3391f0-0627-4f70-ab1b-fd1ab6c758c0 · inbound

LLM Agents for Deliberative Collaboration: A Study on Joint Decision Making Under Partial Observability cites this paper.

LLM Agents for Deliberative Collaboration: A Study on Joint Decision Making Under Partial Observability EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Reference 174

Resolution
verified exact
local_arxiv, observed 2026-07-08T15:05:03.545431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-08T15:03:14.483228Z digest=sha256:bbee37ed53cfd6eaaf38a76711cf4d755abba2505f13e5bec5b19080b54d5512

Observation 5b528198-cced-4cf0-9480-ec9bea17795a · inbound

Who&When Pro: Can LLMs Really Attribute Failures in AI Agents? cites this paper.

Who&When Pro: Can LLMs Really Attribute Failures in AI Agents? EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-14T01:10:44.920887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T01:10:44.920887Z digest=sha256:82665c8d2df2ce63ebd72b633910b8a278764ddd2a16e1c11cd426b419c6e052

Observation f0de9766-8888-461b-9763-29a1d24acac7 · inbound

ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory cites this paper.

ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Reference 105

Resolution
unresolved
no resolver link, observed 2026-07-14T12:26:27.446079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:26:27.446079Z digest=sha256:b8221df4ef1dc24545b698c7c5cbc3da0ecda0e44e9f4312cc03d509274d0371

Observation 8e922246-0a77-497f-8128-b5740718dd89 · inbound

ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory cites this paper.

ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Reference 105

Resolution
unresolved
no resolver link, observed 2026-08-02T07:21:35.065643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:21:35.065643Z digest=sha256:cbe36252987841510b36276ee879d680dfebaa31643950abd9e3c0df6f24f4a0

Observation 44386acd-0f27-472c-9eb1-2b18717fd673 · inbound

Leveraging Trajectory Graphs for Pre-Execution Error Diagnosis in Agentic LLM Systems cites this paper.

Leveraging Trajectory Graphs for Pre-Execution Error Diagnosis in Agentic LLM Systems EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Reference 67

Resolution
unresolved
no resolver link, observed 2026-07-31T00:46:12.910645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:46:12.910645Z digest=sha256:881f1de5bc199597d9b409491ed3c4f6265116f1508a52532de1c88af5b08398

Observation 437c2389-511d-4f3b-ada4-e114328adeaa · inbound

SafeNexus: Discovering and Steering Modality-Universal Safety Neurons in MLLMs cites this paper.

SafeNexus: Discovering and Steering Modality-Universal Safety Neurons in MLLMs EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-03T16:20:30.506056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T16:20:30.506056Z digest=sha256:ba380d0ef91583a3ed231a58dc580b7c63c9cfafc2ebbbea2aa7ad6e9ede3406

Observation ac8bf5ec-4b22-432b-bf15-3e24b956d6e0 · inbound

Long-Horizon Embodied Decision-Making via Multimodal Memory Compression cites this paper.

Long-Horizon Embodied Decision-Making via Multimodal Memory Compression EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T00:13:24.699600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:13:24.699600Z digest=sha256:78290cbc5672a5be4c4963f3f3ccda6480cd106294980d76c9a2628b1abe6754

Observation 69fb0e21-b44b-4472-ac1b-fd15c76213df · inbound

Is Inter-Seed Cross-Play Enough? Evaluating the Robustness of Zero-Shot Coordination Algorithms to Implementation Details cites this paper.

Is Inter-Seed Cross-Play Enough? Evaluating the Robustness of Zero-Shot Coordination Algorithms to Implementation Details EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T15:25:39.444096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:25:39.444096Z digest=sha256:f8b990a88f816c3a659740380930256fcf6ff7d423f0be98d164bffbc7a4a021