Pith. sign in

Paper Citation Record · LEDGER

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

As of 9 August 2026, this Paper Citation Record lists 100 of 112 outbound references and 24 inbound Pith citation observations for arXiv:2506.00123.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.00123 v1

Coverage vector

measured 100 of 112 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:16:15.718768Z

measured 124 of 124 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 24 of 24 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:47:29.605186Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T15:28:34.757769Z

Reference resolution

100 of 112 outbound references displayed

  • verified exact1
  • verified fuzzy16
  • unresolved83
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1ee2d789-c94a-4f9b-a676-0cb0ed3ee942 · outbound

This paper cites GPT-4 Technical Report.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:08.094257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:08.094257Z digest=sha256:cafacf5d42a24283a2385ee221d5c34a28ad25a073ee5f144099e2b500b03d51

Observation 0674a851-2993-401b-9646-6fb3e9800e24 · outbound

This paper cites Do As I Can, Not As I Say: Grounding Language in Robotic Affordances.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Do As I Can, Not As I Say: Grounding Language in Robotic Affordances

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:08.184567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:08.184567Z digest=sha256:13c7a5666eb17970ac0e4cd110bab3b68dd44ce4600b619cd87c7aa352be0e6c

Observation 8518c5f4-2e1c-4473-9f98-374b29f9bc36 · outbound

This paper cites Menick, Sebastian Borgeaud, Andy Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikolaj Binkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, and Karén Simonyan.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Menick, Sebastian Borgeaud, Andy Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikolaj Binkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, and Karén Simonyan

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:08.271878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:08.271878Z digest=sha256:9676cad2a7f1cc988c9aaebecf47769ba89f898d50e4a368ee22845c3ce43601

Observation cb545fed-1e29-4216-b41d-865137d9a172 · outbound

This paper cites Claude 3.5 sonnet.https://www.anthropic.com/news/claude-3-5-sonnet, 2024.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Claude 3.5 sonnet.https://www.anthropic.com/news/claude-3-5-sonnet, 2024

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:08.351375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:08.351375Z digest=sha256:53da35c1e857b36f864fb01c8a140216ca5684cebd2cfced3a86024838bde09e

Observation 0e05a99f-c97b-4d81-8e45-9c47448f95e1 · outbound

This paper cites OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:08.414785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:08.414785Z digest=sha256:cee72247dcac59c4542181137df7fa8159d6056c290a829548fb13e91cd76164

Observation fc79de96-659e-4e20-85b5-93b6aabad814 · outbound

This paper cites Scanqa: 3d question answering for spatial scene understanding.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Scanqa: 3d question answering for spatial scene understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:08.495975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:08.495975Z digest=sha256:256f0118dcf268552256df00742ebaed86fece92b660ac0d66b2d0eadf06d189

Observation e5e7458e-0d66-4212-a353-de4c5255c4d4 · outbound

This paper cites Qwen Technical Report.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Qwen Technical Report

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:08.566118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:08.566118Z digest=sha256:e5768d3d40ad5342629b64ff6298b27f07c93d8b98d6b435af1048c88552cbcd

Observation 3d9a7022-2f3a-407a-8a0b-fd1d9156406f · outbound

This paper cites Qwen2.5-VL Technical Report.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Qwen2.5-VL Technical Report

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:08.626048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:08.626048Z digest=sha256:03b858b20c9afded707cfcb2da136d48bff6f1cabb504dbbb7ab775a03b62465

Observation c32fc78d-0066-4baf-b491-4b4462ee1345 · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:08.701912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:08.701912Z digest=sha256:54f2541c9eb0d3fbd8b2d9ffabc44f3a98a61dbb66309e01f0582c390e202085

Observation 38f2732e-a4a2-4351-b414-6069c7abdd9b · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces RT-1: Robotics Transformer for Real-World Control at Scale

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:08.777108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:08.777108Z digest=sha256:03db3b900ae9d63ffef488ad2f459278c4fc645b4840d968131c015a44fabd7a

Observation a0719468-b015-4571-b5ab-ce5ab0cee77c · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:08.844533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:08.844533Z digest=sha256:502d82c8ab5126f30261b61a5885795a6684134e6e82071495134c76cdb864a4

Observation 696b5cb8-bf54-4657-aabd-a0663a8265c3 · outbound

This paper cites Language models are few-shot learners.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Language models are few-shot learners

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:08.930024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:08.930024Z digest=sha256:96bbb68f8d51d156b833770284ef18c2b2731debbf48750ed77bdf3cf3c02880

Observation 0b5be8d8-f62e-46fc-830e-a9989623dece · outbound

This paper cites Commonsense Reasoning for Legged Robot Adaptation with Vision-Language Models.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Commonsense Reasoning for Legged Robot Adaptation with Vision-Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:09.000071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:09.000071Z digest=sha256:9242867f2e82afd21c76b952094e144f3ff6426454f4dc27f6bbe24cb1166e71

Observation 105cfb30-dcb7-4e70-8984-d9e75ec5ad05 · outbound

This paper cites Scanrefer: 3d object localization in rgb-d scans using natural language.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Scanrefer: 3d object localization in rgb-d scans using natural language

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:09.066166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:09.066166Z digest=sha256:e1b7e8433fb27ecd0dcee37749f646a80a343caef68bbb196ff06af5dc0c3257

Observation 8675d496-772e-46a3-bca9-59a9e49115ab · outbound

This paper cites Eagle 2.5: Boosting long-context post-training for frontier vision-language models.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Eagle 2.5: Boosting long-context post-training for frontier vision-language models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:09.119149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:09.119149Z digest=sha256:0033d7885c9e418acb25433d60e9f740cc0acf695ce03f7a656613928dbf4b47

Observation 1e55fa0c-5216-4a00-976a-ae8b94f5868f · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:09.213009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:09.213009Z digest=sha256:338412c7a2d178852a0e6f0481c2a808beb9d3d425a3cbec385dba9e1048fc70

Observation a4876270-6c15-403f-a4a6-07c8cb24c66d · outbound

This paper cites Sharegpt4v: Improving large multi-modal models with better captions.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Sharegpt4v: Improving large multi-modal models with better captions

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:09.277793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:09.277793Z digest=sha256:c5e297f76f7b857e2b3228cfc4f29b0b445b57ea7d609f7ce698c97d38bca46e

Observation d1407922-9ad2-4d3f-a646-e4f4e64c2731 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:09.351955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:09.351955Z digest=sha256:0c145f9a927c0108c13003a57272f7d8dd56bc80988bdc1a7bce63f180abd6b2

Observation 3984c22a-1ba4-4e61-8e53-46502b5649fb · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:09.427092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:09.427092Z digest=sha256:d883f50609c8e6bbc6543fd42c43b69b1d3bbb63d483ce68d8f6957be0ed0343

Observation faf511ce-2e13-437b-8378-3041049502fe · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:09.508644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:09.508644Z digest=sha256:2602cc2285f29e18d7d6c4d130171008967ef1a17656ba0467848d1fb5b4b9f6

Observation 89df5c1a-f92d-4388-af89-4fa46b24b6bb · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:09.590815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:09.590815Z digest=sha256:719a20ba720a6a39517069e6c753b69af44cc0c65ce6125ecb928404e909bd54

Observation 38816cc3-5d58-4c66-9782-0b54dcd0d5d0 · outbound

This paper cites NaVILA: Legged Robot Vision-Language-Action Model for Navigation.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:09.666040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:09.666040Z digest=sha256:9d64877ee015c64a84da7bbb094289fcce97965c20661f7549488e34bca6fb3f

Observation 75ef2d6e-407d-4877-85d3-9cfbbaa58328 · outbound

This paper cites Spatially-Aware Transformer for Embodied Agents.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Spatially-Aware Transformer for Embodied Agents

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:16:18.467967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:16:09.736172Z digest=sha256:7fd42117b1741f2bd3f9a46cc923be99f1fd0b35616b9559a88d27a1fe325e6f

Observation 80793765-2150-4de9-a05f-3f5f0f266ed7 · outbound

This paper cites Local all-pair correspondence for point tracking.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Local all-pair correspondence for point tracking

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:09.802384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:09.802384Z digest=sha256:a27c84052de3fd619a9097cdb03ebeb47bdce86bb10d3b35cae5f48026d632fa

Observation 8d538287-bf9c-432d-833a-3fc1eeb760c6 · outbound

This paper cites Simple and effective multi-paragraph reading comprehension.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Simple and effective multi-paragraph reading comprehension

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:09.857532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:09.857532Z digest=sha256:12b975dc37203a44756e2c5ee25bd3ffcfd70729c1375422e7eb9a1d1657ef03

Observation d85376bb-fe9a-4889-b1b1-c7529397f013 · outbound

This paper cites Scannet: Richly-annotated 3d reconstructions of indoor scenes.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Scannet: Richly-annotated 3d reconstructions of indoor scenes

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:09.945386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:09.945386Z digest=sha256:c2bba0ad8ee08f4cb967ef5a00108d672961b21a92fac95ef0abd04e1607384b

Observation d9cd709d-1052-43be-89e4-9d466b983348 · outbound

This paper cites Language modeling with gated convolutional networks.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Language modeling with gated convolutional networks

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:10.021970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:10.021970Z digest=sha256:83b7380ebba637aa5419acf6592dbe48d1ff918b7bbd86cd392f0db9e9a34d85

Observation 2d053136-7bcb-43c0-81e8-7dbfb24ed116 · outbound

This paper cites Quar-vla: Vision-language-action model for quadruped robots.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Quar-vla: Vision-language-action model for quadruped robots

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:10.102883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:10.102883Z digest=sha256:c00bb9f2de5231d3f9753eff62bf20fb79b3437bda9845701f03a6428ec5bc81

Observation e06e275e-ba11-47ba-aa0b-f6fc31c2dc26 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces An image is worth 16x16 words: Transformers for image recognition at scale

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:10.177298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:10.177298Z digest=sha256:d14192a850032c03d6d2c0373705daca5c998babee66e708e57c2a7ed1d6643f

Observation 6db3c8b1-01e1-4c50-9129-6fd6c48ccb2a · outbound

This paper cites PaLM-E: An Embodied Multimodal Language Model.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces PaLM-E: An Embodied Multimodal Language Model

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:10.243116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:10.243116Z digest=sha256:45482e1abf8a4592cf1e647ab764b0560470059868f679b276e867329bc2340b

Observation 2a0ae299-5844-406f-b090-fe6419e2ce68 · outbound

This paper cites Centernet: Keypoint triplets for object detection.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Centernet: Keypoint triplets for object detection

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:10.287797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:10.287797Z digest=sha256:73cb94aeef72e4aca64917b1a958ee85f4cab528362b0ed017dfc9a3e433207f

Observation f4629046-9ed0-43c9-80a3-4f6dfb2ca8a7 · outbound

This paper cites Interleave-vla: Enhancing robot manipulation with interleaved image-text instructions.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Interleave-vla: Enhancing robot manipulation with interleaved image-text instructions

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:10.334062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:10.334062Z digest=sha256:b9bc6abf4090c213beb7cdd8639efa1394614efa3547125cbb371dbace5e44ba

Observation a8224738-d488-4f4f-8a18-9997eb7fa620 · outbound

This paper cites Eva: Exploring the limits of masked visual representation learning at scale.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Eva: Exploring the limits of masked visual representation learning at scale

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:10.403215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:10.403215Z digest=sha256:ce6ae4cefe35b24747a4ff1bebf095bef1c33e8f0da7b936fcee3eeeaf22f3f4

Observation 7c64dda6-17a7-4488-b991-c933513eb539 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:10.499411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:10.499411Z digest=sha256:f532268a48d7e171fff0eb80fcc0b689d0520eac664df6e47430f16b03e167ed

Observation a21a4e9b-e31b-4fac-89a6-3bdabb2ecf9c · outbound

This paper cites Rlafford: End-to-end affordance learning for robotic manipulation.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Rlafford: End-to-end affordance learning for robotic manipulation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:10.597240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:10.597240Z digest=sha256:cd146a1004f57e20efde1891d1469096a7b896883a0bff4d3d588fe2ce4eea5c

Observation c49009cd-d63c-43e9-8140-b302563f7a46 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Making the v in vqa matter: Elevating the role of image understanding in visual question answering

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:10.648005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:10.648005Z digest=sha256:4de19291729d501e7f9ed4dbe2bf6cc50240f0ffe87900d81a0656043c4d5dda

Observation 442877b5-7d93-4dc8-9674-df4b5cbf5740 · outbound

This paper cites Stangl, Anhong Guo, Chi Lin, Kristen Grauman, Jiebo Luo, and Jeffrey P.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Stangl, Anhong Guo, Chi Lin, Kristen Grauman, Jiebo Luo, and Jeffrey P

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:10.708199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:10.708199Z digest=sha256:adf6d7fae4cc4747734f0414ce48606e11d351f41e29184e73b92a822490afd8

Observation dee8501e-5fbc-4928-8892-5c0aed834f35 · outbound

This paper cites Girshick.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Girshick

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:10.791941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:10.791941Z digest=sha256:59b7447d85be687212b53fa11e6699807ac1f8d1a281e8633e64fd8669db7fa7

Observation d24edc9a-76f8-4ada-9938-76c194bc2301 · outbound

This paper cites 3d-llm: Injecting the 3d world into large language models.Advances in Neural Information Processing Systems , 36: 20482–20494, 2023.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces 3d-llm: Injecting the 3d world into large language models.Advances in Neural Information Processing Systems , 36: 20482–20494, 2023

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:10.876192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:10.876192Z digest=sha256:d9d2cca764b0ec10059b261b679c65fbbb96b1a987787cca6a3a032acb211c89

Observation ae01a1b4-d01c-473e-8aa0-0758d5d666d8 · outbound

This paper cites Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:10.945513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:10.945513Z digest=sha256:2433f9eb86928e8ce2a36d78f708881e96faa5c78ec0f28e9a391fbf7338d787

Observation 83ed7cf1-98a9-455c-92d6-50f5faa34813 · outbound

This paper cites Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:11.043939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:11.043939Z digest=sha256:37520eb9fdf0debf798c2612dfbadf691943587f00152b9e9d76e1b21ea8ef96

Observation 10c4ad44-4a1e-4e28-a4ca-ca0c0ec860da · outbound

This paper cites Chat-3d v2: Bridging 3d scene and large language models with object identifiers.CoRR, 2023.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Chat-3d v2: Bridging 3d scene and large language models with object identifiers.CoRR, 2023

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:11.134735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:11.134735Z digest=sha256:a4d66c6a658d52fde6cb5dc02e23bb927a331ac4b3bd18ae824bed23a73113c2

Observation cc41a651-4953-4a73-932e-78c486576128 · outbound

This paper cites An Embodied Generalist Agent in 3D World.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces An Embodied Generalist Agent in 3D World

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:11.217497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:11.217497Z digest=sha256:f028a3c892b445d3bc4f13ff14e67f6482090426c8227c1f9a9a83f081b5445c

Observation c08a66b9-b8f2-4040-b0e6-1a299e52ec87 · outbound

This paper cites Inner Monologue: Embodied Reasoning through Planning with Language Models.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Inner Monologue: Embodied Reasoning through Planning with Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:11.289544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:11.289544Z digest=sha256:e63e988afa2b81af19247409a1e71862ffd981f4f46611e4d5c93325066655da

Observation 834ff20c-66be-4c9f-b9db-6b92c79dadcc · outbound

This paper cites Bc-z: Zero-shot task generalization with robotic imitation learning.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Bc-z: Zero-shot task generalization with robotic imitation learning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:11.361561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:11.361561Z digest=sha256:933c97e249bab6ff39f7224cd3089cb615dbe8770f0623e5372845a95a67d8ad

Observation 77e4ec0b-04f0-4385-a4a5-3a983426f210 · outbound

This paper cites RoboBrain: A Unified Brain Model for Robotic Manipulation from Abstract to Concrete.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces RoboBrain: A Unified Brain Model for Robotic Manipulation from Abstract to Concrete

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:11.470944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:11.470944Z digest=sha256:7755a710c392b8ddd350ebba0831ec44d3adf5a4f070025631eec85acea6fa9c

Observation f7dafc7f-ff80-4db9-b147-1bce094aaacd · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Scaling up visual and vision-language representation learning with noisy text supervision

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:11.541136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:11.541136Z digest=sha256:1fd8c55cb7bd58c294a0dabd4a51386d97db3be193d9c1d464b1ed5c53c8ab52

Observation 44b23ff9-3d8f-47d5-a5bf-9b174e55b26c · outbound

This paper cites A diagram is worth a dozen images.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces A diagram is worth a dozen images

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:11.628090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:11.628090Z digest=sha256:a1c9f277190cbb3f59edd534a3ab423b1a44978007597cc32825e7a93f32d687

Observation bf0d064b-94e1-4314-b8cc-85003f6d3a3e · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces OpenVLA: An Open-Source Vision-Language-Action Model

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:11.687388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:11.687388Z digest=sha256:d219e38b1dec5d543ab0e056850b35199510da835765a8456c0334e62629da94

Observation 22f767ef-56a2-4155-a22e-5bbff6497ade · outbound

This paper cites LISA: Reasoning Segmentation via Large Language Model.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces LISA: Reasoning Segmentation via Large Language Model

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:11.756541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:11.756541Z digest=sha256:c0db3870596aaf7baf37e9d5ad73069927e18638f1926a2276f38f89550c51db

Observation fe47318a-98c5-404b-89ef-ab3a454e34fc · outbound

This paper cites Cornernet: Detecting objects as paired keypoints.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Cornernet: Detecting objects as paired keypoints

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:11.826342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:11.826342Z digest=sha256:1e7f49800c3124c11987db1ec930f86e027199e7d71001fa31b7f3ab78a70ed3

Observation 89609035-90f0-40cf-a3e8-4f2913f97ebc · outbound

This paper cites Learning hand-eye coordination for robotic grasping with deep learning and large-scale data collection.The International journal of robotics research, 37(4-5):421–436, 2018.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Learning hand-eye coordination for robotic grasping with deep learning and large-scale data collection.The International journal of robotics research, 37(4-5):421–436, 2018

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:11.899475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:11.899475Z digest=sha256:072e26466b403330959d588c67e369d53d733ada17d234e9c2b00e8be3759e8c

Observation a87c0ae9-e60d-4481-af81-960fba18264f · outbound

This paper cites SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:11.985639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:11.985639Z digest=sha256:81c3a56755921fc18bc2369782d86c906637853bd4a2ff9bcadfea96e897f80a

Observation fb31d28c-6296-4572-9949-15950c5d4bab · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces LLaVA-OneVision: Easy Visual Task Transfer

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:12.062017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:12.062017Z digest=sha256:38c08d6d9d5aeb55815560765fe31d99bdcb077ac4feca610c1072dd0f62bc02

Observation 2d328658-9d72-4246-910b-1096cea908fd · outbound

This paper cites Learning agile skills via adversarial imitation of rough partial demonstrations.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Learning agile skills via adversarial imitation of rough partial demonstrations

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:12.141295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:12.141295Z digest=sha256:623fccece05a6131692a549c7925c1212e985c98ed0eb593cebefc3a8d334ab7

Observation 96670f96-a2ab-47ea-be94-64bf0c89c8af · outbound

This paper cites an unresolved cited work.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Unresolved cited work

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:12.205947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:12.205947Z digest=sha256:6f5d955b767f63e874bea679b84390ccbc9413c46cfce9d22f9116d4187b3d66

Observation e3e139c2-7507-4bfa-abf1-f3de348530d5 · outbound

This paper cites Omnicorpus: A unified multimodal corpus of 10 billion-level images interleaved with text.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Omnicorpus: A unified multimodal corpus of 10 billion-level images interleaved with text

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:12.285277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:12.285277Z digest=sha256:faf0a9e8d0052f9e547044d75049cfe11abd3b7542c92d7cbb7b0f207fe51d36

Observation b3524013-2008-4470-84d9-193b96830b34 · outbound

This paper cites Exploring Plain Vision Transformer Backbones for Object Detection.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Exploring Plain Vision Transformer Backbones for Object Detection

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:12.363276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:12.363276Z digest=sha256:8361bfaf24d91a8bf81e5d47a409a9dc384177410c48c8377d812edea5180951

Observation 2c551f63-1ff8-4ab4-af58-6f087d92ba59 · outbound

This paper cites Robotic Visual Instruction.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Robotic Visual Instruction

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:12.453068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:12.453068Z digest=sha256:779aa3aefdafa5cdb8ff5b34a67ae951f2838ddf0bcb8a7d4b63bf26519f3b95

Observation 8f210c89-75f5-4c59-8221-675fa03e43a8 · outbound

This paper cites Code as policies: Language model programs for embodied control.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Code as policies: Language model programs for embodied control

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:12.537280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:12.537280Z digest=sha256:b4de281667fd0e89b53fd90f755ed2677d63592e3c264e856b799816a833a9eb

Observation 1f13fe24-c711-4692-bfd3-03c1f25cdbde · outbound

This paper cites Vila: On pre-training for visual language models.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Vila: On pre-training for visual language models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:12.611092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:12.611092Z digest=sha256:b8c213562acf4926a92a9672941046ee3e6ce50309134d4206f67a3df0fac890

Observation e8a3dc0b-fc1d-446d-ac2b-f7452a92d593 · outbound

This paper cites Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:12.689514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:12.689514Z digest=sha256:0d848607d170be63006233bbdc5349f32209b4a2c0360ff01c7770ab3a2c3482

Observation 06ca9977-97ca-4c6e-a229-b1cf9a2edc0a · outbound

This paper cites Coarse Correspondences Boost Spatial-Temporal Reasoning in Multimodal Language Model.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Coarse Correspondences Boost Spatial-Temporal Reasoning in Multimodal Language Model

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:12.740878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:12.740878Z digest=sha256:4520fecd0d4d11be9d6a208966e627a70630b66f0db4d093768f6e5fa3015934

Observation 3daa9635-42f0-475b-8fd9-4acffa328232 · outbound

This paper cites Visual instruction tuning.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Visual instruction tuning

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:12.799211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:12.799211Z digest=sha256:3efdb80a45b192042713845edc28df2df177fa9966b0291e70c9e4facc04df14

Observation 36f2259d-718a-4b25-8a9e-45b2f0ae47fc · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces MMBench: Is Your Multi-modal Model an All-around Player?

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:12.881890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:12.881890Z digest=sha256:5a10383ace6c105e9b2501ee3d22738c8593d9cc6792aa745406cf1c852bb1bd

Observation 689a1580-0f37-41e1-a0c9-08676e08dfa2 · outbound

This paper cites OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:12.944251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:12.944251Z digest=sha256:e7ce92be04cdf04645ba1d74d791976e98a63a2dc5983c76b459995bcabd966a

Observation 1a306407-9e2e-46b0-b61d-48280076fd1e · outbound

This paper cites Mminstruct: A high-quality multi-modal instruction tuning dataset with extensive diversity.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Mminstruct: A high-quality multi-modal instruction tuning dataset with extensive diversity

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:22.878623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:16:13.046266Z digest=sha256:d59cf36d58085dcbba9d621ed6c79a5ac4790302fbf565c4a457f6ffd7122352

Observation 9e31acdc-ec99-411c-8518-ebe57bced002 · outbound

This paper cites InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:13.117762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:13.117762Z digest=sha256:34d5abc0916e0a2ad201aefd35bd4848471d2d4f4c979800e4a8dff562e53182

Observation 8ef13d95-c7ca-4dad-b000-c89cc7399d84 · outbound

This paper cites Mono-InternVL: Pushing the Boundaries of Monolithic Multimodal Large Language Models with Endogenous Visual Pre-training.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Mono-InternVL: Pushing the Boundaries of Monolithic Multimodal Large Language Models with Endogenous Visual Pre-training

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:13.209440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:13.209440Z digest=sha256:03a6aa4e18bd9d7b25276ab7dafd63a19304d1c1263bab2ce940d27511875491

Observation 5e872cd2-3e60-48b5-b16e-e9c2b767fd75 · outbound

This paper cites SQA3D: Situated Question Answering in 3D Scenes.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces SQA3D: Situated Question Answering in 3D Scenes

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:13.281354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:13.281354Z digest=sha256:1314b9d3bb61b2d136a092383596bb256a50e055f0da18170d4798f838cb08c0

Observation 3213c8d1-cdf6-4d3c-ad23-d6aa78d04de1 · outbound

This paper cites Where are we in the search for an artificial visual cortex for embodied intelligence? Advances in Neural Information Processing Systems , 36:655–677, 2023.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Where are we in the search for an artificial visual cortex for embodied intelligence? Advances in Neural Information Processing Systems , 36:655–677, 2023

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:22.723933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:16:13.345167Z digest=sha256:9134090e22f4f4fc90ffa11f4ee1f6b1798b959fbcc7f1e32f0ab35d3e0f7be4

Observation 0efa2d26-a987-441b-8041-4ffb0a5eeab3 · outbound

This paper cites Walk these ways: Tuning robot control for generalization with multiplicity of behavior.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Walk these ways: Tuning robot control for generalization with multiplicity of behavior

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:22.514359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:16:13.401055Z digest=sha256:53f731495ec90951cb28cd5758f5b05fc3f31321152582e5c969f8c3bcafaf21

Observation 060f98ff-3b53-4e2c-a68b-e3d12ac5f5be · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Ok-vqa: A visual question answering benchmark requiring external knowledge

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:22.241222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:16:13.458663Z digest=sha256:e838a517322ce3cfe18b0a41ba104b79fcf3ab636b7c26b5778401e39239cc20

Observation 8129ca85-7683-47cb-b8ce-15aea4b9dbb6 · outbound

This paper cites Chartqa: A benchmark for question answering about charts with visual and logical reasoning.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Chartqa: A benchmark for question answering about charts with visual and logical reasoning

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:22.059153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:16:13.513833Z digest=sha256:b511eb0932883121bf9a7e42e6c5d79cc9dff9af7b789a1bb61123802dce39f0

Observation 3251f2c2-46c5-4ac7-92ad-062f19267b9b · outbound

This paper cites Infograph- icvqa.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Infograph- icvqa

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:21.829698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:16:13.613661Z digest=sha256:5d06f55169d6306fe8a8f5a16be87890d46e723f2c474a3d9190e0e880ae83aa

Observation 453ef401-645e-4c86-b12e-a34b569110d1 · outbound

This paper cites QuadrupedGPT: Towards a Versatile Quadruped Agent in Open-ended Worlds.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces QuadrupedGPT: Towards a Versatile Quadruped Agent in Open-ended Worlds

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:13.696485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:13.696485Z digest=sha256:b36dbb9a55fbdb93fe76998ae304ef866b4006e4592101298e2d9039869ec6a3

Observation 93073953-f24b-452a-846b-123bb4f6078a · outbound

This paper cites Embodiedgpt: Vision-language pre-training via embodied chain of thought.Advances in Neural Information Processing Systems, 36:25081–25094, 2023.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Embodiedgpt: Vision-language pre-training via embodied chain of thought.Advances in Neural Information Processing Systems, 36:25081–25094, 2023

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:21.643942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:16:13.764571Z digest=sha256:d3fcf9794adc4d92687003122c19fb8f0efad56d315d387d718be61b39c3d1f1

Observation 20ec6a30-082c-4364-9530-38aabc8e2478 · outbound

This paper cites R3M: A Universal Visual Representation for Robot Manipulation.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces R3M: A Universal Visual Representation for Robot Manipulation

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:13.840242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:13.840242Z digest=sha256:44fc412da49230d5670903860eb7065931c015e556ba62b746a7c97476a440c0

Observation 59d906ba-aa13-49cb-bb2d-79eef539f927 · outbound

This paper cites Gpt-4o system card.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Gpt-4o system card

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:21.324300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:16:13.906541Z digest=sha256:577b3feecf10d65a2942350cc24281852b272de4f989718678cfbdb4cb4153bf

Observation 6fa5edb8-db68-4971-9e85-c132ddab602c · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:13.992330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:13.992330Z digest=sha256:0245e7930a84188ce36b60c82683eb6b0109338753e9c7bb482ca4681bdb34ba

Observation 859eb22a-392a-4190-b6e6-45f594123c8b · outbound

This paper cites GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:14.104148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:14.104148Z digest=sha256:17486df099d792c90f63eef19b9305ca45cc0b88cc19dbe0a4d48175294c324b

Observation f0efc46c-7e1b-425c-8ae2-b32a44a583ca · outbound

This paper cites Improving language understanding by generative pre-training.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Improving language understanding by generative pre-training

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:21.114182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:16:14.201504Z digest=sha256:13f71fe958fbeae2c194eb7ff3bc8b73ee828b0a290ad45ef0a0652b203b0fa7

Observation 5c89e3d8-d4fe-4810-833e-1931d3ffc11c · outbound

This paper cites Language models are unsupervised multitask learners.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Language models are unsupervised multitask learners

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:20.918889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:16:14.275120Z digest=sha256:083d2a1578a0559520c49d6da0b8248cc726110b88e5bdf73351245e1254c565

Observation ee3335e7-0fd7-4dfd-b5a8-50e2e5a0eabb · outbound

This paper cites Learning transferable visual models from natural language supervision.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Learning transferable visual models from natural language supervision

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:20.665790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:16:14.329877Z digest=sha256:403625e0a328b3674b45e189b88d140369de916f4d6023a580ab9fd053227e87

Observation cf7157ea-f21c-4296-8883-60c90dca772f · outbound

This paper cites Real-world robot learning with masked visual pre-training.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Real-world robot learning with masked visual pre-training

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:20.397881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:16:14.436516Z digest=sha256:ca91f961118573865ce1dd46993d0370f73b16d3d9ee707dd0337c00e395df51

Observation 9add7871-7f36-4f34-a621-f36582eb566d · outbound

This paper cites Cliport: What and where pathways for robotic manipulation.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Cliport: What and where pathways for robotic manipulation

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:20.194602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:16:14.584871Z digest=sha256:d89d20996ac18f5747cba87532e50896c3cff428aa41176ca68aede4e9378992

Observation 2a71a9fb-6724-476b-abb4-411f7b0e9b7c · outbound

This paper cites Towards VQA models that can read.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Towards VQA models that can read

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:19.971053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:16:14.692031Z digest=sha256:65fdd61d8f4a2234cedb32c0b859418b1452f43c225efd1611189fdb476f5342

Observation e6059b5a-5b2a-4cda-8596-327f9ed638db · outbound

This paper cites Open-World Object Manipulation using Pre-trained Vision-Language Models.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Open-World Object Manipulation using Pre-trained Vision-Language Models

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:14.782958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:14.782958Z digest=sha256:b94082077c8944c37b878653b577f6fb3834faed6e392dde55772a56a3ca10ce

Observation 74674bf9-e778-4018-b380-b4b587df19cd · outbound

This paper cites SMART: Self-supervised Multi-task pretrAining with contRol Transformers.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces SMART: Self-supervised Multi-task pretrAining with contRol Transformers

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:14.859943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:14.859943Z digest=sha256:f8414eb79bdcacd62bc65f798c1a9d8c2ddf4fd8dece67d1d41d0837fcc9a5ab

Observation 8ffe0a2d-f610-42fa-841c-d5d9ab45293f · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Gemini: A Family of Highly Capable Multimodal Models

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:14.937333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:14.937333Z digest=sha256:00ca8f94f730c03fbad1bd079a227ad89fc657a72c1ca94b9f85c87e1a9d866b

Observation 34992c8b-5c2f-461a-9ebd-97bd33f16688 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:15.034425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:15.034425Z digest=sha256:459d277d269dca83c7251389665b0001c6b0e03a030c6e590d4c40262b6ed1c4

Observation dfa936c2-1361-44c1-aca6-700a74758dfc · outbound

This paper cites Gemini Robotics: Bringing AI into the Physical World.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Gemini Robotics: Bringing AI into the Physical World

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:15.105152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:15.105152Z digest=sha256:315f8153e88e5e27f546646775a3e97b2f9c69f58e829d4f34f08e99dee60223

Observation 61693193-d58e-4177-8a7d-ee10df47b7af · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces LLaMA: Open and Efficient Foundation Language Models

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:15.176031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:15.176031Z digest=sha256:704e195119a7a37f6090ab65e13c1f0a880243eb81e02c115543ee5caa183714

Observation 06719993-6dd7-4eac-95f8-55b4e8a21ec3 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:15.268308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:15.268308Z digest=sha256:86da66fec11138b4f7c4b57718edbe8e9ecf6df121a3dcee379f22c2c3d54e77

Observation a41ee9bc-cd25-455a-8f98-20c48cd04f0e · outbound

This paper cites The all-seeing project: Towards panoptic visual recognition and understanding of the open world.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces The all-seeing project: Towards panoptic visual recognition and understanding of the open world

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:19.827901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:16:15.352827Z digest=sha256:cb1a5df08658e433cbbaf1a019cca0630188e2e2202c4dd684c3f6a7d5533c9b

Observation e1d03a9a-3438-4690-a6a9-9da73d0da9e6 · outbound

This paper cites Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:15.449497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:15.449497Z digest=sha256:ad6cf2865ab3afb64871afb844e0278833926a591388548678bb7fe2e7df7716

Observation e412d259-4bc0-4a4d-8e91-d339d6ca3b4e · outbound

This paper cites Grok-1.5 vision preview.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Grok-1.5 vision preview

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:19.721585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:16:15.510411Z digest=sha256:e322a2dee50ebd492b7306b2f93aa90843c9310d13306317878a9bed47f3a569

Observation eaff1423-9048-4f1f-8ee2-dd20b974265f · outbound

This paper cites Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:15.576937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:15.576937Z digest=sha256:04b0e896cf63ec05378a20164f02550ea02e694aab36fa5694506e79de73dd6a

Observation 239354fb-65f5-4647-9534-f23f80f943f4 · outbound

This paper cites The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision).

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:15.641523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:15.641523Z digest=sha256:2b8993a50aa8443bbb6d5f416fa8ab1b70ccd19d13b8f2e4cef3c1fed2338ae6

Observation 8944cc7b-10b1-4c98-b1bf-5159dde0e1f4 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:15.718768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:15.718768Z digest=sha256:cde915ee248469a1677268a3a8b1d75b4c0e8c5f09db2722874bdafebd4d1334

Pith citing papers

Observation 92387108-7980-4289-8454-b5ce4b3b9582 · inbound

RoboBrain 2.0 Technical Report cites this paper.

RoboBrain 2.0 Technical Report Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:29.605186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:29.605186Z digest=sha256:f0f38366128ca31f4442fdc46d9e6f74f0d628455ade39ad65c4ba4c4d5c4adf

Observation 6253ce52-1ae4-49db-9f13-5e9949c949b3 · inbound

Talk Less, Fly Lighter: Autonomous Semantic Compression for UAV Swarm Communication via LLMs cites this paper.

Talk Less, Fly Lighter: Autonomous Semantic Compression for UAV Swarm Communication via LLMs Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-25T07:40:29.157371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T07:37:38.901354Z digest=sha256:301d50e768794ff43e417627ec920b539c9c857f1e164546a24412d465f6fc6f

Observation cc80ffac-da6b-477b-bd31-f75cdccac3fc · inbound

The high-speed X-ray camera on AXIS: design and performance updates cites this paper.

The high-speed X-ray camera on AXIS: design and performance updates Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T18:48:03.418621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:48:03.418621Z digest=sha256:d2f9ebba1156de170993c0bc8affee6c37b231d0f1d1f596bd2689d6b99db60e

Observation 8684ed19-c732-44f0-b64a-6561ec1fe91b · inbound

Contrastive Representation Regularization for Vision-Language-Action Models cites this paper.

Contrastive Representation Regularization for Vision-Language-Action Models Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-04T12:55:07.932813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:55:07.932813Z digest=sha256:6de6ecbcb63608ee2ec795a0458ad222f9f126aa83b72a0e33243d74217d0bf2

Observation 80a4f563-9ab0-4a4d-96c2-d8e8d525ba01 · inbound

InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy cites this paper.

InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:09:39.833644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T20:09:39.677347Z digest=sha256:a0c48235dcbc067943d58b69ec00673c11be37b41bd58738c3d0699111a92fef

Observation 35c55ac7-b8a4-4f5b-ab32-c4361e29ce64 · inbound

MiMo-Embodied: X-Embodied Foundation Model Technical Report cites this paper.

MiMo-Embodied: X-Embodied Foundation Model Technical Report Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:42:05.714505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T20:40:54.096289Z digest=sha256:3d2b99e40b2978e7db7992a69a4fa35f209e6aa1c1debb0a1352a65fde2387d2

Observation 189a4e30-35d9-47ed-9375-7da26e19f9ba · inbound

Token Warping Helps MLLMs Look from Nearby Viewpoints cites this paper.

Token Warping Helps MLLMs Look from Nearby Viewpoints Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-13T21:08:17.373521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T21:07:55.062113Z digest=sha256:07ccc105f47867c229d106218866c6a894d6d665c8d691bf345dacf410ea1bc1

Observation a7ea5a71-deb2-4e97-9c2b-897a4aad0f53 · inbound

3D-VCD: Hallucination Mitigation in 3D-LLM Embodied Agents through Visual Contrastive Decoding cites this paper.

3D-VCD: Hallucination Mitigation in 3D-LLM Embodied Agents through Visual Contrastive Decoding Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:56:00.254914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T17:51:52.063238Z digest=sha256:439cabd998be42bf418dcb13f1d7cab4fa21d71ba16b2d471b1205d6aff56d71

Observation 0749cb87-7f03-44c0-bac1-e3af4955cb14 · inbound

GeoWorld-VLM: Geometry from World Models for Vision-Language Models cites this paper.

GeoWorld-VLM: Geometry from World Models for Vision-Language Models Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-20T17:58:49.605881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T17:57:00.909897Z digest=sha256:33ee9d6aa20e35012c9a6920ead3246c80de013ee4192aa999f4d46adb393b5b

Observation 37aa5277-49f8-4b04-8a43-0d65279cff87 · inbound

GeoWorld-VLM: Geometry from World Models for Vision-Language Models cites this paper.

GeoWorld-VLM: Geometry from World Models for Vision-Language Models Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-06-30T19:05:00.582050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T19:02:05.125937Z digest=sha256:4bf36a6bbccf5d8b978232fbee9af21a1106f78bdeb6b0969b011ba886b740c9

Observation 319808e9-17f6-4ce6-90d6-8db16d291176 · inbound

SpaceDG: Benchmarking Spatial Intelligence under Visual Degradation cites this paper.

SpaceDG: Benchmarking Spatial Intelligence under Visual Degradation Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-22T06:24:40.387049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T06:24:30.669956Z digest=sha256:fafaa9d159c72dc476ddae2da5613b8c9ca72316313a74acd7617daca7ded087

Observation 63e8f605-6c9f-4107-bda1-eca7581de84b · inbound

SpaceDG: Benchmarking Spatial Intelligence under Visual Degradation cites this paper.

SpaceDG: Benchmarking Spatial Intelligence under Visual Degradation Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:14:57.282353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T17:07:49.005419Z digest=sha256:29030fe49c69d3ec6810550467870fcebda6ea144db5e1d02583cc3605d00654

Observation f8731dac-224d-47d0-9833-45ab6014a0b7 · inbound

Extending Embodied Question Answering from Perception to Decision cites this paper.

Extending Embodied Question Answering from Perception to Decision Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-06-29T21:33:58.950221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T21:30:40.182958Z digest=sha256:8efd53da9263c13d50bc53122af9dec9bc643cbb1bf69f79191ebbc9f0971d23

Observation b4b13735-5d8f-4332-935e-4cce4db729cf · inbound

GEM: Generative Supervision Helps Embodied Intelligence cites this paper.

GEM: Generative Supervision Helps Embodied Intelligence Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:43:28.937015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:944010c74cd96b00d0637e23047f296883bdedee3dfaac756ad06c14997d16cf

Observation cdae8602-7c06-4f94-b73c-6ab9c10f8ddd · inbound

OVO-S-Bench: A Hierarchical Benchmark for Streaming Spatial Intelligence in Multimodal LLMs cites this paper.

OVO-S-Bench: A Hierarchical Benchmark for Streaming Spatial Intelligence in Multimodal LLMs Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:16:27.075845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T11:02:07.122615Z digest=sha256:1489d7aecbd7600ecca4095990ed251efa5bd913988e72be179cbfa4a6b6d620

Observation 8b86d7ef-1904-4510-9c25-e011c7353b06 · inbound

Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models cites this paper.

Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-07-03T06:07:41.154508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T12:55:47.754632Z digest=sha256:7ea9f27ef5d903b11940cb291adf0d4a0b9e690f0bda4ddcbd4d697785095456

Observation ad9c88a9-defe-4f09-9434-36ed67cab158 · inbound

Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models cites this paper.

Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

Reference 58

Resolution
unresolved
no resolver link, observed 2026-07-14T18:07:09.018997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T18:07:09.018997Z digest=sha256:7f7f388c98ed273d0174b40911112b991d413521a43aaae4ae820a75e1b1f8dd

Observation be10cee9-94d9-40db-a399-91baf0cabd54 · inbound

RoboProcessBench: Benchmarking Process-Aware Understanding in Vision-Language Robotic Manipulation cites this paper.

RoboProcessBench: Benchmarking Process-Aware Understanding in Vision-Language Robotic Manipulation Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:28:34.760148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T06:26:32.209719Z digest=sha256:217a202296fdca6bce0b5a5a68be29b15899ed67fb98fac91258c0704462b0ea

Observation ca74376b-ee0f-4340-892e-ef46ed4806a6 · inbound

RoboProcessBench: Benchmarking Process-Aware Understanding in Vision-Language Robotic Manipulation cites this paper.

RoboProcessBench: Benchmarking Process-Aware Understanding in Vision-Language Robotic Manipulation Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T11:43:49.639057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:43:49.639057Z digest=sha256:a150b70cc3062c327b617f20a9f24f6d8e7f0eeedec36652b05f6ed86e750cef

Observation e7ce9633-15b2-48d6-adad-d5b498358984 · inbound

SPARC: Reliable Spatial Annotations from Robot Demonstrations at Scale cites this paper.

SPARC: Reliable Spatial Annotations from Robot Demonstrations at Scale Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:28:34.351453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T06:28:14.694168Z digest=sha256:45fac2d2430e02de56f24f470e7d6fb5848fdfd5a6b22993df3b122ff640d856

Observation 880ac1a5-6849-4c80-8dfe-753e815bd6bb · inbound

RoboPIN: Grounded Embodied Reasoning via Pinned Chain-of-Thought cites this paper.

RoboPIN: Grounded Embodied Reasoning via Pinned Chain-of-Thought Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-06-30T11:24:38.366106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T11:17:48.279808Z digest=sha256:3086610f9bef12a9d2fbfe7f5ef6b8302e2730398340ddba880b7ab7e9ccd5b2

Observation 9d610a94-38fe-468c-b91c-69837c2cc378 · inbound

RoboPIN: Grounded Embodied Reasoning via Pinned Chain-of-Thought cites this paper.

RoboPIN: Grounded Embodied Reasoning via Pinned Chain-of-Thought Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-02T11:20:26.553333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:20:26.553333Z digest=sha256:9031f487ccc9ce1f5517a7e6be2c06c38c2379bca43d27a65f29b43c8742c331

Observation 5ca3a554-d397-484d-9748-597a412c8479 · inbound

ACE-Brain-0.5: A Unified Embodied Foundational Model for Physical Agentic AI cites this paper.

ACE-Brain-0.5: A Unified Embodied Foundational Model for Physical Agentic AI Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

Reference 54

Resolution
unresolved
no resolver link, observed 2026-07-11T19:16:57.396710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:16:57.396710Z digest=sha256:dc465717418cb9957403eb02189d26a3e121ee37a73b65637d425afa38c1065a

Observation e0899ac6-ff62-4be6-bdbd-271e62038c2b · inbound

RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation cites this paper.

RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-01T14:39:39.709300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:39:39.709300Z digest=sha256:c5368ad50093af91729ccee01e1ea3b550f991df7db580c9b42e39b0491e73d8