Pith. sign in

Paper Citation Record · LEDGER

RoboBrain 2.0 Technical Report

As of 15 August 2026, this Paper Citation Record lists 89 of 89 outbound references and 54 inbound Pith citation observations for arXiv:2507.02029.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.02029 v5

Coverage vector

measured 89 of 89 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:47:31.149194Z

measured 143 of 143 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 54 of 54 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T00:36:21.485253Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T12:04:50.478072Z

Reference resolution

89 of 89 outbound references displayed

  • verified exact1
  • verified fuzzy24
  • unresolved64
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 39906c01-6c21-4128-949a-80cee1babd2f · outbound

This paper cites Webdataset: High-performance data loading for deep learning, 2020.

RoboBrain 2.0 Technical Report Webdataset: High-performance data loading for deep learning, 2020

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:26.532740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:26.532740Z digest=sha256:6575b776a61bf37ab9f1260a71f8b51d5f194bf3530c5fb466c5afe35221bd70

Observation 677d2b77-6640-43f0-a663-90604ecedada · outbound

This paper cites Claude sonnet 4.

RoboBrain 2.0 Technical Report Claude sonnet 4

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:26.647128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:26.647128Z digest=sha256:c730e468db92577dbb440bdc0652d33830b6c365bfc68d399379d36fff2775a7

Observation 2d2c5fc0-8b1e-4052-ab4c-0b1b6c11e429 · outbound

This paper cites Scanqa: 3d question answering for spatial scene understanding.

RoboBrain 2.0 Technical Report Scanqa: 3d question answering for spatial scene understanding

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:26.724399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:26.724399Z digest=sha256:442dd8e983d9f5619356cf72dfce27449e77bea7af001817081a6a41680665a2

Observation 244d8de8-3194-4fdd-a734-be9509806f03 · outbound

This paper cites Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning.

RoboBrain 2.0 Technical Report Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:26.803395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:26.803395Z digest=sha256:32e4e2c442139177db49fe3211b00a87a36f640ee2bce0f86c5660ca46ccdb9d

Observation a17216aa-7d8f-4945-8134-3654cc2ca257 · outbound

This paper cites Qwen2.5-VL Technical Report.

RoboBrain 2.0 Technical Report Qwen2.5-VL Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:26.861229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:26.861229Z digest=sha256:b6f3473df7d7de0e0c07a1733ddc6a2ba386da4828399f00f26082e886c79e32

Observation f72dd003-b3e8-4de9-9b7d-1fb15a1c66ad · outbound

This paper cites AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems.

RoboBrain 2.0 Technical Report AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:26.979701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:26.979701Z digest=sha256:53ccb9903563e2f86b86358472aa3ddf8b988d93ee5c49ca055af4f261c8bc89

Observation 0849cc23-5289-46b8-a50a-c9c1a710953a · outbound

This paper cites Sharegpt4video: Improving video understanding and generation with better captions.

RoboBrain 2.0 Technical Report Sharegpt4video: Improving video understanding and generation with better captions

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:27.050956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:27.050956Z digest=sha256:d5bac771c946fd86749207ee3a51aaaf30270d4c992c560eb4142191d4221d1d

Observation 79f26d2a-7a1c-41b2-8ff4-82c0113f6133 · outbound

This paper cites an unresolved cited work.

RoboBrain 2.0 Technical Report Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:27.208513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:27.208513Z digest=sha256:b0d5b6b392376f30f2c88783c99766d7a3623d02e802a2cbd74ba99a5ba60501

Observation ecf0f690-2bb9-40be-8e40-9e22bf23891c · outbound

This paper cites EgoPlan-Bench: Benchmarking Multimodal Large Language Models for Human-Level Planning.

RoboBrain 2.0 Technical Report EgoPlan-Bench: Benchmarking Multimodal Large Language Models for Human-Level Planning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:27.281733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:27.281733Z digest=sha256:3181f249ebf2a72309840720f558ee11402bda9f791fa103b8b3fb402486901d

Observation cc131c5a-11b1-4c06-a342-0008b3b18a14 · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

RoboBrain 2.0 Technical Report Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:27.373818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:27.373818Z digest=sha256:e033fc2e6cd168b1f1f44913ad8d68760beb894563557004c342222c9ee94f71

Observation 5bbf8db0-ef9e-43af-bb5a-fb3b6f9a74cc · outbound

This paper cites Llm agents for education: Advances and applications.arXiv preprint arXiv:2503.11733, 2025.

RoboBrain 2.0 Technical Report Llm agents for education: Advances and applications.arXiv preprint arXiv:2503.11733, 2025

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:27.447799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:27.447799Z digest=sha256:62e96b73f488fdf835117926df33f5afc95145df6aacbea245ae36708d6f6508

Observation 525205e9-feb2-4aad-b8e7-96342ed3f4d9 · outbound

This paper cites Flagscale: A unified meta-framework enabling adaptive heterogeneous computing for the llm ecosystem.https://github.com/FlagOpen/FlagScale, 2024.

RoboBrain 2.0 Technical Report Flagscale: A unified meta-framework enabling adaptive heterogeneous computing for the llm ecosystem.https://github.com/FlagOpen/FlagScale, 2024

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:27.519289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:27.519289Z digest=sha256:57e988f64c5462dfa4977576cf69e1cf77fb55bb655c766802277781bed586e2

Observation b6cfd4bc-253f-4779-a74d-c800be6a9dbc · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

RoboBrain 2.0 Technical Report Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:27.628877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:27.628877Z digest=sha256:36e23e548d17d90a1888db7f638065c2b34851f2f2685bd4f2c0affcd18f2379

Observation 16e40d10-1250-4016-899c-0ada7d423795 · outbound

This paper cites Embspatial-bench: Benchmarking spatial understanding for embodied tasks with large vision-language models.

RoboBrain 2.0 Technical Report Embspatial-bench: Benchmarking spatial understanding for embodied tasks with large vision-language models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:27.780921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:27.780921Z digest=sha256:88652f0f2cb001d1e8537b77abee6bf5cf5f97c0a39c1fe53121cb2cf4ff8667

Observation d40506a7-e4ba-4a11-b952-94aac606ad7c · outbound

This paper cites Blink: Multimodal large language models can see but not perceive.

RoboBrain 2.0 Technical Report Blink: Multimodal large language models can see but not perceive

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:27.854641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:27.854641Z digest=sha256:8c2ea58ada83ef1be1fdcb08b17f81cb63695ee58b89ec4d26b257898d3a4db3

Observation e4ef93e2-9505-4ca9-be1e-d5455f83c727 · outbound

This paper cites Gemini 2.5 pro preview: even better coding performance.https://developers.googleblog.com/en/ gemini-2-5-pro-io-improved-coding-performance/, 2025.

RoboBrain 2.0 Technical Report Gemini 2.5 pro preview: even better coding performance.https://developers.googleblog.com/en/ gemini-2-5-pro-io-improved-coding-performance/, 2025

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:27.976037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:27.976037Z digest=sha256:17de8ccdea83b958e3452e1f512de7c5c6e15e5f3bd3d679d7278f6597a17d83

Observation d50f7c05-705e-4c0b-b809-91381ded576b · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

RoboBrain 2.0 Technical Report DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:28.046364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:28.046364Z digest=sha256:63b1ab2f97a8da06baf0e953f82b3de92681ab54ce3c4e4dc31da1a85760ab86

Observation c67bc056-83d2-4264-83f5-d7a24b974d02 · outbound

This paper cites Seed1.5-VL Technical Report.

RoboBrain 2.0 Technical Report Seed1.5-VL Technical Report

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:28.125248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:28.125248Z digest=sha256:17558524d46fe72949955ff92871b59651e94a5adfd364ca75d0d03fe036eae0

Observation 21c77a9d-7977-47bc-b437-29abdcb4fa9b · outbound

This paper cites Lvis: A dataset for large vocabulary instance segmentation.

RoboBrain 2.0 Technical Report Lvis: A dataset for large vocabulary instance segmentation

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:47:32.278055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:47:28.224406Z digest=sha256:d61a4357d299c2d01c0c5cf0dee2574fc2e4e53d52a0056c9c0407b4cc80d840

Observation 079ee21b-024d-408f-aa54-591abe156fb4 · outbound

This paper cites FlagEvalMM: A Flexible Framework for Comprehensive Multimodal Model Evaluation.

RoboBrain 2.0 Technical Report FlagEvalMM: A Flexible Framework for Comprehensive Multimodal Model Evaluation

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:47:31.781244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:47:28.353756Z digest=sha256:5291e0b6fc5052677d35e854ad8e976f8351060cbfa4b7d0f16087fb80e26e25

Observation 77a87e03-afaf-4699-9570-43234d43b22e · outbound

This paper cites A Comprehensive Survey on Evaluating Large Language Model Applications in the Medical Industry.

RoboBrain 2.0 Technical Report A Comprehensive Survey on Evaluating Large Language Model Applications in the Medical Industry

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:28.421111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:28.421111Z digest=sha256:a4f82617deb30fade7f94f32626cc63fedb852ec06207d95dc304fff8bdfffd6

Observation dfec7ac0-3e6a-4cda-81e7-d2da9b7c947c · outbound

This paper cites GPT-4o System Card.

RoboBrain 2.0 Technical Report GPT-4o System Card

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:28.505007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:28.505007Z digest=sha256:78ed057bb6a94e567c690506cbdd6f1fc94075ddf3e0dfdba2109f9162f7827d

Observation 1f6677a5-f6ce-49cf-a2f9-6ca964fee7fe · outbound

This paper cites Robobrain: A unified brain model for robotic manipulation from abstract to concrete.

RoboBrain 2.0 Technical Report Robobrain: A unified brain model for robotic manipulation from abstract to concrete

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:28.572627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:28.572627Z digest=sha256:0da39771a90eb5365fc973a4768bb10a6de4a91880066d1dc91d7bda79626cc9

Observation 62194061-0669-4ef6-9159-3dd198d0baf1 · outbound

This paper cites Imagic: Text-based real image editing with diffusion models.

RoboBrain 2.0 Technical Report Imagic: Text-based real image editing with diffusion models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:28.632024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:28.632024Z digest=sha256:5467f579e62cbc95a987c228504bcd5cb0e98c640989fc6c3ac11788cb9d5de0

Observation 45924ed9-0e18-4060-a563-080f6458ba87 · outbound

This paper cites AI2-THOR: An Interactive 3D Environment for Visual AI.

RoboBrain 2.0 Technical Report AI2-THOR: An Interactive 3D Environment for Visual AI

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:28.735777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:28.735777Z digest=sha256:bb4aca77e53c1cb6c927c97ba82c8146cf4ba663d1ea20dd6dc6abeab0994fac

Observation 56cb0289-22aa-40af-8685-0e92e1d25712 · outbound

This paper cites Reducing Activation Recomputation in Large Transformer Models.

RoboBrain 2.0 Technical Report Reducing Activation Recomputation in Large Transformer Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:28.846985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:28.846985Z digest=sha256:fc48c4b1d2df5120b448827ecd2921f7b2909a99e0c445ce0467cc1df44efcd2

Observation be4130e0-c720-492f-8576-ab834e2836ea · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annotations.International journal of computer vision, 123:32–73, 2017.

RoboBrain 2.0 Technical Report Visual genome: Connecting language and vision using crowdsourced dense image annotations.International journal of computer vision, 123:32–73, 2017

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:28.932959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:28.932959Z digest=sha256:e0f54502bfcf968a70e97864a85dc569e7c0805c4e4266e346f58b6e9d1de2c3

Observation 6c47d7e7-270a-4983-8b48-7fed382aa73e · outbound

This paper cites The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale.IJCV, 2020.

RoboBrain 2.0 Technical Report The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale.IJCV, 2020

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:47:32.251143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:47:29.014207Z digest=sha256:04c454067e6f0315092ec225313f406b2f210d14dbfc4a5b692f4b846bad64e6

Observation 1c4e674b-8d1d-4997-91ff-05bdece189d2 · outbound

This paper cites Cubify Anything: Scaling Indoor 3D Object Detection.

RoboBrain 2.0 Technical Report Cubify Anything: Scaling Indoor 3D Object Detection

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:29.062663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:29.062663Z digest=sha256:b01125403900c125911bf4c3d5695363837e3cdee4bc697a7056f23dd7485ef6

Observation 93f1fd8f-a9be-4e03-8992-7ad952de31e7 · outbound

This paper cites Energon: Scaling megatron-lm training with data and expert parallelism, 2023.

RoboBrain 2.0 Technical Report Energon: Scaling megatron-lm training with data and expert parallelism, 2023

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:47:32.241849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:47:29.159824Z digest=sha256:9724746f4c8ba9b5c441774b68cbed9e85ea4443c928df198eb533d26fe83c93

Observation 36b5ab3a-f716-4ab1-9df3-cf863982ad9d · outbound

This paper cites DeepSeek-V3 Technical Report.

RoboBrain 2.0 Technical Report DeepSeek-V3 Technical Report

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:29.218019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:29.218019Z digest=sha256:37b96f938a4e48f6a31d0b5dd62f50a5227a7dca2b99e9ef70b6587e72b34638

Observation 5658a705-190b-40ed-a03e-de457d93e25d · outbound

This paper cites Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning.

RoboBrain 2.0 Technical Report Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:29.322899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:29.322899Z digest=sha256:36a8d2f45cfff50704acf7413e951607bf82f4f3101381505ee8ef52d42d88cd

Observation 8c48df8e-bd7d-49b2-96e2-18730c36fc9c · outbound

This paper cites Improved baselines with visual instruction tuning.

RoboBrain 2.0 Technical Report Improved baselines with visual instruction tuning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:29.449571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:29.449571Z digest=sha256:3a3037f2e70d3ba7662d4df5f95145231ea89f5c60400a3451af4cc59a99c95d

Observation 92ea8b59-4c88-4c0a-b179-f028ac161a1e · outbound

This paper cites Grounding dino: Marrying dino with grounded pre-training for open-set object detection.

RoboBrain 2.0 Technical Report Grounding dino: Marrying dino with grounded pre-training for open-set object detection

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:47:32.226010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:47:29.528347Z digest=sha256:330b656485e0e780a76a91462a3e24194c6b80dee48b68f8bc554abb240a9847

Observation 3dc12c3f-bc74-4b62-a834-a1ec1e5d00bc · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

RoboBrain 2.0 Technical Report Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:29.553830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:29.553830Z digest=sha256:51e06d37df9dff674c085995f2f7b265bdf7132e3e2673e3d957ff41d30a5c46

Observation 92387108-7980-4289-8454-b5ce4b3b9582 · outbound

This paper cites Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces.

RoboBrain 2.0 Technical Report Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:29.605186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:29.605186Z digest=sha256:ddcf625cc3227aea8da381da3b18566a997f434b60811f8dacecc912ce0c800e

Observation 9f4eaa05-3c48-4295-867b-5f016dce07cd · outbound

This paper cites GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents.

RoboBrain 2.0 Technical Report GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:29.712258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:29.712258Z digest=sha256:aabe58e832d81d1eb4578e2dd9d20e37d68aa03cf08d92dae12ac96d87e56b5b

Observation b7d85442-4198-47da-8d4a-051b4a91c961 · outbound

This paper cites MMScan: A Multi-Modal 3D Scene Dataset with Hierarchical Grounded Language Annotations.

RoboBrain 2.0 Technical Report MMScan: A Multi-Modal 3D Scene Dataset with Hierarchical Grounded Language Annotations

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:29.800018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:29.800018Z digest=sha256:d74660315061ec57064671fe3e3a1d144363136f3bf7bf46e1afa70c21da95c5

Observation c2252a69-9cda-4174-8228-8eca3d874cff · outbound

This paper cites Sqa3d: Situated question answering in 3d scenes.

RoboBrain 2.0 Technical Report Sqa3d: Situated question answering in 3d scenes

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:47:32.209299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:47:29.963352Z digest=sha256:6b607fba1b5d8743fbb679dc29af66aab279c5c597a86294669ec9fff4160bb9

Observation 3235eac1-294a-40fc-b1e2-14ef618d2217 · outbound

This paper cites Mixed Precision Training.

RoboBrain 2.0 Technical Report Mixed Precision Training

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:30.084151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:30.084151Z digest=sha256:4c233ecaf5fd2bd06d122ce48e6754e035ae9eec44692a7c6493162f9e6937ac

Observation 6dbc3042-abe1-4d06-8e6e-1d41dbed98ad · outbound

This paper cites Ocr-vqa: Visual question answering by reading text in images.

RoboBrain 2.0 Technical Report Ocr-vqa: Visual question answering by reading text in images

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:30.158464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:30.158464Z digest=sha256:292b598cba97d46af11976437eda74c048ed639cb4127317b68046a8ab103ed7

Observation 06a9e879-b5b4-4a56-b5ee-1a82b12ed39b · outbound

This paper cites Efficient large-scale language model training on gpu clusters using megatron-lm.

RoboBrain 2.0 Technical Report Efficient large-scale language model training on gpu clusters using megatron-lm

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:47:32.192373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:47:30.293811Z digest=sha256:8c30da58a111d47facaf6b8c7bcb2bd51a5ac434a0259dd699e5ba5f03088dff

Observation 722dc628-b96a-42c5-82ef-02ce5b32cc0e · outbound

This paper cites Megatron-lm: Training multi-billion parameter language models using model parallelism, 2021.

RoboBrain 2.0 Technical Report Megatron-lm: Training multi-billion parameter language models using model parallelism, 2021

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:47:32.181339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:47:30.452725Z digest=sha256:3e0981686e37a1219cffd56792f3a24ea8387ca709f7c2eaf45984920f43c00c

Observation 66980da7-aeb2-4399-ba2f-9c20f51c687d · outbound

This paper cites GPT-4 Technical Report.

RoboBrain 2.0 Technical Report GPT-4 Technical Report

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:30.603921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:30.603921Z digest=sha256:9a63eaaf696d35b4ea243b797b5945c17530d066360f70ccf818805f30cff828

Observation 93ee9142-f32d-4fbb-834a-00c070996b00 · outbound

This paper cites Gpt-4v(ision) system card.

RoboBrain 2.0 Technical Report Gpt-4v(ision) system card

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:47:32.169931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:47:30.755771Z digest=sha256:083d7a0e872c33af9cc91e2cc8c8d8583978694d536a39b13acd7dc914a562ff

Observation 434bef0a-bcf3-4392-ae1b-d79ef4c3c70b · outbound

This paper cites SpaceR: Reinforcing MLLMs in Video Spatial Reasoning.

RoboBrain 2.0 Technical Report SpaceR: Reinforcing MLLMs in Video Spatial Reasoning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:30.915731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:30.915731Z digest=sha256:4955a5e520887ba19c2f5c5e79aca79aa159d6f6e8c8b46b5f810f0182b49c45

Observation 98422b66-815e-4ab5-a764-b357fc839249 · outbound

This paper cites Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0.

RoboBrain 2.0 Technical Report Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:30.976655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:30.976655Z digest=sha256:d35ceb2ac6b1b4a76e00265e7c110487825ba776d0171fa766cf2a58997437cd

Observation 63d5768d-ce84-43c8-a18b-371b80610a31 · outbound

This paper cites Unidepthv2: Universal monocular metric depth estimation made simpler.arXiv, 2025.

RoboBrain 2.0 Technical Report Unidepthv2: Universal monocular metric depth estimation made simpler.arXiv, 2025

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:47:32.152636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:47:30.980578Z digest=sha256:7e62c7669571cc558ec4449a15b104555271624e7bbe5cd591bce6b95f37dabe

Observation 3fc90b62-ee1e-4288-a698-b003809ff63d · outbound

This paper cites Cuda memory management, 2023.

RoboBrain 2.0 Technical Report Cuda memory management, 2023

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:47:32.142246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:47:30.984067Z digest=sha256:550f8dc960cc882cab56d8195ac25b03f739a9c7bfb31d5716388e1b2fa9dcc2

Observation 63904301-a55f-459e-978d-2f56f8d39dae · outbound

This paper cites Qwen2.5-vl: Multimodal llms from alibaba, 2025.

RoboBrain 2.0 Technical Report Qwen2.5-vl: Multimodal llms from alibaba, 2025

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:47:32.132077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:47:30.987819Z digest=sha256:c4f9820f1db6a4639c0e1d99712aba96a07cb2600e1811c81d47ddece200c21c

Observation 9a1d25a4-97f9-4bb7-89d2-fb32017536ff · outbound

This paper cites Paco: Parts and attributes of common objects.

RoboBrain 2.0 Technical Report Paco: Parts and attributes of common objects

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:47:32.122618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:47:30.991348Z digest=sha256:18a617268d276ce16157d082639c9c1741f6ff75c87d8ce3e989f982ff7ca85f

Observation efccb56a-0994-459d-8329-e462582d17b7 · outbound

This paper cites Sam 2: Segment anything in images and videos.ICLR, 2025.

RoboBrain 2.0 Technical Report Sam 2: Segment anything in images and videos.ICLR, 2025

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:47:32.111805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:47:30.995693Z digest=sha256:d58b7e46b983e5b42186c5203f19e0fea6bb602f04a4a8438d2144c9b08583d0

Observation 7d82e8a4-157c-4daa-95df-f7e611df51ae · outbound

This paper cites Sat: Spatial aptitude training for multimodal language models.

RoboBrain 2.0 Technical Report Sat: Spatial aptitude training for multimodal language models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:30.998852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:30.998852Z digest=sha256:3ec8bc6f1ee67dd98c8506825b7e7619f0eed9cfe537cf433102e6adb4df5465

Observation a0a5c116-98b8-404e-95fa-a049e3bef303 · outbound

This paper cites A-okvqa: A benchmark for visual question answering using world knowledge.

RoboBrain 2.0 Technical Report A-okvqa: A benchmark for visual question answering using world knowledge

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:31.002436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:31.002436Z digest=sha256:fec078eed000b8edaae54a2885e44f4acaa53e51a47581b3a79e3a9fe6d2a26a

Observation 1c2c044b-5aef-42ff-879d-e3859db2d95b · outbound

This paper cites Robovqa: Multimodal long-horizon reasoning for robotics.

RoboBrain 2.0 Technical Report Robovqa: Multimodal long-horizon reasoning for robotics

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:47:32.095280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:47:31.005433Z digest=sha256:dd9a0c281fadbae22d92e04a6eb2c824d1178994ecd691ea18b6dd350f353ca0

Observation 47522736-b458-4690-b46a-4353ca7b761f · outbound

This paper cites Hybridflow: A flexible and efficient rlhf framework.

RoboBrain 2.0 Technical Report Hybridflow: A flexible and efficient rlhf framework

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:31.008891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:31.008891Z digest=sha256:ec246a7bc10bc913b8eed9b4819535b3300d09bed3287fa09720f206e9c1ca40

Observation fd1a35aa-1e42-47d4-aab2-bd4d01433f9c · outbound

This paper cites Emu edit: Precise image editing via recognition and generation tasks.

RoboBrain 2.0 Technical Report Emu edit: Precise image editing via recognition and generation tasks

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:31.013713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:31.013713Z digest=sha256:5e90de44a8d772ba71d7b6d191bb2d363a3dd1fa82a33af9769ae0b244abec2e

Observation 9dcd3b63-22cc-4889-9413-54d1c3612abb · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

RoboBrain 2.0 Technical Report Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:31.016748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:31.016748Z digest=sha256:1aa2f6ae44334ec4c43c49d705cfb9f026ff49b37d12fa338b852016fbaadd26

Observation 75e5e9dd-0aca-4622-bf7b-ed12be3f80a2 · outbound

This paper cites Robospatial: Teaching spatial understanding to 2d and 3d vision-language models for robotics.

RoboBrain 2.0 Technical Report Robospatial: Teaching spatial understanding to 2d and 3d vision-language models for robotics

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:31.020517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:31.020517Z digest=sha256:da31453f9739c2835ca23c54b4c13760fadb2c50bea77306c369e9bcc9141d1d

Observation 8e506e8c-a861-404e-a997-9e38d5b02cf5 · outbound

This paper cites MultiModalQA: Complex Question Answering over Text, Tables and Images.

RoboBrain 2.0 Technical Report MultiModalQA: Complex Question Answering over Text, Tables and Images

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:31.024566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:31.024566Z digest=sha256:716c35ed66d734b5e45e2dee726726302221be2c1d3913c70c4e52f89f83b90b

Observation dc5661de-b82b-4453-97b1-4facdefd8cd9 · outbound

This paper cites RoboOS: A Hierarchical Embodied Framework for Cross-Embodiment and Multi-Agent Collaboration.

RoboBrain 2.0 Technical Report RoboOS: A Hierarchical Embodied Framework for Cross-Embodiment and Multi-Agent Collaboration

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:31.028205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:31.028205Z digest=sha256:e0efc94b582b3cd28019b7a0538c7f21de346eece94ec2b8c2575e99c81e4ee6

Observation 94ff3ac6-ef97-4b1c-af79-3ed3e0d694df · outbound

This paper cites Reason-rft: Reinforcement fine-tuning for visual reasoning.arXiv preprint arXiv:2503.20752, 2025.

RoboBrain 2.0 Technical Report Reason-rft: Reinforcement fine-tuning for visual reasoning.arXiv preprint arXiv:2503.20752, 2025

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:31.031498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:31.031498Z digest=sha256:6f05abfa852db8a837b7a22aef70036cb1fe734e502d4f91b9d3b1852aac0a58

Observation 1ecd60d6-94d8-4d48-b674-fe45f3c604b2 · outbound

This paper cites Video understanding with large language models: A survey.IEEE Transactions on Circuits and Systems for Video Technology, 2025.

RoboBrain 2.0 Technical Report Video understanding with large language models: A survey.IEEE Transactions on Circuits and Systems for Video Technology, 2025

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:31.034168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:31.034168Z digest=sha256:eb28f3022a64ec292a793722cf5debe62563f1bb92d8da9c5a77271936e7db50

Observation be8963c2-8f60-44b4-96df-23fb6d9a322e · outbound

This paper cites Gemini Robotics: Bringing AI into the Physical World.

RoboBrain 2.0 Technical Report Gemini Robotics: Bringing AI into the Physical World

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:31.038680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:31.038680Z digest=sha256:47cf2d100fe056ecedf956f84018f6b8a086308e666968b63f92d0284d58959f

Observation cac9d90b-871c-408d-b7af-81af1a63183d · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

RoboBrain 2.0 Technical Report Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:31.043230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:31.043230Z digest=sha256:cb1bb57068811f92429389497257c47da6dd0bdec8a1eb925f93be96e4fa21b7

Observation 7f133fbf-58a4-4744-bb1d-709678cb675d · outbound

This paper cites Qwq-32b: Embracing the power of reinforcement learning, March 2025.

RoboBrain 2.0 Technical Report Qwq-32b: Embracing the power of reinforcement learning, March 2025

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:31.047827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:31.047827Z digest=sha256:a5040c1d8184cd5caaff89130d61e881ddd94017f4417cc9eb732ff99092a449

Observation d475f2c7-393b-4a55-a68f-386d731bd9ff · outbound

This paper cites Cambrian-1: A fully open, vision-centric exploration of multimodal llms.NeurIPS, 2024.

RoboBrain 2.0 Technical Report Cambrian-1: A fully open, vision-centric exploration of multimodal llms.NeurIPS, 2024

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:31.052452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:31.052452Z digest=sha256:6b701fe25c4e29f62822f98d1bffd6b4499b6463d993c0e70d7db0810633fa41

Observation 50253336-2da7-445b-9c93-ff970d64e1fb · outbound

This paper cites verl: Volcano engine reinforcement learning for llms, 2024.

RoboBrain 2.0 Technical Report verl: Volcano engine reinforcement learning for llms, 2024

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:47:32.054021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:47:31.060947Z digest=sha256:c6900cb09f8d91765a14aa51d3e0e5f6d37204ef755d2623235d2583af70c199

Observation 913dd2e6-7038-44db-8f2c-675451c9a2b9 · outbound

This paper cites Rio: 3d object instance re-localization in changing indoor environments.

RoboBrain 2.0 Technical Report Rio: 3d object instance re-localization in changing indoor environments

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:47:32.044195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:47:31.068156Z digest=sha256:99bd8e0bd612c3149b1d2be643c3d024fb9bd9f710b5ebd6dfefb413585670d0

Observation 355ff154-54f6-468d-9d2a-cf1a9b7c12af · outbound

This paper cites HAQ: Hardware-Aware Automated Quantization with Mixed Precision.

RoboBrain 2.0 Technical Report HAQ: Hardware-Aware Automated Quantization with Mixed Precision

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:31.071952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:31.071952Z digest=sha256:91a3617c307e7a73767ba22cdf9c81d1b040df2c27e1f4d18da3307bc3aca3a8

Observation fb068da7-b4c8-4806-8f81-2d16971dd52a · outbound

This paper cites GUI Agents with Foundation Models: A Comprehensive Survey.

RoboBrain 2.0 Technical Report GUI Agents with Foundation Models: A Comprehensive Survey

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:31.078436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:31.078436Z digest=sha256:1e9526bc662b5ac2a08d2253e8964df617a83d9b88c99d3ddb0a7215f1265659

Observation 9b6cb17f-e9e5-48a1-9047-bc052a47239d · outbound

This paper cites Internvideo2: Scaling foundation models for multimodal video understanding.

RoboBrain 2.0 Technical Report Internvideo2: Scaling foundation models for multimodal video understanding

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:31.082918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:31.082918Z digest=sha256:dd266d70f87155cc557dafbd363b4f26c626d4fc6b420dc5e2abea504abbdeb1

Observation 2b6ac7c0-0881-4f2b-babd-4425db39a016 · outbound

This paper cites Qwen3 Technical Report.

RoboBrain 2.0 Technical Report Qwen3 Technical Report

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:31.085978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:31.085978Z digest=sha256:a358bd0c7d3a2b7c27c3120af1a3db16dccd2f0164abea48073ae52cc9efe0bf

Observation 58be3216-1d8d-45c0-ab0e-5d2d615eeecf · outbound

This paper cites Magma: A foundation model for multimodal ai agents.

RoboBrain 2.0 Technical Report Magma: A foundation model for multimodal ai agents

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:31.092418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:31.092418Z digest=sha256:d2ef2d01dc0d4a56223a4c12ee5332d58fec9bc4c039e19bd36032a58f7e26a2

Observation bfbd99d3-e43e-460a-b0bb-7c491d0731b7 · outbound

This paper cites Thinking in space: How multimodal large language models see, remember, and recall spaces.

RoboBrain 2.0 Technical Report Thinking in space: How multimodal large language models see, remember, and recall spaces

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:31.100133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:31.100133Z digest=sha256:6626ffc1d1754117306849e13c0f09c7caf5d76637fe39bab311927763af03bb

Observation 8d4193a7-446a-441a-9e79-8037f04ca699 · outbound

This paper cites Modeling context in referring expressions.

RoboBrain 2.0 Technical Report Modeling context in referring expressions

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:47:32.013979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:47:31.103154Z digest=sha256:e100a088c2d0720cd4d187ab1e0b34c2785ae69deff7ff3aaf85ae3e971f275e

Observation 467bcb34-25e3-48b1-a861-3e2ead12dff4 · outbound

This paper cites Robopoint: A vision-language model for spatial affordance prediction for robotics,.

RoboBrain 2.0 Technical Report Robopoint: A vision-language model for spatial affordance prediction for robotics,

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:47:32.005019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:47:31.106194Z digest=sha256:c02ffa2a2f3a7dd4500273941456445e919d71bfecde911e069cf85217d36bdf

Observation 77eb9349-8346-4a05-9614-cfb18c4f8996 · outbound

This paper cites Embodied-Reasoner: Synergizing Visual Search, Reasoning, and Action for Embodied Interactive Tasks.

RoboBrain 2.0 Technical Report Embodied-Reasoner: Synergizing Visual Search, Reasoning, and Action for Embodied Interactive Tasks

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:31.113444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:31.113444Z digest=sha256:61ef56ca61b313d2c059e577eedcbc0f43f1547c903560f80572bf25e12d7c03

Observation 9cbe1e21-b169-43ae-a279-94f17ba1a9e2 · outbound

This paper cites Recognize anything: A strong image tagging model.

RoboBrain 2.0 Technical Report Recognize anything: A strong image tagging model

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:47:31.995896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:47:31.116414Z digest=sha256:3ecd5568a496cb2b65d7363bda195f98ae6850ffde15330392b0470881cb5c71

Observation 1c3f9caa-788f-4c8b-aae9-61c6d5482610 · outbound

This paper cites Code-as-Monitor: Constraint-aware Visual Programming for Reactive and Proactive Robotic Failure Detection.

RoboBrain 2.0 Technical Report Code-as-Monitor: Constraint-aware Visual Programming for Reactive and Proactive Robotic Failure Detection

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:31.119288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:31.119288Z digest=sha256:8a0ae36bfd23515248e095f1da271ae5ddc364b3d0382ca78a9b3a92521c371a

Observation 4c8f8020-3653-4925-9c84-bdfc2bee5e80 · outbound

This paper cites Roborefer: Towards spatial referring with reasoning in vision-language models for robotics.

RoboBrain 2.0 Technical Report Roborefer: Towards spatial referring with reasoning in vision-language models for robotics

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:31.124480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:31.124480Z digest=sha256:fc5a1a1f30483c2047a0ac7e421a0ce52bee2130aea27c8e256f8c9830aa59ea

Observation 9ce09e21-fdc9-4c29-839f-f1f87cea34fd · outbound

This paper cites Large language model (llm) for telecommunications: A comprehensive survey on principles, key techniques, and opportunities.

RoboBrain 2.0 Technical Report Large language model (llm) for telecommunications: A comprehensive survey on principles, key techniques, and opportunities

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:47:31.986318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:47:31.127650Z digest=sha256:1c05f5a71743047c6b55beb7c1ab5a61ab207b7aaf4096c7aba4f767b9bd6183

Observation 152ebfed-9f71-45d7-862a-04f50438a944 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

RoboBrain 2.0 Technical Report InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:31.131012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:31.131012Z digest=sha256:4037a6db0e0ba9b63dc9aaad86d52b2fd8a32cf381c75dd0d78891ad581bcc2a

Observation baa29bac-ba56-488b-9d9a-b09757409be3 · outbound

This paper cites Please point out the orange box,.

RoboBrain 2.0 Technical Report Please point out the orange box,

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:47:31.976540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:47:31.133782Z digest=sha256:8c9947fa8b6e838ed950f918c1d6402daab0d04c4ff1ce25fecb03ea031514c9

Observation e041b661-cad0-4a2b-a86a-5917f5656bc7 · outbound

This paper cites an unresolved cited work.

RoboBrain 2.0 Technical Report Unresolved cited work

Reference 86

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:47:31.966825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:47:31.138713Z digest=sha256:a89c5f6534c5fb35f80a490caaa2538ee90a8a7499425d51b31285190215b957

Observation f38b44b1-b4f7-48c1-b0cc-271732b62447 · outbound

This paper cites Never use variable names as the action arguments, use the value instead.

RoboBrain 2.0 Technical Report Never use variable names as the action arguments, use the value instead

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:47:31.956781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:47:31.142243Z digest=sha256:01cb36cb7e7406d9362b7e3b960a2061c31946a18e570c50b32b9d94c21debf5

Observation 6b57e33b-0328-4d46-861d-ece54d711147 · outbound

This paper cites If no tool call is needed, use final_answer tool to return your answer.

RoboBrain 2.0 Technical Report If no tool call is needed, use final_answer tool to return your answer

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:47:31.946710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:47:31.145508Z digest=sha256:5b033d396318e32c55fe7b4d1c6fd4d85f70c6c4156e2d510f541b74ffa59e8d

Observation bfa14531-053c-4059-b7db-43deb12b9c8f · outbound

This paper cites # Now Begin! If you solve the task correctly, you will receive a reward of $1,000,000.

RoboBrain 2.0 Technical Report # Now Begin! If you solve the task correctly, you will receive a reward of $1,000,000

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:47:31.936002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:47:31.149194Z digest=sha256:f36f9cd3038e8d3f746d1cbffda65341e0585d20b63dd2e6aa3dc9b6242723d2

Observation 3fd2d4a8-16fb-4e9f-8587-06bfe8fdcf83 · outbound

This paper cites RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics.

RoboBrain 2.0 Technical Report RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:31.110057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:31.110057Z digest=sha256:64d939d450f3e599cda6a6921e460959c1b2c67ef3d18fe2d59433bbcab7b8ff

Pith citing papers

Observation 1e619059-6be7-4144-bfef-f0c04d9be0bc · inbound

Volume-Distance-Ratio Asymptote and Spacetime Inextendibility for FLRW Spacetimes cites this paper.

Volume-Distance-Ratio Asymptote and Spacetime Inextendibility for FLRW Spacetimes RoboBrain 2.0 Technical Report

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T04:36:34.405669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:36:34.405669Z digest=sha256:4e581287e56c0251e88727469eebf70dfd6170bd52aab1a8ce96431bdc2ca8ce

Observation a8104a18-b6bd-4ad5-b0a8-412a6d027c3a · inbound

Robix: A Unified Model for Robot Interaction, Reasoning and Planning cites this paper.

Robix: A Unified Model for Robot Interaction, Reasoning and Planning RoboBrain 2.0 Technical Report

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-05T12:59:09.771943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:59:09.771943Z digest=sha256:7d4e9b52fe6aec4458aa231fbc8a978f19d6cdf63b39b95ca068abff9498938f

Observation 721b8a4e-f81f-432b-8093-86e6492b33a4 · inbound

Contrastive Representation Regularization for Vision-Language-Action Models cites this paper.

Contrastive Representation Regularization for Vision-Language-Action Models RoboBrain 2.0 Technical Report

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T12:55:08.254972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:55:08.254972Z digest=sha256:7c57c99b5c4efd93e862adbc670be33558a55d1b8e8fdfe1265ff7d63c3808bd

Observation 5d270656-4c34-4b04-988e-df566723113e · inbound

InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy cites this paper.

InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy RoboBrain 2.0 Technical Report

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:09:40.028188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-14T20:09:39.677347Z digest=sha256:72154446a603b460e9b1cdcab5e2afde4b97927e617c58bd12daa1ffcd102e79

Observation c6fb03d1-eca1-4379-84f1-a1909f7b60b0 · inbound

DeepThinkVLA: Enhancing Reasoning Capability of Vision-Language-Action Models cites this paper.

DeepThinkVLA: Enhancing Reasoning Capability of Vision-Language-Action Models RoboBrain 2.0 Technical Report

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-18T03:10:48.948121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-18T03:09:09.713822Z digest=sha256:8862e60e6755414d9f390f0a731a8e0ad1ce4be11a5535cc6519c0943fea63f8

Observation 312253bc-4d71-48ed-bf49-7c0cd97e49b6 · inbound

MiMo-Embodied: X-Embodied Foundation Model Technical Report cites this paper.

MiMo-Embodied: X-Embodied Foundation Model Technical Report RoboBrain 2.0 Technical Report

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:42:05.697672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-17T20:40:54.096289Z digest=sha256:0a2eb08e24453cfe63ba27331e421c8bc80004868993354a8e73f73b41fecf13

Observation 1310b23c-6acf-40e2-a448-6ee664d4bf35 · inbound

RoboCOIN: An Open-Sourced Bimanual Robotic Data Collection for Integrated Manipulation cites this paper.

RoboCOIN: An Open-Sourced Bimanual Robotic Data Collection for Integrated Manipulation RoboBrain 2.0 Technical Report

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:20:12.127420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-17T20:16:16.133013Z digest=sha256:96be87f423da3d5538b06937fc7c7cca60044383a5369dea5e51a946ce231db7

Observation d3cd3b4f-b244-4860-961c-2e6fc1eb8c58 · inbound

Beyond Description: Cognitively Benchmarking Fine-Grained Action for Embodied Agents cites this paper.

Beyond Description: Cognitively Benchmarking Fine-Grained Action for Embodied Agents RoboBrain 2.0 Technical Report

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T20:42:56.199932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:42:56.199932Z digest=sha256:9a65265ec0c61ce893af19cf69486c434a09f494a560bff42511a3a2539cff7c

Observation 59b7420e-52ac-42cf-b44c-0629262cf52b · inbound

Scaling Up AI-Generated Image Detection with Generator-Aware Prototypes cites this paper.

Scaling Up AI-Generated Image Detection with Generator-Aware Prototypes RoboBrain 2.0 Technical Report

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:08:36.084109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T22:06:34.356607Z digest=sha256:dabe349eed79a68a79b73a4baa2b340715ed76999d93baf2807ac566d1487cae

Observation 1e942416-964a-4ecf-a20c-2f39435d7dca · inbound

Towards Spatial Trace with Reasoning in Vision-Language Models for Robotics cites this paper.

Towards Spatial Trace with Reasoning in Vision-Language Models for Robotics RoboBrain 2.0 Technical Report

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-03T16:27:34.144120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:27:34.144120Z digest=sha256:32c024c6926f08a07674d5cdc9bb45a2ef2dbe31a34d99c78d36243af4279c1f

Observation 40d4743d-ee15-45b4-8475-052663c11788 · inbound

VLM4VLA: Revisiting Vision-Language-Models in Vision-Language-Action Models cites this paper.

VLM4VLA: Revisiting Vision-Language-Models in Vision-Language-Action Models RoboBrain 2.0 Technical Report

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T12:30:33.865225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:30:33.865225Z digest=sha256:2c0d90c3d22648b3d5d195d43a6ace10b47c0a2b8341f5a6f2a8f1a28f7e59f5

Observation 3500d414-05ab-4a48-b91e-94189105986f · inbound

VIEW2SPACE: Studying Multi-View Visual Reasoning from Sparse Observations cites this paper.

VIEW2SPACE: Studying Multi-View Visual Reasoning from Sparse Observations RoboBrain 2.0 Technical Report

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-13T23:42:44.515158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:42:44.515158Z digest=sha256:a15dbc08cfec5884e43484a1d88bec7dd2b0f6fecd472a4dddd4bb09ade07a91

Observation 3c464154-1166-491c-bd68-a4bb32f19043 · inbound

Token Warping Helps MLLMs Look from Nearby Viewpoints cites this paper.

Token Warping Helps MLLMs Look from Nearby Viewpoints RoboBrain 2.0 Technical Report

Reference 89

Resolution
verified exact
arxiv_id, observed 2026-05-13T21:08:17.391514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T21:07:55.062113Z digest=sha256:fef88cbef74304cb68eed078792eb4f293f26cf0fde1ba31f45c2fed4bbb94bd

Observation 792e59bc-433d-4cc6-84e5-57b6a9705f94 · inbound

RoboAgent: Chaining Basic Capabilities for Embodied Task Planning cites this paper.

RoboAgent: Chaining Basic Capabilities for Embodied Task Planning RoboBrain 2.0 Technical Report

Reference 104

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:15:57.143333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T18:15:08.727921Z digest=sha256:7b766539f6b5c123cbdd8f3098b9366a8aef4c1b1d89905c0681d77b3796e4da

Observation 116bf40b-0e40-460e-9ac7-9c301b178980 · inbound

OpenVLThinkerV2: A Generalist Multimodal Reasoning Model for Multi-domain Visual Tasks cites this paper.

OpenVLThinkerV2: A Generalist Multimodal Reasoning Model for Multi-domain Visual Tasks RoboBrain 2.0 Technical Report

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:30:58.595352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T17:09:33.968186Z digest=sha256:b546f218cb8a2d86e260f313501181308ffa31bab73aed19804bdf2f952d4df9

Observation 184893ee-00ef-4250-97fa-b4bcc3484863 · inbound

AssemLM: A Spatial Reasoning Multimodal Large Language Model for Robotic Assembly cites this paper.

AssemLM: A Spatial Reasoning Multimodal Large Language Model for Robotic Assembly RoboBrain 2.0 Technical Report

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:41:02.016734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T18:24:42.800233Z digest=sha256:ce80782fa0b6ed01204dcf103b191b4dd473c01cf7e17b368de948cf30e94764

Observation 7c981598-7bfe-4828-8115-cb08b43cab03 · inbound

AssemLM: A Spatial Reasoning Multimodal Large Language Model for Robotic Assembly cites this paper.

AssemLM: A Spatial Reasoning Multimodal Large Language Model for Robotic Assembly RoboBrain 2.0 Technical Report

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-12T23:37:01.875768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T23:37:01.875768Z digest=sha256:a2dbafe191a78a24be592933e30e55340a9a4c0da5af2634e7d665d2cada1cbf

Observation c96ee93e-b8a2-4904-8d92-d2025b1fee69 · inbound

Unmasking the Illusion of Embodied Reasoning in Vision-Language-Action Models cites this paper.

Unmasking the Illusion of Embodied Reasoning in Vision-Language-Action Models RoboBrain 2.0 Technical Report

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:56:28.939398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T04:27:18.284698Z digest=sha256:47dd33278ba5830fd7e0dc7ef758e0b80f7dc1e15eeb3a87cc9cd19416949e35

Observation 8865ec66-acc0-40a1-b058-ada5ceb91e39 · inbound

Long-Horizon Manipulation via Trace-Conditioned VLA Planning cites this paper.

Long-Horizon Manipulation via Trace-Conditioned VLA Planning RoboBrain 2.0 Technical Report

Reference 59

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:41:38.906730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-09T21:10:03.554504Z digest=sha256:734913ed37fc2ad6a4d6b9659430efbe78feab9b195357456a47f11bf49b20b4

Observation 562395d8-55e6-4094-aced-4014a526ab31 · inbound

Assistance Without Interruption: A Benchmark and LLM-based Framework for Non-Intrusive Human-Robot Assistance cites this paper.

Assistance Without Interruption: A Benchmark and LLM-based Framework for Non-Intrusive Human-Robot Assistance RoboBrain 2.0 Technical Report

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:51:06.874405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-09T14:51:10.807902Z digest=sha256:48e79bbeca802e8d18bc49971344616a471fb1b4b69ee9e8091ca30a92d27c47

Observation be5652f5-c3a7-465d-8802-17a00bfcf6de · inbound

RoboEvolve: Co-Evolving Planner-Simulator for Robotic Manipulation with Limited Data cites this paper.

RoboEvolve: Co-Evolving Planner-Simulator for Robotic Manipulation with Limited Data RoboBrain 2.0 Technical Report

Reference 71

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T17:57:33.123620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-14T17:54:50.325820Z digest=sha256:e8eca657d3d0d806419154151a40658988de11900ffdb55db843038756ef64ff

Observation cc82a7c8-12e2-42ec-a883-59a68d12ed19 · inbound

Rethinking VLM Representation for VLA Initialization cites this paper.

Rethinking VLM Representation for VLA Initialization RoboBrain 2.0 Technical Report

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:24:00.048577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-29T22:21:29.733181Z digest=sha256:6a2c42eea2e0420157dfd54ae0ce83ec67ae9fed650c51d5174e04f0f06339ca

Observation def6574d-f7bc-4516-84fa-b5d4a983d7fc · inbound

Extending Embodied Question Answering from Perception to Decision cites this paper.

Extending Embodied Question Answering from Perception to Decision RoboBrain 2.0 Technical Report

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-06-29T21:33:58.975304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-29T21:30:40.182958Z digest=sha256:0a832e23c7ee2714279f673f312bb097dd3addfaefc07494f3bf787e8f23c40b

Observation 789c2ad1-2173-4197-8139-3b88c54c1d28 · inbound

Grounded 3D-Aware Spatial Vision-Language Modeling cites this paper.

Grounded 3D-Aware Spatial Vision-Language Modeling RoboBrain 2.0 Technical Report

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:13:15.373936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-29T08:08:36.012761Z digest=sha256:08df472d4ad56c727cccfe98fd78c97c7ff2af9da04d7e7410c84804a82f6f39

Observation 3f16f2a2-b1dc-499f-8ce2-581e53067d84 · inbound

Token Predictors Are Not Planners: Building Physically Grounded Causal Reasoners cites this paper.

Token Predictors Are Not Planners: Building Physically Grounded Causal Reasoners RoboBrain 2.0 Technical Report

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-01T23:26:23.011643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-28T14:19:40.492920Z digest=sha256:324d5083c7708263953633b7c244bde95b92abc2660c2b57f54037e6180f6743

Observation 6e5147c4-8e9e-4803-9239-d81c15592bd9 · inbound

Thinking with Imagination: Agentic Visual Spatial Reasoning with World Simulators cites this paper.

Thinking with Imagination: Agentic Visual Spatial Reasoning with World Simulators RoboBrain 2.0 Technical Report

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:06:56.199501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-28T02:25:30.998989Z digest=sha256:5b6cc74b617722b55a0ec9867db09430ba40aa6b65cbad1d6eabf85315d1a337

Observation 321c1fc5-bf7f-4f15-a02b-395bb2b7be90 · inbound

Two Bridges, One Pathway: From VLMs to Generalizable VLAs with Embodied Trajectory-Coupled Data cites this paper.

Two Bridges, One Pathway: From VLMs to Generalizable VLAs with Embodied Trajectory-Coupled Data RoboBrain 2.0 Technical Report

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:57:26.661631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T18:31:03.548002Z digest=sha256:96d464e933eb1f74b6e7d029ac2e085a2b71ab4f716b1bb90ee4425d08350937

Observation 1c86762d-7507-4838-a49b-2c4e8bb7adef · inbound

Kwai Keye-VL-2.0 Technical Report cites this paper.

Kwai Keye-VL-2.0 Technical Report RoboBrain 2.0 Technical Report

Reference 77

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T04:27:36.986423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-27T13:53:10.352603Z digest=sha256:f1cf74832a52f52c9d25c418e74a80abf391ea08d09370ae327e8308312a96f7

Observation b280ef3b-d1f4-4507-b32d-f76ad2c8f552 · inbound

Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models cites this paper.

Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models RoboBrain 2.0 Technical Report

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-07-03T06:07:41.253641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T12:55:47.754632Z digest=sha256:a3baa7cc829bcf4315ccf5232c6690c001e21c82ceaeddcc3c4d16424137e9e6

Observation accaf290-34ed-4811-a996-24dba529a17e · inbound

Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models cites this paper.

Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models RoboBrain 2.0 Technical Report

Reference 87

Resolution
unresolved
no resolver link, observed 2026-07-14T18:07:09.018997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T18:07:09.018997Z digest=sha256:c1f62d2be7d4d395bfd5283ca85bd3bc4e434d78e03c353f0277d5b1d30fb51b

Observation d14bc912-963f-4384-bef1-2007f85d81ef · inbound

RoboProcessBench: Benchmarking Process-Aware Understanding in Vision-Language Robotic Manipulation cites this paper.

RoboProcessBench: Benchmarking Process-Aware Understanding in Vision-Language Robotic Manipulation RoboBrain 2.0 Technical Report

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:28:34.762558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T06:26:32.209719Z digest=sha256:59af8624871fc62788859f02c6e51d73c6457a02794528d22edb6c1d25f914bc

Observation 8e09e1a5-fbbc-4c8b-96aa-c59eb4f48b7d · inbound

RoboProcessBench: Benchmarking Process-Aware Understanding in Vision-Language Robotic Manipulation cites this paper.

RoboProcessBench: Benchmarking Process-Aware Understanding in Vision-Language Robotic Manipulation RoboBrain 2.0 Technical Report

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T11:43:49.545830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:43:49.545830Z digest=sha256:774760b4551a865c6ef1c655d5e1374149b38d9fe3b378450d537ca5e1b2cbd7

Observation 67811327-e148-4b8a-a6c2-c92a09affd0d · inbound

SPARC: Reliable Spatial Annotations from Robot Demonstrations at Scale cites this paper.

SPARC: Reliable Spatial Annotations from Robot Demonstrations at Scale RoboBrain 2.0 Technical Report

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:28:34.372338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T06:28:14.694168Z digest=sha256:6928598824bed7d66b8013f5867d7969b4c60d8b035ca09443bf6a1a69902c4b

Observation 5afc2cf3-1a49-4f3e-b658-38abde83644e · inbound

RoboPIN: Grounded Embodied Reasoning via Pinned Chain-of-Thought cites this paper.

RoboPIN: Grounded Embodied Reasoning via Pinned Chain-of-Thought RoboBrain 2.0 Technical Report

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-06-30T11:24:38.342268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T11:17:48.279808Z digest=sha256:77346fd6009cc8f6bdd655d97f271d34457d3246c45c7882d48def44c46e856c

Observation 7be93e99-e912-46d8-a1e9-be24969da321 · inbound

RoboPIN: Grounded Embodied Reasoning via Pinned Chain-of-Thought cites this paper.

RoboPIN: Grounded Embodied Reasoning via Pinned Chain-of-Thought RoboBrain 2.0 Technical Report

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T11:20:27.100594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:20:27.100594Z digest=sha256:02865d0f6d3e0d58e251d59defc53fb7fa6023d3915cbbfb79ef6617c83985e7

Observation b0a6796d-6f6e-498f-896d-ebb3505f7740 · inbound

Vesta: A Generalist Embodied Reasoning Model cites this paper.

Vesta: A Generalist Embodied Reasoning Model RoboBrain 2.0 Technical Report

Reference 109

Resolution
verified exact
arxiv_id, observed 2026-07-04T04:29:35.627244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-26T16:55:12.518255Z digest=sha256:e109425c093b80eab18d366abb40215f256b78c2eb3fcd955b54e47102ededd5

Observation c588d71b-7b20-48a4-92ad-439f99dc3459 · inbound

LIBERO-Safety: A Comprehensive Benchmark for Physical and Semantic Safety in Vision-Language-Action Models cites this paper.

LIBERO-Safety: A Comprehensive Benchmark for Physical and Semantic Safety in Vision-Language-Action Models RoboBrain 2.0 Technical Report

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T11:29:51.478385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-26T07:55:28.884473Z digest=sha256:c71b5bbb9d777c4ba0d0a67c20098b1055def4475bbbcd4d553b8dc99e0619b7

Observation 24cb4fcb-10bf-48fd-8a4a-65e2b9372917 · inbound

LIBERO-Safety: A Comprehensive Benchmark for Physical and Semantic Safety in Vision-Language-Action Models cites this paper.

LIBERO-Safety: A Comprehensive Benchmark for Physical and Semantic Safety in Vision-Language-Action Models RoboBrain 2.0 Technical Report

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T19:33:54.753630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-29T04:41:28.564104Z digest=sha256:3c1acd1e040eb635deaa1796c1868f02c9d84ca5501d19b8ba1af138687ec23e

Observation 4c17d25a-fd0e-4a60-90f7-ddad61284316 · inbound

Advancing Omnimodal Embodied Agents from Isolated Skills to Everyday Physical Autonomy cites this paper.

Advancing Omnimodal Embodied Agents from Isolated Skills to Everyday Physical Autonomy RoboBrain 2.0 Technical Report

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:49:52.044905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-26T04:52:51.524022Z digest=sha256:0f04101bf2cfb21ca89ef812a472a56b692d9a89f3faeae4b26f52706178b2ba

Observation 06fed410-4c3c-41b4-9306-01dfa57cda84 · inbound

FutureNav: Unified World-Action Modeling for Vision-and-Language Navigation cites this paper.

FutureNav: Unified World-Action Modeling for Vision-and-Language Navigation RoboBrain 2.0 Technical Report

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-06-30T15:44:49.418964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T05:05:56.065261Z digest=sha256:fdfc0c7f24ce4aa8966bfd54ce48a39b551f6e02470060aab7815119f2f1ff9d

Observation 897b8aa1-69fd-47f7-8154-e21fb5a12df2 · inbound

OmniView-Space: Reinforcing Spatial Reasoning via Multi-Perspective Spatial Mapping cites this paper.

OmniView-Space: Reinforcing Spatial Reasoning via Multi-Perspective Spatial Mapping RoboBrain 2.0 Technical Report

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-07-02T14:17:02.454959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-02T14:16:39.649823Z digest=sha256:e611ad1ed69e696228201e0ded425131cd4b8750dd8b8ad84e2cc8c674e87eaf

Observation eddaea49-160f-4f46-b2b9-b8d3487ef1b8 · inbound

ACE-Brain-0.5: A Unified Embodied Foundational Model for Physical Agentic AI cites this paper.

ACE-Brain-0.5: A Unified Embodied Foundational Model for Physical Agentic AI RoboBrain 2.0 Technical Report

Reference 55

Resolution
unresolved
no resolver link, observed 2026-07-11T19:16:57.396710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:16:57.396710Z digest=sha256:f24bc68facc167ae37e76520d81c333e0b3bf56296b338579b4be15245f0ccb8

Observation c4030680-8200-4dd6-a557-2cc703beefc7 · inbound

Diagnosing Semantic Handoff Failures in Agent-Orchestrated Vision-Language-Action Skill Composition cites this paper.

Diagnosing Semantic Handoff Failures in Agent-Orchestrated Vision-Language-Action Skill Composition RoboBrain 2.0 Technical Report

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-08T12:04:50.479862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-08T12:00:14.336135Z digest=sha256:8d95df6d5ab66eaee6ec1e3a8cac5ea937f728099afaabb10596fad5f4a489d0

Observation 08bcabca-1051-4af6-b540-b0d21096afe0 · inbound

Diagnosing Semantic Handoff Failures in Agent-Orchestrated Vision-Language-Action Skill Composition cites this paper.

Diagnosing Semantic Handoff Failures in Agent-Orchestrated Vision-Language-Action Skill Composition RoboBrain 2.0 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T08:21:36.948882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:21:36.948882Z digest=sha256:9657ba5424ec49d432a7c2c4485adbf24612672e03f5200c8c5aeb11cf491ab7

Observation bd05024f-bb2d-4fd8-a1c9-7448c1edee98 · inbound

UniLM-Nav: A Unified Framework for Zero-Shot Last-Mile Navigation cites this paper.

UniLM-Nav: A Unified Framework for Zero-Shot Last-Mile Navigation RoboBrain 2.0 Technical Report

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-07-08T02:44:27.771049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-08T02:41:24.989404Z digest=sha256:754df5eca0bbfab8d7dac9cb932babe5acda72ae1d4d768bfaefeca1d436e1b5

Observation 98b34350-2ead-4960-abfc-0934a5747fc6 · inbound

UniLM-Nav: A Unified Framework for Zero-Shot Last-Mile Navigation cites this paper.

UniLM-Nav: A Unified Framework for Zero-Shot Last-Mile Navigation RoboBrain 2.0 Technical Report

Reference 55

Resolution
unresolved
no resolver link, observed 2026-07-14T16:00:02.584798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:00:02.584798Z digest=sha256:c446ed361a3cf39902a34bd35fd198cc585859d9f082cdcec245448434bf1e4e

Observation 3b298bf2-7407-4c2f-bb5f-bab3dc6a40ab · inbound

ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory cites this paper.

ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory RoboBrain 2.0 Technical Report

Reference 74

Resolution
unresolved
no resolver link, observed 2026-07-14T12:26:27.446079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:26:27.446079Z digest=sha256:5b2a5165909cb0884ed276b0daae84b966245dfa191469a4bfdafd20402c7017

Observation 275781c7-e1be-445d-8f89-fd43418dc941 · inbound

ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory cites this paper.

ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory RoboBrain 2.0 Technical Report

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-02T07:21:31.686911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:21:31.686911Z digest=sha256:14f4bd89ef6e4d28abb421806588a8e2fca23db8018035ea79c31d00467d6a2f

Observation 8e10392d-9963-4043-9964-8d0bf5d46972 · inbound

UniETP: Unifying Environments for Generalizable Embodied Task Planning cites this paper.

UniETP: Unifying Environments for Generalizable Embodied Task Planning RoboBrain 2.0 Technical Report

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T16:20:28.787487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:20:28.787487Z digest=sha256:9b23e69248e804872e8e2d40f9a59f59b613bcd12cb075a980491bf54c140f0d

Observation 0695f7e4-7f7f-49ba-9bf2-497e60fab566 · inbound

RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation cites this paper.

RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation RoboBrain 2.0 Technical Report

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-01T14:39:41.142227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:39:41.142227Z digest=sha256:45b69142e237ff1151b32852c5e26789d9854fab1a4648e113262415540888a6

Observation 57720f11-e530-450d-82e3-d0b31f305f76 · inbound

Towards General Language-Conditioned Latent Safety Filters cites this paper.

Towards General Language-Conditioned Latent Safety Filters RoboBrain 2.0 Technical Report

Reference 246

Resolution
unresolved
no resolver link, observed 2026-08-04T00:49:36.759144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:49:36.759144Z digest=sha256:dae3b3d4f7238b2fd9ac17c529abb950f6484c3d9a00dde21037366e5a417629

Observation 29606749-d338-4124-8d6e-841411590730 · inbound

SpatialAfford: Teaching Compact VLMs Where to Look and Where to Ground for Affordance cites this paper.

SpatialAfford: Teaching Compact VLMs Where to Look and Where to Ground for Affordance RoboBrain 2.0 Technical Report

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-05T00:54:48.781637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T00:54:48.781637Z digest=sha256:d1b87c949687b30c0ddc58166e470ba175701f8b99dc2ac28e46328e3b07d4f6

Observation a1f67bc6-4803-4003-9aae-34437e4bcdcd · inbound

Beyond Flat Policies: Hierarchical Post-Training for Embodied Agents in Robotic Manipulation cites this paper.

Beyond Flat Policies: Hierarchical Post-Training for Embodied Agents in Robotic Manipulation RoboBrain 2.0 Technical Report

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T19:38:41.751687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T19:38:41.751687Z digest=sha256:255741b7354a61829d2deef8178dfa76278818c78cef507a981c40cf323ec17e

Observation d0722474-d9b0-4ecd-8677-e418b507e6aa · inbound

Compiling and Benchmarking Task-State Horizons for Embodied Agents cites this paper.

Compiling and Benchmarking Task-State Horizons for Embodied Agents RoboBrain 2.0 Technical Report

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T00:36:21.485253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:36:21.485253Z digest=sha256:7a1b20189fc17df3ef4736632b3254484f505f9b7305186a5288f7c0cf9b97d6