Pith. sign in

Paper Citation Record · LEDGER

PaLM-E: An Embodied Multimodal Language Model

As of 4 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 100 inbound Pith citation observations for arXiv:2303.03378.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2303.03378 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T22:29:29.631351Z

measured 143 of 143 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 100 of 267 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T18:25:37.332342Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-11T00:27:51.904332Z

Reference resolution

43 of 43 outbound references displayed

  • verified exact28
  • verified fuzzy2
  • unresolved0
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch12

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8d5980c1-8d64-4b0f-8f14-a23bce35f2a3 · outbound

This paper cites Do As I Can, Not As I Say: Grounding Language in Robotic Affordances.

PaLM-E: An Embodied Multimodal Language Model Do As I Can, Not As I Say: Grounding Language in Robotic Affordances

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-10T22:29:29.801274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T22:29:29.631351Z digest=sha256:d15763aa307a1217e932682a52abc81232d1aa11cae7075c01e02835e82d91c2

Observation 02f15836-2537-456b-b3bf-1d45accaf8a4 · outbound

This paper cites Flamingo: a Visual Language Model for Few-Shot Learning.

PaLM-E: An Embodied Multimodal Language Model Flamingo: a Visual Language Model for Few-Shot Learning

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:22:31.147880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T22:29:29.631351Z digest=sha256:1e6c3cfdc7b57448ab47e98837d00eb0e009172ad425cb5c566cf1e88efb6b20

Observation 303298d0-28e9-4451-b9d2-c4fde0a67946 · outbound

This paper cites On the Opportunities and Risks of Foundation Models.

PaLM-E: An Embodied Multimodal Language Model On the Opportunities and Risks of Foundation Models

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-05-10T22:29:29.750349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T22:29:29.631351Z digest=sha256:03a6a6de716822bf2f8589cd251e1adc4f8cdd378b96c573fd6493de7918343e

Observation ea8564a8-7678-4399-beb8-d57b5879c0f0 · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

PaLM-E: An Embodied Multimodal Language Model RT-1: Robotics Transformer for Real-World Control at Scale

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:41:14.065009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T22:29:29.631351Z digest=sha256:a4deca5f2e7cbc2b319c1377f5bea551413800a3a746f0c3fd9d2734737fca42

Observation 86941646-4b35-4482-a8e6-5315bddc62d8 · outbound

This paper cites D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al.

PaLM-E: An Embodied Multimodal Language Model D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:29:30.118388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T22:29:29.631351Z digest=sha256:21f1a97bdc9ed917e2b49716e95e540c7b5450d591ab0955000caf3d0f043e10

Observation 811630a8-0bf9-44f0-8603-2461f295d179 · outbound

This paper cites All You May Need for VQA are Image Captions.

PaLM-E: An Embodied Multimodal Language Model All You May Need for VQA are Image Captions

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:29:29.812541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T22:29:29.631351Z digest=sha256:af786f33c211ed746db32775b4fb44e8ee00c500871ae11912f8a11bec4c54d3

Observation 24c19ffd-ab1b-4d7f-a102-0260ab345d82 · outbound

This paper cites PaLI: A Jointly-Scaled Multilingual Language-Image Model.

PaLM-E: An Embodied Multimodal Language Model PaLI: A Jointly-Scaled Multilingual Language-Image Model

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:29:06.753688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T22:29:29.631351Z digest=sha256:4459f9c7d335374f91dac1eb6d8d2a52dd338cb1c1334971ca18a4480108cdbd

Observation 45ef8d62-45a7-419a-9967-2137931745d9 · outbound

This paper cites PaLM: Scaling Language Modeling with Pathways.

PaLM-E: An Embodied Multimodal Language Model PaLM: Scaling Language Modeling with Pathways

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:45:07.766206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T22:29:29.631351Z digest=sha256:429434415d49376e4c2f648ed0d9b8618a988f431187445773f35263be95446d

Observation 23729b9b-4725-47c5-9564-2da4060c3202 · outbound

This paper cites Scaling Vision Transformers to 22 Billion Parameters.

PaLM-E: An Embodied Multimodal Language Model Scaling Vision Transformers to 22 Billion Parameters

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:29:29.833939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T22:29:29.631351Z digest=sha256:cd13b224bfbc78091003152d076e268a2f2dbd525319b125e19fa09dbf3842ba

Observation ae7593c5-f0e6-43be-8c37-4bf43db9b423 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

PaLM-E: An Embodied Multimodal Language Model BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-10T22:29:29.839743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T22:29:29.631351Z digest=sha256:605db8ccc17a8d0246d90ba3f3beb9018699bb88122b09fb37d0149e07bafc16

Observation 1d081a66-bd2b-4288-8371-238113aa6fba · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

PaLM-E: An Embodied Multimodal Language Model An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-10T22:29:29.844548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T22:29:29.631351Z digest=sha256:66887114be2336184e16f77647e1baeafeefc44d7b4c049b40d0d33b27d18bfc

Observation 7d1a0ea5-566c-4df5-b07e-4a56bc74b0ec · outbound

This paper cites Improving alignment of dialogue agents via targeted human judgements.

PaLM-E: An Embodied Multimodal Language Model Improving alignment of dialogue agents via targeted human judgements

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-14T17:54:02.323766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T22:29:29.631351Z digest=sha256:7418ed6e17d1f020dde6dea03683595baf2e7ba6e329868b84999764efa5f361

Observation 297c9fa1-139f-4d70-a3a4-47dcc2a7a35f · outbound

This paper cites Instruction-driven history-aware policies for robotic manipulations.

PaLM-E: An Embodied Multimodal Language Model Instruction-driven history-aware policies for robotic manipulations

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:29:29.858550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T22:29:29.631351Z digest=sha256:5604493feb45ed209fece8448c56b8c6dfe9c3afdf32428089312258c18ebef8

Observation 74c66787-ce03-4df3-a046-8fa008fb051d · outbound

This paper cites Language Models are General-Purpose Interfaces.

PaLM-E: An Embodied Multimodal Language Model Language Models are General-Purpose Interfaces

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:29:29.866806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T22:29:29.631351Z digest=sha256:e2faa9b74313ea5fa56ed057e48822acaf2c101a422ef79c1103bbaa42e774cf

Observation 82328e8e-2e44-473b-9336-2e13dd6899f1 · outbound

This paper cites Visual Language Maps for Robot Navigation.

PaLM-E: An Embodied Multimodal Language Model Visual Language Maps for Robot Navigation

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:29:29.899114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T22:29:29.631351Z digest=sha256:4fd7a6498d0b1f9836a41b646bdd9ae6e0246bdeaa34c302be371c44ee273c2f

Observation 89eb5b12-26de-4f9c-82e2-b75251f261a7 · outbound

This paper cites VIMA: General Robot Manipulation with Multimodal Prompts.

PaLM-E: An Embodied Multimodal Language Model VIMA: General Robot Manipulation with Multimodal Prompts

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:29:29.904362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T22:29:29.631351Z digest=sha256:acfbd10248ef38521f8a52f28b1fed034f52de2451c8ffa3a6bb63ab72462e4a

Observation eb9b95c6-a477-40d2-9bfd-f07c4b0835f2 · outbound

This paper cites Large Language Models are Zero-Shot Reasoners.

PaLM-E: An Embodied Multimodal Language Model Large Language Models are Zero-Shot Reasoners

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T19:04:13.904485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T22:29:29.631351Z digest=sha256:b41810548addbdcb506f353023225d49313ed8b6d1ea0568c3b598837c0589f9

Observation a953ebc0-911b-4701-b30b-b78fe8da32c8 · outbound

This paper cites The Power of Scale for Parameter-Efficient Prompt Tuning.

PaLM-E: An Embodied Multimodal Language Model The Power of Scale for Parameter-Efficient Prompt Tuning

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:34:05.826717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T22:29:29.631351Z digest=sha256:2e5098e28845aeb948d92de07c2307411c58a03151a224b18ddd924835de792c

Observation 01278255-d868-4db2-b19b-803b333fe507 · outbound

This paper cites Solving Quantitative Reasoning Problems with Language Models.

PaLM-E: An Embodied Multimodal Language Model Solving Quantitative Reasoning Problems with Language Models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-12T22:43:59.458883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T22:29:29.631351Z digest=sha256:ed3086c191416e8d87df4efc5b0aafc59144903a00f0cdd64c67665b9719cf8b

Observation ca722207-024a-4b52-81a1-09397919ae31 · outbound

This paper cites VisualBERT: A Simple and Performant Baseline for Vision and Language.

PaLM-E: An Embodied Multimodal Language Model VisualBERT: A Simple and Performant Baseline for Vision and Language

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T23:59:38.201352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T22:29:29.631351Z digest=sha256:b854013758000547e414e9b9ef0a8bcc814bc716713f82f5a4becc7590f4840f

Observation f885576d-2d45-4e59-adb5-288dda9c0720 · outbound

This paper cites TrOCR: Transformer-based Optical Character Recognition with Pre-trained Models.

PaLM-E: An Embodied Multimodal Language Model TrOCR: Transformer-based Optical Character Recognition with Pre-trained Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:29:29.944797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T22:29:29.631351Z digest=sha256:1a0e289beae1be00d950a9783c5b89f5db4a158843161d614afc1c9c137fd213

Observation 733b2ae1-ab8b-4aa2-bb54-607b9a731b36 · outbound

This paper cites Pre-Trained Language Models for Interactive Decision-Making.

PaLM-E: An Embodied Multimodal Language Model Pre-Trained Language Models for Interactive Decision-Making

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:29:29.954072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T22:29:29.631351Z digest=sha256:d930d413c9e49de5b30a510d2ccb49b84c645aa8ada9093f59053d6d1907e2a7

Observation b1a2fde3-bc68-4418-8d14-22627f5d2edf · outbound

This paper cites Code as Policies: Language Model Programs for Embodied Control.

PaLM-E: An Embodied Multimodal Language Model Code as Policies: Language Model Programs for Embodied Control

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-15T00:38:03.084841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T22:29:29.631351Z digest=sha256:63a36ec0ce4b4548fff6f63e34b8dda55799adb83d1dbcb15ede21312654e298

Observation fcf4203a-294f-4086-a5e6-9c0a2bab3cab · outbound

This paper cites Pretrained Transformers as Universal Computation Engines.

PaLM-E: An Embodied Multimodal Language Model Pretrained Transformers as Universal Computation Engines

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:29:29.969876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T22:29:29.631351Z digest=sha256:9dd4bc361cf4e420329e9c06e7e43158e8c38cc2042ada7315f8e413d7203583

Observation f845b85c-d44c-44d9-a7c8-bca9235d2992 · outbound

This paper cites Language Conditioned Imitation Learning over Unstructured Data.

PaLM-E: An Embodied Multimodal Language Model Language Conditioned Imitation Learning over Unstructured Data

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:29:29.974642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T22:29:29.631351Z digest=sha256:9a07a111cbbdfff3d59450ed8439a2389903e9ab6b3901359c2e3214be1b8418

Observation bba9eada-1ce2-4d47-a1d4-bb4401214cfd · outbound

This paper cites Interactive Language: Talking to Robots in Real Time.

PaLM-E: An Embodied Multimodal Language Model Interactive Language: Talking to Robots in Real Time

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:29:29.983883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T22:29:29.631351Z digest=sha256:3998dc8d45e3b9fbd0c932db05ef85be01faea2313a7fa65bf7ab6bc8d1cb97c

Observation 92fbd9ce-f5de-4aa8-bc4b-917fd279b3e1 · outbound

This paper cites Do Embodied Agents Dream of Pixelated Sheep: Embodied Decision Making using Language Guided World Modelling.

PaLM-E: An Embodied Multimodal Language Model Do Embodied Agents Dream of Pixelated Sheep: Embodied Decision Making using Language Guided World Modelling

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:29:29.990431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T22:29:29.631351Z digest=sha256:c5dda33b7cbc4a7b02f3e03f828041c7ac3ba2ec6c74641ba6421eb5d8b48ba9

Observation 706073ba-a5e9-47a0-8e05-3bf21917435d · outbound

This paper cites Formal Mathematics Statement Curriculum Learning.

PaLM-E: An Embodied Multimodal Language Model Formal Mathematics Statement Curriculum Learning

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:29:29.994704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T22:29:29.631351Z digest=sha256:7cb5b350d4de33defe43890356b6f6b28b64a252213017bf720c8ffb03e5b923

Observation 6f5579bb-94b4-4f69-bf05-f04ed0a8dff1 · outbound

This paper cites A Generalist Agent.

PaLM-E: An Embodied Multimodal Language Model A Generalist Agent

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T06:24:50.082909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T22:29:29.631351Z digest=sha256:36aaa55a2d59ff6791c2f11682a6636254051b4150417756a5342d716702e4ea

Observation 30516c5a-679c-47c5-afd6-0be2ed571833 · outbound

This paper cites TokenLearner: What Can 8 Learned Tokens Do for Images and Videos?.

PaLM-E: An Embodied Multimodal Language Model TokenLearner: What Can 8 Learned Tokens Do for Images and Videos?

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:29:30.012482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T22:29:29.631351Z digest=sha256:18e191d53fc66f5879c78f39683162c87a5b18654f3957fb6145c459e6cee54f

Observation 6879b259-3c74-4613-a2c8-59e00d13d838 · outbound

This paper cites LM-Nav: Robotic Navigation with Large Pre-Trained Models of Language, Vision, and Action.

PaLM-E: An Embodied Multimodal Language Model LM-Nav: Robotic Navigation with Large Pre-Trained Models of Language, Vision, and Action

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:29:30.017202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T22:29:29.631351Z digest=sha256:0736eaeafbe86c41e922da33778ae14913bc7cfaf04cb36c8831dbac0e7c2cfe

Observation 543ae5af-f3c4-4947-a350-7707bfbd69d3 · outbound

This paper cites Skill Induction and Planning with Latent Language.

PaLM-E: An Embodied Multimodal Language Model Skill Induction and Planning with Latent Language

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:29:30.028289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T22:29:29.631351Z digest=sha256:63a35e5861f183aeb8775fee585940d95547e86a722b87022ae760e300ffd697

Observation 2ec82c29-58e0-45e9-a046-338b31e8d5c0 · outbound

This paper cites Perceiver-Actor: A Multi-Task Transformer for Robotic Manipulation.

PaLM-E: An Embodied Multimodal Language Model Perceiver-Actor: A Multi-Task Transformer for Robotic Manipulation

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:29:30.039748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T22:29:29.631351Z digest=sha256:633496fa1633aaed87045dc800fed66b9d7f097612a138002118d42085c33a79

Observation c0bf4e8f-a860-45aa-924c-6c0bcf3661b8 · outbound

This paper cites ProgPrompt: Generating Situated Robot Task Plans using Large Language Models.

PaLM-E: An Embodied Multimodal Language Model ProgPrompt: Generating Situated Robot Task Plans using Large Language Models

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-17T04:22:17.799460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T22:29:29.631351Z digest=sha256:173bc03e0981150e13bfb2cf5d1f2a1c48248fd7a7a3011c12440e89cbae1ef7

Observation 52901c5e-8564-413a-a4dd-a12d540e7d93 · outbound

This paper cites LaMDA: Language Models for Dialog Applications.

PaLM-E: An Embodied Multimodal Language Model LaMDA: Language Models for Dialog Applications

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:29:30.059000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T22:29:29.631351Z digest=sha256:00781781f3c951aa1b493b2fe1ad51856422d29e36776a10b0f96d61212ba8d0

Observation 4e8bf568-b553-4fd3-a621-682da4214420 · outbound

This paper cites Describe, Explain, Plan and Select: Interactive Planning with Large Language Models Enables Open-World Multi-Task Agents.

PaLM-E: An Embodied Multimodal Language Model Describe, Explain, Plan and Select: Interactive Planning with Large Language Models Enables Open-World Multi-Task Agents

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-16T03:27:40.929615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T22:29:29.631351Z digest=sha256:0f77d487d9036980023a3f2e998393995cec4795f761d67ebd989b4f701dfe42

Observation 06fe88a1-4e83-40ed-bd8b-8643fd594637 · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

PaLM-E: An Embodied Multimodal Language Model Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-10T22:29:30.103778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T22:29:29.631351Z digest=sha256:4186f0c67db459c2b87664a7f87ddee72eea9f494b6681a72e5e4036834484ee

Observation 93ea694c-e438-49af-ac54-7457bd0ed1e9 · outbound

This paper cites Robotic Skill Acquisition via Instruction Augmentation with Vision-Language Models.

PaLM-E: An Embodied Multimodal Language Model Robotic Skill Acquisition via Instruction Augmentation with Vision-Language Models

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:29:30.111909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T22:29:29.631351Z digest=sha256:00910c4a6e459ddea00f0cc26b891a096daa7e5a03f4125f0a576c2377594885

Observation 57b594b3-7414-4806-b7b3-42a2e27569e3 · outbound

This paper cites PIGLeT: Language Grounding Through Neuro-Symbolic Interaction in a 3D World.

PaLM-E: An Embodied Multimodal Language Model PIGLeT: Language Grounding Through Neuro-Symbolic Interaction in a 3D World

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:29:29.784246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T22:29:29.631351Z digest=sha256:b93f3f682bbf4173d8eb6e10fec2d0ba27e04b9a7ac4857431eaad60a5111d6b

Observation 6e493240-fcb5-40d5-ad5f-86acf5448797 · outbound

This paper cites Hierarchical Task Learning from Language Instructions with Unified Transformers and Self-Monitoring.

PaLM-E: An Embodied Multimodal Language Model Hierarchical Task Learning from Language Instructions with Unified Transformers and Self-Monitoring

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:29:29.794554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T22:29:29.631351Z digest=sha256:d21147acd85040fd4bcb4434e4d3b4ee8378a8af23ba8f6eeae8056427240637

Observation 9a2d297a-b69b-4ecc-89c3-bcd07579daa5 · outbound

This paper cites full mixture.

PaLM-E: An Embodied Multimodal Language Model full mixture

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:29:30.134622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T22:29:29.631351Z digest=sha256:83e49e36b4074f3957066dfa35e89e2892c206607e896d133a33115f0692d5b1

Observation 034efaa3-7e5f-4051-8f70-3ebb075075f8 · outbound

This paper cites an unresolved cited work.

PaLM-E: An Embodied Multimodal Language Model Unresolved cited work

Reference 42

Resolution
malformed identifier
raw_fallback, observed 2026-05-10T22:29:30.127437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T22:29:29.631351Z digest=sha256:7062020b2db70c4d00c89105c3d5a96d86be1d266ad4eb3e2f1678258b7ee2f3

Observation faa2a7d9-cf9a-4528-91b6-0d9534f22801 · outbound

This paper cites an unresolved cited work.

PaLM-E: An Embodied Multimodal Language Model Unresolved cited work

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:29:29.680084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T22:29:29.631351Z digest=sha256:51d0c1f9eb2f316ebeaf35be8e8ba5e117324147cc53588c6c0a96cfa0ddb88c

Pith citing papers

Observation 6a8c0286-ff1f-4f13-ae9c-071fae1588f2 · inbound

MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action cites this paper.

MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action PaLM-E: An Embodied Multimodal Language Model

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-14T01:17:58.747224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T01:17:58.678036Z digest=sha256:f390f8e2ebd47788ada5378c47a39d2083fa4b99e7e9c35eeb97763eae12446c

Observation c1f45a0d-ffc9-4e1d-9d69-a49c7aa7edb0 · inbound

Language Models can Solve Computer Tasks cites this paper.

Language Models can Solve Computer Tasks PaLM-E: An Embodied Multimodal Language Model

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-17T12:17:26.720074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T12:17:26.602361Z digest=sha256:b362d40dd9060731088937d49a2f1c544d375294b85e564c0b163424f24c38a7

Observation 4a3ee7e5-e8bb-49e5-81a3-7dfc0eccc039 · inbound

A Survey of Large Language Models cites this paper.

A Survey of Large Language Models PaLM-E: An Embodied Multimodal Language Model

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-05-10T22:46:40.806798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T22:46:39.268353Z digest=sha256:993073afb5b6095b3d307fb14f0210e795dd3e45c3ee5359a8ef0ed362ec9a16

Observation 8b7de1f7-e58e-4c01-af4f-9a3c173f2770 · inbound

Visual Instruction Tuning cites this paper.

Visual Instruction Tuning PaLM-E: An Embodied Multimodal Language Model

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-11T08:22:03.701283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T08:22:03.403362Z digest=sha256:68ea544534e937832f5cade5628446adede2fc70862eba913d852a59f6822ff9

Observation 9a4f7279-e83b-4ab4-986b-a61b6cf0eac1 · inbound

MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models cites this paper.

MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models PaLM-E: An Embodied Multimodal Language Model

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:29:30.136237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T20:37:01.617345Z digest=sha256:d0dec1bdbbaa0b80f8b9a033c4c90481ddfd1f381b5e63d974625860d7f8f6b5

Observation 2844c17a-77aa-4896-baf2-ddf06f5b48cf · inbound

LLM+P: Empowering Large Language Models with Optimal Planning Proficiency cites this paper.

LLM+P: Empowering Large Language Models with Optimal Planning Proficiency PaLM-E: An Embodied Multimodal Language Model

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-14T18:36:18.410581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T18:36:18.342052Z digest=sha256:4aff4e4c50e61f355aa3164cdb4f38b3666bed11b7bf85ce4f0c9a55c34568f1

Observation efad6fc9-822f-42e5-a7a5-99b56c91b13e · inbound

The Internal State of an LLM Knows When It's Lying cites this paper.

The Internal State of an LLM Knows When It's Lying PaLM-E: An Embodied Multimodal Language Model

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-16T00:08:36.592913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T00:08:36.530560Z digest=sha256:6bcc225e9352b5cade977161e74eabc6db55cea2e3f5c9ad30cd44c1e8a338a4

Observation 9e89c9d5-617c-40b2-8dd6-4055bd7caac9 · inbound

mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality cites this paper.

mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality PaLM-E: An Embodied Multimodal Language Model

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-24T09:04:15.068093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-24T09:02:31.211260Z digest=sha256:d25e64254deed4e90304daf670053e4040500c5420a8db6face21c7758a3b285

Observation e7edaf90-8665-4772-8823-72cc880c046a · inbound

LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model cites this paper.

LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model PaLM-E: An Embodied Multimodal Language Model

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-15T08:41:04.846297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T08:41:04.743886Z digest=sha256:a835d0290b6a9b1870bbe266ff730ae7655173a5837cd7b0f9499a96e17bd8b1

Observation f43a8c3e-77df-480f-a432-95bbaf898ee0 · inbound

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models cites this paper.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models PaLM-E: An Embodied Multimodal Language Model

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-17T09:55:35.587553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:f3a8473a9d342d7ffa8de49528ec66d2b8b9bb16273d13f1d4595c23f7f41273

Observation a61b9212-9f65-4561-84eb-0a62b84de972 · inbound

Voyager: An Open-Ended Embodied Agent with Large Language Models cites this paper.

Voyager: An Open-Ended Embodied Agent with Large Language Models PaLM-E: An Embodied Multimodal Language Model

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:29:30.136237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T13:11:40.995345Z digest=sha256:606b0303952e75c6363a9565f2431e3f82219a175d2ff71e80d563263ac3fb61

Observation 59b95322-c39d-4de3-93f7-441576791ead · inbound

Ghost in the Minecraft: Generally Capable Agents for Open-World Environments via Large Language Models with Text-based Knowledge and Memory cites this paper.

Ghost in the Minecraft: Generally Capable Agents for Open-World Environments via Large Language Models with Text-based Knowledge and Memory PaLM-E: An Embodied Multimodal Language Model

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-15T18:37:20.942500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T18:37:20.889942Z digest=sha256:446ad6d1a6a5e9f1a129b411b94595237c604dddc06bafb467c9edf739fdae3c

Observation 48e82d2d-e45e-47ff-959b-2efdcf5126df · inbound

ReWOO: Decoupling Reasoning from Observations for Efficient Augmented Language Models cites this paper.

ReWOO: Decoupling Reasoning from Observations for Efficient Augmented Language Models PaLM-E: An Embodied Multimodal Language Model

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-15T18:15:55.694905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T18:15:55.525596Z digest=sha256:114394d9a8bade2a10c92a081822c5c8de6c9256baf8c00fa1d4811c6d245263

Observation c385e7b0-7ee2-41d2-b470-94c63f61d571 · inbound

AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration cites this paper.

AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration PaLM-E: An Embodied Multimodal Language Model

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-05-24T08:29:11.372361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-24T08:27:35.798991Z digest=sha256:df843adc4a5e2627f8409d2198cf540b7702e2eef0ee769b376f4ae480e0363a

Observation 83f65583-5e0c-4be7-8785-73654365fcce · inbound

MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models cites this paper.

MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models PaLM-E: An Embodied Multimodal Language Model

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:29:30.136237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T20:25:33.854923Z digest=sha256:3db9feffeb02f2d9425578adf02c2021edbecc83208b8da73958fc5cb64a32d0

Observation b59ebb8b-5cb6-462e-afaf-277bde92bbc8 · inbound

A Survey on Multimodal Large Language Models cites this paper.

A Survey on Multimodal Large Language Models PaLM-E: An Embodied Multimodal Language Model

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-16T02:56:42.331731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T02:56:41.658658Z digest=sha256:37077719632ac2558af7348573e3aef28e33d1233632afbcd829c1665327cf74

Observation 8b5d5616-a38c-47be-88a7-e57303cb0b83 · inbound

Kosmos-2: Grounding Multimodal Large Language Models to the World cites this paper.

Kosmos-2: Grounding Multimodal Large Language Models to the World PaLM-E: An Embodied Multimodal Language Model

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-12T05:19:48.014398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T05:19:47.907355Z digest=sha256:35e0cd4102655d619901199fe48b459fd30d784aa8784fac4dd23f40206a0ef6

Observation 67047b39-235c-4e3a-b88e-505211a4a1c2 · inbound

Secrets of RLHF in Large Language Models Part I: PPO cites this paper.

Secrets of RLHF in Large Language Models Part I: PPO PaLM-E: An Embodied Multimodal Language Model

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-17T13:17:42.937746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:17:42.848578Z digest=sha256:be1d4dd92166e3af834346a359c878b531e430ddd350f46100798da12acc00f4

Observation ba69446c-6dd3-4e1a-bc10-8135129201cc · inbound

VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models cites this paper.

VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models PaLM-E: An Embodied Multimodal Language Model

Reference 68

Resolution
verified exact
local_arxiv, observed 2026-05-13T08:57:22.542159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T08:57:22.299028Z digest=sha256:8e7d33de3d7e2d34ecd88c94bcde60359a00966249c9b06bab375a3548e1bdf1

Observation f54abbe3-1a4b-47c5-a506-a21f0472364b · inbound

A Comprehensive Overview of Large Language Models cites this paper.

A Comprehensive Overview of Large Language Models PaLM-E: An Embodied Multimodal Language Model

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-19T20:28:39.081971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T20:28:38.900026Z digest=sha256:9e4393cd9e2149dd5fb673cea14d158f8fb5a5897930b3c91d3e3c3c74ffb272

Observation af5bd3bc-25e3-4141-971c-53f902e070d2 · inbound

OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models cites this paper.

OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models PaLM-E: An Embodied Multimodal Language Model

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-14T01:52:01.409552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T01:52:01.163900Z digest=sha256:1c7a78a63870c7169e4f585894726a89e9e775dd73013eb5521a3e0ef4d61058

Observation 2362ebfb-3aeb-48ed-ac34-cfa67633ce2d · inbound

DragNUWA: Fine-grained Control in Video Generation by Integrating Text, Image, and Trajectory cites this paper.

DragNUWA: Fine-grained Control in Video Generation by Integrating Text, Image, and Trajectory PaLM-E: An Embodied Multimodal Language Model

Reference 204

Resolution
verified exact
local_arxiv, observed 2026-05-20T13:03:57.952986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-20T13:03:57.828598Z digest=sha256:9485b7981513defcad417b940ed0d11ef8c5ea117e6a3ed4c337b11a30ba5f41

Observation 5e1b54e4-d41e-4a1f-a8f9-af89238c1717 · inbound

Cognitive Architectures for Language Agents cites this paper.

Cognitive Architectures for Language Agents PaLM-E: An Embodied Multimodal Language Model

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-16T19:33:44.317111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T19:33:44.146134Z digest=sha256:960078ecabcc4dfbafd01701bf1a712feebfea22d8edca18fb6195247bba5e53

Observation 40e88d1b-ac75-47ea-81db-2c4a6e036608 · inbound

Aligning Large Multimodal Models with Factually Augmented RLHF cites this paper.

Aligning Large Multimodal Models with Factually Augmented RLHF PaLM-E: An Embodied Multimodal Language Model

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-15T17:58:17.836952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:58:17.699042Z digest=sha256:d3bc5909e55dd65577139a5772be46b8c77ee481b31c4ba0c5604b3cc2c89441

Observation f0331151-8237-4e7d-9eca-407cad275eaa · inbound

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition cites this paper.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition PaLM-E: An Embodied Multimodal Language Model

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-17T13:48:48.724699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:4c1a4c676c6cb8615184caaef760238cb17c3afa397a14bfecf8510f1e880e3d

Observation 918f790b-79ea-458c-ab0b-98910dfd3274 · inbound

Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution cites this paper.

Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution PaLM-E: An Embodied Multimodal Language Model

Reference 94

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:12:31.442707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T08:12:30.984870Z digest=sha256:80f4d3d568d3e1b6b9932feeed1e871d087e41b2ae9f2362c55ff03b3a7709a0

Observation 9ab82333-e7c2-47d0-8aa8-daffb8a7dd87 · inbound

The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision) cites this paper.

The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision) PaLM-E: An Embodied Multimodal Language Model

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-15T23:26:06.272791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T23:26:06.183574Z digest=sha256:869afde72b876e62be95d131c9edb079761c289748aa7eb1f55cfc5e854396f8

Observation 9f6c2646-4c75-4abf-aa30-060dd56f3344 · inbound

Reinforcement Learning with Foundation Priors: Let the Embodied Agent Efficiently Learn on Its Own cites this paper.

Reinforcement Learning with Foundation Priors: Let the Embodied Agent Efficiently Learn on Its Own PaLM-E: An Embodied Multimodal Language Model

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-24T06:44:02.584833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-24T06:40:00.328012Z digest=sha256:40fc08914e6a06477e2dfa6cea4934088b4e0015c2f798743419dc978d12afcd

Observation 44dac76c-e983-42fc-85d3-767f562ab7d4 · inbound

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation cites this paper.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation PaLM-E: An Embodied Multimodal Language Model

Reference 200

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T20:06:44.706232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:991888d330dedd292de62655f772886c2d9c5ff1ef3b59a3538cf395a3ceb045

Observation eee8fc93-af35-490e-9edc-4edc446f0197 · inbound

Learning Interactive Real-World Simulators cites this paper.

Learning Interactive Real-World Simulators PaLM-E: An Embodied Multimodal Language Model

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-16T02:15:18.405183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T02:15:18.265190Z digest=sha256:20c97d4173d3f7a02895b245ea2c46aee2f893797d311cd7c48b5161781683a3

Observation cd92feae-8d32-4127-81e4-3287c6ac2597 · inbound

Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models cites this paper.

Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models PaLM-E: An Embodied Multimodal Language Model

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-16T05:54:59.162016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T05:54:58.940428Z digest=sha256:03e473658c8f3785d6e73bb7950f163a7f640dc27217f1062dc69219bf824d99

Observation cc147c8e-ef71-4204-9c81-5b2527ea353c · inbound

TD-MPC2: Scalable, Robust World Models for Continuous Control cites this paper.

TD-MPC2: Scalable, Robust World Models for Continuous Control PaLM-E: An Embodied Multimodal Language Model

Reference 105

Resolution
metadata mismatch
local_arxiv, observed 2026-05-14T17:27:35.982838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-14T17:27:35.733800Z digest=sha256:0f349e57561ca22e5ba2195c42d1fa270b9e04b3f230f0115853a7d90a473ed0

Observation 0fde5446-ce46-40d3-8080-6026e75ffaf5 · inbound

Vision-Language Foundation Models as Effective Robot Imitators cites this paper.

Vision-Language Foundation Models as Effective Robot Imitators PaLM-E: An Embodied Multimodal Language Model

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T21:44:27.655058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T21:44:27.562453Z digest=sha256:cff25b07ef81deec37568cef16f0357932b35d5241003e64d5be2dace10789c5

Observation 7d18a581-6c1c-4952-a4fc-95a30b3e9649 · inbound

Remember what you did so you know what to do next cites this paper.

Remember what you did so you know what to do next PaLM-E: An Embodied Multimodal Language Model

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-25T08:25:34.254592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-25T08:20:54.753641Z digest=sha256:9522e600187accf3067a3e77c91889d73acc422358f3f82f687615b93f82e96b

Observation adf8a310-159f-45f3-8dcb-db23fc9b56e4 · inbound

CogVLM: Visual Expert for Pretrained Language Models cites this paper.

CogVLM: Visual Expert for Pretrained Language Models PaLM-E: An Embodied Multimodal Language Model

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T15:46:06.522386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T15:46:06.334088Z digest=sha256:ebd4f729b7837932db936d009a29213284b2611d05339404bb7e601fddd65b08

Observation 31e8a08f-f84d-4c03-b1cf-cf0b0883dbbc · inbound

Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation cites this paper.

Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation PaLM-E: An Embodied Multimodal Language Model

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-13T16:32:05.702046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-13T16:32:05.538507Z digest=sha256:d96661331ba1b28517ca04febfae502ec37e43d2703254d0cce1eb97f0c994e6

Observation d2bb2690-eed2-4093-ae7c-c962639163d2 · inbound

VideoPoet: A Large Language Model for Zero-Shot Video Generation cites this paper.

VideoPoet: A Large Language Model for Zero-Shot Video Generation PaLM-E: An Embodied Multimodal Language Model

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T17:51:05.666225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:51:05.465548Z digest=sha256:ce7cff73e45c4d78778f1fb215fc3fb9bb8741b996e94a55dbc75c0f952a6346

Observation 88468e8d-7784-4899-9b32-48d9f7184fd7 · inbound

InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks cites this paper.

InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks PaLM-E: An Embodied Multimodal Language Model

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-13T22:46:09.876897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T22:46:09.693156Z digest=sha256:10bdb4c3648d9308bc4a480343d5933f06381e4287d5f19550098b74a91d3136

Observation 9042e4e2-01a4-45e9-aed2-683fc4b0a119 · inbound

Agent AI: Surveying the Horizons of Multimodal Interaction cites this paper.

Agent AI: Surveying the Horizons of Multimodal Interaction PaLM-E: An Embodied Multimodal Language Model

Reference 119

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T14:25:59.340568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-18T14:25:58.876978Z digest=sha256:8bee6627d5f94235434ae38cd6f9f113b943015fa9d6cee1ea3a228ed0bd5ba0

Observation 26584fc9-8f8d-4ebc-b1d5-9d7fb50f6475 · inbound

InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model cites this paper.

InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model PaLM-E: An Embodied Multimodal Language Model

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-17T05:30:27.719250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T05:30:27.667126Z digest=sha256:59d24383b1998a3a10de6268938a895ea35ed2e866da3e62e2d7e719ad21e0b1

Observation 1263e5be-8c9d-43e1-9995-f521a508bcc0 · inbound

NaVid: Video-based VLM Plans the Next Step for Vision-and-Language Navigation cites this paper.

NaVid: Video-based VLM Plans the Next Step for Vision-and-Language Navigation PaLM-E: An Embodied Multimodal Language Model

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-18T04:55:20.445697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T04:55:20.362512Z digest=sha256:571d79ce71aac42b5737e94660ff9b1258140eeff5cfc658100e70668dd3f651

Observation a03fb367-6ac8-4016-8ed1-8efcaae7ce79 · inbound

RT-H: Action Hierarchies Using Language cites this paper.

RT-H: Action Hierarchies Using Language PaLM-E: An Embodied Multimodal Language Model

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-17T06:53:27.735933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T06:53:27.642020Z digest=sha256:736e30c0f6f645eb360a92307ceaa0fe2585020690e431a328130bcfdbbf79c6

Observation 02ce5c85-46d4-4d83-ab83-ad289e7fc8c3 · inbound

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training cites this paper.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training PaLM-E: An Embodied Multimodal Language Model

Reference 28

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T04:09:36.320359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:11c78d9a23fa6996a3722d00ea491d02684697da182f45094c1f8f8b07aa00da

Observation 79465bc0-9128-4873-b79b-f6626afb40a9 · inbound

3D-VLA: A 3D Vision-Language-Action Generative World Model cites this paper.

3D-VLA: A 3D Vision-Language-Action Generative World Model PaLM-E: An Embodied Multimodal Language Model

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T18:18:27.245035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-13T18:18:27.211034Z digest=sha256:5fb508cf0068fddc11e2a198f7a5a4d4444b102e25fe810610b64ee26308d024

Observation cbf2d667-6f10-4697-9837-1669c1c76fc6 · inbound

Capabilities of Gemini Models in Medicine cites this paper.

Capabilities of Gemini Models in Medicine PaLM-E: An Embodied Multimodal Language Model

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-15T17:13:23.006080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-15T17:13:22.759147Z digest=sha256:24667b2f681ee12cd7d6b280ff0a750239954d90dd9c77e14eefc75396c9c007

Observation 3b6380f8-77d7-4507-b2c2-0912d0d8779c · inbound

The Platonic Representation Hypothesis cites this paper.

The Platonic Representation Hypothesis PaLM-E: An Embodied Multimodal Language Model

Reference 233

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T06:03:56.507115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-15T06:03:56.328012Z digest=sha256:e70066adce4f5fa5939262c7e5849a00e92a964ed066643df8ddb613a3dcecc6

Observation 922d9af2-9fe3-4c01-b43d-e3fc3250ffeb · inbound

Octo: An Open-Source Generalist Robot Policy cites this paper.

Octo: An Open-Source Generalist Robot Policy PaLM-E: An Embodied Multimodal Language Model

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-11T00:26:15.703429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T00:26:15.163358Z digest=sha256:a83daaec89595f4c1933d04bff6205b00a30b8cfba0332ff43b18df830f212e1

Observation 9470930f-2e22-40e9-ae66-9da1e7b4f620 · inbound

OpenVLA: An Open-Source Vision-Language-Action Model cites this paper.

OpenVLA: An Open-Source Vision-Language-Action Model PaLM-E: An Embodied Multimodal Language Model

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:29:30.136237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T14:46:35.942338Z digest=sha256:f719d846c8aca4a22d796082d836c5b832546ac1b0b7b5d2f56a858192eccec0

Observation 4e9a8100-73c2-4ba1-a044-33f7c5769e73 · inbound

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output cites this paper.

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output PaLM-E: An Embodied Multimodal Language Model

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-17T10:46:28.583353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T10:46:28.447347Z digest=sha256:8b81013d2a51f62c758165f05745737eb955599f79943cb668c48124faf6a792

Observation ea921a89-baf1-462c-9244-d255d3a4832d · inbound

LongVILA: Scaling Long-Context Visual Language Models for Long Videos cites this paper.

LongVILA: Scaling Long-Context Visual Language Models for Long Videos PaLM-E: An Embodied Multimodal Language Model

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:51:25.482472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T03:51:25.396887Z digest=sha256:8e9b2eff16a1372eb5d4759b78a73031591427ef13e40ca14eec997f7de59bdb

Observation 5113b5f0-e972-4b78-ba35-7313b2af6d6f · inbound

MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans? cites this paper.

MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans? PaLM-E: An Embodied Multimodal Language Model

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-16T07:59:32.841139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T07:59:32.638758Z digest=sha256:e15c675770974cab10c1a2b195f44f2813404a7d5ca9f2419592554c67b6a3f0

Observation be0fd2e6-b2ad-4a4a-91e6-cfa436c0366a · inbound

RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation cites this paper.

RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation PaLM-E: An Embodied Multimodal Language Model

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T07:46:30.666368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T07:46:30.553633Z digest=sha256:ac4ccc6710622ba4031632a6309528cea659ca50bf88303a47ed2544beb946a6

Observation 22071d76-cdc5-4a55-8986-348c5d881b6c · inbound

OS-ATLAS: A Foundation Action Model for Generalist GUI Agents cites this paper.

OS-ATLAS: A Foundation Action Model for Generalist GUI Agents PaLM-E: An Embodied Multimodal Language Model

Reference 59

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T09:29:27.343773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-13T09:29:27.173784Z digest=sha256:6214f611841477d995d25900452bef5a6e2afadca4a4db72c2c3599f2443b049

Observation 8408a01b-1c64-4b77-9d0d-1d97d06d5472 · inbound

EMMA: End-to-End Multimodal Model for Autonomous Driving cites this paper.

EMMA: End-to-End Multimodal Model for Autonomous Driving PaLM-E: An Embodied Multimodal Language Model

Reference 191

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T05:08:54.460154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-15T05:08:54.368109Z digest=sha256:086789fb71f9b6acbfdca8940afc6cdb1836d9c3b3ae2d0c827d390bb6a2d476

Observation b201e70e-56da-42ed-9ee4-7f5f261b8aea · inbound

$\pi_0$: A Vision-Language-Action Flow Model for General Robot Control cites this paper.

$\pi_0$: A Vision-Language-Action Flow Model for General Robot Control PaLM-E: An Embodied Multimodal Language Model

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:29:30.136237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T12:38:24.425784Z digest=sha256:5ab793fada3d1cd4363bff85ccbcee424462ea8ae0d4467e3363a30c1f4e4bb3

Observation 51625069-9ce4-4537-82c0-84b0ce89bc62 · inbound

VeriGraph: Scene Graphs for Execution Verifiable Robot Planning cites this paper.

VeriGraph: Scene Graphs for Execution Verifiable Robot Planning PaLM-E: An Embodied Multimodal Language Model

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-05-23T17:13:13.766619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T17:12:31.645346Z digest=sha256:18872fc49874c2ff7247880fefc5b68a4ee2cad1a9547f705da6f67f55e9b866

Observation e8b8ebcd-a10e-49d7-94e3-054448bdaab3 · inbound

SOWing Information: Cultivating Contextual Coherence with MLLMs in Image Generation cites this paper.

SOWing Information: Cultivating Contextual Coherence with MLLMs in Image Generation PaLM-E: An Embodied Multimodal Language Model

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-05-23T08:37:44.695709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T08:37:13.529654Z digest=sha256:dde01828f6753757e83be36f266e324ba37db283abaf8418c36d6d982abb2dac

Observation 160e633c-f8ef-4b2a-b949-64ab1338d218 · inbound

Motion-aware Contrastive Learning for Temporal Panoptic Scene Graph Generation cites this paper.

Motion-aware Contrastive Learning for Temporal Panoptic Scene Graph Generation PaLM-E: An Embodied Multimodal Language Model

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-05-23T07:25:28.521741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-23T07:23:51.435139Z digest=sha256:6aa84a80623911d27295016f3b108e25196885a282060a328463e44ad6dbe43d

Observation e1ae76c8-380a-4300-a750-309566723e12 · inbound

TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies cites this paper.

TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies PaLM-E: An Embodied Multimodal Language Model

Reference 68

Resolution
verified exact
local_arxiv, observed 2026-05-15T18:27:22.879453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-15T18:27:22.760982Z digest=sha256:fad8626f1225b7d08f81df08be588b36835c0030799016b284a63f2df782c0f9

Observation c63f7160-8da2-4290-9278-66676e7af971 · inbound

Modality-Inconsistent Continual Learning of Multimodal Large Language Models cites this paper.

Modality-Inconsistent Continual Learning of Multimodal Large Language Models PaLM-E: An Embodied Multimodal Language Model

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-23T06:52:40.192933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-23T06:50:12.315919Z digest=sha256:8510ce90b8004d9d5d2c891eb7acaa6182d848592be4482630b1428b99ce600b

Observation a54155e5-b5b8-4145-9dc9-ca23203faeab · inbound

FAST: Efficient Action Tokenization for Vision-Language-Action Models cites this paper.

FAST: Efficient Action Tokenization for Vision-Language-Action Models PaLM-E: An Embodied Multimodal Language Model

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-11T08:52:31.964936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:b48c9e6e218eccae50ce4d5de90b033d7cca164026445788e9442640760442c9

Observation aff7c028-1642-48aa-8402-795b5b24490f · inbound

WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs cites this paper.

WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs PaLM-E: An Embodied Multimodal Language Model

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-17T05:53:26.281814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T05:53:26.066674Z digest=sha256:6042852f50cfe5c6929c318cd0fbcaf25399049338fbae1317297de78267098e

Observation 2a3ad65f-e7ed-476b-b96b-4df071b5b615 · inbound

Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models cites this paper.

Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models PaLM-E: An Embodied Multimodal Language Model

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T22:53:37.344281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-15T22:53:37.120692Z digest=sha256:6c4b13ddf0a101ade3dcde76ea70e4f38ef785df620c770a0be8e0d5334bb0e7

Observation a8f50b03-d4bf-46d0-9fec-17c88edbaddf · inbound

HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model cites this paper.

HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model PaLM-E: An Embodied Multimodal Language Model

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-15T22:00:49.182851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T22:00:48.667428Z digest=sha256:fe84f1a4cce8274ce292462a6f3d60536be43ac012cffd17280d531a5ea58975

Observation a785531e-2b06-42e6-a472-203fb1be99aa · inbound

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots cites this paper.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots PaLM-E: An Embodied Multimodal Language Model

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:29:30.136237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:eb57b7e98df8d9460179be5f6276f8d6a904c533545c27c1d5773001f5fcf589

Observation d5e232ec-da20-4e3c-8254-d6b11cb46119 · inbound

CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models cites this paper.

CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models PaLM-E: An Embodied Multimodal Language Model

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-16T05:21:45.257773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T05:21:44.903048Z digest=sha256:0d43cedd1c9c5a32a271d591f0b9e139d32a0e2962b700a969dd79e290ad9277

Observation fc9d53e5-909f-454f-8b1a-a95d15c0e966 · inbound

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization cites this paper.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization PaLM-E: An Embodied Multimodal Language Model

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-22T18:05:00.999789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:da04c981c321e66f1b736253014cab9189186685309197f2ac10b4f77db7f70f

Observation ef462909-7f9e-4948-abf1-d7b69473c166 · inbound

NORA: A Small Open-Sourced Generalist Vision Language Action Model for Embodied Tasks cites this paper.

NORA: A Small Open-Sourced Generalist Vision Language Action Model for Embodied Tasks PaLM-E: An Embodied Multimodal Language Model

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T15:53:29.447866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T15:53:29.412890Z digest=sha256:82fdc233d29ec116c7efddd85b38238fd2f1ded9aae3e202834515c38710541f

Observation e0a948c7-efcf-45aa-bd55-25a2851e2e3b · inbound

GraspVLA: a Grasping Foundation Model Pre-trained on Billion-scale Synthetic Action Data cites this paper.

GraspVLA: a Grasping Foundation Model Pre-trained on Billion-scale Synthetic Action Data PaLM-E: An Embodied Multimodal Language Model

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:55:52.324344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T20:55:52.109166Z digest=sha256:b50e0e3c2894845be92fd8e868f2b3e9fc67fa85b316a6566746b2f2fd8828df

Observation 2e5f4096-288d-4977-8568-a16440b0a44a · inbound

Multi-SpatialMLLM: Multi-Frame Spatial Understanding with Multi-Modal Large Language Models cites this paper.

Multi-SpatialMLLM: Multi-Frame Spatial Understanding with Multi-Modal Large Language Models PaLM-E: An Embodied Multimodal Language Model

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-25T08:40:33.079258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-25T08:38:02.944230Z digest=sha256:62374ed420e47be7b15f818e427cf0ab6a27446166f5199d68719b99f7ffbe3e

Observation 14921338-3d23-4ef3-b191-c22200b59d40 · inbound

Sensorimotor Self-Recognition in Multimodal Large Language Model-Driven Robots cites this paper.

Sensorimotor Self-Recognition in Multimodal Large Language Model-Driven Robots PaLM-E: An Embodied Multimodal Language Model

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-19T13:32:19.259827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T13:29:34.151546Z digest=sha256:fcb05a03f46bf804c7b09145d623bd4ca65c78847f19c9957e1e381a6cbbda39

Observation 0f3717bb-0979-4d6a-b905-db5ae4c5c2e4 · inbound

From Standalone LLMs to Integrated Intelligence: A Survey of Compound Al Systems cites this paper.

From Standalone LLMs to Integrated Intelligence: A Survey of Compound Al Systems PaLM-E: An Embodied Multimodal Language Model

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-19T11:52:16.449091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T11:49:36.574471Z digest=sha256:1cab025e50af7f57eef4a88d318f18af39d012052132b0790771b907cfdf627d

Observation 084ee922-7067-481e-a22f-0448d7797dec · inbound

DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge cites this paper.

DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge PaLM-E: An Embodied Multimodal Language Model

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-16T15:42:41.427827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T15:42:41.363422Z digest=sha256:a7b9839830d96fe08f478d9829a66b135b46716911bee13ec860b8bc0744a4b9

Observation 2924e9c3-14e0-4bd0-bc68-52495b83f842 · inbound

A Careful Examination of Large Behavior Models for Multitask Dexterous Manipulation cites this paper.

A Careful Examination of Large Behavior Models for Multitask Dexterous Manipulation PaLM-E: An Embodied Multimodal Language Model

Reference 26

Resolution
metadata mismatch
local_arxiv, observed 2026-05-25T04:32:56.634193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-25T04:32:56.397350Z digest=sha256:5f13f8d9df6a31ee2839a3a5ed46daefb8975e20f31c7fb90e23344ced83ffc8

Observation 62ed886b-dfc4-4a56-a692-7e364c738bcd · inbound

GR-3 Technical Report cites this paper.

GR-3 Technical Report PaLM-E: An Embodied Multimodal Language Model

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-17T08:04:12.532207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T08:04:12.433863Z digest=sha256:39ca84d2aba2e5284564eeb7c15c02bae96b53e63c220266bd1c5bd4ec4e7ff4

Observation 681e7877-35cb-47b5-a6f1-9e5dd7c70425 · inbound

Training-Free Multimodal Large Language Model Orchestration cites this paper.

Training-Free Multimodal Large Language Model Orchestration PaLM-E: An Embodied Multimodal Language Model

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-19T00:12:54.110441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T00:12:39.834892Z digest=sha256:cd69d895c42163f3b93a345e85d6d9ac69020365bea2e493c3c527ebe157935d

Observation fef614d7-d952-4f64-9233-86413224dec1 · inbound

Training-Free Multimodal Large Language Model Orchestration cites this paper.

Training-Free Multimodal Large Language Model Orchestration PaLM-E: An Embodied Multimodal Language Model

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-25T08:05:30.703364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-25T08:02:15.950975Z digest=sha256:7d1242d056fa6d6432b6434bc5dacda6fc109dbd7ddf57244856038ad305dc49

Observation e1baf8b1-648a-4ad7-8800-a2793e56f74f · inbound

Unsupervised Hallucination Detection by Inspecting Reasoning Processes cites this paper.

Unsupervised Hallucination Detection by Inspecting Reasoning Processes PaLM-E: An Embodied Multimodal Language Model

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T18:25:37.332342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T18:25:37.332342Z digest=sha256:5bbefc188a9a2db13eb9b06daaa46d8ca32145fffed374e0c8a8d86d28ff58a7

Observation a4df57e5-3c09-46bf-8fae-56d98aa526de · inbound

A Vision Toward Energy-Efficient Domain-Specific Artificial Intelligence Models and Agents cites this paper.

A Vision Toward Energy-Efficient Domain-Specific Artificial Intelligence Models and Agents PaLM-E: An Embodied Multimodal Language Model

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T08:12:10.058079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:12:10.058079Z digest=sha256:234925da8544a782ede6604125023876b3041de37ce2a13b10a2b298b8dd404a

Observation a1aaa130-4a6f-4434-be1a-76188a9e8cde · inbound

A Collaborative Reasoning Framework for Anomaly Diagnostics in Underwater Robotics cites this paper.

A Collaborative Reasoning Framework for Anomaly Diagnostics in Underwater Robotics PaLM-E: An Embodied Multimodal Language Model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T00:03:22.907790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:03:22.907790Z digest=sha256:97e248de51ac208c69d6f3178500ad11c28594176c49fd395db6e73595e41053

Observation dba930d4-e738-403e-9386-61288e5c2f41 · inbound

AsyncVLA: Asynchronous Flow Matching for Vision-Language-Action Models cites this paper.

AsyncVLA: Asynchronous Flow Matching for Vision-Language-Action Models PaLM-E: An Embodied Multimodal Language Model

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-17T21:30:18.269867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T21:28:18.630934Z digest=sha256:b036d5a52bc7c78a5fcb734e9304d0f5893265130f06c0ef711843170deaa71b

Observation 4ecf3f31-478a-4501-b09e-5d764a27f447 · inbound

SPEAR-1: Scaling Beyond Robot Demonstrations via 3D Understanding cites this paper.

SPEAR-1: Scaling Beyond Robot Demonstrations via 3D Understanding PaLM-E: An Embodied Multimodal Language Model

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:22:04.617134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T20:21:12.375936Z digest=sha256:789fa12c551e54f96849ddb9be9b773e1073699e90a2d026d40bb5d83e5b30cb

Observation 505b478b-7ba9-40dd-ae48-5b4240396403 · inbound

Large Video Planner Enables Generalizable Robot Control cites this paper.

Large Video Planner Enables Generalizable Robot Control PaLM-E: An Embodied Multimodal Language Model

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-16T21:28:33.993855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T21:26:32.048309Z digest=sha256:cb6aaf6102990bf79463fceebb9abb8901c43059e1fdc2ea5f38c778951beaaa

Observation 97332b31-f0c2-4758-9b7f-837ed0c264a2 · inbound

PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic Manipulation cites this paper.

PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic Manipulation PaLM-E: An Embodied Multimodal Language Model

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-16T15:08:01.718261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T15:05:21.907878Z digest=sha256:c022013621827f2d7dab813250cf296cd95aca8b123857a446d1fe1dde1bbedc

Observation 96cde638-0e7c-4008-8640-8864aa6f6ee1 · inbound

Real2Sim via Active Perception with Behavior Trees Automatically Generated by VLMs cites this paper.

Real2Sim via Active Perception with Behavior Trees Automatically Generated by VLMs PaLM-E: An Embodied Multimodal Language Model

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-21T15:10:16.665533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T15:07:54.930273Z digest=sha256:cd7c0a5b2da811e5c63011ca1a7bb608db47e3186dff9004e16877e866e8e2f6

Observation 42483006-231d-482f-a6f1-8f5d8d567429 · inbound

ST-BiBench: Benchmarking Multi-Stream Multimodal Coordination in Bimanual Embodied Tasks for MLLMs cites this paper.

ST-BiBench: Benchmarking Multi-Stream Multimodal Coordination in Bimanual Embodied Tasks for MLLMs PaLM-E: An Embodied Multimodal Language Model

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-16T06:00:40.640861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T06:00:02.043029Z digest=sha256:b7eab16f2eb60dbec07f3e0075ac1c0dd325e874e4aa3dead79d04bcd5d51ad7

Observation b5e77811-fd14-4ebd-b11c-f058f01278b2 · inbound

Robust Policy Optimization to Prevent Catastrophic Forgetting cites this paper.

Robust Policy Optimization to Prevent Catastrophic Forgetting PaLM-E: An Embodied Multimodal Language Model

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-16T05:37:24.325187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T05:33:42.965249Z digest=sha256:bc5e42d775383325f43d8e8c21a5a6af61a92e7588ed145a3d0ce05d344aa418

Observation 415cdc25-37c9-4baa-85be-fdaec00b3489 · inbound

World Action Models are Zero-shot Policies cites this paper.

World Action Models are Zero-shot Policies PaLM-E: An Embodied Multimodal Language Model

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-11T16:18:15.891584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T16:18:15.003371Z digest=sha256:00ff73e62cf6801635103589f6569a420de49323ccc757238075ab64ad7a80c5

Observation 35e096b6-a059-4961-9b99-0070b986c192 · inbound

UniLACT: Depth-Aware RGB Latent Action Learning for Vision-Language-Action Models cites this paper.

UniLACT: Depth-Aware RGB Latent Action Learning for Vision-Language-Action Models PaLM-E: An Embodied Multimodal Language Model

Reference 27

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T20:20:17.549476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T20:18:31.988002Z digest=sha256:ff4658c5eedcac8cc7ec4a97b9ea9197fb7d81614b185d23663137b981f5e64d

Observation 22460ac1-2ede-49e5-a9f0-d060d797912f · inbound

PhysMem: Scaling Test-Time Memory for Embodied Physical Reasoning cites this paper.

PhysMem: Scaling Test-Time Memory for Embodied Physical Reasoning PaLM-E: An Embodied Multimodal Language Model

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-15T20:06:33.859979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T20:05:30.324309Z digest=sha256:7328dd93cc8fd5d403dc22247b29cf286ead4fd3150e866831f91f17cf932a42

Observation 2ed83f21-211d-4ce4-be14-f74dfadc1992 · inbound

Learning from Trials and Errors: Reflective Test-Time Planning for Embodied LLMs cites this paper.

Learning from Trials and Errors: Reflective Test-Time Planning for Embodied LLMs PaLM-E: An Embodied Multimodal Language Model

Reference 1977

Resolution
unresolved
no resolver link, observed 2026-08-02T21:10:58.087934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:10:58.087934Z digest=sha256:e40b3aa9265c98ffba27e1f8165a5cd42d1d32912e6e35221b474b90ba643f44

Observation 4dd14148-f670-4c3e-9c2e-7dec83809345 · inbound

ReCoN-Ipsundrum: An Inspectable Recurrent Persistence Loop Agent with Affect-Coupled Control and Mechanism-Linked Consciousness Indicator Assays cites this paper.

ReCoN-Ipsundrum: An Inspectable Recurrent Persistence Loop Agent with Affect-Coupled Control and Mechanism-Linked Consciousness Indicator Assays PaLM-E: An Embodied Multimodal Language Model

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-02T20:29:41.309509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:29:41.309509Z digest=sha256:96271defe653b537600cc06188661f6726d23b04c6471e1f88a59f9ed95cda61

Observation 33caab44-b41a-4789-a319-954c6c0c8a1b · inbound

Mema: Memory-Augmented Adapter for Enhanced Vision-Language Understanding cites this paper.

Mema: Memory-Augmented Adapter for Enhanced Vision-Language Understanding PaLM-E: An Embodied Multimodal Language Model

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-15T17:56:25.088912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:53:29.860496Z digest=sha256:480034751605b72117318d632308292ec1f0f5b93e0e01884eb16ad08a566a58

Observation 000115f5-a140-47c1-affb-70c41af7f58c · inbound

Vision Language Models Cannot Reason About Physical Transformation cites this paper.

Vision Language Models Cannot Reason About Physical Transformation PaLM-E: An Embodied Multimodal Language Model

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-15T13:27:51.848177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T13:27:51.848177Z digest=sha256:1172d80dd5832ff5d20144297febf60b0a165805bbd7557caa650de150a8ab68

Observation d982ae49-7132-4484-9f3f-0c3dc35cb62f · inbound

AR-VLA: True Autoregressive Action Expert for Vision-Language-Action Models cites this paper.

AR-VLA: True Autoregressive Action Expert for Vision-Language-Action Models PaLM-E: An Embodied Multimodal Language Model

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-15T12:50:37.156650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T12:50:02.670780Z digest=sha256:b17314dbd453379dabf39642cc5ceebe45c96c8cd2255a28887863da1c2f8812

Observation 861dd294-8eba-463c-a0ed-3e00aef39b82 · inbound

Tactile Modality Fusion for Vision-Language-Action Models cites this paper.

Tactile Modality Fusion for Vision-Language-Action Models PaLM-E: An Embodied Multimodal Language Model

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T18:13:07.837052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:13:07.837052Z digest=sha256:3207b18de1f2853d65b840f09352b0c3231596228254b17500c965544eb0172f

Observation 62100a11-86c7-447f-aacb-8e84e3bed7e2 · inbound

LongTail Driving Scenarios with Reasoning Traces: The KITScenes LongTail Dataset cites this paper.

LongTail Driving Scenarios with Reasoning Traces: The KITScenes LongTail Dataset PaLM-E: An Embodied Multimodal Language Model

Reference 19

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T00:33:23.342498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T00:30:56.156113Z digest=sha256:fa28152fe0c67119bab5343c591fbddfa5a61bfb5a06b978d5903e24f1b62754

Observation 25e4232e-34be-4b18-a4c8-de634a0d41ef · inbound

CoME-VL: Scaling Complementary Multi-Encoder Vision-Language Learning cites this paper.

CoME-VL: Scaling Complementary Multi-Encoder Vision-Language Learning PaLM-E: An Embodied Multimodal Language Model

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-13T20:33:17.134640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T20:28:30.864143Z digest=sha256:753743fb46ac6f4a0d2198afe557202f766bf5d3fcebacd68b1018b51cb0cd09

Observation 82a924bc-291f-43bb-aaa9-f2f969001b82 · inbound

ROSClaw: A Hierarchical Semantic-Physical Framework for Heterogeneous Multi-Agent Collaboration cites this paper.

ROSClaw: A Hierarchical Semantic-Physical Framework for Heterogeneous Multi-Agent Collaboration PaLM-E: An Embodied Multimodal Language Model

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-10T22:55:48.696691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T19:30:30.463475Z digest=sha256:18388da286fd752663ca0c9d975d6162c6c36878dd52f0c5ea85d20662f05ea5

Observation 2f5b8bcc-fcbe-4302-ae9f-4450cc79759e · inbound

CoEnv: Driving Embodied Multi-Agent Collaboration via Compositional Environment cites this paper.

CoEnv: Driving Embodied Multi-Agent Collaboration via Compositional Environment PaLM-E: An Embodied Multimodal Language Model

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-05-10T22:55:52.463513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T19:27:12.286456Z digest=sha256:d294129096e95cff0da921a13e8351c174d986293396e544f07cdbef005cc05c