Pith. sign in

Paper Citation Record · LEDGER

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents

As of 12 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 0 inbound Pith citation observations for arXiv:2412.10410.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.10410 v1

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T20:41:45.249886Z

measured 65 of 65 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

65 of 65 outbound references displayed

  • verified exact1
  • verified fuzzy18
  • unresolved46
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f50f39f2-e052-4655-b5da-a369fe0d3d19 · outbound

This paper cites write newline.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:44.943729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:44.943729Z digest=sha256:e64e7b9931f16077a393c03b79f4619ebc50c490a216ad57d8828f88343d3dda

Observation 28369869-51de-4e3f-b658-9b9c64591727 · outbound

This paper cites Multi-objective latent space optimization of generative molecular design models.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Multi-objective latent space optimization of generative molecular design models

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:46.001240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T20:41:44.949822Z digest=sha256:be6ca8c2330f693bc5325c71a5ac87470f8a3bbb7c025a7821e31b5048631102

Observation 14e1e029-8980-4bd3-a986-8f2bf97615fc · outbound

This paper cites Imitating Interactive Intelligence.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Imitating Interactive Intelligence

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:44.960832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:44.960832Z digest=sha256:50efa8b71f1b1f934a3c0b41dbe07680f933b0b5f214b7b8c3e01a68909eac9e

Observation 218ff899-930d-4e8f-a422-d73ca4dcb064 · outbound

This paper cites An optimistic perspective on offline reinforcement learning.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents An optimistic perspective on offline reinforcement learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:44.965648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:44.965648Z digest=sha256:75c97177c8f0b011ded891679b4bd801de62d4622d55653f023fb556885e8e41

Observation e7a7730c-45b8-4919-b357-713eb95f3b6b · outbound

This paper cites OPAL: Offline Primitive Discovery for Accelerating Offline Reinforcement Learning.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents OPAL: Offline Primitive Discovery for Accelerating Offline Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:44.970264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:44.970264Z digest=sha256:56579fd377151ed38e60d8d694f1f79e6c0365dcfb50c66f2a6292715b988a2c

Observation 5edea55d-0b14-4aca-a234-78064cce8fe4 · outbound

This paper cites Flamingo: a Visual Language Model for Few-Shot Learning.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Flamingo: a Visual Language Model for Few-Shot Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:44.976022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:44.976022Z digest=sha256:5b607efd7c2aa5d0b4867515e807c9b6fb7a5496ca153398d32b0c385fc2b7d7

Observation 8744f31c-5e5a-4fa7-9fa7-7aaeed675001 · outbound

This paper cites Fixing a broken elbo.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Fixing a broken elbo

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:44.981536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:44.981536Z digest=sha256:9002e52b33e705741511d23d1d024c6360862e9ccac0214604fe2c582ab86e01

Observation 69ca765b-5454-4b5d-9ab9-0dba13e472d5 · outbound

This paper cites Hindsight Experience Replay.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Hindsight Experience Replay

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:44.985917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:44.985917Z digest=sha256:5e78f1cebf86e258eede34d1d6b574c6a98ca5defef57ee95a5103c9fbc77e2c

Observation 7c984fe0-9b1a-41e1-9940-fc603d54ba1f · outbound

This paper cites Agent57: Outperforming the atari human benchmark.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Agent57: Outperforming the atari human benchmark

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:45.972738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T20:41:44.990370Z digest=sha256:8f07c5b8d1fc089d26978ba5795a643fae883884c9857d2bfadb9a5870519e20

Observation 06676397-6534-4398-8f3b-c47ad9508a58 · outbound

This paper cites Video PreTraining (VPT): Learning to Act by Watching Unlabeled Online Videos.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Video PreTraining (VPT): Learning to Act by Watching Unlabeled Online Videos

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:44.993876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:44.993876Z digest=sha256:7014271202e7cc78cc71e1398301baf1854e46d33902631d0c937612e0e92996

Observation 500ce3f9-1759-4886-9837-0aaeff9386bb · outbound

This paper cites The arcade learning environment: An evaluation platform for general agents.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents The arcade learning environment: An evaluation platform for general agents

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:44.999032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:44.999032Z digest=sha256:f5cff42327458e1439785ea6b8bf872a7d0a69ada056d32e1f37c0bf4c73eb5c

Observation 2cf34361-3637-4dd8-8133-e34fa7774007 · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents RT-1: Robotics Transformer for Real-World Control at Scale

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.003192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.003192Z digest=sha256:4bf7320706efcce73c2076538bb132e8e6276cdf5f1191a5a0a73a1633492594

Observation 1732c799-21f6-4b38-b370-1f9ef5864090 · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.013831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.013831Z digest=sha256:5e26fbc7c8495bda6f96c7af91d146e9274d8e470fa9930351cd5b2452278b0c

Observation 8da9e833-c68d-4395-b509-f4bd10d0700b · outbound

This paper cites Language Models are Few-Shot Learners.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Language Models are Few-Shot Learners

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.021602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.021602Z digest=sha256:39915d206cb0b2031c822bee73e00b499bd87ec6013c6203647e8457ef95d356

Observation 70444ac8-ee54-4e16-946a-323d725855cd · outbound

This paper cites Open-world multi-task control through goal-aware representation learning and adaptive horizon prediction.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Open-world multi-task control through goal-aware representation learning and adaptive horizon prediction

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:45.950291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T20:41:45.026106Z digest=sha256:b8ce34be41ad5d3ae501462a7c56059ffdecec0f8ca34f89bb273520d8f02707

Observation f24ea4a2-6aa5-4c12-90fb-e4beb55d40b1 · outbound

This paper cites Groot: Learning to follow instructions by watching gameplay videos.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Groot: Learning to follow instructions by watching gameplay videos

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:45.937746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T20:41:45.030662Z digest=sha256:9e6f1f2b02fa369d6dbe2ba4208e8691d19032abd77780e14b0fbfb937d762ae

Observation 83b80df0-fc92-467f-b030-94ed50058ad0 · outbound

This paper cites Abbeel, A.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Abbeel, A

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:45.923865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T20:41:45.034180Z digest=sha256:4731ce5778fb4bc3fdc653a8a934a3717a156666d49bf847d36dc04b49325445

Observation 18ad4ab7-8f55-4320-bc0c-aa46bd7932f0 · outbound

This paper cites Transformer-xl: Attentive language models beyond a fixed-length context.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Transformer-xl: Attentive language models beyond a fixed-length context

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.042834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.042834Z digest=sha256:189c65fb3355a57fd567160dceaea4062352b1cd55660193cae2a6fed27745eb

Observation 4f7cf927-0b2e-4ad3-abc6-6678d4951c82 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.046948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.046948Z digest=sha256:ea50f0dd79b92b452643adf177badec83382f77669e7765e69c1c3482354fff4

Observation bc4d0dcc-4b45-404c-bb45-3bccd86d4211 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.049894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.049894Z digest=sha256:d90b19484a4017e06386dd44f076188327f6bc77633742ea13e5caf834331b4c

Observation 2c6542bb-2217-4c5d-8b02-0beca396f1c0 · outbound

This paper cites One-shot imitation learning.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents One-shot imitation learning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:45.911107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T20:41:45.053932Z digest=sha256:528ee4b6fbbefd2360ab7cabb975451be294606149b7b878b4d138efb8289658

Observation d26f0261-4177-45a8-8f12-8f61498869d1 · outbound

This paper cites Implicit Deep Latent Variable Models for Text Generation.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Implicit Deep Latent Variable Models for Text Generation

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-11T20:41:45.530666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T20:41:45.057282Z digest=sha256:9ad2679e8578e12bc41d8213570fcbf5891e5057e86c696a43e71cccf702e374

Observation 293ec0b5-6d0c-43aa-8f62-93c0d783b7d5 · outbound

This paper cites Relay Policy Learning: Solving Long-Horizon Tasks via Imitation and Reinforcement Learning.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Relay Policy Learning: Solving Long-Horizon Tasks via Imitation and Reinforcement Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.061059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.061059Z digest=sha256:db31ce0aa36bf3f5dab08a32f82c1efd2b989bdb64287a29c22881dca5b595a0

Observation 13f6abe7-be3c-4463-8429-2a2b75e76ab5 · outbound

This paper cites Guss, Brandon Houghton, Nicholay Topin, Phillip Wang, Cayden Codel, Manuela M.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Guss, Brandon Houghton, Nicholay Topin, Phillip Wang, Cayden Codel, Manuela M

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.064991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.064991Z digest=sha256:7fefe96bde4b8713ffde726c4a9bb8197949584605a335dff0640af59696187b

Observation 02419236-e53b-41a7-8067-4f6d8b769539 · outbound

This paper cites Vln-bert: A recurrent vision-and-language bert for navigation.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Vln-bert: A recurrent vision-and-language bert for navigation

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:45.891844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T20:41:45.068745Z digest=sha256:77bf845a8255117adb7edaee72badc05506ec8060e15f19af31a372bb3c660e7

Observation 14dc7441-235e-4a76-a095-56bdd444ded0 · outbound

This paper cites An Embodied Generalist Agent in 3D World.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents An Embodied Generalist Agent in 3D World

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.074200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.074200Z digest=sha256:d21a6d20ec5843348514df91aed78feccfb12fde79970de0e4c4be372faebff0

Observation 5fafca65-a57c-47b7-ab54-a0b2091c12a5 · outbound

This paper cites Vid2Robot: End-to-end Video-conditioned Policy Learning with Cross-Attention Transformers.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Vid2Robot: End-to-end Video-conditioned Policy Learning with Cross-Attention Transformers

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.078763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.078763Z digest=sha256:add4e7591845db1fb3019bd74fd13931a7be038f729d9e6a5c9935c51d3bee02

Observation 61b1740d-0afc-4bfd-9177-f8d85f0c2d28 · outbound

This paper cites BC-Z: Zero-Shot Task Generalization with Robotic Imitation Learning.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents BC-Z: Zero-Shot Task Generalization with Robotic Imitation Learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.083397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.083397Z digest=sha256:dfef86f1aea5b4850ea58a8c4954c313ffb18599669bec70fa28725cdbaf4c3d

Observation 0352c213-bfdc-4930-bc33-46e87ae24b0f · outbound

This paper cites The malmo platform for artificial intelligence experimentation.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents The malmo platform for artificial intelligence experimentation

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:45.878744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T20:41:45.088113Z digest=sha256:ebcd003d97cddc103db1179222f6493ed657126cad6b56468f5d5d2ef56185d7

Observation 1b0a3218-9245-42d3-960d-a904105392fc · outbound

This paper cites Auto-Encoding Variational Bayes.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Auto-Encoding Variational Bayes

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.092438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.092438Z digest=sha256:f97472c90878caae333bef501c0822d32befdcd34194d92d88a3a9fe8ba07779

Observation 3f602080-664e-463e-a314-8bd7086f9af2 · outbound

This paper cites Multi-game decision transformers.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Multi-game decision transformers

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.096021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.096021Z digest=sha256:a4d2a586a4fc1a5b658291baa948d06ff01d0d0fbcb6be1f12274f2153723aee

Observation 3daa7d74-c326-499c-82d9-52b066f61ea0 · outbound

This paper cites Evaluating Real-World Robot Manipulation Policies in Simulation.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Evaluating Real-World Robot Manipulation Policies in Simulation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.100715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.100715Z digest=sha256:e6a710f0f49d04e01ddce38d1143ae3327674a20e1bc739d452a7946b99cbc22

Observation 158adb09-896f-45fb-9640-3dd60c27624c · outbound

This paper cites STEVE-1: A Generative Model for Text-to-Behavior in Minecraft.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents STEVE-1: A Generative Model for Text-to-Behavior in Minecraft

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.105155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.105155Z digest=sha256:96049c310a01fbd678c2c43cc218d1ebb712e1d288583874b0babd192d719f51

Observation 4bb96c3a-44ef-4fd2-8a33-8135c6256bc9 · outbound

This paper cites Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.116529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.116529Z digest=sha256:c69f440784c25cf637dcbf027b103d5d74a5a81f0a4230234d77bf1254120a2b

Observation 151077ff-6bb6-4603-adb4-be779ff5523e · outbound

This paper cites Language Conditioned Imitation Learning over Unstructured Data.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Language Conditioned Imitation Learning over Unstructured Data

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.120926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.120926Z digest=sha256:2310df7d9e46786ae4e2a75189b31b3143f1533dc9d64a3b8cd1884195c8bca2

Observation 8b7ab2c4-e7d8-40eb-acd9-3de1f078aaf5 · outbound

This paper cites Learning latent plans from play.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Learning latent plans from play

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:45.854733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T20:41:45.124914Z digest=sha256:92d0faa12f46f9cdd303f21690cc9554db4d58b4a0ff2d07719d67af13e5cb9c

Observation 8f0b4f59-1616-491c-9527-c1e1dc268d6d · outbound

This paper cites Learning latent plans from play.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Learning latent plans from play

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:45.841767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T20:41:45.128533Z digest=sha256:d69d539ffc5357b9659a60bfeb4b15b0b04a86723ab881de1691944aeab995ec

Observation 85dae5c4-4642-4d51-9bf8-e5e2390ec52a · outbound

This paper cites Learning latent plans from play.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Learning latent plans from play

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:45.827698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T20:41:45.132613Z digest=sha256:76187782860b43b24b91dbba3d2a97f2a6b23b17af968b3352459a84dff454f7

Observation 53c39c27-ffb2-4a5c-a38b-200f01ec7b37 · outbound

This paper cites Interactive language: Talking to robots in real time.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Interactive language: Talking to robots in real time

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.136889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.136889Z digest=sha256:5a95be3adad3a2ba19be81da0a0a529b13b04427c33fb325ca367c27b36f3b78

Observation 1c85f61f-d6cc-400c-9029-c29b0d88694e · outbound

This paper cites ZSON: Zero-Shot Object-Goal Navigation using Multimodal Goal Embeddings.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents ZSON: Zero-Shot Object-Goal Navigation using Multimodal Goal Embeddings

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.140755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.140755Z digest=sha256:0830886dbe2bc832663f9b43fa97f58f8c33492458f24c15820dc23deff35a29

Observation ee0692f7-bc58-4de7-9fd4-0870847b5b14 · outbound

This paper cites What matters in language conditioned robotic imitation learning over unstructured data.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents What matters in language conditioned robotic imitation learning over unstructured data

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:45.807648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T20:41:45.145567Z digest=sha256:0130aadaa1cc44d0a803ecdd0355faacdbea3883dc3e25fe7b9a4b140be41a1d

Observation f5b071b7-f4f6-432b-978a-2fe318553e90 · outbound

This paper cites Rusu, Joel Veness, Marc G.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Rusu, Joel Veness, Marc G

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.149036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.149036Z digest=sha256:4891a03171a3be451505f783f85080a3a9380df25fce7a89b28cd3db8efc8b0d

Observation 58d8e0f2-56fc-4a86-b104-0c53e31d9b77 · outbound

This paper cites Goal representations for instruction following: A semi-supervised language interface to control.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Goal representations for instruction following: A semi-supervised language interface to control

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:45.787144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T20:41:45.152603Z digest=sha256:f7e2ca54600e46d4194e5dc79e9f9f464e00d08cbdf1c131a127c03b00a07860

Observation f89879f9-987a-4a0c-99f1-cc009537a33d · outbound

This paper cites Octo: An open-source generalist robot policy.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Octo: An open-source generalist robot policy

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.156190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.156190Z digest=sha256:a0e9afef2dc473404374c583a75c92685e3c40172b11f6b4fe6aaeb24ac5eb91

Observation 152c4e05-a8ed-4815-983d-9f22f37cce05 · outbound

This paper cites Open X-Embodiment: Robotic Learning Datasets and RT-X Models.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Open X-Embodiment: Robotic Learning Datasets and RT-X Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.159279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.159279Z digest=sha256:ccbf2f403f12cca028abf662f37111a3cbd17a81bef398a4abda1884b2c7dbb5

Observation 0bb8bfe7-c9e8-4958-aef8-a31f933bc552 · outbound

This paper cites Conditional Variational Autoencoder for Neural Machine Translation.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Conditional Variational Autoencoder for Neural Machine Translation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.162510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.162510Z digest=sha256:6bd73928f5f9a70bc7a3d56da97238b62448fbe7b1b16b4970c655bc828309b8

Observation 84e935d3-5426-47a3-b218-c758916988de · outbound

This paper cites Episodic transformer for vision-and-language navigation.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Episodic transformer for vision-and-language navigation

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:45.766547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T20:41:45.166517Z digest=sha256:798fa9efe77bf4434545c50515dc307698145e323c4bb2bfccb9d09448d5df83

Observation 87dd3bdf-20cf-41ee-badf-f7e0df4ffe25 · outbound

This paper cites Accelerating reinforcement learning with learned skill priors.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Accelerating reinforcement learning with learned skill priors

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:45.749383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T20:41:45.169963Z digest=sha256:f5bf0e0057ae675839a534e10bdb2b059c3fa3e0264b8fb7eb1546d1a3a78dae

Observation b0412b2c-a674-4c5b-812e-05dd181154cb · outbound

This paper cites Scaling Instructable Agents Across Many Simulated Worlds.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Scaling Instructable Agents Across Many Simulated Worlds

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.173544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.173544Z digest=sha256:a444deac5e489567464dab900d805e8f8d7575f1fd4b4687c86d13890ded1c21

Observation ea1ff3cd-f5d9-45a3-b93d-d96a13eb9d6c · outbound

This paper cites Improving language understanding by generative pre-training.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Improving language understanding by generative pre-training

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.177377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.177377Z digest=sha256:05856294e844782d7b16e571918fc672f63b2ca4728a407ca7bc54c313bb5b83

Observation 35cf3de0-21f3-4e86-848e-2cd6bc0f9f24 · outbound

This paper cites Language models are unsupervised multitask learners.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Language models are unsupervised multitask learners

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.181411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.181411Z digest=sha256:2c2151dd12175578ee76c16c0d5e0e6b442c60a8c7d319282cbbb601c7c4dd3e

Observation 4c021881-6925-4123-a7b0-f3c5c08f3fee · outbound

This paper cites Learning transferable visual models from natural language supervision.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Learning transferable visual models from natural language supervision

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.186191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.186191Z digest=sha256:819d726fbedd5da254ba97cdd5762a94dbced6bae3de29dfa98cee83bf10f8e4

Observation 407263d1-4278-4042-b2b2-63265791bec9 · outbound

This paper cites A Generalist Agent.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents A Generalist Agent

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.191304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.191304Z digest=sha256:f05eab0f77cab5da0b443688d7c8dc813839dcdd4b37a801273c884b453f72b0

Observation e417dc5c-f1c3-4089-a3b8-ec46ce4657ab · outbound

This paper cites Habitat: A platform for embodied ai research.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Habitat: A platform for embodied ai research

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.195401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.195401Z digest=sha256:1697492d8c398869138d54c3954664bae0419aa49bc65602a4ba7057a4433b30

Observation c31d09c1-597f-45b4-a07a-4debc3ede1aa · outbound

This paper cites Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.198736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.198736Z digest=sha256:48a50c098c9dc3cfc98c82cd8e15f1ca7e813c4fc6e506b2818b41256ff12d53

Observation f5f0f9ab-0452-4a05-859e-96f9a2094da0 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Gemini: A Family of Highly Capable Multimodal Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.203599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.203599Z digest=sha256:abd228f9d933c6be070a83959c2e665a7a944745a64841c9b1c9d383fb986f71

Observation 3b518a75-3d8d-4552-bceb-8f71998c5f2b · outbound

This paper cites Describe, explain, plan and select: interactive planning with llms enables open-world multi-task agents.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Describe, explain, plan and select: interactive planning with llms enables open-world multi-task agents

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:45.706557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T20:41:45.213308Z digest=sha256:c4d5ec3f31eecde574b0149affdc96eaf294c5612c1cf5c3261705fc4dfbc7dc

Observation e80437eb-59ba-462c-83a0-4a437aeed3ba · outbound

This paper cites JARVIS-1: Open-World Multi-task Agents with Memory-Augmented Multimodal Language Models.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents JARVIS-1: Open-World Multi-task Agents with Memory-Augmented Multimodal Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.217706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.217706Z digest=sha256:5549d41cc228a8e6e8b982af82bd403b384023e36250bf255760c002dcf7dfab

Observation 3b12a3f0-374d-4c8a-91ef-fbd89547d5b3 · outbound

This paper cites Xskill: Cross embodiment skill discovery.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Xskill: Cross embodiment skill discovery

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:45.695643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T20:41:45.223484Z digest=sha256:45ad64fdbc8d10a3e69845aa465d7eb5b4519da72637380ec58a1beac969fbb5

Observation cdb9dadc-fcbe-496d-b6d0-89fe6d0c414c · outbound

This paper cites Meta-World: A Benchmark and Evaluation for Multi-Task and Meta Reinforcement Learning.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Meta-World: A Benchmark and Evaluation for Multi-Task and Meta Reinforcement Learning

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.228028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.228028Z digest=sha256:f143b57f478ddcfc1275a7542608ee4f7fef47be6e6ff8214692e8c2a0663f37

Observation a5e56cbd-a5b0-421b-bd47-db18e53ca0e6 · outbound

This paper cites Learning to drive by watching youtube videos: Action-conditioned contrastive policy pretraining.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Learning to drive by watching youtube videos: Action-conditioned contrastive policy pretraining

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:45.685235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T20:41:45.232668Z digest=sha256:21fa241cf2e7ae2c783331074c2f6d8b5a1d1cacf0682e1e4869649853cb8092

Observation 6a3966b8-a3b1-4ae6-b34e-daebfa857382 · outbound

This paper cites Hierarchical Task Learning from Language Instructions with Unified Transformers and Self-Monitoring.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Hierarchical Task Learning from Language Instructions with Unified Transformers and Self-Monitoring

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.236842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.236842Z digest=sha256:fb4732615596a1708ae98d0381c02b45d53bdfe7c85b7fb08e0937043f0675db

Observation 09bb8d68-c8de-434a-bcf7-6a0cf4209639 · outbound

This paper cites @esa (Ref.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents @esa (Ref

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.240583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.240583Z digest=sha256:0b433bc4d35388cb4ca7c19ae401625bc04164d9be737176150ad50b913ca6c8

Observation 83a22687-78ff-4f75-b957-fa99edf90c46 · outbound

This paper cites an unresolved cited work.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Unresolved cited work

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.245534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.245534Z digest=sha256:aa9c94b360c4041c4f83fe8993fe2d03b1ddc7506cf1dcdd74a9248f94f3bb9a

Observation 6b90dc01-faa4-4423-835c-a6af1e088068 · outbound

This paper cites MCU: An Evaluation Framework for Open-Ended Game Agents.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents MCU: An Evaluation Framework for Open-Ended Game Agents

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.249886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.249886Z digest=sha256:daa922ebfda54841a14dbbdbefc3e9ed852ea201e98e48b9a2b68263bd1d174c

Pith citing papers

No inbound Pith citation observations are available.