Pith. sign in

Paper Citation Record · LEDGER

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents

As of 21 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 0 inbound Pith citation observations for arXiv:2412.10410.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.10410 v1

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T20:41:45.249886Z

measured 65 of 65 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

65 of 65 outbound references displayed

  • verified exact1
  • verified fuzzy18
  • unresolved46
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f50f39f2-e052-4655-b5da-a369fe0d3d19 · outbound

This paper cites write newline.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:44.943729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:44.943729Z digest=sha256:1123fd701437f4463696ded47f9b4a0d9d740c66e80a838301efa23eecf0040d

Observation 28369869-51de-4e3f-b658-9b9c64591727 · outbound

This paper cites Multi-objective latent space optimization of generative molecular design models.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Multi-objective latent space optimization of generative molecular design models

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:46.001240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T20:41:44.949822Z digest=sha256:49dcbcffe05121a6173e8ec53d967dd25480836dbf74f1ddd6c69d34e4de8ab7

Observation 14e1e029-8980-4bd3-a986-8f2bf97615fc · outbound

This paper cites Imitating Interactive Intelligence.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Imitating Interactive Intelligence

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:44.960832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:44.960832Z digest=sha256:c70c03c5074ada93073c7a6bccf51296f25af1d044c95e278d12b2d44487960b

Observation 218ff899-930d-4e8f-a422-d73ca4dcb064 · outbound

This paper cites An optimistic perspective on offline reinforcement learning.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents An optimistic perspective on offline reinforcement learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:44.965648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:44.965648Z digest=sha256:451d37788454f732531e4cc95c766f0a6dc8955dde5b2167c45afefb26953345

Observation e7a7730c-45b8-4919-b357-713eb95f3b6b · outbound

This paper cites OPAL: Offline Primitive Discovery for Accelerating Offline Reinforcement Learning.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents OPAL: Offline Primitive Discovery for Accelerating Offline Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:44.970264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:44.970264Z digest=sha256:2818174ab6d3083d6c341a5cbcec5fc3e660967b09be032f13b4e25a33b23771

Observation 5edea55d-0b14-4aca-a234-78064cce8fe4 · outbound

This paper cites Flamingo: a Visual Language Model for Few-Shot Learning.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Flamingo: a Visual Language Model for Few-Shot Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:44.976022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:44.976022Z digest=sha256:040c55465bb368e53f2ee9d760875212488463221fe0582e331a0ed1eed62e5c

Observation 8744f31c-5e5a-4fa7-9fa7-7aaeed675001 · outbound

This paper cites Fixing a broken elbo.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Fixing a broken elbo

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:44.981536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:44.981536Z digest=sha256:bab0218c1a777e198d12bd2f063e15e79e4a279b752845c2ae955b4c3cb7ac95

Observation 69ca765b-5454-4b5d-9ab9-0dba13e472d5 · outbound

This paper cites Hindsight Experience Replay.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Hindsight Experience Replay

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:44.985917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:44.985917Z digest=sha256:0b1fd6e7635f6a092abbbf109a1d47a2ad07d4d474435a10fd43e77dc5c09ba7

Observation 7c984fe0-9b1a-41e1-9940-fc603d54ba1f · outbound

This paper cites Agent57: Outperforming the atari human benchmark.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Agent57: Outperforming the atari human benchmark

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:45.972738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T20:41:44.990370Z digest=sha256:f4a3b7c7b67d0ab528f7da88806b121b4cd1604f54562e4becf0db0f4098119b

Observation 06676397-6534-4398-8f3b-c47ad9508a58 · outbound

This paper cites Video PreTraining (VPT): Learning to Act by Watching Unlabeled Online Videos.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Video PreTraining (VPT): Learning to Act by Watching Unlabeled Online Videos

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:44.993876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:44.993876Z digest=sha256:3e5a2be05b2c3ab4dcd11c4e6eecb3fffe9626f67920d8827e8fbde352d6e740

Observation 500ce3f9-1759-4886-9837-0aaeff9386bb · outbound

This paper cites The arcade learning environment: An evaluation platform for general agents.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents The arcade learning environment: An evaluation platform for general agents

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:44.999032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:44.999032Z digest=sha256:7bdee09da08dbef078a04745dbeeb533a1aa327e57ad42be62bcbcaec61d5cc6

Observation 2cf34361-3637-4dd8-8133-e34fa7774007 · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents RT-1: Robotics Transformer for Real-World Control at Scale

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.003192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.003192Z digest=sha256:f3ac8f5f7d2edc38cdc666cf7cd6fa390d599da7e532b0edb7f6c017b1c735eb

Observation 1732c799-21f6-4b38-b370-1f9ef5864090 · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.013831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.013831Z digest=sha256:b06814c89075246a2381f9cc01a7959f974b114f804cd79c0b9f7d83ba5d1427

Observation 8da9e833-c68d-4395-b509-f4bd10d0700b · outbound

This paper cites Language Models are Few-Shot Learners.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Language Models are Few-Shot Learners

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.021602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.021602Z digest=sha256:0d0a967515e9aac7a760693160bc71668b696217c881086f5ab8895183d224e3

Observation 70444ac8-ee54-4e16-946a-323d725855cd · outbound

This paper cites Open-world multi-task control through goal-aware representation learning and adaptive horizon prediction.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Open-world multi-task control through goal-aware representation learning and adaptive horizon prediction

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:45.950291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T20:41:45.026106Z digest=sha256:2638888cbfaf6ddfac8aef4690b0268e06e8bd7d2f115d1a60dfb4166305c42c

Observation f24ea4a2-6aa5-4c12-90fb-e4beb55d40b1 · outbound

This paper cites Groot: Learning to follow instructions by watching gameplay videos.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Groot: Learning to follow instructions by watching gameplay videos

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:45.937746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T20:41:45.030662Z digest=sha256:0964f65848dacb437feb7873dd8c7e62e86169aedbe2e074ad1083bc2898edf0

Observation 83b80df0-fc92-467f-b030-94ed50058ad0 · outbound

This paper cites Abbeel, A.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Abbeel, A

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:45.923865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T20:41:45.034180Z digest=sha256:004244c24009b2b6cf88e1ce942b52c4bf34228073bba7340efac59dc258990f

Observation 18ad4ab7-8f55-4320-bc0c-aa46bd7932f0 · outbound

This paper cites Transformer-xl: Attentive language models beyond a fixed-length context.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Transformer-xl: Attentive language models beyond a fixed-length context

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.042834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.042834Z digest=sha256:1ba74136a5169e16c9f1c0a72ca7d7680ec1f78e78de0f7d60d7e6205cf6c49e

Observation 4f7cf927-0b2e-4ad3-abc6-6678d4951c82 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.046948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.046948Z digest=sha256:5b7760c83e51833ba7b9b419606df8ab9a436f8b8cbdbb863ffdb1d64c665586

Observation bc4d0dcc-4b45-404c-bb45-3bccd86d4211 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.049894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.049894Z digest=sha256:7a11e0fdca853fba28bcfe13460bc33d05809878c6890e447182342ac0b1e2fa

Observation 2c6542bb-2217-4c5d-8b02-0beca396f1c0 · outbound

This paper cites One-shot imitation learning.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents One-shot imitation learning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:45.911107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T20:41:45.053932Z digest=sha256:02612827bf6f644f5e07343526ad4fb94e38d2c1145ee04a0294e64c28570210

Observation d26f0261-4177-45a8-8f12-8f61498869d1 · outbound

This paper cites Implicit Deep Latent Variable Models for Text Generation.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Implicit Deep Latent Variable Models for Text Generation

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-11T20:41:45.530666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T20:41:45.057282Z digest=sha256:0e97897943de49aebce4252a3869f65f9ae5c3641925657153c7ffba7e1fe6d8

Observation 293ec0b5-6d0c-43aa-8f62-93c0d783b7d5 · outbound

This paper cites Relay Policy Learning: Solving Long-Horizon Tasks via Imitation and Reinforcement Learning.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Relay Policy Learning: Solving Long-Horizon Tasks via Imitation and Reinforcement Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.061059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.061059Z digest=sha256:37a4bc1ac66c5be2cf8deb127da7e9a69b0a65d7d647399ca33e1d81975711a1

Observation 13f6abe7-be3c-4463-8429-2a2b75e76ab5 · outbound

This paper cites Guss, Brandon Houghton, Nicholay Topin, Phillip Wang, Cayden Codel, Manuela M.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Guss, Brandon Houghton, Nicholay Topin, Phillip Wang, Cayden Codel, Manuela M

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.064991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.064991Z digest=sha256:f08555ebe2db9bf7e9d95b8625cc00b24a4de55c4cd0c365c85453b5eaa7e9a7

Observation 02419236-e53b-41a7-8067-4f6d8b769539 · outbound

This paper cites Vln-bert: A recurrent vision-and-language bert for navigation.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Vln-bert: A recurrent vision-and-language bert for navigation

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:45.891844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T20:41:45.068745Z digest=sha256:a73f68e7258f7c329eaf53d55dce86d0eb689731462beb7ec806dc2a98920890

Observation 14dc7441-235e-4a76-a095-56bdd444ded0 · outbound

This paper cites An Embodied Generalist Agent in 3D World.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents An Embodied Generalist Agent in 3D World

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.074200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.074200Z digest=sha256:a20c831d9393454d05505379d99c45b71bc123f1f3fc3955bd6ec46a119c9798

Observation 5fafca65-a57c-47b7-ab54-a0b2091c12a5 · outbound

This paper cites Vid2Robot: End-to-end Video-conditioned Policy Learning with Cross-Attention Transformers.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Vid2Robot: End-to-end Video-conditioned Policy Learning with Cross-Attention Transformers

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.078763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.078763Z digest=sha256:a2594ec6909700b1809c7f35a8845dcff72c7762565867da0a12420f78d3f4fa

Observation 61b1740d-0afc-4bfd-9177-f8d85f0c2d28 · outbound

This paper cites BC-Z: Zero-Shot Task Generalization with Robotic Imitation Learning.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents BC-Z: Zero-Shot Task Generalization with Robotic Imitation Learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.083397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.083397Z digest=sha256:4c3e26c05c772afbbd97a5febd86ef3a2d7b8318057c0e7edad4207fbe6915f3

Observation 0352c213-bfdc-4930-bc33-46e87ae24b0f · outbound

This paper cites The malmo platform for artificial intelligence experimentation.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents The malmo platform for artificial intelligence experimentation

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:45.878744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T20:41:45.088113Z digest=sha256:c34d36adf5bf2f3130e1532586d36034c8bf1f9659cbbf2f1a4f2cfb2229e473

Observation 1b0a3218-9245-42d3-960d-a904105392fc · outbound

This paper cites Auto-Encoding Variational Bayes.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Auto-Encoding Variational Bayes

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.092438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.092438Z digest=sha256:fd20583dfabc1e503347ce585183dbbc4a93c1f7da6b5a773bdd26ed7d2a38e9

Observation 3f602080-664e-463e-a314-8bd7086f9af2 · outbound

This paper cites Multi-game decision transformers.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Multi-game decision transformers

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.096021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.096021Z digest=sha256:b91ef0255e915e8cf02aa34efd198bd74374241381cae3bbb757f3f73713a838

Observation 3daa7d74-c326-499c-82d9-52b066f61ea0 · outbound

This paper cites Evaluating Real-World Robot Manipulation Policies in Simulation.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Evaluating Real-World Robot Manipulation Policies in Simulation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.100715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.100715Z digest=sha256:2136ecb597dbd851fb675455574e93649dddf8f3fdd378f74e4608cc5aa515ac

Observation 158adb09-896f-45fb-9640-3dd60c27624c · outbound

This paper cites STEVE-1: A Generative Model for Text-to-Behavior in Minecraft.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents STEVE-1: A Generative Model for Text-to-Behavior in Minecraft

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.105155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.105155Z digest=sha256:b0e894e2c3ec3ba5c4d4da9d96185d10f36c6e36f181ce98da5ec79f08b44cc9

Observation 4bb96c3a-44ef-4fd2-8a33-8135c6256bc9 · outbound

This paper cites Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.116529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.116529Z digest=sha256:7bd72a608c6373741966f3c582e1160c32c05ba19600f1d60919706faefb7905

Observation 151077ff-6bb6-4603-adb4-be779ff5523e · outbound

This paper cites Language Conditioned Imitation Learning over Unstructured Data.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Language Conditioned Imitation Learning over Unstructured Data

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.120926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.120926Z digest=sha256:15b1ccbaad009963e6d9af0efc3b4e93123b97cf811500caefaf02b2b8abc302

Observation 8b7ab2c4-e7d8-40eb-acd9-3de1f078aaf5 · outbound

This paper cites Learning latent plans from play.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Learning latent plans from play

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:45.854733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T20:41:45.124914Z digest=sha256:55794d9ddc77cd9492aa26977f22491f75c9874762080d36430aa8c5adf04579

Observation 8f0b4f59-1616-491c-9527-c1e1dc268d6d · outbound

This paper cites Learning latent plans from play.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Learning latent plans from play

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:45.841767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T20:41:45.128533Z digest=sha256:b978d116b37e286a76441c134d9e87b984c6393735c1559df97ad1cb54a368c4

Observation 85dae5c4-4642-4d51-9bf8-e5e2390ec52a · outbound

This paper cites Learning latent plans from play.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Learning latent plans from play

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:45.827698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T20:41:45.132613Z digest=sha256:86fe2f7909eeb250ba54a6985b511619401ba636e118ffd789b07bd8a7d8e9c9

Observation 53c39c27-ffb2-4a5c-a38b-200f01ec7b37 · outbound

This paper cites Interactive language: Talking to robots in real time.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Interactive language: Talking to robots in real time

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.136889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.136889Z digest=sha256:1dd1577b513c1f78f41fb1a6a4884cd9cc28a98daa4157504a61bc336491062a

Observation 1c85f61f-d6cc-400c-9029-c29b0d88694e · outbound

This paper cites ZSON: Zero-Shot Object-Goal Navigation using Multimodal Goal Embeddings.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents ZSON: Zero-Shot Object-Goal Navigation using Multimodal Goal Embeddings

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.140755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.140755Z digest=sha256:b71da95f488b755a4e1726229775e76087c6e7f5c9acd1ee5ac3b8a7a98b29b0

Observation ee0692f7-bc58-4de7-9fd4-0870847b5b14 · outbound

This paper cites What matters in language conditioned robotic imitation learning over unstructured data.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents What matters in language conditioned robotic imitation learning over unstructured data

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:45.807648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T20:41:45.145567Z digest=sha256:c6c6ca50f7146093e15470aee725277a8100193a49ed03ad259047c2e77d8c11

Observation f5b071b7-f4f6-432b-978a-2fe318553e90 · outbound

This paper cites Rusu, Joel Veness, Marc G.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Rusu, Joel Veness, Marc G

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.149036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.149036Z digest=sha256:78fe324d6f040bcc871ddc874c95c90ebc909f18b11b2d8fa84e23679a4ebf68

Observation 58d8e0f2-56fc-4a86-b104-0c53e31d9b77 · outbound

This paper cites Goal representations for instruction following: A semi-supervised language interface to control.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Goal representations for instruction following: A semi-supervised language interface to control

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:45.787144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T20:41:45.152603Z digest=sha256:f7808045508a92f5a1ff5b15abcb21a20aa74205bd0021c6de492da347ebe009

Observation f89879f9-987a-4a0c-99f1-cc009537a33d · outbound

This paper cites Octo: An open-source generalist robot policy.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Octo: An open-source generalist robot policy

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.156190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.156190Z digest=sha256:55f9df7ee54570bf4b1984fb1fd571ad8202f2f9d7424fe7a19cafec7f71880d

Observation 152c4e05-a8ed-4815-983d-9f22f37cce05 · outbound

This paper cites Open X-Embodiment: Robotic Learning Datasets and RT-X Models.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Open X-Embodiment: Robotic Learning Datasets and RT-X Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.159279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.159279Z digest=sha256:e8a1f382c2a8820fccbab0b2f23bdc51c88699be85d33d7d8332637b0a0082ad

Observation 0bb8bfe7-c9e8-4958-aef8-a31f933bc552 · outbound

This paper cites Conditional Variational Autoencoder for Neural Machine Translation.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Conditional Variational Autoencoder for Neural Machine Translation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.162510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.162510Z digest=sha256:bb783ea2de338a3bfee711313b04f3af865d8eba68835ed0764208ef9916ee93

Observation 84e935d3-5426-47a3-b218-c758916988de · outbound

This paper cites Episodic transformer for vision-and-language navigation.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Episodic transformer for vision-and-language navigation

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:45.766547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T20:41:45.166517Z digest=sha256:27a0c32b3b93602dca672c01eebd474f761da37ca2b6f0aa463c3f2aec168fda

Observation 87dd3bdf-20cf-41ee-badf-f7e0df4ffe25 · outbound

This paper cites Accelerating reinforcement learning with learned skill priors.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Accelerating reinforcement learning with learned skill priors

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:45.749383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T20:41:45.169963Z digest=sha256:d5492feca2e238c81835abe252aeab902fca4510cd5bed48494e22f3c9c04f96

Observation b0412b2c-a674-4c5b-812e-05dd181154cb · outbound

This paper cites Scaling Instructable Agents Across Many Simulated Worlds.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Scaling Instructable Agents Across Many Simulated Worlds

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.173544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.173544Z digest=sha256:3d896bc2756039b9d042d799a79f5f2dfa5759a13e1cbbbf03f2a8a4f5e0cbdc

Observation ea1ff3cd-f5d9-45a3-b93d-d96a13eb9d6c · outbound

This paper cites Improving language understanding by generative pre-training.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Improving language understanding by generative pre-training

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.177377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.177377Z digest=sha256:f99acd28eff213da6e6271d581fabfa831bdb176c7aa6ef5cf54b876b35b406d

Observation 35cf3de0-21f3-4e86-848e-2cd6bc0f9f24 · outbound

This paper cites Language models are unsupervised multitask learners.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Language models are unsupervised multitask learners

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.181411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.181411Z digest=sha256:125bb1d4ec3fe3ec852156aef1169bbfb8459de27059037dba3ef787000afbd8

Observation 4c021881-6925-4123-a7b0-f3c5c08f3fee · outbound

This paper cites Learning transferable visual models from natural language supervision.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Learning transferable visual models from natural language supervision

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.186191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.186191Z digest=sha256:fb9020e2c59df894f960b3927b4ffb381ecf95e13c3d2487ece9261f01c49b05

Observation 407263d1-4278-4042-b2b2-63265791bec9 · outbound

This paper cites A Generalist Agent.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents A Generalist Agent

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.191304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.191304Z digest=sha256:c7d705be64625b3b26beeec6c82f93506a2f8700f1e2d2644f725df9dfcc999c

Observation e417dc5c-f1c3-4089-a3b8-ec46ce4657ab · outbound

This paper cites Habitat: A platform for embodied ai research.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Habitat: A platform for embodied ai research

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.195401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.195401Z digest=sha256:fd9dce21d93404eaf8ecdaea0a02e43d1db11242d8d0f0079aae03e9bc4ee3d1

Observation c31d09c1-597f-45b4-a07a-4debc3ede1aa · outbound

This paper cites Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.198736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.198736Z digest=sha256:6779c1a95c00cbfe14e810d400d6343750dd98a6af64fcee16984f9063f91c0c

Observation f5f0f9ab-0452-4a05-859e-96f9a2094da0 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Gemini: A Family of Highly Capable Multimodal Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.203599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.203599Z digest=sha256:a9b95e5b94776d017d12c3c4c688a5037088d82dec22a8e411878f2fd8dc8c22

Observation 3b518a75-3d8d-4552-bceb-8f71998c5f2b · outbound

This paper cites Describe, explain, plan and select: interactive planning with llms enables open-world multi-task agents.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Describe, explain, plan and select: interactive planning with llms enables open-world multi-task agents

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:45.706557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T20:41:45.213308Z digest=sha256:122344d19299e5949042ba6e69ff1645f262aa31e6061b077df0af2a436c0b06

Observation e80437eb-59ba-462c-83a0-4a437aeed3ba · outbound

This paper cites JARVIS-1: Open-World Multi-task Agents with Memory-Augmented Multimodal Language Models.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents JARVIS-1: Open-World Multi-task Agents with Memory-Augmented Multimodal Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.217706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.217706Z digest=sha256:d86a08188a710a5875e71c075a5d09d461bb9793bd447466ef16c5708b64321b

Observation 3b12a3f0-374d-4c8a-91ef-fbd89547d5b3 · outbound

This paper cites Xskill: Cross embodiment skill discovery.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Xskill: Cross embodiment skill discovery

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:45.695643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T20:41:45.223484Z digest=sha256:ebd3c73a5c51941eabf08ad7ff2b64a6c9baf7e98c6c77d59f31a1ce06582c48

Observation cdb9dadc-fcbe-496d-b6d0-89fe6d0c414c · outbound

This paper cites Meta-World: A Benchmark and Evaluation for Multi-Task and Meta Reinforcement Learning.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Meta-World: A Benchmark and Evaluation for Multi-Task and Meta Reinforcement Learning

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.228028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.228028Z digest=sha256:92802388170bc4ea65815c5f98caeaf6f83c8dc801ac120e8adeb3236dc9e006

Observation a5e56cbd-a5b0-421b-bd47-db18e53ca0e6 · outbound

This paper cites Learning to drive by watching youtube videos: Action-conditioned contrastive policy pretraining.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Learning to drive by watching youtube videos: Action-conditioned contrastive policy pretraining

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:45.685235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T20:41:45.232668Z digest=sha256:c1f52761b046c68ccc3bad0351f08d9cb3545cbf450daef10a3c2773ff744718

Observation 6a3966b8-a3b1-4ae6-b34e-daebfa857382 · outbound

This paper cites Hierarchical Task Learning from Language Instructions with Unified Transformers and Self-Monitoring.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Hierarchical Task Learning from Language Instructions with Unified Transformers and Self-Monitoring

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.236842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.236842Z digest=sha256:3fe3b1b04a0f8e66e7dccbcbad31875cd4232080c092e57374aba6990676848a

Observation 09bb8d68-c8de-434a-bcf7-6a0cf4209639 · outbound

This paper cites @esa (Ref.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents @esa (Ref

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.240583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.240583Z digest=sha256:5d9ce47400e39824b87a2f40c3c3ed0ba07e413656b562613b9002a71ff31078

Observation 83a22687-78ff-4f75-b957-fa99edf90c46 · outbound

This paper cites an unresolved cited work.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Unresolved cited work

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.245534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.245534Z digest=sha256:6c43104515daa9cd8f2aec7193ec182468cfbc52410f2bf7e148e3e6a7178a1c

Observation 6b90dc01-faa4-4423-835c-a6af1e088068 · outbound

This paper cites MCU: An Evaluation Framework for Open-Ended Game Agents.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents MCU: An Evaluation Framework for Open-Ended Game Agents

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.249886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.249886Z digest=sha256:ec7a5bf3975102b4591cde0232947e166b2af3ebc7424f2a4c3b7712e56fcd70

Pith citing papers

No inbound Pith citation observations are available.