Pith. sign in

Paper Citation Record · LEDGER

JARVIS-1: Open-World Multi-task Agents with Memory-Augmented Multimodal Language Models

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 24 inbound Pith citation observations for arXiv:2311.05997.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.05997 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 24 of 24 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:15:12.317283Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

6
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 7f5ae896-6dc5-4d83-a399-e905b05d9976 · inbound

AppAgent: Multimodal Agents as Smartphone Users cites this paper.

AppAgent: Multimodal Agents as Smartphone Users JARVIS-1: Open-World Multi-task Agents with Memory-Augmented Multimodal Language Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-17T10:16:43.859010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-17T10:16:43.364787Z digest=sha256:9059fe397846d3a8299e937089be8da6e251b4665f1bac3da7c40c9442b3825e

Observation 15526ff7-f592-46b4-97f6-d0f44ac0f41c · inbound

A Survey on the Memory Mechanism of Large Language Model based Agents cites this paper.

A Survey on the Memory Mechanism of Large Language Model based Agents JARVIS-1: Open-World Multi-task Agents with Memory-Augmented Multimodal Language Models

Reference 159

Resolution
verified exact
arxiv_id, observed 2026-05-15T07:21:39.639525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T07:21:39.440092Z digest=sha256:616509ed0df2d3cb30388991f2db73e864a8f4d64009022658ed70ed55434b2b

Observation 4a599350-0594-40a3-b514-7e728e33be38 · inbound

STEVE-Audio: Expanding the Goal Conditioning Modalities of Embodied Agents in Minecraft cites this paper.

STEVE-Audio: Expanding the Goal Conditioning Modalities of Embodied Agents in Minecraft JARVIS-1: Open-World Multi-task Agents with Memory-Augmented Multimodal Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T04:54:54.559169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:54:54.559169Z digest=sha256:38caa4e809c3de16c45d15b27222ad1ffd19af27babf2f0b4ff21888784fcae2

Observation 824b9835-d3f1-4f75-a8c2-726ef302f0a7 · inbound

Preference Goal Tuning: Post-Training as Latent Control for Frozen Policies cites this paper.

Preference Goal Tuning: Post-Training as Latent Control for Frozen Policies JARVIS-1: Open-World Multi-task Agents with Memory-Augmented Multimodal Language Models

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-23T08:22:44.348038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-23T08:20:05.898025Z digest=sha256:2eb123360d7c5919417814df0847afa8b4650b683445ac74a7c9eecea7652023

Observation 35c7445e-d855-4984-94bf-e9b6e8129581 · inbound

WiS Platform: Enhancing Evaluation of LLM-Based Multi-Agent Systems Through Game-Based Analysis cites this paper.

WiS Platform: Enhancing Evaluation of LLM-Based Multi-Agent Systems Through Game-Based Analysis JARVIS-1: Open-World Multi-task Agents with Memory-Augmented Multimodal Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T22:33:53.250733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:33:53.250733Z digest=sha256:8955fc1db2afad580c96209b429ea43eff5e083948d8b6ebd36e6a2c8ae57ad8

Observation 0a25bac1-aa92-4b0c-99f3-632e7f9f7e47 · inbound

TeamCraft: A Benchmark for Multi-Modal Multi-Agent Systems in Minecraft cites this paper.

TeamCraft: A Benchmark for Multi-Modal Multi-Agent Systems in Minecraft JARVIS-1: Open-World Multi-task Agents with Memory-Augmented Multimodal Language Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T20:53:35.431771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:53:35.431771Z digest=sha256:8f2935d327c1e0a7e7d4676e210f0d72a3e854cb03362ccae7f74a94c5dc4c3c

Observation e80437eb-59ba-462c-83a0-4a437aeed3ba · inbound

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents cites this paper.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents JARVIS-1: Open-World Multi-task Agents with Memory-Augmented Multimodal Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.217706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.217706Z digest=sha256:d86a08188a710a5875e71c075a5d09d461bb9793bd447466ef16c5708b64321b

Observation a8f09fce-0f1c-40b5-abb3-a032e7d84de0 · inbound

MineStudio: A Streamlined Package for Minecraft AI Agent Development cites this paper.

MineStudio: A Streamlined Package for Minecraft AI Agent Development JARVIS-1: Open-World Multi-task Agents with Memory-Augmented Multimodal Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T04:51:58.309631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T04:51:58.309631Z digest=sha256:cca4c40ca5e887d92e01d955f02c0d72df700ace71b2816245bd444b87e82aeb

Observation e35f2d17-2504-4a29-bd3d-1928b2a34646 · inbound

FaGeL: Fabric LLMs Agent empowered Embodied Intelligence Evolution with Autonomous Human-Machine Collaboration cites this paper.

FaGeL: Fabric LLMs Agent empowered Embodied Intelligence Evolution with Autonomous Human-Machine Collaboration JARVIS-1: Open-World Multi-task Agents with Memory-Augmented Multimodal Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T23:27:21.422487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:27:21.422487Z digest=sha256:07c13794c94c6f0c470e7ff282b24fcbc2ae881b63090d8cb788aac4bdc5a9cf

Observation f02204f5-6697-41a8-9983-636805a31499 · inbound

Plancraft: an evaluation dataset for planning with LLM agents cites this paper.

Plancraft: an evaluation dataset for planning with LLM agents JARVIS-1: Open-World Multi-task Agents with Memory-Augmented Multimodal Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:35.187048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:35.187048Z digest=sha256:0fcacaf6ecd23727b2b57910ed318b300ee8306649f20c9db54361e581160706

Observation b5b10156-0b3d-4669-b251-8dedbf674037 · inbound

Embodied VideoAgent: Persistent Memory from Egocentric Videos and Embodied Sensors Enables Dynamic Scene Understanding cites this paper.

Embodied VideoAgent: Persistent Memory from Egocentric Videos and Embodied Sensors Enables Dynamic Scene Understanding JARVIS-1: Open-World Multi-task Agents with Memory-Augmented Multimodal Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:54.804278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:54.804278Z digest=sha256:3e3b946b79451327eddc963e32b4c0e982a3f2940f46c715f7530124d749f685

Observation 4089d60e-4f4e-4217-b525-d4723ac00e63 · inbound

Visual Language Models as Operator Agents in the Space Domain cites this paper.

Visual Language Models as Operator Agents in the Space Domain JARVIS-1: Open-World Multi-task Agents with Memory-Augmented Multimodal Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:43.847257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:43.847257Z digest=sha256:e76ed023e02eca0e85765b061a34b736d3a4a2fe51cf3a639ab2b1d3b1be727d

Observation 7d1ff9b9-edd1-4000-a3c4-76583f6e71b5 · inbound

Conditional Multi-Stage Failure Recovery for Embodied Agents cites this paper.

Conditional Multi-Stage Failure Recovery for Embodied Agents JARVIS-1: Open-World Multi-task Agents with Memory-Augmented Multimodal Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:28.360265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:28.360265Z digest=sha256:17c9ccc17f0315572c07b506a6f12bb3f9c6e276afa57cfb2cd0923ff75c5ed9

Observation e154e997-5481-41c7-97ac-58f74b2233a0 · inbound

CrafterDojo: A Suite of Foundation Models for Building Open-Ended Embodied Agents in Crafter cites this paper.

CrafterDojo: A Suite of Foundation Models for Building Open-Ended Embodied Agents in Crafter JARVIS-1: Open-World Multi-task Agents with Memory-Augmented Multimodal Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T17:15:12.317283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:15:12.317283Z digest=sha256:03a355a2361128a980be088fe56aef8629bf962ae8764189f9905621a139b2f4

Observation 3a0e4777-3c75-4ded-818a-f0f47d392161 · inbound

SoK: Agentic Skills -- Beyond Tool Use in LLM Agents cites this paper.

SoK: Agentic Skills -- Beyond Tool Use in LLM Agents JARVIS-1: Open-World Multi-task Agents with Memory-Augmented Multimodal Language Models

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-14T23:19:31.125901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T23:19:31.024268Z digest=sha256:c47cd6c70c7ca5b70fa6b08176682a9ab87b592fd0b8e9c5927d0516c3621620

Observation 739e99f3-24a0-4af3-9f5c-d42b8a4c71a9 · inbound

GameWorld: Towards Standardized and Verifiable Evaluation of Multimodal Game Agents cites this paper.

GameWorld: Towards Standardized and Verifiable Evaluation of Multimodal Game Agents JARVIS-1: Open-World Multi-task Agents with Memory-Augmented Multimodal Language Models

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:46:07.697516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T17:57:36.038091Z digest=sha256:cb6d40c44f046053403597f3afceeba814214a9fbb3dc43121c47dcba2a9fcb4

Observation 52b44462-e38f-410a-be05-86e6d43ddfeb · inbound

Dream-Cubed: Controllable Generative Modeling in Minecraft by Training on Billions of Cubes cites this paper.

Dream-Cubed: Controllable Generative Modeling in Minecraft by Training on Billions of Cubes JARVIS-1: Open-World Multi-task Agents with Memory-Augmented Multimodal Language Models

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:31:02.641583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T01:36:23.517745Z digest=sha256:41bf3a41f01d31044c70388e753a006611e9df380237f0f417354ce2307bcade

Observation 7c497296-ea37-4c23-811e-5fa4a4a2e53d · inbound

A Comprehensive Survey on Agent Skills: Taxonomy, Techniques, and Applications cites this paper.

A Comprehensive Survey on Agent Skills: Taxonomy, Techniques, and Applications JARVIS-1: Open-World Multi-task Agents with Memory-Augmented Multimodal Language Models

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:20:57.243431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-11T01:47:39.926540Z digest=sha256:637e2f1d1b6ffcef9e8cafab4b249e679fe1b7d3f64e6e9db808861d9f132ff1

Observation 3693cabc-54e7-4662-ad09-4ad65c97457a · inbound

A Comprehensive Survey on Agent Skills: Taxonomy, Techniques, and Applications cites this paper.

A Comprehensive Survey on Agent Skills: Taxonomy, Techniques, and Applications JARVIS-1: Open-World Multi-task Agents with Memory-Augmented Multimodal Language Models

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-20T23:19:15.144913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T23:15:44.550045Z digest=sha256:8f0edf5ba862911afc2670b0e9e54d5ed178e47b52902c82ba0f4c2558267a00

Observation fe68687f-7002-4018-be4d-a9fef70d153a · inbound

A Comprehensive Survey on Agent Skills: Taxonomy, Techniques, and Applications cites this paper.

A Comprehensive Survey on Agent Skills: Taxonomy, Techniques, and Applications JARVIS-1: Open-World Multi-task Agents with Memory-Augmented Multimodal Language Models

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-06-30T23:25:07.322279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T23:23:42.883286Z digest=sha256:8a46c19faebae6ab9f8c4c70057e4df5d335ffbca4e4a67abc2cb928ef87700d

Observation 9a54de7f-832a-45f2-a6a5-5807ba3488a9 · inbound

Do Vision-Language-Models show human-like logical problem-solving capability in point and click puzzle games? cites this paper.

Do Vision-Language-Models show human-like logical problem-solving capability in point and click puzzle games? JARVIS-1: Open-World Multi-task Agents with Memory-Augmented Multimodal Language Models

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T02:02:05.547051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-13T01:58:39.476408Z digest=sha256:eacd9af1f519128429ee24f474646b6bd3217ab79adda18925b51b04ecb25588

Observation 38091461-5cb6-4fdf-9737-c8aeff9360c9 · inbound

SPIKE: An Adaptive Dual Controller Framework for Cost-Efficient Long-Horizon Game Agents cites this paper.

SPIKE: An Adaptive Dual Controller Framework for Cost-Efficient Long-Horizon Game Agents JARVIS-1: Open-World Multi-task Agents with Memory-Augmented Multimodal Language Models

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:33:12.266648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T10:32:26.668583Z digest=sha256:08d2e2929e64f1c5e0f2c55aae320450cfb9c15aaaddb71a1a4098641f3ed037

Observation 0ce76651-5009-4389-bde0-20a88a50eefc · inbound

HiMe: Hierarchical Embodied Memory for Long-Horizon Vision-Language-Action Control cites this paper.

HiMe: Hierarchical Embodied Memory for Long-Horizon Vision-Language-Action Control JARVIS-1: Open-World Multi-task Agents with Memory-Augmented Multimodal Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-12T02:24:21.020383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T02:24:21.020383Z digest=sha256:039a3b1fbefb9a636f44b1f2ac39de5fecd38ce988143df38a2760b3ba6bdfc9

Observation 1a894190-d2c8-4030-ae46-267294b59629 · inbound

Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills cites this paper.

Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills JARVIS-1: Open-World Multi-task Agents with Memory-Augmented Multimodal Language Models

Reference 259

Resolution
unresolved
no resolver link, observed 2026-08-04T19:45:35.319727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:45:35.319727Z digest=sha256:7afbf4edbec35f2088f8c2e9f25fc7a5d55c33978bd02d0e94081028cced4b04