Pith. sign in

REVIEW 9 cited by

ALFRED: A Benchmark for Interpreting Grounded Instructions for Everyday Tasks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1912.01734 v2 pith:YIE6UMCM submitted 2019-12-03 cs.CV cs.AIcs.CLcs.LGcs.RO

classification cs.CVcs.AIcs.CLcs.LGcs.RO
keywords alfredlanguagetasksbenchmarkdirectivesinstructionsactioncoffee
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present ALFRED (Action Learning From Realistic Environments and Directives), a benchmark for learning a mapping from natural language instructions and egocentric vision to sequences of actions for household tasks. ALFRED includes long, compositional tasks with non-reversible state changes to shrink the gap between research benchmarks and real-world applications. ALFRED consists of expert demonstrations in interactive visual environments for 25k natural language directives. These directives contain both high-level goals like "Rinse off a mug and place it in the coffee maker." and low-level language instructions like "Walk to the coffee maker on the right." ALFRED tasks are more complex in terms of sequence length, action space, and language than existing vision-and-language task datasets. We show that a baseline model based on recent embodied vision-and-language tasks performs poorly on ALFRED, suggesting that there is significant room for developing innovative grounded visual language understanding models with this benchmark.

Discussion (0). Sign in to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Masked Diffusion Language Models are Strong and Steerable Text-Based World Models for Agentic RL

    cs.AI 2026-05 conditional novelty 7.0 of 10

    Masked diffusion language models, not larger autoregressive LLMs, are the better building block for text-based world models in agentic RL, improving rollout fidelity, diversity, and downstream task success.

  2. Conditional Multi-Stage Failure Recovery for Embodied Agents

    cs.CL 2025-07 conditional novelty 6.0 of 10

    A conditional four-stage chain-prompting method for failure recovery improves success on the TEACH embodied-agent benchmark from 24.9% to 36.5% with the same plan and executor.

  3. Embodied Web Agents: Bridging Physical-Digital Realms for Integrated Agent Intelligence

    cs.AI 2025-06 conditional novelty 6.0 of 10

    A new benchmark integrates AI2-THOR, Google Street View, and functional websites to test agents that must combine physical actions with online information retrieval.

  4. IndoorWorld: Integrating Physical Task Solving and Social Simulation in A Heterogeneous Multi-Agent Environment

    cs.MA 2025-06 conditional novelty 6.0 of 10

    IndoorWorld is a new multi-agent environment that combines physical task solving with social interaction, and its experiments show effects of collaboration, resource competition, and layout on agent behavior.

  5. RoboGPT-R1: Enhancing Robot Task Planning with Reinforcement Learning

    cs.AI 2025-10 conditional novelty 5.0 of 10

    A 3B VLM trained with SFT plus GRPO and an LCS-based reward reaches 55.3% on EmbodiedBench's EB-ALFRED, beating GPT-4o-mini and the 7B REBP planner.

  6. Embodied Spatial Intelligence: from Implicit Scene Modeling to Spatial Reasoning

    cs.RO 2025-08 conditional novelty 4.0 of 10

    The thesis demonstrates that combining implicit 3D scene representations with LLM-based reasoning, using text as an interface, yields strong performance on robotic perception and spatial language tasks.

  7. Learning to Reason for Factuality

    cs.CL 2025-08 unverdicted novelty 4.0 of 10

    An online RL reward balancing factual precision, detail, and relevance is reported to reduce reasoning-LLM hallucination by 23.1 points, but the attached manuscript is an unrelated paper.

  8. VerifyLLM: LLM-Based Pre-Execution Task Plan Verification for Robots

    cs.RO 2025-07 reject novelty 4.0 of 10

    An LLM plus LTL-based verification module that reorders, inserts, and removes steps in household robot plans, reporting reduced ordering errors but with weak experimental support.

  9. Analyze-Prompt-Reason: A Collaborative Agent-Based Framework for Multi-Image Vision-Language Reasoning

    cs.CV 2025-08 conditional novelty 3.0 of 10

    A prompt-engineered Claude 3.7, guided by GPT-4o-generated prompts and few-shot examples, reaches near-ceiling accuracy on most of the 18 MIRAGE multi-image reasoning tasks.

Pith tools