REVIEW 9 cited by
ALFRED: A Benchmark for Interpreting Grounded Instructions for Everyday Tasks
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We present ALFRED (Action Learning From Realistic Environments and Directives), a benchmark for learning a mapping from natural language instructions and egocentric vision to sequences of actions for household tasks. ALFRED includes long, compositional tasks with non-reversible state changes to shrink the gap between research benchmarks and real-world applications. ALFRED consists of expert demonstrations in interactive visual environments for 25k natural language directives. These directives contain both high-level goals like "Rinse off a mug and place it in the coffee maker." and low-level language instructions like "Walk to the coffee maker on the right." ALFRED tasks are more complex in terms of sequence length, action space, and language than existing vision-and-language task datasets. We show that a baseline model based on recent embodied vision-and-language tasks performs poorly on ALFRED, suggesting that there is significant room for developing innovative grounded visual language understanding models with this benchmark.
Forward citations
Cited by 9 Pith papers
-
Masked Diffusion Language Models are Strong and Steerable Text-Based World Models for Agentic RL
Masked diffusion language models, not larger autoregressive LLMs, are the better building block for text-based world models in agentic RL, improving rollout fidelity, diversity, and downstream task success.
-
Conditional Multi-Stage Failure Recovery for Embodied Agents
A conditional four-stage chain-prompting method for failure recovery improves success on the TEACH embodied-agent benchmark from 24.9% to 36.5% with the same plan and executor.
-
Embodied Web Agents: Bridging Physical-Digital Realms for Integrated Agent Intelligence
A new benchmark integrates AI2-THOR, Google Street View, and functional websites to test agents that must combine physical actions with online information retrieval.
-
IndoorWorld: Integrating Physical Task Solving and Social Simulation in A Heterogeneous Multi-Agent Environment
IndoorWorld is a new multi-agent environment that combines physical task solving with social interaction, and its experiments show effects of collaboration, resource competition, and layout on agent behavior.
-
RoboGPT-R1: Enhancing Robot Task Planning with Reinforcement Learning
A 3B VLM trained with SFT plus GRPO and an LCS-based reward reaches 55.3% on EmbodiedBench's EB-ALFRED, beating GPT-4o-mini and the 7B REBP planner.
-
Embodied Spatial Intelligence: from Implicit Scene Modeling to Spatial Reasoning
The thesis demonstrates that combining implicit 3D scene representations with LLM-based reasoning, using text as an interface, yields strong performance on robotic perception and spatial language tasks.
-
Learning to Reason for Factuality
An online RL reward balancing factual precision, detail, and relevance is reported to reduce reasoning-LLM hallucination by 23.1 points, but the attached manuscript is an unrelated paper.
-
VerifyLLM: LLM-Based Pre-Execution Task Plan Verification for Robots
An LLM plus LTL-based verification module that reorders, inserts, and removes steps in household robot plans, reporting reduced ordering errors but with weak experimental support.
-
Analyze-Prompt-Reason: A Collaborative Agent-Based Framework for Multi-Image Vision-Language Reasoning
A prompt-engineered Claude 3.7, guided by GPT-4o-generated prompts and few-shot examples, reaches near-ceiling accuracy on most of the 18 MIRAGE multi-image reasoning tasks.
Discussion (0). Sign in to comment.