Pith. sign in

REVIEW 4 cited by

Blindfold Baselines for Embodied QA

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1811.05013 v1 pith:IWH5J3G3 submitted 2018-11-12 cs.CV cs.AIcs.CLcs.LG

classification cs.CVcs.AIcs.CLcs.LG
keywords blindfoldagentansweringbaselinebaselinesembodiedembodiedqaenvironment
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We explore blindfold (question-only) baselines for Embodied Question Answering. The EmbodiedQA task requires an agent to answer a question by intelligently navigating in a simulated environment, gathering necessary visual information only through first-person vision before finally answering. Consequently, a blindfold baseline which ignores the environment and visual information is a degenerate solution, yet we show through our experiments on the EQAv1 dataset that a simple question-only baseline achieves state-of-the-art results on the EmbodiedQA task in all cases except when the agent is spawned extremely close to the object.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Blind text-only likelihood models match or exceed CLIP on many compositionality benchmarks because positives and negatives differ systematically in length, plausibility, or image style.

  2. Interactive Language Learning by Question Answering

    cs.CL 2019-08 conditional novelty 6.0 of 10

    QAit turns question answering into an interactive text-game task, and the paper's baselines show current agents cannot generalize beyond memorized games, while humans can.

  3. VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering

    cs.CV 2019-08 conditional novelty 6.0 of 10

    VideoNavQA pairs 101,000 template-generated questions with near-optimal navigation videos in House3D, and adapted VQA models beat language-only baselines by about 8 accuracy points.

  4. Walking with MIND: Mental Imagery eNhanceD Embodied QA

    cs.CV 2019-08 conditional novelty 6.0 of 10

    A mental imagery module that predicts future views and treats them as short-term subgoals improves an embodied agent's navigation and question-answering accuracy in simulation.

Pith tools