REVIEW 4 cited by
Blindfold Baselines for Embodied QA
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We explore blindfold (question-only) baselines for Embodied Question Answering. The EmbodiedQA task requires an agent to answer a question by intelligently navigating in a simulated environment, gathering necessary visual information only through first-person vision before finally answering. Consequently, a blindfold baseline which ignores the environment and visual information is a degenerate solution, yet we show through our experiments on the EQAv1 dataset that a simple question-only baseline achieves state-of-the-art results on the EmbodiedQA task in all cases except when the agent is spawned extremely close to the object.
Forward citations
Cited by 4 Pith papers
-
A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks
Blind text-only likelihood models match or exceed CLIP on many compositionality benchmarks because positives and negatives differ systematically in length, plausibility, or image style.
-
Interactive Language Learning by Question Answering
QAit turns question answering into an interactive text-game task, and the paper's baselines show current agents cannot generalize beyond memorized games, while humans can.
-
VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering
VideoNavQA pairs 101,000 template-generated questions with near-optimal navigation videos in House3D, and adapted VQA models beat language-only baselines by about 8 accuracy points.
-
Walking with MIND: Mental Imagery eNhanceD Embodied QA
A mental imagery module that predicts future views and treats them as short-term subgoals improves an embodied agent's navigation and question-answering accuracy in simulation.
Discussion (0). Continue with ORCID to comment.