REVIEW 3 cited by
WebQA: Multihop and Multimodal QA
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Scaling Visual Question Answering (VQA) to the open-domain and multi-hop nature of web searches, requires fundamental advances in visual representation learning, knowledge aggregation, and language generation. In this work, we introduce WebQA, a challenging new benchmark that proves difficult for large-scale state-of-the-art models which lack language groundable visual representations for novel objects and the ability to reason, yet trivial for humans. WebQA mirrors the way humans use the web: 1) Ask a question, 2) Choose sources to aggregate, and 3) Produce a fluent language response. This is the behavior we should be expecting from IoT devices and digital assistants. Existing work prefers to assume that a model can either reason about knowledge in images or in text. WebQA includes a secondary text-only QA task to ensure improved visual performance does not come at the cost of language understanding. Our challenge for the community is to create unified multimodal reasoning models that answer questions regardless of the source modality, moving us closer to digital assistants that not only query language knowledge, but also the richer visual online world.
Forward citations
Cited by 3 Pith papers
-
POQD: Performance-Oriented Query Decomposer for Multi-vector retrieval
POQD uses an LLM-based optimizer to search the query-decomposition prompt together with RAG generator training, improving multi-vector retrieval and QA accuracy over fixed decomposition baselines.
-
Reconstructing Content with Collaborative Attention for Universal Multimodal Representation Learning
CoCoA forces an MLLM to reconstruct masked text through a single EOS token, improving multimodal embedding quality on MMEB-V1 and matching MoCa at 3B with far less pretraining data.
-
Analyze-Prompt-Reason: A Collaborative Agent-Based Framework for Multi-Image Vision-Language Reasoning
A prompt-engineered Claude 3.7, guided by GPT-4o-generated prompts and few-shot examples, reaches near-ceiling accuracy on most of the 18 MIRAGE multi-image reasoning tasks.
Discussion (0). Continue with ORCID to comment.