Pith. sign in

REVIEW 6 cited by

SituatedQA: Incorporating Extra-Linguistic Contexts into QA

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2109.06157 v1 pith:WWCXMAXT submitted 2021-09-13 cs.CL

classification cs.CL
keywords answersquestionssituatedqacontextsexistingextra-linguisticquestionasked
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Answers to the same question may change depending on the extra-linguistic contexts (when and where the question was asked). To study this challenge, we introduce SituatedQA, an open-retrieval QA dataset where systems must produce the correct answer to a question given the temporal or geographical context. To construct SituatedQA, we first identify such questions in existing QA datasets. We find that a significant proportion of information seeking questions have context-dependent answers (e.g., roughly 16.5% of NQ-Open). For such context-dependent questions, we then crowdsource alternative contexts and their corresponding answers. Our study shows that existing models struggle with producing answers that are frequently updated or from uncommon locations. We further quantify how existing models, which are trained on data collected in the past, fail to generalize to answering questions asked in the present, even when provided with an updated evidence corpus (a roughly 15 point drop in accuracy). Our analysis suggests that open-retrieval QA benchmarks should incorporate extra-linguistic context to stay relevant globally and in the future. Our data, code, and datasheet are available at https://situatedqa.github.io/ .

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DailyQA: A Benchmark to Evaluate Web Retrieval Augmented LLMs Based on Capturing Real-World Changes

    cs.IR 2025-05 conditional novelty 6.0 of 10

    DailyQA, automatically built from Wikipedia revision logs, shows that LLMs with web retrieval answer only about half of time-sensitive questions correctly, and that reranking documents outperforms both snippets and ti...

  2. CORG: Generating Answers from Complex, Interrelated Contexts

    cs.CL 2025-04 conditional novelty 6.0 of 10

    CORG, a graph-based context grouping framework, improves disambiguated answer recall on QA with distracting, ambiguous, counterfactual, and duplicated contexts, and reaches performance comparable to per-document proce...

  3. MTPChat: A Multimodal Time-Aware Persona Dataset for Conversational Agents

    cs.CL 2025-02 conditional novelty 6.0 of 10

    MTPChat adds explicit date stamps and synthetic earlier responses to multimodal persona dialogues, defines two temporal retrieval tasks, and reports modest gains from a gated fusion module.

  4. ToolComp: A Multi-Tool Reasoning & Process Supervision Benchmark

    cs.CL 2025-01 conditional novelty 6.0 of 10

    ToolComp, a human-verified multi-tool reasoning benchmark with per-step labels, shows process-supervised reward models outperform outcome-supervised ones in ranking tool-use trajectories.

  5. Concept Incongruence: An Exploration of Time and Death in Role Playing

    cs.CL 2025-05 conditional novelty 5.0 of 10

    LLMs asked to role-play dead historical figures rarely abstain from answering post-death questions, and their factual accuracy drops due to poorly encoded death states and role-playing-induced shifts in temporal repre...

  6. TelcoLM: collecting data, adapting, and benchmarking language models for the telecommunication domain

    cs.CL 2024-12 conditional novelty 5.0 of 10

    For Llama-2-7B on telecommunications tasks, instruction tuning on a telco-generated dataset suffices; continuing pretraining on raw telco text adds little (max +0.03 accuracy).

Pith tools