Pith. sign in

REVIEW 10 cited by

RealTime QA: What's the Answer Right Now?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2207.13332 v2 pith:WO32PFNP submitted 2022-07-27 cs.CL

classification cs.CL
keywords realtimeresultsanswergpt-3informationretrievalansweringapplications
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We introduce REALTIME QA, a dynamic question answering (QA) platform that announces questions and evaluates systems on a regular basis (weekly in this version). REALTIME QA inquires about the current world, and QA systems need to answer questions about novel events or information. It therefore challenges static, conventional assumptions in open-domain QA datasets and pursues instantaneous applications. We build strong baseline models upon large pretrained language models, including GPT-3 and T5. Our benchmark is an ongoing effort, and this paper presents real-time evaluation results over the past year. Our experimental results show that GPT-3 can often properly update its generation results, based on newly-retrieved documents, highlighting the importance of up-to-date information retrieval. Nonetheless, we find that GPT-3 tends to return outdated answers when retrieved documents do not provide sufficient information to find an answer. This suggests an important avenue for future research: can an open-domain QA system identify such unanswerable cases and communicate with the user or even the retrieval module to modify the retrieval results? We hope that REALTIME QA will spur progress in instantaneous applications of question answering and beyond.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. BESPOKE: Benchmark for Search-Augmented Large Language Model Personalization via Diagnostic Feedback

    cs.CL 2025-09 conditional novelty 7.0 of 10

    BESPOKE provides 2,870 real user history sessions, 150 user-authored queries with gold information needs, and fine-grained human feedback, enabling evaluation and diagnosis of personalization in search-augmented LLMs.

  2. Will It Still Be True Tomorrow? Multilingual Evergreen Question Classification to Improve Trustworthy QA

    cs.CL 2025-05 conditional novelty 7.0 of 10

    EverGreenQA and EG-E5 provide a multilingual, human-labeled evergreen question classifier that improves self-knowledge estimation and QA dataset curation.

  3. Memorization vs. Reasoning: Updating LLMs with New Knowledge

    cs.CL 2025-04 conditional novelty 7.0 of 10

    KUP and MCT: a new benchmark and training method showing LLMs can memorize post-cutoff knowledge updates but fail to reason over them in indirect tests.

  4. Evaluating List Construction and Temporal Understanding capabilities of Large Language Models

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A new benchmark shows LLMs give incomplete lists and inaccurate time intervals for temporal list questions, and retrieval helps only partly.

  5. DailyQA: A Benchmark to Evaluate Web Retrieval Augmented LLMs Based on Capturing Real-World Changes

    cs.IR 2025-05 conditional novelty 6.0 of 10

    DailyQA, automatically built from Wikipedia revision logs, shows that LLMs with web retrieval answer only about half of time-sensitive questions correctly, and that reranking documents outperforms both snippets and ti...

  6. MedBrowseComp: Benchmarking Medical Deep Research and Computer Use

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A new benchmark of more than 1,000 multi-hop medical browsing questions shows that even the best deep-research and computer-use AI agents answer fewer than half correctly.

  7. MRAG: A Modular Retrieval Framework for Time-Sensitive Question Answering

    cs.CL 2024-12 conditional novelty 6.0 of 10

    A modular, training-free retrieval framework that decomposes time-sensitive questions into semantic content and temporal constraints, then ranks evidence by combined semantic and symbolic temporal scores, outperforms ...

  8. AntiLeakBench: Preventing Data Contamination by Automatically Constructing Benchmarks with Updated Real-World Knowledge

    cs.CL 2024-12 conditional novelty 6.0 of 10

    AntiLeakBench automatically constructs QA benchmarks from knowledge updated after each model's cutoff, and its experiments suggest that pre-cutoff evaluation overstates LLM ability.

  9. Language hooks: a modular framework for augmenting LLM reasoning that decouples tool usage from the model and its prompt

    cs.CL 2024-12 conditional novelty 6.0 of 10

    Language hooks interleave conditional external programs with an LLM's sentence-by-sentence generation, reaching competitive multi-tool QA and arithmetic performance without model fine-tuning or task-specific tool demo...

  10. NewsEdits 2.0: Learning the Intentions Behind Updating News

    cs.CL 2024-11 conditional novelty 6.0 of 10

    NewsEdits 2.0 introduces an edit-intention taxonomy and text-based models that predict factual updates in news revisions, enabling LLMs to abstain from answering with outdated facts at near-oracle accuracy.

Pith tools