Pith. sign in

REVIEW 13 cited by

Researchy Questions: A Dataset of Multi-Perspective, Decompositional Questions for LLM Web Agents

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.17896 v1 pith:6K2DTJ2H submitted 2024-02-27 cs.CL cs.AI

classification cs.CLcs.AI
keywords questionsansweringlikemulti-perspectiveresearchybenchmarkschallengingdataset
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
abstract

Existing question answering (QA) datasets are no longer challenging to most powerful Large Language Models (LLMs). Traditional QA benchmarks like TriviaQA, NaturalQuestions, ELI5 and HotpotQA mainly study ``known unknowns'' with clear indications of both what information is missing, and how to find it to answer the question. Hence, good performance on these benchmarks provides a false sense of security. A yet unmet need of the NLP community is a bank of non-factoid, multi-perspective questions involving a great deal of unclear information needs, i.e. ``unknown uknowns''. We claim we can find such questions in search engine logs, which is surprising because most question-intent queries are indeed factoid. We present Researchy Questions, a dataset of search engine queries tediously filtered to be non-factoid, ``decompositional'' and multi-perspective. We show that users spend a lot of ``effort'' on these questions in terms of signals like clicks and session length, and that they are also challenging for GPT-4. We also show that ``slow thinking'' answering techniques, like decomposition into sub-questions shows benefit over answering directly. We release $\sim$ 100k Researchy Questions, along with the Clueweb22 URLs that were clicked.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Large-Scale Study of Relevance Assessments with Large Language Models: An Initial Look

    cs.IR 2024-11 conditional novelty 7.0 of 10

    Fully automatic LLM relevance judgments rank TREC 2024 retrieval runs as well as human judgments, and human-in-the-loop assistance provides no clear extra benefit.

  2. Relational Priors as Convergence Pressure in LLM-Based Multi-Agent Systems

    cs.CL 2026-08 conditional novelty 6.0 of 10

    Explicit positive or negative relationship cues in multi-agent prompts act mainly as convergence pressure, increasing agreement without reliably improving answer correctness.

  3. LocalSearchBench: Benchmarking Agentic Search in Real-World Local Life Services

    cs.AI 2025-12 conditional novelty 6.0 of 10

    LocalSearchBench—1.3M merchant records and 900 multi-hop local-life QA tasks across 9 Chinese cities—shows the best reasoning agent reaches only 35.6% correctness.

  4. Neural Prioritisation for Web Crawling

    cs.IR 2025-06 conditional novelty 6.0 of 10

    Prioritizing the crawl frontier with a neural quality estimator substantially improves early harvest rate and search effectiveness for natural language queries compared to breadth-first crawling.

  5. Support Evaluation for the TREC 2024 RAG Track: Comparing Human versus LLM Judges

    cs.CL 2025-04 conditional novelty 6.0 of 10

    GPT-4o and human judges agree on RAG citation support labels 56% (from scratch) and 72% (with post-editing) of the time, with run-level correlations above 0.79 Kendall's tau.

  6. The Great Nugget Recall: Automating Fact Extraction and RAG Evaluation with Large Language Models

    cs.IR 2025-04 conditional novelty 6.0 of 10

    A fully automatic LLM-based nugget evaluation for RAG systems matches human assessments at the run level on TREC 2024, with stronger agreement when only nugget assignment is automated.

  7. FreshStack: Building Realistic Benchmarks for Evaluating Retrieval on Technical Documents

    cs.IR 2025-04 conditional novelty 6.0 of 10

    A framework that builds realistic, refreshable IR and RAG benchmarks on niche technical topics, with five datasets and large measured headroom for retrieval models.

  8. Challenges in Trustworthy Human Evaluation of Chatbots

    cs.HC 2024-12 conditional novelty 6.0 of 10

    Open chatbot leaderboards like Chatbot Arena can have their model rankings moved by several positions with only 10% low-quality or adversarial votes.

  9. Initial Nugget Evaluation Results for the TREC 2024 RAG Track with the AutoNuggetizer Framework

    cs.IR 2024-11 conditional novelty 6.0 of 10

    An automatic LLM-based nugget evaluation (AutoNuggetizer) correlates strongly with manual NIST evaluation at the run level (Kendall's tau = 0.783 over 21 topics), though topic-level agreement is weaker.

  10. An Adversary-Resistant Multi-Agent LLM System via Credibility Scoring

    cs.MA 2025-05 conditional novelty 5.0 of 10

    A credibility-scoring framework for multi-agent LLM systems, learning agent trustworthiness on the fly and weighting outputs accordingly, improves accuracy under adversarial conditions in some benchmarks.

  11. Chatbot Arena Meets Nuggets: Towards Explanations and Diagnostics in the Evaluation of LLM Responses

    cs.IR 2025-04 conditional novelty 5.0 of 10

    Nugget-based fact-recall scores on Search Arena battles correlate with human preferences, with about 54% agreement, and offer finer-grained diagnostic signals than raw win/loss votes.

  12. Not Worth Another Token: Marginal Value Estimation for Efficient Deep Research Agents

    cs.AI 2026-08 conditional novelty 4.0 of 10

    Stage placement, not the scoring rule, dominates pruning effectiveness in deep research agents; early post-retrieval pruning cuts token usage by up to 73% with modest quality loss.

  13. A Vision for Geo-Temporal Deep Research Systems: Towards Comprehensive, Transparent, and Reproducible Geo-Temporal Information Synthesis

    cs.CL 2025-06 conditional novelty 4.0 of 10

    A research agenda calling for geo-temporal reasoning in deep research systems, with no experiments or system implementation.

Pith tools