REVIEW 13 cited by
Researchy Questions: A Dataset of Multi-Perspective, Decompositional Questions for LLM Web Agents
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
abstract
Existing question answering (QA) datasets are no longer challenging to most powerful Large Language Models (LLMs). Traditional QA benchmarks like TriviaQA, NaturalQuestions, ELI5 and HotpotQA mainly study ``known unknowns'' with clear indications of both what information is missing, and how to find it to answer the question. Hence, good performance on these benchmarks provides a false sense of security. A yet unmet need of the NLP community is a bank of non-factoid, multi-perspective questions involving a great deal of unclear information needs, i.e. ``unknown uknowns''. We claim we can find such questions in search engine logs, which is surprising because most question-intent queries are indeed factoid. We present Researchy Questions, a dataset of search engine queries tediously filtered to be non-factoid, ``decompositional'' and multi-perspective. We show that users spend a lot of ``effort'' on these questions in terms of signals like clicks and session length, and that they are also challenging for GPT-4. We also show that ``slow thinking'' answering techniques, like decomposition into sub-questions shows benefit over answering directly. We release $\sim$ 100k Researchy Questions, along with the Clueweb22 URLs that were clicked.
Forward citations
Cited by 13 Pith papers
-
A Large-Scale Study of Relevance Assessments with Large Language Models: An Initial Look
Fully automatic LLM relevance judgments rank TREC 2024 retrieval runs as well as human judgments, and human-in-the-loop assistance provides no clear extra benefit.
-
Relational Priors as Convergence Pressure in LLM-Based Multi-Agent Systems
Explicit positive or negative relationship cues in multi-agent prompts act mainly as convergence pressure, increasing agreement without reliably improving answer correctness.
-
LocalSearchBench: Benchmarking Agentic Search in Real-World Local Life Services
LocalSearchBench—1.3M merchant records and 900 multi-hop local-life QA tasks across 9 Chinese cities—shows the best reasoning agent reaches only 35.6% correctness.
-
Neural Prioritisation for Web Crawling
Prioritizing the crawl frontier with a neural quality estimator substantially improves early harvest rate and search effectiveness for natural language queries compared to breadth-first crawling.
-
Support Evaluation for the TREC 2024 RAG Track: Comparing Human versus LLM Judges
GPT-4o and human judges agree on RAG citation support labels 56% (from scratch) and 72% (with post-editing) of the time, with run-level correlations above 0.79 Kendall's tau.
-
The Great Nugget Recall: Automating Fact Extraction and RAG Evaluation with Large Language Models
A fully automatic LLM-based nugget evaluation for RAG systems matches human assessments at the run level on TREC 2024, with stronger agreement when only nugget assignment is automated.
-
FreshStack: Building Realistic Benchmarks for Evaluating Retrieval on Technical Documents
A framework that builds realistic, refreshable IR and RAG benchmarks on niche technical topics, with five datasets and large measured headroom for retrieval models.
-
Challenges in Trustworthy Human Evaluation of Chatbots
Open chatbot leaderboards like Chatbot Arena can have their model rankings moved by several positions with only 10% low-quality or adversarial votes.
-
Initial Nugget Evaluation Results for the TREC 2024 RAG Track with the AutoNuggetizer Framework
An automatic LLM-based nugget evaluation (AutoNuggetizer) correlates strongly with manual NIST evaluation at the run level (Kendall's tau = 0.783 over 21 topics), though topic-level agreement is weaker.
-
An Adversary-Resistant Multi-Agent LLM System via Credibility Scoring
A credibility-scoring framework for multi-agent LLM systems, learning agent trustworthiness on the fly and weighting outputs accordingly, improves accuracy under adversarial conditions in some benchmarks.
-
Chatbot Arena Meets Nuggets: Towards Explanations and Diagnostics in the Evaluation of LLM Responses
Nugget-based fact-recall scores on Search Arena battles correlate with human preferences, with about 54% agreement, and offer finer-grained diagnostic signals than raw win/loss votes.
-
Not Worth Another Token: Marginal Value Estimation for Efficient Deep Research Agents
Stage placement, not the scoring rule, dominates pruning effectiveness in deep research agents; early post-retrieval pruning cuts token usage by up to 73% with modest quality loss.
-
A Vision for Geo-Temporal Deep Research Systems: Towards Comprehensive, Transparent, and Reproducible Geo-Temporal Information Synthesis
A research agenda calling for geo-temporal reasoning in deep research systems, with no experiments or system implementation.
Discussion (0). Continue with ORCID to comment.