{"id":"59c02571-41d4-493e-8d7a-414ad05e013b","arxiv_id":"2508.17580","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":7.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":2,"one_line_summary":"Unsolved Stack Exchange questions can serve as a dynamic benchmark: the best tested model passes validator screening on only 15% of 500 questions.","lead":"This paper introduces UQ, a set of 500 genuinely unsolved questions taken from Stack Exchange, and tests whether AI models can answer them. It reports that the strongest model passes automated validation on only 15% of these questions, offering a new way to benchmark AI on real, open problems instead of exams.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Validator precision is uncalibrated: without ground truth, the 15% pass rate may measure agreement with LLM validators rather than correctness, so UQ's central claim is not yet established.","rationale":"Good-faith reading: the paper proposes a genuinely novel asynchronous benchmark with community verification, and releasing the dataset and platform is a concrete contribution. The concern is not that unsolved questions are illegitimate in principle; it is that the headline quantitative claim cannot be interpreted until the validator's operating characteristics are measured. This is an internal-validity question, not a disagreement with consensus. The reader flagged the same assumption—that validator pass implies correctness—so I agree. I do not move the verdict because the supplied full text is garbled and the missing analyses could exist in the unreadable sections; the appropriate status remains unverified. The concrete test would settle it: a calibration set of questions that were unsolved at training time but have since received accepted answers, run through the exact validator pipeline, plus blind expert review of a random sample of actual passes. Absent that, the 15% figure is uninterpretable as evidence about model ability on unsolved questions.","tokens_in":69722,"tokens_out":4224,"duration_ms":48034,"concrete_test":"Build a calibration set from Stack Exchange questions that were unsolved at model training time but have since received accepted, community-verified answers. Run the exact UQ-Validator pipeline on model-generated answers for those questions, treating the accepted answers as ground truth, and report precision and recall at the deployed threshold, both overall and per question category. Additionally, take a random sample of 50 validator-passing candidates from the actual UQ benchmark and have two independent domain experts blindly judge correctness, reporting inter-annotator agreement and the false-positive rate. If precision at the deployed threshold is not high or expert agreement is low, the 15% pass rate cannot be read as a difficulty measurement.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that UQ is 'difficult and realistic by construction' and that 'the top model passes UQ-validation on only 15% of questions'—turns on what UQ-validation actually measures. Because every UQ question is unsolved, there is no ground-truth answer against which the validator's positive calls can be checked. The abstract's 'compound validation strategies that leverage the generator-validator gap' therefore must supply a proxy for correctness. If that proxy is built on LLM judges or self-consistency, shared failure modes—sycophancy, style preference, hallucination, reward hacking—can make a wrong answer pass or a right answer fail. A low pass rate is equally consistent with genuine difficulty, validator over-strictness, or questions that are unanswerable or ill-posed despite curation. The paper's 'preliminary human verification' is existence evidence only: without a random sample of passes reviewed blindly by independent experts, or a calibration set of questions with known accepted answers, the precision of UQ-Validators is unknown. If precision is low, the 15% figure measures validator agreement rather than model ability, and the claimed paradigm's usefulness is unsupported. The curation filters ('well-defined and difficult') are themselves LLM-judge outputs, so the 'difficult and realistic by construction' claim inherits the same circularity.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a new evaluation paradigm, UQ, in which language models are assessed on 500 unsolved questions curated from Stack Exchange. The authors contribute three artifacts: UQ-Dataset and its collection pipeline (rule-based filters, LLM judges, and human review), UQ-Validators (compound validation strategies built on the generator-validator gap), and UQ-Platform for asynchronous community verification. The central empirical claim is that the top model passes UQ-validation on only about 15% of questions, with preliminary human verification confirming some of those passes as correct answers. The paper argues that because the questions are unsolved and naturally arising, the benchmark is 'difficult and realistic by construction' and offers a path toward evaluating frontier models on open-ended, real-world problems.","tokens_in":70008,"tokens_out":2144,"duration_ms":24602,"significance":"If the claims are established, UQ would be a genuinely useful complement to static benchmarks: it has ecological validity, an asynchronous community-verification mechanism, and a concrete public artifact. The paper is also commendable for releasing the dataset and platform, and for making a falsifiable quantitative claim (the 15% pass rate) that can be re-measured. However, the significance currently rests on an unvalidated proxy for correctness: since the questions are unsolved, there is no ground truth against which the validator precision can be checked, and the reported pass rate may partly reflect agreement with the validator rather than model ability. The strength of the contribution therefore depends on additional calibration evidence that is not yet reported.","major_comments":[{"comment":"The headline result, 'The top model passes UQ-validation on only 15% of questions,' is reported without any uncertainty estimate or statistical framing. For a 500-question testbed, even a simple 95% confidence interval would be informative, and the manuscript should also report the raw counts of passes, fails, and abstentions. Without this, readers cannot distinguish a true 15% ability level from a noisy estimate.","section":"Abstract"},{"comment":"Validator precision is uncalibrated, and this is load-bearing. Because every UQ question is unsolved by construction, 'passing UQ-validation' is defined by the paper's own compound validator strategies; there is no external ground truth. The abstract states that 'preliminary human verification has already identified correct answers among those that passed,' but that is existence evidence only. The paper needs a blind, independent review of a random sample of validator passes and validator fails, with inter-annotator agreement reported, to estimate precision and recall. Without such calibration, the 15% figure may measure agreement with the validator rather than genuine problem-solving ability.","section":"UQ-Validators and Abstract"},{"comment":"The claim that UQ is 'difficult and realistic by construction' is an assertion, not a measured property. The curation pipeline uses rule-based filters, LLM judges, and human review to ensure questions are 'well-defined and difficult,' but no validation statistics are reported for these filters (e.g., agreement between LLM-judge screening and human review, or pass-rate comparisons against existing benchmarks). Since the curation filters are themselves LLM-judge outputs, the 'difficult' claim inherits the same circularity as the validator: it needs an independent anchor such as human difficulty ratings or a comparison set of solved questions with known answer distributions.","section":"Abstract / Contribution (1)"},{"comment":"The generator-validator gap is the central methodological assumption of UQ-Validators, but the paper provides no evidence that this gap is a reliable correctness signal on unsolved questions. If the generator and validator share failure modes (sycophancy, style preference, hallucination, or reward hacking), a wrong answer can pass and a right answer can fail. The manuscript should specify each validator strategy, report the agreement and disagreement rates among them, and show that disagreement correlates with human judgments on a calibration set drawn from questions with known or later-confirmed answers.","section":"Abstract / generator-validator gap"},{"comment":"The paper does not report how the same models perform on existing solved benchmarks under the same evaluation protocol. Without such a baseline, the statement that the 15% pass rate demonstrates 'difficulty' is not quantified: a low pass rate is equally consistent with validator over-strictness or with questions that are unanswerable or ill-posed despite curation. At minimum, the authors should report pass rates on a matched set of solved Stack Exchange questions processed through the same validator pipeline.","section":"Results / comparison baselines"}],"minor_comments":[{"comment":"The manuscript contains many typographical and rendering artifacts in the provided text, including garbled characters and repeated fragments; the authors should carefully proofread the final version.","section":"Throughout"},{"comment":"Several tables appear without clear captions or explicit units; for example, the pass-rate tables should state the number of questions per category and the number of validator strategies applied.","section":"Tables"},{"comment":"The paper would benefit from an explicit comparison to prior benchmark-construction efforts that use human-LLM verification pipelines, such as human feedback for difficult evaluations, to clarify the novelty of the generator-validator gap approach.","section":"Related work"}],"recommendation":"major_revision","confidential_remarks":"The core idea is attractive and the artifacts are potentially useful, but the central 15% claim cannot be assessed without validator calibration and uncertainty quantification. I would be willing to look at a revised version that adds a calibration set, a blinded human-review sample, and a matched baseline; if those additions are infeasible, the manuscript's central claim should be substantially weakened."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"UQ is worth a serious referee, but the headline number needs more support before it becomes a result. The idea — evaluate models on questions no human has answered, score them asynchronously with validator screening and community verification — is a real departure from static exam-style benchmarks and real-user logs. The 500-question Stack Exchange dataset, the public release, and the framing around the difficulty-realism tension are useful contributions for evaluation researchers. If the full manuscript delivers what the abstract promises, this is the kind of artifact that shapes benchmark design.\n\nI could not check the full text because the supplied file is badly corrupted; my read rests on the abstract and the reader's report. That is not a strike against the authors, but it limits how hard I can push on method details. The reader's instinct is right: the central 15% pass rate is only as meaningful as UQ-Validators' precision. There is no ground truth for unsolved questions, so a pass is defined by the compound validator plus a small human check. If the LLM judges share failure modes or reward style, the pass rate can reflect validator agreement rather than correct problem solving. A low pass rate is also consistent with over-strict validators or ill-posed questions. The abstract says human review confirmed some correct answers, but that is existence evidence; without a blind random sample of passes, precision and recall are unknown. The curation filters that promise well-defined and difficult questions are themselves LLM-judge outputs, so the 'difficult by construction' claim inherits the same circularity.\n\nThose concerns are addressable, not fatal. Add a calibration set of questions with publicly accepted answers, report validator precision/recall and inter-annotator agreement, present confidence intervals on the 15%, and show a random subset of passes and failures to independent experts. The authors already ship the dataset and platform, so they are positioned to do this.\n\nWho benefits: benchmark builders, capability trackers, and anyone studying contamination-resistant evaluation. I would bring it to a reading group; I would cite it if the validation analysis holds. Recommendation: send to peer review, and assign a referee who cares about measurement validity, not just scale.","headline":"A genuinely new benchmarking paradigm on unsolved questions, but the 15% pass rate needs calibrated validators before the central claim is established.","tokens_in":70575,"tokens_out":2919,"would_cite":true,"duration_ms":32581,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that evaluating language models on genuinely unsolved real-world questions yields a benchmark that is both hard and realistic, and that the top model still fails most of it, passing only 15%.","keywords":["unsolved questions","language model evaluation","benchmark","generator-validator gap","Stack Exchange","asynchronous evaluation","community verification","frontier models"],"falsifier":"Collect a cohort of UQ questions that later receive human-accepted answers on the original sites, and score validator-passing model answers on those questions: if the pass rate on later-solved questions is no higher than on questions that remain unsolved, the validator signal is not tracking correctness.","tokens_in":69505,"feed_emoji":"❓","tokens_out":3340,"duration_ms":34328,"temperature":0.7,"pith_summary":"The paper proposes evaluating language models not on exam-style benchmarks with known answers but on genuinely unanswered questions drawn from real human askers. It introduces UQ, a 500-question testbed collected from Stack Exchange and spanning CS theory, mathematics, sci-fi, and history, together with a validation protocol that combines generated candidate answers, LLM judges, and human review. Its central finding is that the strongest model passes UQ-validation on only 15% of questions, while preliminary human verification has already confirmed some correct answers among the passes. If the paradigm works, it gives the field a benchmark whose difficulty cannot be gamed by training-data exposure and whose solutions carry direct real-world value.","feed_headline":"Frontier models pass only 15% of unsolved questions","feed_subtitle":"A 500-question benchmark drawn from real unanswered posts asks whether AI can solve problems nobody has solved yet.","key_machinery":"The load-bearing mechanism is the generator-validator gap: candidate answers are produced by a generator model and then screened by separate validator models, composed into compound strategies, so that the checkers do not simply share the generator's blind spots and can pre-select candidates for expert human review. Around this sit three components: UQ-Dataset, the filtered collection pipeline that turns raw unsolved Stack Exchange posts into well-defined hard questions; UQ-Validators, the compound validation strategies that supply the evaluation signal when no ground-truth answer exists; and UQ-Platform, an open platform for asynchronous expert verification. The 15% pass rate is the output of the validator stage, with human verification reserved for the surviving candidates.","core_discovery":"The central claim is that unsolved real-world questions form a viable evaluation paradigm that resolves the difficulty-realism tension: they are hard by nature and realistic because they arise from people actually seeking answers. Concretely, the paper curates 500 unsolved Stack Exchange questions, filters them through rule-based checks, LLM judges, and human review to ensure they are well-defined and difficult, and evaluates models asynchronously with validator-assisted screening followed by community verification. The headline result is that the top model passes UQ-validation on only 15% of questions, and early human verification has found correct answers among those that passed. The authors frame this as a path for assessing frontier models on open-ended challenges where success pushes the frontier of human knowledge.","pith_inferences":["If UQ were scaled up with automated tracking of which validator-passing answers later receive community acceptance, the benchmark could double as a live measure of whether validator quality tracks true correctness; this is an extension the paper leaves implicit.","The same generator-validator gap mechanism could be ported to domains without a built-in community, such as internal enterprise question queues, where 'unsolved' status is defined by the absence of a trusted answer.","A testable extension is to compare model pass rates on questions that later receive human-written accepted answers against those that remain unsolved: if the rates do not differ, the validator signal is measuring plausibility rather than correctness."],"forward_implications":["If the paradigm holds, benchmark difficulty can be sourced from human ignorance rather than constructed by test designers, so future models cannot memorize their way to high scores.","A model that solves an UQ question produces an answer that is already known to be wanted by a real asker, so progress on the benchmark translates into tangible utility.","The validation pipeline, rather than a static answer key, becomes the scoring instrument; as validators improve, the same question set can be re-scored asynchronously.","The 15% pass rate becomes a baseline for measuring whether future frontier models are actually expanding the space of answerable questions."],"supporting_citations":[],"fun_headline_variants":["AI passes just 15% on 500 real unsolved questions","Unsolved question benchmark: top AI at 15% pass rate","AI passes only 15% of real unsolved Stack Exchange queries","New benchmark uses unsolved questions – AI scores 15%","Top model passes 15% of unsolved questions in UQ test"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The benchmark assumes that an unsolved Stack Exchange question is well-defined and answerable, and that a candidate answer passing the validator screening is actually correct rather than merely persuasive.","fun_headline_variants_meta":{"raw":{"variants":["AI passes just 15% on 500 real unsolved questions","Unsolved question benchmark: top AI at 15% pass rate","AI passes only 15% of real unsolved Stack Exchange queries","New benchmark uses unsolved questions – AI scores 15%","Top model passes 15% of unsolved questions in UQ test"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000329,"raw_usage":{"total_tokens":1871,"prompt_tokens":1015,"completion_tokens":856,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":631,"completion_tokens_details":{"reasoning_tokens":765}},"tokens_in":631,"tokens_out":856,"duration_ms":7506,"temperature":1.0,"reasoning_tokens":765,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T17:03:32.288095+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Collect a cohort of UQ questions that later receive human-accepted answers on the original sites, and score validator-passing model answers on those questions: if the pass rate on later-solved questions is no higher than on questions that remain unsolved, the validator signal is not tracking correctness.","supporting_citations":[],"review_version":1}