{"id":"8c48627f-764e-4953-8773-47cda27a35aa","arxiv_id":"2508.11988","paper_version":2,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"The full text introduces FutureX, a live contamination-free evaluation benchmark for LLM agents on future prediction tasks, but it does not match the submitted abstract about event-based micro-expression analysis.","lead":"The submission's metadata describe an event-camera micro-expression dataset, but the full text is a different paper, 'FutureX', on a live benchmark for LLM future prediction. Because the manuscript does not match its title or abstract, the advertised findings cannot be assessed from this submission.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The leaderboard depends on an unaudited LLM answer-extraction pipeline: the reported >97% is a fetch-success rate, not a label-accuracy rate, so ground-truth noise could be driving the reported rankings.","rationale":"The reader's verdict is already UNVERDICTED, and this stress-test does not move it: the advertised micro-expression paper is absent from the submitted full text, so the abstract's claims cannot be checked at all. Judged on its own terms, the FutureX text's central empirical claim is that its leaderboard reflects genuine agent forecasting ability. That claim rests on the automated answer-acquisition pipeline producing correct, unbiased ground truth. The paper asserts a >97% acquisition success rate but never reports extraction accuracy, label-error rates, or any audit that would rule out systematic mistakes. Since every score, ranking, and factor-analysis coefficient is computed against these labels, even a small error rate could alter close rankings or shift domain-level conclusions. This is precisely the weakest assumption the reader identified, and it is load-bearing. The proposed human-audit test would settle the question: if label noise is small and uniform, the concern does not land; if it is large or correlated with model or website, the leaderboard cannot be trusted as a measure of forecasting skill. No other concern is more central: website-selection bias is secondary to labeling integrity, and the paper's own admitted limitations (e.g., human/model question sets not aligned) are explicit caveats rather than hidden failure points.","tokens_in":36131,"tokens_out":5727,"duration_ms":66155,"concrete_test":"Randomly sample 200–300 resolved events from the July 20–August 3 window, stratified by event type (single-choice, multi-choice, ranking, numeric) and source category. For each event, reconstruct the exact prompt and have two independent human annotators re-derive the answer from archived source-page captures at the resolution-date crawl times (14:00/16:00/18:00/20:00), without seeing the pipeline label. Measure inter-annotator agreement and the pipeline's label error rate (where the pipeline answer differs from both annotators). Then recompute the per-model overall scores with corrected labels and compare the leaderboard. If the top-5 ordering or the reported margins change by more than a few points, the original rankings are not robust to label noise; if the error rate is below 1% and uniformly distributed across domains and answer formats, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"FutureX's central claim—that its leaderboard measures agent forecasting ability—requires that the automatically extracted ground-truth answers are correct and unbiased. Section 3.2.4 reports a >97% 'answer acquisition success rate' in the stable version, but acquisition success is not label accuracy: it counts events for which some answer was scraped, not events for which the scraped answer is correct. The pipeline's only quality control for extraction failures is manual review and prompt tweaking, and no error rate, confusion matrix, or independent audit is reported. If extraction errors are non-negligible (say >3–5%) or correlate with answer format, website, or domain (e.g., numeric ranking pages vs. narrative news pages), then the scores of all 25 models inherit that noise, and the reported leaderboard (Grok-4 > Gemini-2.5-flash Deep Research > GPT-o4-mini) and the factor analysis in Section 4.4 are not trustworthy as measures of forecasting skill. This assumption is load-bearing because every downstream result—difficulty-tier findings, domain comparisons, planning/search analyses, and the human-expert comparison—is computed against these labels. The paper provides no evidence that label noise is small or unbiased.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript, as provided in the body, presents FutureX, a live benchmark for evaluating LLM agents on future prediction. The pipeline collects future-event questions from 195 curated websites, runs 25 models (base LLMs, think-and-search LLMs, open- and closed-source deep research agents) at each event's start date, then automatically scrapes the resolved outcome and scores the stored predictions. The authors report overall scores and analyses across four difficulty tiers, eleven domains, and a factor regression, plus three out-of-benchmark case studies on financial forecasting, fake-website vulnerability, and real-time information retrieval. The stated ambitions are that FutureX is the largest and most diverse live future-prediction benchmark, that it is contamination-free by design, and that the automated answer-acquisition pipeline supports reliable evaluation with a reported success rate above 97%. An additional human-expert comparison is presented as a rough benchmark of human performance. As submitted, there is a major front-matter discrepancy: the abstract and title describe a different paper on event-based facial micro-expression analysis, while the full text describes FutureX.","tokens_in":36321,"tokens_out":6153,"duration_ms":62203,"significance":"If the benchmark's validity holds, FutureX would be a valuable community resource. A live, prospective benchmark for agent future prediction addresses the known pitfalls of backtesting and retrieval contamination, and the paper's scale (25 models, daily updates, 1,272 events over two weeks) is substantially larger than prior work such as FutureBench. The open-ended and high-volatility event types, the case studies on adversarial websites, and the planning/search analyses are useful and go beyond simple leaderboard reporting. However, the significance is contingent on two key assumptions: that the automatically extracted ground-truth answers are correct and unbiased, and that the hand-chosen scoring parameters do not drive the reported rankings. The paper provides no direct evidence for the first assumption, and only weak internal evidence for the second. The front-matter mismatch also prevents the manuscript from being assessed as a coherent submission in its present form.","major_comments":[{"comment":"The reported >97% answer acquisition success rate is not an answer-accuracy rate; it counts events for which a scraped answer was obtained, not events for which the extracted answer is correct. The pipeline uses Seed1.5-Thinking to extract the precise answer, and extraction errors are handled by manual review and prompt tweaking, but no error rate, confusion matrix, or independent audit is reported. Because every downstream result (overall leaderboard, difficulty-tier comparisons, domain analyses, factor regression, human comparison) is computed against these labels, systematic extraction errors correlated with domain or answer format would directly corrupt the reported rankings. The authors should provide a manually validated sample of extracted answers with accuracy broken down by event type and website category, and should demonstrate that label noise is small and uncorrelated with model identity or search behavior.","section":"§3.2.4"},{"comment":"The manuscript's front matter identifies it as 'Exploring Spatial-Temporal Dynamics in Event-based Facial Micro-Expression Analysis' (arXiv:2508.11988, cs.CV), but the full text is 'FutureX: An Advanced Live Benchmark for LLM Agents in Future Prediction' (cs.AI). This is not a typographical issue but an inconsistency in the paper's core identity: the abstract, title, and application domain do not match the content. The authors must correct the title, abstract, and metadata to match the body, or submit the intended manuscript. As submitted, the document cannot be assessed as a coherent paper.","section":"Title and Abstract vs. Full Text"},{"comment":"The evaluation protocol introduces several hand-chosen parameters with no sensitivity analysis: the tier weights (10%/20%/30%/40%), the 80% partial credit for ranking overlap, the 1-standard-deviation tolerance for numerical predictions, and the 7-day window for computing that standard deviation. The overall leaderboard and the factor-analysis conclusions depend on these choices. The authors should show that the ranking of models and the main qualitative findings (e.g., Grok-4 first, search-augmented models better on harder tiers) are stable under reasonable perturbations of these parameters, such as equal tier weights, partial credit between 70% and 90%, tolerance between 0.5 and 1.5 standard deviations, and windows between 3 and 14 days. Without such robustness checks, the quantitative comparisons may partly reflect metric choices.","section":"§3.4.3 and §4.1"},{"comment":"The claim that 'our difficulty tiers accurately reflect the complexity of the events' is supported by observing that model performance declines across the four tiers. This is a circular consistency check: the tiers were defined by event type and volatility, and the same model scores are used both to validate the tiers and to report the benchmark's main results. Independent validation is needed, for example human-rated difficulty, calibration against prediction-market prices, or item-response-theory analysis of the event bank. Without external evidence, the tier ordering and the tier-weighted overall score should be treated as a convention rather than a validated difficulty scale.","section":"§4.2, Finding 1"}],"minor_comments":[{"comment":"The submitted text contains multiple garbled or missing tokens, such as 'the������� ��� ���� �������' in Section 1 and '�' placeholders throughout Section 3. These appear to be rendering or conversion artifacts from the PDF-to-text process. The authors should provide a clean, complete version of the manuscript so that all content is legible.","section":"Global rendering"},{"comment":"The Related Work text states FutureBench includes a 'very small number of events (≥30)' while Table 1 reports 42 events. The text should be made consistent with the table.","section":"§2 vs. Table 1"},{"comment":"The authors already note that the human comparison uses a different question set from the model evaluation and label the comparison as 'rough indicators.' Given this limitation, the subsequent statement that 'humans significantly outperform LLM agents on Level 1, Level 3, and Level 4' should be rephrased as a preliminary observation, or the evaluation should be repeated on a matched question set.","section":"§4.2, Finding 5"},{"comment":"The footnote explaining that GPT-o1, GPT-o3, and GPT-o4-mini were excluded because they frequently refused to make predictions is useful, but the authors should also state whether any of the other 25 models had systematic refusal or missing-prediction patterns that could bias the missingness analysis in §3.4.2.","section":"§3.2.3"},{"comment":"In the event-filtering description, the text says between July 15 and July 22, 294 unsuitable events were dropped as '(both harmful and subjective events).' The phrase 'both' is ambiguous because the preceding sentence lists three categories; the authors should clarify whether the 294 figure covers harmful and subjective only, or also some trivial events.","section":"§3.2.2"}],"recommendation":"major_revision","confidential_remarks":"The front-matter mismatch (micro-expression abstract vs. FutureX body) is unusual and may indicate a submission-system error. The editor should verify with the authors that the intended submission is the FutureX paper; if so, the metadata must be corrected. The benchmark is developed primarily by ByteDance and uses ByteDance's Seed1.5-Thinking for answer extraction and Gemini-2.5-pro for plan-quality assessment; these self-dependencies are disclosed but should be scrutinized for potential bias. The central technical concern, however, is not the affiliation but the absence of any label-accuracy audit for the automatically extracted ground truth. That issue, plus the lack of sensitivity analysis for the scoring parameters, justifies major revision rather than acceptance at this stage."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things upfront. First, the submission is not a coherent paper: the title, abstract, and author list are about event-based facial micro-expression analysis, but the full text is a completely different manuscript, a live LLM forecasting benchmark called FutureX. That is a desk-reject-level structural problem on its own. Second, the FutureX text, taken on its own merits, is a serious and potentially useful benchmark, but it has a load-bearing validity gap that the paper does not address.\n\nWhat is actually new in the FutureX content: a live, multi-domain forecasting benchmark that combines automated event curation from 195 web sources (filtered from 2,008), daily/weekly updates, four difficulty tiers, evaluation of 25 models including deep research agents, and several out-of-benchmark case studies (fake-website injection, real-time search, comparison with Wall Street analysts). The design principles—contamination-free, live, forward-looking—are well motivated, and the comparison table with ForecastBench and FutureBench is informative. The difficulty tiers are a reasonable stratification, and the case studies are a genuine addition beyond the leaderboard.\n\nWhere the soft spots are, in proportion: the biggest is the identity mismatch; no editor should send this to review as is. Next is the answer-extraction pipeline. The paper claims a >97% answer acquisition success rate, but that is a fetch-success rate, not a label-accuracy rate. The stress test note is right: extraction errors are not quantified, and if they are systematic (e.g., numeric ranking pages vs. narrative news), the entire leaderboard and the factor analysis inherit that noise. The paper provides no error analysis, confusion matrix, or independent audit. That is a serious weakness for any claim that the rankings measure forecasting skill. Also, the human expert comparison is acknowledged by the authors as rough, but the acknowledgment is buried; the claimed human-vs-model gaps should be treated as indicative only. The scoring weights (10/20/30/40) and numerical tolerance are hand-chosen; not fatal, but they deserve ablation. Finally, the PDF has many garbled, corrupted passages, which is consistent with a hastily assembled manuscript.\n\nWho is this for? The FutureX benchmark, if properly cleaned and audited, would be of real interest to the LLM-agent evaluation community. It fills a gap that ForecastBench and FutureBench only partially cover. But as submitted, no reader can responsibly evaluate the advertised micro-expression contribution, and the FutureX contribution is not yet evidence-backed enough to act on.\n\nMy recommendation: desk reject this submission with a clear note about the manuscript mismatch. Encourage the authors to resubmit the FutureX paper separately, with an explicit label-accuracy audit for the answer-extraction pipeline and a more careful human-comparison protocol. If those are fixed, it deserves serious peer review.","headline":"This submission cannot be reviewed as submitted: the title and abstract describe a micro-expression paper while the full text is an entirely different future-prediction benchmark, and that benchmark itself needs an audit of its answer-extraction pipeline before its leaderboard can be trusted.","tokens_in":36909,"tokens_out":1959,"would_cite":false,"duration_ms":21505,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A live, contamination-free benchmark for AI future forecasting is feasible, and it shows search-and-reasoning agents in the lead while humans keep the edge on hard events.","keywords":["future prediction","LLM agents","live benchmark","data contamination","deep research agents","forecasting evaluation","web search agents","event prediction"],"falsifier":"Take a random sample of resolved events from one week and hand-check the pipeline's extracted answer against the archived content of the source page on the resolution date, then recompute each model's score using only the corrected answers; if extraction accuracy is materially below the claimed 97%, or if correcting errors changes the leaderboard, the benchmark's validity claim fails. A second decisive check is to run a trivial baseline that predicts the current observed value for every open-ended numerical event and compare it with the agent scores on the volatility-based metric.","tokens_in":35879,"feed_emoji":"🔮","tokens_out":9608,"duration_ms":95130,"temperature":0.7,"pith_summary":"FutureX is a live benchmark that evaluates LLM agents by asking them to predict real, time-stamped future events—such as stock prices, rankings, sports outcomes, and data releases—before the answers exist anywhere. The paper's claim is that this forward-looking design is the only methodologically sound way to measure forecasting ability, because any retrospective test leaks the outcome into the model's training data and into search results. The benchmark runs on a semi-automated daily pipeline: roughly 500 events per week are curated from 195 websites, 25 models answer on each event's start date, and after the resolution date an automated crawler retrieves the ground truth and scores every stored prediction. The paper reports that search-and-reasoning models dominate the leaderboard, with Grok-4 first, while a base model wins the easy retrieval tiers, and that even the best agents trail human experts on the hardest open-ended events. The work matters because it proposes a standard for gauging whether autonomous agents can do the analytical forecasting that human professionals do in finance, politics, and economics—and it documents how far they currently fall short.","feed_headline":"Live benchmark scores AI agents on future events before they happen","feed_subtitle":"FutureX adds fresh questions daily so answers can't leak into training data—yet top agents still trail human experts.","key_machinery":"The load-bearing mechanism is the fully automated, closed-loop daily pipeline with four stages. Event database construction starts from 2,008 candidate websites collected by the AIME agent, filters them through LLM checks and manual review to 195 sources. Future event daily curation turns these sources into roughly 500 weekly prediction questions via prediction-market crawling and randomized question templates, then filters out harmful, subjective, and binary events. Agent daily prediction runs 25 models on 70-100 events per day with a 30-minute cap per question. Answer daily acquisition crawls each resolved event's source up to four times daily and uses the Seed1.5-Thinking model to extract the answer, achieving a claimed acquisition success rate above 97%. The validity argument is carried by the four-tier difficulty stratification (Basic, Wide Search, Deep Search, Super Agent) with tier-specific metrics—0-1 accuracy, F1 for multi-choice, set-overlap partial credit for rankings, and volatility-adjusted tolerance for numerical forecasts—and by the one-week evaluation delay that makes it impossible to optimize against recent feedback.","core_discovery":"The central discovery, stated on the paper's own terms, is that future prediction can be turned into a rigorous, contamination-proof agent evaluation: because ground truth does not exist at prediction time, the benchmark is \"contamination-impossible by design.\" On the first two weeks of operation (1,272 events, July 20 to August 3), model performance declines monotonically across the four difficulty tiers, search and reasoning become the decisive capabilities at the harder tiers, and the top model (Grok-4) beats both open- and closed-source deep research agents while remaining below human expert scores on three of the four tiers. The paper additionally shows that most deep research agents can be steered toward a false outcome by a fabricated webpage, that agents retrieve resolved \"past\" outcomes far better than they predict unresolved ones, and that no tested model beats professional sell-side analysts on more than about a third of financial forecasting tasks.","pith_inferences":["Because the benchmark is live by design, its contamination protection applies at prediction time but erodes as resolved answers accumulate on the public web and enter future training corpora; the daily randomization and fresh event supply are what actually preserve it over time.","A persistence baseline that simply repeats the current observed value for numerical events would test whether the volatility-based scoring rewards genuine forecasting or recency tracking; the paper does not report such a baseline.","The human-expert comparison used a different question set from the model runs; re-running the same 300 questions through both would give a cleaner estimate of the human-model gap.","The fake-website attack protocol, with its LLM-generated pages and iterative feedback loop, could be packaged as a reusable adversarial robustness test for research agents generally."],"forward_implications":["A contamination-free, live evaluation of forecasting agents is feasible at scale and on a daily schedule, not just for static knowledge benchmarks.","Search and reasoning, not model size or internal knowledge, are the capabilities that separate agents on future prediction; the largest gains appear exactly where the benchmark becomes open-ended.","The four-tier scoring with 10/20/30/40 weights gives a stable leaderboard, and the factor analysis (R-squared = 0.418) supports treating difficulty and domain as first-order drivers of performance.","No agent in the study is yet a reliable financial forecaster: even the best models beat professional sell-side analysts on at most 37.5% of revenue and 32.3% of EPS predictions.","Deep research agents are vulnerable to adversarial web content; three of four tested closed-source agents were consistently deceived by fabricated pages, so robustness to misinformation must be part of any deployment claim."],"supporting_citations":[{"why":"The prior live forecasting benchmark (prediction-market events, vanilla LLMs, multiple choice) that FutureX expands in scale, domain coverage, and agent evaluation.","marker":"[15]"},{"why":"The pitfalls study that motivates the prospective design by showing retrospective forecasting evaluation suffers temporal leakage and retrieval contamination.","marker":"[22]"},{"why":"GAIA, the real-world-question agent benchmark whose blend of reasoning, multi-modality, and web browsing FutureX extends to open-ended future events.","marker":"[18]"},{"why":"BrowseComp, the inverted-question search benchmark that sets the persistence and multi-hop retrieval standard compared against in FutureX's search analysis.","marker":"[33]"},{"why":"Seed1.5-Thinking, the LLM that performs website screening, event filtering, and answer extraction, on whose reliability the >97% answer-acquisition claim rests.","marker":"[27]"},{"why":"The autonomous multi-agent system AIME, used to collect the initial pool of 2,008 candidate websites.","marker":"[28]"},{"why":"SmolAgent, the open-source deep research framework whose visible planning memory grounds the planning-quality factor analysis.","marker":"[25]"},{"why":"Gemini Deep Research, the closed-source deep research agent whose refusal to cite a fabricated page anchors the fake-website robustness case study.","marker":"[7]"}],"fun_headline_variants":["Event cameras lift micro-expression accuracy to 51% from 23%","New event-based dataset for micro-expressions: SNN hits 51%","Event cameras reveal micro-expressions better than RGB in new study","Micro-expression analysis: event-based data outperforms RGB by 28%","Spiking neural networks achieve 51% on event-based micro-expressions"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the automated pipeline—LLM-based event curation and answer extraction, with a claimed success rate above 97%—produces accurate, unbiased ground truth, and that the 195 curated websites give a balanced and representative sample of forecastable events; if extraction is systematically wrong or the sources are skewed, the scores do not measure forecasting ability.","fun_headline_variants_meta":{"raw":{"variants":["Event cameras lift micro-expression accuracy to 51% from 23%","New event-based dataset for micro-expressions: SNN hits 51%","Event cameras reveal micro-expressions better than RGB in new study","Micro-expression analysis: event-based data outperforms RGB by 28%","Spiking neural networks achieve 51% on event-based micro-expressions"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000703,"raw_usage":{"total_tokens":3152,"prompt_tokens":906,"completion_tokens":2246,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":522,"completion_tokens_details":{"reasoning_tokens":2151}},"tokens_in":522,"tokens_out":2246,"duration_ms":16704,"temperature":1.0,"reasoning_tokens":2151,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T17:26:55.543795+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a random sample of resolved events from one week and hand-check the pipeline's extracted answer against the archived content of the source page on the resolution date, then recompute each model's score using only the corrected answers; if extraction accuracy is materially below the claimed 97%, or if correcting errors changes the leaderboard, the benchmark's validity claim fails. A second decisive check is to run a trivial baseline that predicts the current observed value for every open-ended numerical event and compare it with the agent scores on the volatility-based metric.","supporting_citations":[],"review_version":1}