{"id":"c8d4edca-115e-4ba8-93c8-e73a4befce18","arxiv_id":"2607.20441","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"English news context systematically biases LLM predictions on Ukraine territorial markets toward Russian capture, and the bias originates in the text, not the model.","lead":"Using 111 Polymarket prediction markets on Ukraine, the authors show that giving LLMs English news context systematically pushes their territorial forecasts toward Russian capture, and these pushes are wrong 64–72% of the time—even for a model that already knows the real outcomes. This introduces a calibrated way to measure how much information ecosystems bias AI forecasts, relevant to anyone relying on LLMs for conflict understanding and decisions.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Push-accuracy headline compares to 50% without a base-rate baseline; in NO-heavy territorial markets the 64–72% 'wrong' rate may simply reflect the low frequency of upward price moves, so Table 1 does not yet show English text causes the error.","rationale":"I focused on the headline statistic because it is the paper's strongest quantitative claim and the one most likely to be cited. The 64–72% wrong-push rate is presented as evidence that English news induces directional error, invariant to model knowledge. But the test used—binomial vs 50%—implicitly assumes that, absent the push, upward and downward price moves are equally likely. The paper's own data contradict this: 62/65 territorial markets resolved NO, and the market reference is biased +3.5 pp toward capture. Under a low base rate of upward moves, a random upward push already has accuracy below 50%; the observed accuracies near 28–36% are exactly what the base-rate alone would predict. This is not a disagreement with consensus; it is an internal statistical issue: the null hypothesis is misspecified. The contaminated-model control, a key piece of evidence, inherits the same problem—if the subset is defined by D−A > 0, a model with outcome knowledge can still have upward-push accuracy near the base rate because short-horizon price direction differs from final resolution. I therefore do not think Table 1 establishes that the error originates in the text. The bias-shift tests (D−A) are better specified and may support a weaker version of the claim, but they do not justify the 'wrong 64–72%' language. The reader's concern about output probabilities as beliefs is real, but it is less crisply testable; the base-rate issue can be settled by reanalysis of the released dataset. For these reasons the paper should remain conditional, pending reanalysis of Table 1 with a proper base-rate or matched-push baseline. My recommendation is unchanged relative to the reader's verdict: CONDITIONAL, with the push-accuracy analysis as the key required revision.","tokens_in":14261,"tokens_out":10094,"duration_ms":108259,"concrete_test":"Recompute Table 1 with a proper baseline. For each model: (1) restrict to instances where D−A > 0; (2) compute the base rate R = P(sign(p_horizon − p_current) > 0) over those same instances; (3) test the D push accuracy against R (one-sided exact binomial), not against 0.5; (4) also compute the push accuracy for A and B in the same subset, defined as sign(p_model − p_current) > 0, to see whether D's upward pushes are worse than other conditions' upward pushes. If D's push accuracy is not significantly lower than R or than A's, Table 1's 'wrong 64–72%' is an artifact of the NO-heavy market and cannot be cited as evidence that English news induces bias. All needed data are already in the released dataset, so this is a pure reanalysis.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claim—English context pushes are wrong 64–72% of the time (Abstract; §4.2; Table 1)—rests on a binomial test against 50%. But the appropriate null is the base rate of upward price moves over the 7-day horizon in the selected subset. Territorial markets are overwhelmingly NO (62/65 resolve NO; §4.1) and the market itself has a +3.5 pp pro-capture bias, so the unconditional probability that sign(p_horizon − p_current) > 0 may be far below 50%. If the base rate is ~28–30%, the observed push accuracies (27.9–36.3%) are exactly what any upward-predicting rule would produce; the binomial p<10^-6 only shows that upward pushes are wrong, not that English news makes them wrong. The contaminated-model control is uninformative for this claim: a model with outcome knowledge will also have upward-push accuracy near the base rate if it makes any upward predictions, because short-horizon price direction is not the same as final outcome. To attribute the error to the text, the authors must compare D's upward-push accuracy to (a) the base rate in the same instances, (b) the upward-push accuracy of A or B, or (c) a matched set of downward pushes. The paper reports none of these. The bias-shift (D−A) and MAE results are separate and may survive, but the headline 'wrong 64–72%' is not supported by Table 1 as analyzed.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a method for quantifying information-ecosystem bias in LLM predictions, using Polymarket price trajectories as an external calibration reference. On 111 Ukraine-related markets (~93,000 predictions), five information conditions (A: blind; B: +chart; C: +English news; D: full English context; DUA: D + Ukrainian military sources) are run on three clean models plus a contaminated model whose training data include realized outcomes. The main claims are that English-language context systematically shifts territorial predictions toward Russian capture, that such pro-capture pushes are wrong 64–72% of the time (binomial p<10^-6), that a contaminated model shows the same push-error rate and therefore the bias originates in the text, and that adding Ukrainian military-analytical sources reduces the directional bias while MAE gains are partial and model-dependent.","tokens_in":14600,"tokens_out":6518,"duration_ms":72941,"significance":"If the central claims hold, the paper offers a novel, externally anchored way to quantify the cost of framing in LLM world models, with practical implications for multilingual RAG and forecasting. The design has genuine strengths: market-level clustering for the bias-shift tests, Bonferroni correction, permutation tests, a diplomatic-market placebo, per-horizon analyses, a contaminated-model control, and an honest limitations section. The released dataset of predictions and reasoning traces is a useful resource. However, the headline push-accuracy result currently rests on an inappropriate null hypothesis, and the attribution to 'English news text' is clouded by the composition of condition D. These issues are load-bearing and require reanalysis before the main quantitative claim is accepted.","major_comments":[{"comment":"The headline 'wrong 64–72%' compares upward/pro-capture push accuracy against a 50% binomial null. The appropriate baseline is the unconditional rate of upward price moves in the same market/horizon subset. With 62/65 territorial markets resolving NO and Polymarket carrying a +3.5 pp pro-capture bias, the base rate of sign(p_horizon − p_current) > 0 may be well below 50%. If it is ~28–30%, the observed 27.9–36.3% accuracies are close to what any upward-predicting rule would produce. The binomial p<10^-6 only shows that upward pushes are often wrong, not that English text creates the error. The contaminated-model comparison does not fix this: a model that knows final outcomes can still have low short-horizon direction accuracy because 7-day price direction is not the same as final binary resolution. The paper should report (a) the base rate of upward price moves in the same instances, (b)","section":"§4.2 / Table 1"},{"comment":"The paper states that 'all tests use market-level aggregates with cluster-robust inference,' but the push-accuracy p-values are binomial tests across individual predictions, despite ~44% overlap between adjacent cutoffs and intra-cluster correlations up to 0.141. This violates the stated unit-of-independence assumption and makes the p<10^-6 values unreliable. These tests are also explicitly excluded from the Bonferroni family. Please report a market-clustered test (e.g., cluster bootstrap or market-level accuracy means) and either include push accuracy in the multiple-testing family or justify the exclusion.","section":"Appendix B / Table 1"},{"comment":"The main claim is that 'English news' systematically biases predictions, but the push-accuracy table appears to be for condition D, which bundles English news with a price chart, a war map, and Polymarket trader comments. The introduction defines C as A + English news blocks, yet no C push accuracy is reported in Table 1. The conclusion that 'the bias originates primarily in the text' therefore conflates the full English-language ecosystem with news text. If the claim is about news, report condition C; if it is about the full ecosystem, revise the abstract and discussion. The contaminated model is run only in conditions A–D, so the source attribution cannot separate text from the other components of D.","section":"§3.1 / §4.2"},{"comment":"The measurement instrument assumes that 'the output probability is the induced belief,' justified by two LLM theory papers. This is load-bearing: if output probabilities partly reflect prompt formatting, instruction-following, or reasoning heuristics rather than text-induced beliefs, the measured 'framing cost' is not what is claimed. The contaminated model rules out model ignorance but not instrument infidelity. Add a validation study — for example, compare LLM probabilities under C/D against a human belief-elicitation benchmark, or show that prompt/order perturbations do not change the measured push error rates.","section":"Section 1 / §4.2"}],"minor_comments":[{"comment":"The abstract reports 'wrong 64–72%' while Table 1 reports accuracies of 27.9–36.3%. Stating the accuracy and its complement explicitly would avoid confusion.","section":"Abstract / Table 1"},{"comment":"The entry for 'advance into' reads '770∞'; this appears to be a formatting error for '77 / 0' and should be corrected.","section":"Appendix L, Table 9"},{"comment":"The bullet list has a typo: 'Aprovides' should be 'A provides'.","section":"§3.1"},{"comment":"The contaminated model is called 'Gemini 3.1 Pro Preview' in the text and 'Pro 3.1*' in tables; define this abbreviation at first use.","section":"§3.2 / Appendix J"},{"comment":"'Bias 2 accounts for only 2–9%' should use a proper superscript or spell out 'bias squared' for clarity.","section":"Appendix G"},{"comment":"The limitations section notes that DUA vs D MAE improvement is not significant and that MAE effects are underpowered; this should be reflected in the abstract's claims about 'accuracy gains'.","section":"Limitations"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses a politically sensitive topic, but my recommendation is based solely on the statistical and identification issues. The ethical and political framing sections are unusual for a CS methods paper but do not affect the verdict. If the authors can reanalyze the push accuracy against the appropriate base rate and cluster-level inference, the contribution could become publishable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe paper from arXiv:2607.20441 is worth knowing about because it tries something genuinely new: using LLMs as instruments to extract the beliefs a news corpus implies, and using prediction-market price trajectories as a calibration reference. That's distinct from both framing classification and LLM forecasting benchmarks. The authors build an ablation ladder (A→B→C→D→DUA), collect ~93k predictions across four models, and include a contaminated model that knows actual outcomes. The dataset and reasoning traces are a real asset.\n\nThe core finding—that adding English news context shifts territorial predictions toward Russian capture by about 1.4–2.4 pp and degrades MAE relative to a no-change baseline—is plausible and largely survives scrutiny, though the effects are modest and underpowered for two of the three clean models. The placebo (diplomatic markets) and the per-horizon analysis are nice touches.\n\nThe soft spots are serious. The headline \"wrong 64–72% of the time\" is not supported by the analysis as presented. The binomial test compares against 50% without considering the base rate of upward price moves. Territorial markets are overwhelmingly NO, so upward moves over a 7-day horizon are likely rare. If the unconditional probability of an upward move is around 30%, then observed push accuracies of 28–36% are exactly what any upward-predicting rule would produce. The paper never reports the base rate, and the contaminated model does not fix that: a model with outcome knowledge can still have low upward-push accuracy if it makes any upward predictions, because the 7-day price direction is not the same as the final resolution. To attribute the error to the text, the authors must compare push accuracy to the base rate in the same instances, to the push accuracy of condition A or B, or to a matched set of downward pushes. As it stands, Table 1 doesn't show that English news causes the errors.\n\nThe contaminated-model control is also weaker than it looks: it uses a different model (Gemini 3.1 Pro Preview) than the clean models (Gemini 2.5 Flash/Pro, GPT-5-mini), so architecture and knowledge are confounded. And the exclusion of push-accuracy tests from the Bonferroni family is a red flag.\n\nThe \"output probability is the induced belief\" assumption is cited but not validated. If LLM probabilities respond to formatting or instruction artifacts, the measured framing cost is not a clean readout of the text's belief.\n\nBottom line: the bias-shift and MAE results are worth taking seriously, and the method is a genuine contribution. But the headline number is not earned, and the control needs rework. I'd send it to peer review, expecting major revision. The right referee would ask for the base-rate analysis and a tighter control before the push-accuracy claim can stand.","headline":"A clever measurement idea that overreaches in its headline: the push-accuracy claim ignores base rates, but the bias-shift and MAE results are worth a careful look.","tokens_in":15096,"tokens_out":6300,"would_cite":true,"duration_ms":62553,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"English-language news context systematically biases LLM territorial predictions toward Russian capture, and these biased pushes are wrong 64–72% of the time, a distortion the paper argues originates in the sources, not the models.","keywords":["LLM beliefs","prediction markets","information bias","framing cost","territorial predictions","ablation study","belief elicitation","Ukraine conflict"],"falsifier":"A direct calibration experiment: have a panel of financially incentivized human forecasters read the same English news blocks and give probability updates under the same conditions; if human updates do not reproduce the 64–72% wrong-push rate relative to the market reference, then LLM output probabilities are not faithful readouts of text-induced beliefs, and the measured framing cost would be an artifact of the model rather than a property of the information ecosystem.","tokens_in":14154,"feed_emoji":"🎯","tokens_out":8972,"duration_ms":85942,"temperature":0.7,"pith_summary":"The paper tries to establish a calibrated way to measure how far the beliefs an information ecosystem induces deviate from an external reference. It treats an LLM as an instrument that reads out, in probability form, the beliefs implied by the text it is given, and uses prediction-market prices — anchored at resolution by real outcomes — as the calibration scale. Applied to 111 Ukraine-related markets and roughly 93,000 predictions from four models, the paper finds that adding English-language news context systematically shifts territorial predictions toward Russian capture, and those shifts are wrong 64–72% of the time. A contaminated model that already knows the actual outcomes shows the same error rate, which the paper takes as evidence that the bias lives in the text, not in the models. Supplementing the context with Ukrainian military-analytical sources reduces the bias across all clean models, though absolute-error gains are partial and model-dependent.","feed_headline":"News-pushed AI capture calls are wrong 64-72%","feed_subtitle":"Even a model that already knows outcomes shows the same failure rate — the bias is in the sources, not the models.","key_machinery":"The measuring instrument is an ablation ladder: the same model predicts under progressively richer information contexts — starting from bare market data, then adding a price chart, then English news articles, then the full English-language ecosystem, and finally augmented with Ukrainian military-analytical sources. The difference between conditions, expressed in percentage points relative to the prediction-market price trajectory, is the 'framing cost' of each text source. The decisive diagnostic is the push-error rate: when a context shift moves a prediction toward capture, does the later price path confirm it? The use of a contaminated model that already knows the outcomes serves as a cont","core_discovery":"The paper's central claim is that the English-language information ecosystem about the war in Ukraine carries a measurable pro-capture bias, and that this bias propagates through any model that processes it. The evidence is a push-error rate: when the addition of English news context moves an LLM's probability toward Russian territorial capture, that movement is later contradicted by the outcome-anchored market reference in 64–72% of cases across all four models, with binomial p values below 10^-6. The contaminated-model control — a model whose training data includes the outcomes — shows the same failure rate, which the authors argue isolates the text corpus as the source of the distortion r","pith_inferences":["One testable extension would be to apply the same push-error metric directly to news outlets, ranking them by how often their inclusion shifts model forecasts toward capture and then proves wrong; the paper's data hint that the most-cited analytical source correlates with worse directional accuracy.","The contaminated-model result does not separate two possible mechanisms — offense-dominant framing within the text versus exclusion of mitigating sources; a cleaner ablation would compare English-only context with Ukrainian-only context (without the English mix) to quantify the relative contribution of each.","If the instrument-fidelity assumption holds, the method could be used as a general 'belief calibration' audit for any text corpus on any question with a prediction market, not just war outcomes.","The paper's ethical discussion suggests the bias may influence human policy and public opinion; the method could theoretically be extended to measure human belief update rates on the same texts, though that goes beyond the current study."],"forward_implications":["English-language LLM forecasting of territorial conflicts will systematically overstate the attacking side's success unless the information diet is broadened.","Adding sources from the affected side (e.g., its military-analytical ecosystem) can reduce the directional bias, but the benefit varies by model and is not guaranteed to improve absolute error.","The bias is a property of the corpus, so any downstream system that consumes the same English news text will inherit it.","Model selection becomes a strategic lever: conservative reasoning may be preferable to deep reasoning when the text itself carries a directional bias.","The method offers a general, probability-calibrated measure of 'framing cost' that can be applied to other conflicts or contested topics."],"fun_headline_variants":["English news pushes LLM war calls wrong 64-72%","Bias in sources, not models: 64-72% error on capture calls","News context drives 64-72% of capture predictions into error","LLM bias originates in text: 64-72% wrong on capture","Pro-capture news bias survives outcome-aware LLM"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that the probability an LLM outputs is a direct, faithful measure of the belief the input text induces — an assumption adopted from prior work but not validated against an external belief measure such as human judgment.","fun_headline_variants_meta":{"raw":{"variants":["English news pushes LLM war calls wrong 64-72%","Bias in sources, not models: 64-72% error on capture calls","News context drives 64-72% of capture predictions into error","LLM bias originates in text: 64-72% wrong on capture","Pro-capture news bias survives outcome-aware LLM"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000811,"raw_usage":{"total_tokens":3398,"prompt_tokens":753,"completion_tokens":2645,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":497,"completion_tokens_details":{"reasoning_tokens":2551}},"tokens_in":497,"tokens_out":2645,"duration_ms":16901,"temperature":1.0,"reasoning_tokens":2551,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T14:14:04.932879+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A direct calibration experiment: have a panel of financially incentivized human forecasters read the same English news blocks and give probability updates under the same conditions; if human updates do not reproduce the 64–72% wrong-push rate relative to the market reference, then LLM output probabilities are not faithful readouts of text-induced beliefs, and the measured framing cost would be an artifact of the model rather than a property of the information ecosystem.","supporting_citations":[],"review_version":1}