{"id":"e9266e14-de92-4657-a21f-149da00e21fd","arxiv_id":"2608.06549","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"TradeVerse is a longitudinal trade-negotiation benchmark from WTO minutes where current LLMs show a Western-country advantage in respondent identification, over-predict product codes, and generate generic final statements.","lead":"This paper introduces TradeVerse, a new benchmark built from 1,170 WTO trade-dispute records, and tests six language models on predicting product codes, identifying the responding country, and writing a country's final response. The models name respondents well overall but are systematically worse for non-Western members, over-predict product categories, and produce fluent but content-thin statements.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"ST-2 'full anonymization' is not validated; product-origin and measure-specific cues (e.g., 'China-made electric vehicles', 'presidential decree') can leak the respondent, so the Western-bias result may be an artifact.","rationale":"The strongest claim is conditional on the anonymization being complete. The paper's own example shows that even after masking, the identity of the respondent is recoverable from substantive details (the Turkish EV tariff decree). The authors' only evidence against identity-cue reliance is that ST-1 and ST-2 accuracies are nearly identical; but if both conditions leak, equality proves nothing. The statistical inconsistency between Table 4 and Table 8 independently weakens the same result: the main text's universal p<0.01 is contradicted by the appendix's own p-values. The dataset reconstruction pipeline is also unvalidated, but the masking question is the narrowest and most decisive point of failure for the headline geopolitical-bias finding. A targeted leakage audit can settle it. I do not see grounds to reject outright: the corpus is authentic, the tasks are well-defined, and the release is promised; the appropriate verdict remains conditional on the audit and on reconciling the statistics. This matches the reader's conditional verdict, so no change is needed.","tokens_in":14570,"tokens_out":6706,"duration_ms":65061,"concrete_test":"Run a leakage audit on a random sample of 200 Task-2 ST-2 inputs. First, paraphrase away all country-denominated product origins and unique measure references (e.g., 'China-made electric vehicles' -> 'imported electric vehicles'; 'presidential decree' -> 'government measure'), keeping the argumentative structure and all placeholder tokens unchanged. Second, rerun the six models on this sanitized subset and compare accuracy and the Western-minus-non-Western gap against the original ST-2 subset. If accuracy or the gap drops by more than 5 percentage points, the residual cues caused the headline result. As a complementary check, recompute the chi-square p-values for all twelve Table 4 comparisons from the released model predictions; if they match Table 8 rather than Table 4, the main text's 'p<0.01 everywhere' claim must be corrected before the bias finding is interpreted.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Task 2's ST-2 condition carries the paper's strongest interpretive claim: accuracy stays near 90% and the Western/non-Western gap persists when all country names are masked, which the authors read as evidence that models use substantive longitudinal reasoning rather than surface identity cues (Section 4.3, Tables 3-4). The paper never demonstrates that masking removes all identity-correlated information. The prompt in A.3 instructs models to use 'specific trade measures, tariff bound rates, presidential decrees, product sectors, dates, and legal arguments'; these are exactly the features most correlated with the respondent's identity. In the representative concern (Figure 2 and Figure 8), the dispute is about a 'presidential decree imposing a 40% additional tariff only on electric vehicles originating from China.' Even with every country name replaced by a placeholder, a model can answer 'Türkiye' from the presidential decree plus the EV tariff narrative, because this matches Türkiye's publicly known 2023 measure. Product-origin adjectives ('China-made electric vehicles'), national programs, and unique legal frameworks are country-specific residue that the described masking may not remove. The appendix's own EU-ablation (Table 8) shows that when EU samples are removed, DeepSeek-V4-Pro's ST-2 gap is not significant (p=0.0597) and several comparisons hover near 0.01-0.06, contradicting the main text's claim that 'every difference is significant at p<0.01.' Unless a leakage audit rules out such cues, the near-identical ST-1/ST-2 accuracy does not show models ignore identity; it may show enough country-correlated substance remains for identification.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript introduces TradeVerse, a benchmark constructed from 1,170 World Trade Organization Specific Trade Concerns reconstructed as longitudinal multi-party dialogues. It defines three tasks: (1) predicting the HS chapters of the products under discussion, (2) identifying the responding country from transcripts with either only the respondent masked (ST-1) or all country names masked (ST-2), and (3) generating the respondent's final statement. The paper evaluates six contemporary LLMs and reports high overall respondent-identification accuracy, a Western/non-Western accuracy gap that persists under ST-2 masking, over-prediction in HS-code identification, and moderate fluency but low specificity in generated statements. It claims to be the first benchmark for longitudinal political trade negotiations and releases the dataset, extraction pipeline, and evaluation code.","tokens_in":14839,"tokens_out":5791,"duration_ms":49240,"significance":"The benchmark addresses a genuinely underserved area: authentic institutional negotiation texts with a longitudinal structure, and all ground-truth labels are recovered directly from WTO proceedings without manual annotation. The release of the dataset, extraction pipeline, and evaluation code is a concrete and useful contribution. If the Western/non-Western gap under anonymization is real and robust, it is an important measurement of geopolitical bias in LLMs on authentic institutional text. The paper also provides useful evidence that current LLMs over-predict product categories and produce fluent but nonspecific diplomatic statements. However, the central interpretive claim about the persistence of the bias under full anonymization is not yet established, because the masking procedure does not demonstrably remove country-correlated substantive cues.","major_comments":[{"comment":"The claim that accuracy remains near 90% when all country names are masked, indicating that models rely on substance rather than identity, is not supported by the design of the masking. The prompt in Appendix A.3 (Task 2B) explicitly directs the model to use 'specific trade measures, tariff bound rates, presidential decrees, product sectors, dates, and legal arguments,' and the representative concern in Figures 2 and 8 shows that a 'presidential decree imposing a 40% additional tariff only on electric vehicles originating from China' uniquely identifies Türkiye's 2023 measure even when country names are replaced by placeholders. Country-specific residue such as 'China-made electric vehicles', national programs, and unique legal frameworks therefore remains in the 'anonymized' transcripts. The stability of accuracy between ST-1 and ST-2 does not establish that models use longitudinal substance rather than identity-correlated surface cues. The authors should provide a leakage audit: identify and remove, or at least measure, residual country-specific n-grams, and report Task 2 accuracy on a subset of concerns whose measures are not uniquely attributable to a single member.","section":"Section 4.3, Task 2 (ST-2); Appendix A.3, Task 2B"},{"comment":"Tables 4 and 8 report incompatible numbers for the same model/setting combinations. For example, DeepSeek-V4-Pro under ST-2 appears as Western 96.98%, non-Western 88.21%, χ²=24.09, p=1×10⁻⁶ in Table 4, but as Western 93.53%, non-Western 88.21%, χ²=3.55, p=0.0597 in Table 8; similar discrepancies occur for every model. The main text (Section 5) claims 'every difference is significant at p < 0.01', but the appendix's EU-ablation results include p=0.0597 (DeepSeek ST-2), p=0.0278 (GLM ST-2), and p=0.0061 (GLM ST-1). The authors need to clarify which table reports the full corpus and which reports the EU-ablation, reconcile the numbers, and correct the significance claims in Section 5 and the Conclusion.","section":"Table 4 vs. Table 8; Section 5"},{"comment":"The only description of the reconstruction is that the authors 'parsed these documents to organise countries and their roles and corresponding statements', and the paper does not report any validation of this parsing. The respondent and raiser role labels are the ground truth for Task 2 and the input structure for Task 3, so parsing errors propagate directly into all benchmark scores. At minimum, the paper should report the precision of the role-attribution parser on a manually labeled random sample of meetings, together with any filtering rules used to resolve ambiguous or missing role assignments.","section":"Section 3, Data collection"},{"comment":"The statement that the Western/non-Western gap 'remains when all participant identities are masked and cannot be attributed to surface cues' is stronger than the evidence, because the masking procedure leaves country-correlated substantive cues in place (see the major comment on Section 4.3). The conclusion should be reworded to claim only that the gap persists under the paper's masking procedure, pending the leakage audit.","section":"Section 6, Conclusion"}],"minor_comments":[{"comment":"The sentence 'for the first task, we evaluate the models’ performance by measuring Accuracy' is inconsistent with Table 2, which reports precision, recall, and F1; it appears that 'second task' is meant.","section":"Section 4.2"},{"comment":"The benchmark name is spelled inconsistently as 'TradeVerse', 'TRADEVERSE', and 'TradeVerseis' (e.g., in the abstract); please use one consistent spelling.","section":"Throughout"},{"comment":"Table 1 reports 1,170 concerns and 6,933 meeting records, while Appendix A.1 reports a 'cleaned corpus' of 1,101 concerns and 6,418 records; the paper should state which corpus is used for each table and why the counts differ.","section":"Table 1 vs. Appendix A.1"},{"comment":"The caption for Table 8 does not mention that it is the EU-ablation, so it initially appears to be a duplicate of Table 4; the ablation status should be in the caption and referenced explicitly from Section 5.","section":"Table 8 caption"},{"comment":"The figure caption in the appendix ends with 'shown abridged in All statements are verbatim from the WTO proceedings, with the boilerplate opening omitted.', and the figure reference is missing; please complete the sentence and add the figure number.","section":"Figure 8 caption"}],"recommendation":"major_revision","confidential_remarks":"To the editor: The benchmark is potentially a good fit for the journal, and the dataset release is valuable. The two central quantitative claims—the persistence of the Western/non-Western gap under full anonymization and the universal p<0.01 significance—rest on unresolved leakage and table inconsistencies. I would not reject the paper; a revised version with a leakage audit and reconciled statistics could be publishable. I also note that the manuscript cites several 2026 technical reports that may not be publicly available, and the authors should verify that all citations are verifiable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read this one if you care about LLM evaluation on institutional text. The benchmark itself is a real contribution: reconstructing 1,170 WTO trade concerns into structured multi-round dialogues with annotation-free labels is useful, and the three tasks (HS-code prediction, masked respondent ID, statement generation) are a sensible combination. The corpus stats look plausible, and the appendix does some honest work—the history-ablation and the EU-removal analysis are the kind of thing that makes me trust the authors' intent. So the raw material is good.\n\nThat said, the paper currently has problems that keep it from being acceptable as-is. Most glaring: Table 4 in the main text and Table 8 in the appendix report different percentages and p-values for the same model/setting combinations. The main text says every Western/non-Western difference is significant at p<0.01, but the appendix's own numbers show DeepSeek-V4-Pro at p=0.0597 for ST-2 and several others above 0.01. You cannot have two versions of the same result, and you cannot make a stronger claim than your own appendix supports. That needs to be reconciled before the bias finding is credible.\n\nSecond, the masking claim. The paper reads the near-identical ST-1/ST-2 accuracy as evidence that models use substance, not surface identity. But masking only replaces country names, demonyms, and agencies. It leaves \"China-made electric vehicles,\" \"presidential decree,\" tariff percentages, product sectors, and legal arguments—exactly the features the appendix prompt tells the model to use. The representative example (China vs. Türkiye on EV tariffs) is identifiable from the measure alone. So the stability under masking does not show what they claim. The stress-test note is right, and it lands on the paper's own example. A leakage audit—e.g., a simple classifier trained on the masked text—would settle it. Also, the parsing of the colored Word documents is described in one sentence; if role attribution is noisy, every downstream task suffers.\n\nMinor but relevant: no baselines (random, majority-class, or simple lexical models) for Task 2, and the dataset/code release is promised but no link appears anywhere. Those are easy fixes.\n\nVerdict: the resource deserves referee time, but the current preprint cannot be published as-is. I'd ask for (1) reconciled results tables, (2) a leakage audit for the masking, (3) baselines, and (4) actual artifact links. Fix those and this becomes a solid within-subfield contribution; leave them and the geopolitics finding is just an artifact in waiting.","headline":"Genuinely new longitudinal trade-negotiation benchmark, but the headline bias result currently rests on an unvalidated masking assumption and the paper's own numbers don't agree between main text and appendix.","tokens_in":15396,"tokens_out":1911,"would_cite":false,"duration_ms":20724,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"TradeVerse reconstructs 1,170 WTO trade disputes into a three-task longitudinal benchmark and finds that LLMs identify Western responding countries far more accurately than non-Western ones, even when all country names are masked.","keywords":["longitudinal negotiation","WTO trade concerns","LLM benchmark","respondent identification","geopolitical bias","HS code prediction","statement generation","diplomatic text"],"falsifier":"A targeted audit would settle it: take a random sample of reconstructed concerns, check the parsed role labels against the original meeting minutes, and re-run Task 2 on a version of the transcripts from which every product-origin phrase (e.g., ‘China-made electric vehicles’), demonym, and agency name has been removed. If manual inspection finds role misassignment above a few percent, or if the Western-minus-non-Western gap shrinks to near zero when such cues are stripped, the central bias finding is not robust.","tokens_in":14347,"feed_emoji":"🌐","tokens_out":5616,"duration_ms":50004,"temperature":0.7,"pith_summary":"TradeVerse is a benchmark built from 1,170 World Trade Organization “specific trade concern” disputes, reconstructed from 6,933 meeting minutes into multi-party dialogues that unfold over months or years. The paper argues this is the first benchmark for longitudinal political negotiation, and it defines three tasks: predicting product categories (HS chapters), identifying the responding country from anonymized transcripts, and generating the respondent’s final statement. Across six current LLMs, the paper finds that respondent identification is accurate overall (roughly 90 percent) but systematically better for Western members than for the rest of the membership, with the gap persisting under full anonymization. It also reports that models over-predict product categories (high recall, low precision) and generate final statements that are fluent but share little specific content with the actual interventions. If the benchmark is sound, it gives a reusable measurement of how institutional and geopolitical structure enters LLM reasoning on authentic text.","feed_headline":"90% accuracy on WTO respondents — with a Western lean","feed_subtitle":"Six LLMs identify responding countries accurately, but a Western edge survives full name masking.","key_machinery":"The central object is the TradeVerse corpus: each trade concern is represented as a chronological sequence of rounds, each round containing statements by raiser, supporter, and respondent, with labels (HS chapters, respondent identity, final statement) recovered directly from WTO proceedings rather than from human annotation. The machinery that carries the argument is the pairing of two anonymization settings with a Western/non-Western partition of respondents; this pairing lets the paper separate substance-based inference from identity-cue leakage. The history-ablation on the statement-generation task, comparing full history against the final round alone, is the complementary mechanism that shows the longitudinal information drives generation-quality gains.","core_discovery":"The paper claims that longitudinal structure carries real signal that LLMs can exploit, and that a stable geopolitical skew in respondent identification exists even when all names are masked. The core evidence is the comparison between two anonymization settings: masking only the respondent versus masking every country. Accuracy stays nearly identical across these two settings, while the Western-minus-non-Western gap remains (usually 5 to 14 percentage points) and sometimes widens under full masking. The paper interprets this as models relying on the substance of the dispute—measures, products, and legal arguments—rather than on surface names, while their internal knowledge still reflects a bias toward Western respondents.","pith_inferences":["If the Western bias is driven by training-data imbalance, the same masking design could be reused as a screening tool: any LLM deployed in trade or legal analysis could be tested for respondent-specific accuracy gaps before deployment.","A natural extension the paper leaves implicit is to turn Task 3 into a multi-agent negotiation setting, where separate models play raiser, supporter, and respondent, and to test whether simulated coalitions reproduce the structure of real WTO disputes.","Another testable extension is to control for respondent frequency, since the EU is the most frequent respondent; the paper’s EU-ablation still shows the gap, but a per-country accuracy curve would separate frequency effects from geographic ones.","A multilingual version of the benchmark could reveal whether the Western advantage is a property of English-language training distributions or a more general geopolitical skew."],"forward_implications":["A model that performs well on TradeVerse must track a concern across multiple meetings rather than answer from a single document, so the benchmark directly measures longitudinal reasoning.","The persistent Western/non-Western gap under full masking indicates that LLMs carry internally learned knowledge about trade-dispute participants, which matters for any downstream use in institutional decision support.","High recall with low precision on HS chapter prediction shows that current LLMs hedge, listing plausible product categories rather than committing to a precise set.","The history-ablation result (five of six models improve with full history) supports the benchmark’s premise that longitudinal context carries signal.","Because all labels come from official proceedings without manual annotation, TradeVerse is reusable as a live evaluation suite for future models."],"supporting_citations":[{"why":"Establishes the WTO as a governance and dispute-settlement setting, justifying the corpus source.","marker":"Lang and Scott [2009]"},{"why":"LongBench, the long-context benchmark whose scope excludes longitudinal political text and serves as a contrast for the novelty claim.","marker":"Bai et al. [2023]"},{"why":"Babilong, another long-context reasoning benchmark that TradeVerse distinguishes from its longitudinal negotiation setting.","marker":"Kuratov et al. [2024]"},{"why":"UNBench, a political statement-generation benchmark without longitudinal negotiation structure.","marker":"Liang et al. [2026]"},{"why":"UNSC-Bench, a role-playing vote-prediction baseline that does not follow negotiations over repeated meetings.","marker":"Nangia et al. [2026]"},{"why":"TradeGov, the closest existing trade dataset, which is legal question-answering rather than longitudinal negotiation.","marker":"Mahajan [2025]"},{"why":"Supplies the Western/non-Western country classification used for the bias analysis.","marker":"Pritchard and Wallace [2011]"},{"why":"BERTScore, the semantic similarity metric used to evaluate generated final statements in Task 3.","marker":"Zhang et al. [2019]"}],"fun_headline_variants":["TradeVerse benchmark: LLMs decode WTO disputes, but lean Western","LLMs hit 90% on WTO respondent ID, still Western-biased","Longitudinal trade benchmark reveals LLMs' Western bias","LLMs read multi-round WTO disputes—and inherit a Western bias"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the pipeline that converts WTO Word documents into structured dialogues assigns roles (raiser, supporter, respondent) correctly and that masking removes every identity cue; if role attribution is noisy, or masking leaves identifying content, the Western-bias and Task 2 results become artifacts.","fun_headline_variants_meta":{"raw":{"variants":["TradeVerse benchmark: LLMs decode WTO disputes, but lean Western","LLMs hit 90% on WTO respondent ID, still Western-biased","Longitudinal trade benchmark reveals LLMs' Western bias","LLMs read multi-round WTO disputes—and inherit a Western bias"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001681,"raw_usage":{"total_tokens":6645,"prompt_tokens":907,"completion_tokens":5738,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":523,"completion_tokens_details":{"reasoning_tokens":5664}},"tokens_in":523,"tokens_out":5738,"duration_ms":40994,"temperature":1.0,"reasoning_tokens":5664,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T04:16:35.233640+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A targeted audit would settle it: take a random sample of reconstructed concerns, check the parsed role labels against the original meeting minutes, and re-run Task 2 on a version of the transcripts from which every product-origin phrase (e.g., ‘China-made electric vehicles’), demonym, and agency name has been removed. If manual inspection finds role misassignment above a few percent, or if the Western-minus-non-Western gap shrinks to near zero when such cues are stripped, the central bias finding is not robust.","supporting_citations":[],"review_version":1}