{"id":"e040f110-cf85-477b-8b64-22ba5577e6ea","arxiv_id":"2607.07469","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":4,"one_line_summary":"A 21-configuration LLM judge ensemble (7 models × 3 prompts) validates synthetic e-commerce attribute labels at 95.2% agreement with human experts across 4 languages and 12,726 products.","lead":"The paper builds a multilingual e-commerce benchmark of 12,726 products and validates synthetic labels using a 21-LLM judge ensemble that reaches 95.2% agreement with human experts. A smart generalist might read this to learn whether multi-model voting can replace expensive human annotation at scale.","discovery_kind":"unclear","skeptic_critique":{"model":"glm-5.2","headline":"The κ=0.92 headline is largely tautological: for ~89% of the dataset, the ground truth IS the arena majority vote, so agreement is 100% by construction. The only independently measured accuracy is on ~946 disagreement cases (83.1%) and 400 triaged agreement cases (97%).","rationale":"The reader correctly identified the most load-bearing concern: the ground truth is partially constructed by the same system being evaluated, and the 400-sample triage is too small to fully validate the extrapolation to ~11,780 agreement cases. My analysis confirms this and sharpens it: the circularity is more severe than a sampling concern — for ~89% of the dataset, the 'human expert evaluation' in Table 3 is literally the arena's own output, making the 95.2% agreement figure largely tautological rather than measured. The only independently measured accuracies are 83.1% on disagreement cases and 97% (CI [95.3%, 98.7%]) on 400 triaged agreement cases. The paper is honest about this in its Limitations section, and the framework is defensible for production use with awareness of residual error. However, the headline claim as stated in the abstract — 'agrees with human experts at Cohen's κ = 0.92' — overstates what was directly measured. The reader's CONDITIONAL verdict with MODERATE confidence is appropriate: the contribution is useful, the pipeline is practical, but the central accuracy claim should be understood as an estimate that depends on an extrapolation from a small sample, not a directly measured quantity across the full dataset. I recommend UNCHANGED because the reader already captured this concern accurately in the rationale, even if the weakest_assumption field framed it primarily as a sampling issue rather than emphasizing the more fundamental circularity in the ground truth construction.","tokens_in":16033,"tokens_out":5584,"duration_ms":374806,"concrete_test":"Compute Cohen's κ between the arena majority vote and human ground truth on ONLY the ~946 disagreement cases (where humans independently determined labels). This isolates the genuinely measured agreement from the constructed agreement. If this κ is substantially below 0.92 (which the 83.1% accuracy on these cases suggests it would be), it confirms that the headline figure is inflated by the circular ground truth construction. Additionally, expand the triage to 1,000+ agreement cases stratified proportionally by the five agreement levels in Table 7 (rather than the current 100 unanimous / 300 mixed split) to obtain a representative overturn rate with a tighter confidence interval.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that the 21-judge ensemble achieves 95.2% agreement (κ=0.92) with human experts. However, the ground truth construction in §4 creates a circularity: when the majority vote matches the original synthetic label (~11,780 of 12,726 cases, ~89%), the sample is accepted as correct without human review. The 'human expert evaluation' that the arena is compared against is, for these cases, literally the arena's own output. This means the 95.2% figure is not a directly measured quantity — it is an estimate that combines (a) constructed 100% agreement on ~89% of the data, adjusted downward by a 3% overturn rate extrapolated from 400 samples, with (b) genuinely measured agreement on ~946 disagreement cases where humans actually determined the ground truth. On those disagreement cases, the arena's accuracy is only 83.1% (786/946 corrected, per Table 4). The 400-sample triage (Appendix F) is the sole independent evidence for the agreement cases, and its 3% overturn rate has a 95% CI of [1.3%, 4.7%]; at the upper bound, the true accuracy on agreement cases would be 95.3%, pulling the overall figure down to ~94.4%. Furthermore, the triage stratification (100 unanimous, 300 mixed) does not match the dataset's agreement-level distribution (Table 7: 4,612 unanimous, 5,346 very-high, 1,398 high, 1,198 medium, 172 low), so the extrapolated overturn rate may not be representative. The paper is transparent about this in its Limitations section, but the abstract and Table 3 present κ=0.92 as a measured agreement with human experts without this caveat.","agreement_with_reader":"agree"},"referee_report":{"model":"glm-5.2","summary":"The paper presents SynthAVE, a multilingual benchmark for e-commerce attribute value extraction comprising 12,726 products across 4 languages, 229 product categories, and 792 attributes. To validate synthetic labels at scale, the authors introduce a multi-LLM arena framework in which 21 judge configurations (7 model families × 3 prompts) independently evaluate each sample via majority voting. The paper reports that the ensemble achieves 95.2% agreement with human experts (Cohen's κ = 0.92) and demonstrates cost-effective data cleaning ($0.023 per product). The dataset and per-judge predictions are to be publicly released.","tokens_in":16331,"tokens_out":1331,"duration_ms":221553,"significance":"The paper addresses a practical and important problem: scalable quality assurance for synthetic labels in industrial e-commerce settings. The multi-LLM arena framework with model and prompt diversity is a reasonable and well-motivated design. The cost analysis ($290.50 for 267,246 API calls) and the public release of the dataset with per-judge predictions are concrete strengths. The stratified triage of agreement cases (Appendix F) provides falsifiable evidence for the disagreement-based annotation strategy. The per-class and per-agreement-level breakdowns (Tables 5, 7) are useful for practitioners assessing where the framework is reliable.","major_comments":[{"comment":"§4 (Ground Truth Establishment) and Table 3: The headline κ=0.92 (95.2% agreement) is presented as the arena's agreement with human experts, but the ground truth construction creates a partial circularity. For the ~89% of samples where the arena majority vote matches the original synthetic label, the sample is auto-accepted as correct without human review. The arena's agreement on these cases is 100% by construction, not by measurement. The only independently measured accuracy comes from (a) 946 disagreement cases where humans determined ground truth (83.1% correction rate, Table 4) and (b) 400 triaged agreement cases (3% overturn rate, Appendix F). The paper should explicitly decompose the 95.2% figure into these components so readers understand what is measured versus estimated. As stated, the abstract claim that the ensemble 'agrees with human experts at κ=0.92' overstates the degree,","section":null},{"comment":"§4 and Appendix F: The 400-sample triage is the sole independent evidence for the ~11,326 unreviewed agreement cases, but its stratification (100 unanimous, 300 mixed) does not match the dataset's agreement-level distribution (Table 7: 4,612 unanimous, 5,346 very-high, 1,398 high, 1,198 medium, 172 low). The 3% overturn rate with 95% CI [1.3%, 4.7%] is extrapolated to all agreement cases, but if error rates differ systematically across agreement levels not represented proportionally in the triage sample, the overall accuracy estimate could shift. The paper should either (i) weight the triage results by the actual agreement-level distribution or (ii) explicitly acknowledge this stratification mismatch and its potential impact on the headline accuracy claim.","section":null},{"comment":"Table 4 and §5: The arena's accuracy on disagreement cases is reported as 83.1% error correction (786/946), but Table 4 also reports 'Arena Accuracy: 95.0%' and 'Precision: 98.0%, Recall: 95.2%, F1: 96.6%' for bad-label detection. These metrics appear to be computed against the partially constructed ground truth (including auto-accepted agreement cases), which inflates precision and recall by construction. The paper should clarify which metrics are computed against independently verified ground truth versus the full (partially constructed) ground truth, and report the arena's performance on disagreement cases separately and clearly.","section":null}],"minor_comments":[{"comment":"Abstract: 'nabling' should be 'enabling' (also appears in the contributions bullet in §1).","section":null},{"comment":"Table 3 caption: 'LLM Arena agreement with human evaluation by language' — given the circularity concern, 'human evaluation' should be qualified (e.g., 'human-verified ground truth (see §4)').","section":null},{"comment":"§3: The dataset is described as spanning '2,607 distinct product-attribute combinations' (§3) but '2,598 product type-attribute combinations' in the Ethics Statement. These should be reconciled.","section":null},{"comment":"Appendix F, Table 16: The '95% CI' column header appears but the CI values are listed in the 'OVERTURN' column format. Formatting should be clarified.","section":null},{"comment":"§4: The qualification criteria (general competency, task-specific competency, self-consistency, instruction adherence) are described conceptually but no specific thresholds or evaluation results are provided. A brief note on which models met or failed these criteria would strengthen the section.","section":null},{"comment":"Table 7: The 'VERY HIGH (>85%)' row shows 9,958 products, but the main text (§5) says 'very high agreement (>85%) yields 98.8%' — the product count in Table 7 appears to combine very-high and unanimous cases (9,958 + 4,612 = 14,570 > 12,726). This should be checked.","section":null},{"comment":"Figure 2 caption references 'Appendix C' for other languages, but the appendix figures (Figures 3–9) are labeled by language individually. Cross-references could be clearer.","section":null},{"comment":"§5: 'No individual judge configuration outperformed the majority vote ensemble (95.0%)' — the parenthetical says 95.0% but Table 3 reports 95.2%. These should be consistent.","section":null}],"recommendation":"major_revision","confidential_remarks":"The circularity concern is the central issue. The paper is transparent about the disagreement-based annotation strategy in §4 and the Limitations section, but the abstract and Table 3 present the κ=0.92 figure without qualification. A revision that decomposes the accuracy claim into measured vs. estimated components, weights the triage by actual agreement-level distributions, and adjusts the abstract language would resolve the concern. The dataset release and cost analysis are genuine contributions worth preserving. The paper fits the journal's scope if these issues are addressed."},"author_rebuttal":{"model":"glm-5.2","summary":"We thank the referee for a careful and constructive review. The three major comments all identify a legitimate concern about the distinction between measured and estimated accuracy in our ground truth construction. We agree that the manuscript needs to be more transparent about this distinction and will revise accordingly. Below we address each comment point by point.","responses":[{"response":"The referee is correct that the 95.2% figure conflates two distinct sources: directly measured agreement on the 1,346 human-reviewed samples (946 disagreement cases + 400 triaged agreement cases) and estimated agreement on the ~11,380 unreviewed agreement cases. We agree this decomposition should be explicit in the paper. In the revision, we will add a clear breakdown showing that (a) 946 disagreement cases were fully human-annotated, (b) 400 agreement cases were independently triaged yielding a 3.0% overturn rate (95% CI [1.3%, 4.7%]), and (c) the remaining ~10,980 agreement cases are estimated at ~97% accuracy based on the triage extrapolation. The headline κ=0.92 is computed against the full ground truth, which includes the auto-accepted cases, so it does reflect a partially constructed ground truth rather than purely measured agreement. We will revise the abstract and §4 to state this more precisely, replacing language like 'agrees with human experts at κ=0.92' with wording that distinguishes measured validation on the human-reviewed subset from estimated accuracy on the auto-accepted subset. We will also add a sentence in the abstract acknowledging the estimation component.","revision_made":"yes","referee_comment":"§4 (Ground Truth Establishment) and Table 3: The headline κ=0.92 (95.2% agreement) is presented as the arena's agreement with human experts, but the ground truth construction creates a partial circularity. For the ~89% of samples where the arena majority vote matches the original synthetic label, the sample is auto-accepted as correct without human review. The arena's agreement on these cases is 100% by construction, not by measurement. The paper should explicitly decompose the 95.2% figure into these components so readers understand what is measured versus estimated."},{"response":"This is a fair point. The triage sample was deliberately stratified to over-represent mixed-agreement cases (where errors are more likely) rather than to mirror the natural agreement-level distribution. This means the 3.0% overturn rate is an upper bound on the true error rate among all agreement cases, since the natural distribution is heavily skewed toward unanimous and very-high agreement cases (which have lower error rates). In the revision, we will add a weighted estimate: applying the observed per-stratum overturn rates (0% for unanimous, 4.0% for mixed) to the actual agreement-level distribution from Table 7 yields a weighted overturn rate of approximately 2.6%, which is lower than the unweighted 3.0%. This actually strengthens our headline accuracy claim slightly (estimated accuracy ~97.4% rather than ~97.0%). We will include this weighted calculation in Appendix F and add an explicit acknowledgment of the stratification mismatch in §4, noting that the unweighted 3.0% figure is conservative. We cannot, however, produce per-sub-level breakdowns (e.g., very-high vs. high vs. medium) because the triage sample only distinguished unanimous from mixed, not the finer granularity in Table 7. We will state this limitation explicitly.","revision_made":"yes","referee_comment":"§4 and Appendix F: The 400-sample triage is the sole independent evidence for the ~11,326 unreviewed agreement cases, but its stratification (100 unanimous, 300 mixed) does not match the dataset's agreement-level distribution (Table 7: 4,612 unanimous, 5,346 very-high, 1,398 high, 1,198 medium, 172 low). The 3% overturn rate with 95% CI [1.3%, 4.7%] is extrapolated to all agreement cases, but if error rates differ systematically across agreement levels not represented proportionally in the triage sample, the overall accuracy estimate could shift. The paper should either (i) weight the triage results by the actual agreement-level distribution or (ii) explicitly acknowledge this stratification mismatch and its potential impact on the headline accuracy claim."},{"response":"The referee is correct that the precision, recall, and F1 metrics in Table 4 are computed against the full ground truth, which includes auto-accepted agreement cases. This means that true positives (correctly identified bad labels) include both human-verified corrections from the 946 disagreement cases and auto-accepted corrections from agreement cases where the arena changed the label. The latter are not independently verified, so the precision and recall figures do benefit from the construction. In the revision, we will (i) add a note to Table 4 clarifying that these metrics are computed against the full (partially constructed) ground truth, (ii) add a separate row or sub-table reporting the arena's performance on the 946 independently verified disagreement cases only, and (iii) note that the 83.1% error correction rate is the key metric computed purely against human-verified ground truth. We will also clarify that the 95.0% 'Arena Accuracy' figure in Table 4 is the post-cleaning dataset accuracy (i.e., the proportion of all 12,726 labels that are correct after applying the arena), which is distinct from the 95.2% agreement-with-humans figure in Table 3, though the two are closely related. The distinction is subtle but worth making explicit.","revision_made":"yes","referee_comment":"Table 4 and §5: The arena's accuracy on disagreement cases is reported as 83.1% error correction (786/946), but Table 4 also reports 'Arena Accuracy: 95.0%' and 'Precision: 98.0%, Recall: 95.2%, F1: 96.6%' for bad-label detection. These metrics appear to be computed against the partially constructed ground truth (including auto-accepted agreement cases), which inflates precision and recall by construction. The paper should clarify which metrics are computed against independently verified ground truth versus the full (partially constructed) ground truth, and report the arena's performance on disagreement cases separately and clearly."}],"tokens_in":16014,"tokens_out":1396,"duration_ms":251157,"standing_objections":[]},"desk_editor":{"model":"glm-5.2","letter":"The headline κ=0.92 is not a directly measured quantity across the full dataset, and the stress-test note is right to flag this. For ~89% of the 12,726 products, the ground truth IS the arena majority vote — when the arena agreed with the original synthetic label, no human looked at it. The 95.2% figure combines constructed 100% agreement on those ~11,300 cases (adjusted by a 3% overturn rate extrapolated from 400 triaged samples) with genuinely measured agreement on ~946 disagreement cases (where accuracy is 83.1%). The abstract and Table 3 present κ=0.92 as measured agreement with human experts without this caveat, which overstates the evidence. The 400-sample triage is the sole independent check on agreement cases, and its stratification (100 unanimous, 300 mixed) doesn't match the dataset's actual agreement-level distribution (Table 7: 4,612 unanimous, 9,958 very-high, etc.), so the extrapolated 3% overturn rate may not be representative. At the upper CI bound (4.7%), overall accuracy drops to ~94.4% — still decent, but not what's advertised. The paper is transparent about this in the Limitations section, which I appreciate, but the framing in the abstract and results table should match that honesty. That said, the circularity concern is somewhat mitigated by the fact that the arena is being compared against the *original synthetic label*, not against itself. The arena's majority vote agreeing with a separately generated label is a weaker form of circularity than the arena validating its own output directly. The 946 disagreement cases where humans actually adjudicated provide real, independent evidence — and the 83.1% correction rate there is a genuinely measured quantity. The triage results (0% overturn on unanimous cases, 3% on mixed) are also real evidence, just limited in scope. What's genuinely new and useful: the multilingual benchmark (12,726 products, 4 languages, 229 categories, 792 attributes) with human-validated labels is a concrete artifact the community can use. The cost analysis ($0.023/product, $290.50 total) is practical and reproducible. The per-class and per-agreement-level breakdowns (Tables 5, 7) are honest about where the ensemble fails — low-agreement cases at 65.1% accuracy, INCORRECT labels hardest to catch. The within-family agreement clustering shown in Figure 2 is a useful empirical finding about judge independence. The constituent techniques (LLM-as-judge, majority voting, synthetic data) are not new, but the specific 21-judge configuration and its measured properties across languages are a contribution to applied NLP. The judge independence assumption is questionable — Figure 2 shows within-family κ of 0.80-0.90, meaning the 21 judges are not truly independent. This inflates the ensemble's apparent diversity. The paper acknowledges model families cluster but doesn't grapple with what this means for the majority vote's statistical properties. This is a minor-to-moderate concern: the ensemble still works empirically, but the theoretical justification is weaker than presented. This paper is for applied ML teams building data validation pipelines in e-commerce or similar domains. A researcher working on LLM-as-judge methodology will find the empirical configuration useful but won't find theoretical novelty. The benchmark itself has value if the labels hold up under broader sampling. I'd recommend a serious referee. The main revision ask: reframe the κ=0.92 as an estimate (not a measurement), add the construction caveat to the abstract and Table 3, and either expand the triage sample or explicitly bound the headline claim with the CI. The judge independence issue deserves a paragraph, not just a figure caption. If the authors make those framing fixes, this is a solid applied contribution.","headline":"The circularity concern is real but manageable; the paper's practical contribution is a useful multilingual benchmark and a cost-effective validation recipe.","tokens_in":16934,"tokens_out":876,"would_cite":false,"duration_ms":190812,"reading_group":"no","serious_thinker":"no","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"glm-5.2","headline":"21 LLM judges vote on labels, match humans 95% of the time","keywords":["LLM-as-judge","synthetic data validation","majority voting ensemble","attribute value extraction","e-commerce","multilingual","data quality assurance","inter-rater agreement"],"falsifier":"Find that the 3% overturn rate from the 400 sampled agreement cases substantially underestimates the true error rate among the ~11,300 unreviewed agreement cases—particularly if certain attribute types, product categories, or languages have error profiles that the stratified sample did not capture.","tokens_in":16255,"feed_emoji":"🗳️","tokens_out":1225,"duration_ms":211294,"temperature":0.7,"pith_summary":"This paper claims that a panel of 21 LLM judge configurations—7 model families each given 3 different prompts—can validate synthetic e-commerce product labels at near-human reliability through majority voting, at roughly two cents per product. The authors built a dataset of 12,726 products across 4 languages and 229 categories, generated synthetic attribute labels (CORRECT, INCORRECT, or UNKNOWN) for each, and then used the 21-judge panel to audit those labels. The ensemble's majority vote agrees with human experts at Cohen's κ = 0.92 (95.2% agreement), while individual judges only reach Fleiss' κ = 0.76 among themselves. The key mechanism is diversity-driven error cancellation: no single judge configuration outperforms the ensemble, and unanimous agreement across all 21 judges yields 100% accuracy in the human-checked sample. The authors use a disagreement-based annotation strategy where human review is triggered only when the panel contradicts the original synthetic label, reducing human effort to the ~8% of cases where disagreement occurs. Expert triage of 400 agreement cases found a 3% overturn rate (0% for unanimous cases). The total validation cost was $290.50 for the full dataset, and the framework corrects 83% of synthetic labeling errors.","feed_headline":"21 LLM judges vote on labels, match humans 95% of the time","feed_subtitle":"A diverse ensemble of 21 LLM configurations validates synthetic e-commerce labels at $0.02 per product, matching expert agreement at κ=0.92.","key_machinery":"The multi-LLM arena: 7 model families (Claude, Nova, GPT, Mistral, DeepSeek, Qwen, Gemma) × 3 prompt variants = 21 independent judge configurations, each evaluating every product. Final labels are determined by majority vote across all 21. The disagreement-based annotation strategy routes only cases where the panel contradicts the original synthetic label to human review, while agreement cases are accepted without human inspection.","core_discovery":"The central finding is that diversity among LLM judges—across both model families and prompt strategies—produces an ensemble whose majority vote substantially outperforms any individual judge. The authors interpret moderate inter-judge agreement (Fleiss' κ = 0.76) combined with high ensemble-human agreement (Cohen's κ = 0.92) as evidence that judges with different biases cancel each other's errors. Unanimous 21-judge agreement was never overturned by human review in the sampled cases, while low-agreement cases (<50% consensus) dropped to 65% accuracy, suggesting that agreement level itself serves as a calibrated confidence signal. The framework also functions as a data-cleaning tool: the raw","pith_inferences":["The framework is validated only on a three-way classification task with relatively objective ground truth (attribute value verification). Whether the same ensemble agreement signals transfer to subjective tasks—such as assessing product review quality, sentiment nuance, or recommendation relevance—remains untested and is a natural next experiment.","The 21-judge configuration is a specific point in a design space. A systematic study of how ensemble accuracy scales with the number of judge configurations (e.g., 3, 5, 7, 11, 15, 21) would reveal whether there is a diminishing-returns threshold, which would make the approach more accessible to budget-constrained teams.","The finding that INCORRECT labels are the hardest for both the synthetic generator and the arena suggests an asymmetry in how LLMs handle contradiction versus confirmation. Investigating whether this asymmetry persists across task domains could reveal a structural limitation of LLM-as-judge approaches."],"forward_implications":["If majority-vote LLM ensembles reliably approximate human judgment on verification tasks, the cost bottleneck for creating labeled training and evaluation datasets shifts from human annotation to API costs, which are orders of magnitude cheaper.","The finding that agreement level correlates with accuracy (unanimous = 100%, low-agreement = 65%) suggests a natural triage mechanism: confident cases bypass human review entirely, while split decisions can be routed to human auditors, creating an adaptive human-in-the-loop pipeline.","The result that model-family diversity matters more than prompt diversity for ensemble performance implies that adding more models from different providers yields diminishing but real returns, while adding more prompts to the same model saturates faster.","If the 3% overturn rate from the 400-sample triage generalizes, the released dataset carries a known residual error rate that downstream users must account for when using it as ground truth for benchmarking."],"fun_headline_variants":["21 diverse LLM judges ensemble-match human experts at 95% agreement","LLM judge diversity cancels bias: ensemble hits 95% human agreement","Unanimous 21-judge LLM panels never overturned by human review","Voting across 21 LLM configs matches human label quality at 95%","Disagreement among LLM judges flags low-confidence labels for review"],"cache_read_input_tokens":0,"weakest_assumption_plain":"The ground truth validation relies on a disagreement-based strategy where samples are accepted without human review when the LLM majority vote matches the original synthetic label. The reported 97% accuracy for these unreviewed agreement cases is estimated from only 400 stratified samples (100 per language), and 6 'unsure' cases were excluded from the overturn-rate calculation. If the remaining ~11,300 unreviewed agreement cases have a systematically different error profile, ","fun_headline_variants_meta":{"raw":{"variants":["21 diverse LLM judges ensemble-match human experts at 95% agreement","LLM judge diversity cancels bias: ensemble hits 95% human agreement","Unanimous 21-judge LLM panels never overturned by human review","Voting across 21 LLM configs matches human label quality at 95%","Disagreement among LLM judges flags low-confidence labels for review","21-LLM ensemble agreement signals label confidence, matches humans","Multi-model LLM voting validates synthetic labels at 95% accuracy","Diverse LLM judges outperform individuals via majority vote","Judge agreement level calibrates synthetic label reliability","Seven model families, three prompts, one reliable label ensemble","Low-consensus LLM panels drop to 65% accuracy, flagging bad labels","Synthetic e-commerce labels validated by 21-LLM majority vote"]},"model":"glm-5.2","effort":"low","cost_usd":0.0,"raw_usage":{"total_tokens":1502,"prompt_tokens":569,"completion_tokens":933,"prompt_tokens_details":null},"tokens_in":569,"tokens_out":933,"duration_ms":46397,"temperature":1.0,"reasoning_tokens":739,"cache_read_input_tokens":0,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-09T09:48:15.380532+00:00","model_set":{"reader":"glm-5.2"},"falsifier":"Find that the 3% overturn rate from the 400 sampled agreement cases substantially underestimates the true error rate among the ~11,300 unreviewed agreement cases—particularly if certain attribute types, product categories, or languages have error profiles that the stratified sample did not capture.","supporting_citations":[],"review_version":1}