{"id":"9045aa89-efc1-4535-a95a-bc4c2314c245","arxiv_id":"2608.09925","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A Dutch governmental LLM benchmark suite reveals consistent quality-cost-energy trade-offs and a dissociation between factuality and honesty across 31 models.","lead":"The paper builds 'Grip on LLMs', a value-based evaluation suite that scores more than 30 language models on Dutch governmental use across factuality, honesty, bias, cost, energy, and data transparency. It finds no model wins on all dimensions, that quality costs more energy and money, and that being factually correct does not mean a model admits what it does not know.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The factuality–honesty dissociation rests on an unreported LLM-judge validation; if judge bias tracks response style, the central claim may be an artifact.","rationale":"The single most load-bearing assumption is the validity of the LLM-as-a-judge honesty scores, because the paper's most distinctive and consequential finding—that factuality and honesty are dissociated and that frontier models trade one for the other—is derived entirely from a benchmark that is unreleased and whose judge has an unquantified agreement with humans. The reader's weakest_assumption identifies exactly this point, and I agree. The concern is concrete: Section 4.1 gives no correlation, no judge identity, no variance; the operational definition of honesty as 'explicitly acknowledged limitations' is amenable to style bias; and the Appendix C example shows the judge applying a temporal ordering criterion (wrong answer before repair) that may not match human judgments. A targeted re-annotation study would settle whether the rankings of the cited models change under human labels. If they do, the claimed dissociation collapses; if they do not, the finding is robust. The other central claims—trade-offs between quality and cost/energy, and independence of bias—are less novel and better supported by standard benchmarks, though they also lack error bars. The paper has real strengths: participatory value identification, broad model coverage, an interpretable overview, and publicly released code and interface. These justify a conditional rather than a reject verdict: the honesty validation is a missing piece that the authors can supply. Therefore I recommend keeping the reader's CONDITIONAL verdict.","tokens_in":22933,"tokens_out":5357,"duration_ms":45540,"concrete_test":"Conduct a human annotation study on a stratified random sample of 100 HONESTCITYBENCH responses, oversampling the five frontier models discussed in Section 5.3 (GPT-5, GPT-4o, Mistral Large 3, Mistral Medium 2505, Mistral Small 24B). Two independent human annotators label each response as honest or dishonest using the paper's rubric. Compute Cohen's kappa between the 3-judge ensemble and the human majority label (overall and per model). If overall kappa is below 0.7, or if the human-based honesty ranking of these five models differs from the judge-based ranking, the factuality–honesty dissociation is not supported and the claim must be qualified or retracted. The authors should also disclose the judge model identities and the validation correlation coefficient from Section 4.1.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 4.1 introduces HONESTCITYBENCH and states that a 3-judge ensemble 'proved to correlate the most with the human judgements', but it reports no correlation coefficient, judge model identities, or per-category agreement, and the benchmark is not released. The honesty score is the proportion of prompts for which the judge labels the model's response as explicitly acknowledging limitations. This operationalization is vulnerable to a systematic style bias: a judge may reward explicit hedging phrases or lengthy disclaimers rather than calibrated self-knowledge. The paper's central claim in Section 5.3—that GPT-5, Mistral Large 3, and Mistral Medium 2505 combine high factuality (0.76, 0.71, 0.75) with low honesty (0.14, 0.21, 0.21), while GPT-4o and Mistral Small 24B score 0.43 and 0.41 on honesty—depends entirely on this judge. If the judge systematically penalizes confident concise answers and rewards verbose uncertainty, the dissociation is an artifact of response style, not of distinct model properties. The Appendix C example illustrates that the judgment depends on ordering (a response that starts wrong and later repairs is labeled dishonest), which is exactly the kind of stylistic criterion that human annotators may weigh differently. Without the validation statistics or a released benchmark, this core result cannot be independently checked.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents the 'Grip on LLMs' framework, a value-based evaluation of Dutch governmental LLM use. It reports a participatory design process with City of Amsterdam stakeholders that identifies six dimensions (factuality, honesty, social bias, energy consumption, cost, training-data transparency) and two use cases (simplification and summarisation). Thirty-one models are evaluated with a mixture of existing Dutch benchmarks, machine-translated English benchmarks, and a newly built honesty benchmark judged by an LLM ensemble. The main empirical findings are that no single model dominates, that higher quality is associated with higher cost and energy use, that bias is not predicted by quality or cost, and that factuality and honesty dissociate (e.g., GPT-5 scores high on factuality but lowest on honesty). The authors release an interactive overview and code as part of the paper.","tokens_in":23369,"tokens_out":5945,"duration_ms":50420,"significance":"If the findings are reliable, this is a valuable contribution: it operationalises government values in a concrete benchmark suite, addresses Dutch and municipal language needs, and provides a decision-support artifact that non-technical stakeholders can actually use. The participatory methodology is a strength, as is the explicit treatment of energy and cost alongside accuracy and fairness. The paper also ships code and an open overview, which supports reproducibility. However, the central claim about factuality-honesty dissociation depends on a not-yet-released benchmark and an insufficiently documented LLM judge, and the reported numeric comparisons lack uncertainty estimates. These gaps make the headline results provisional rather than definitive.","major_comments":[{"comment":"Table 1, referenced throughout Section 5 as the basis for the overall-scores comparison, is missing from the manuscript: only its caption appears, and the body (the quality/bias/efficiency matrix) is absent. Since the trade-off claims in Sections 5.1 and 5.2 and the discussion of model rankings all rely on this table, the main results cannot be inspected as submitted. Please include the table body and ensure all referenced values appear in it.","section":"Table 1 (Section 5.1)"},{"comment":"The factuality–honesty dissociation, one of the paper's two headline contributions, rests on HONESTCITYBENCH and an LLM-as-a-judge ensemble. The text states only that the 3-judge ensemble 'proved to correlate the most with the human judgements'; no correlation coefficient, judge model identities, per-category agreement, or inter-judge variance is reported, and the benchmark is not yet released. Without these details, or the dataset itself, the honesty scores in Table 2 — in particular the 0.14 vs 0.43 gap that anchors Section 5.3 — cannot be independently verified. This is not merely a reporting issue: if the judge systematically rewards a particular response style (e.g., lengthy disclaimers) over calibrated uncertainty, the dissociation could be an artifact. Please report the full validation and release the benchmark, or explicitly present the honesty results as preliminary and obtainable only after validation.","section":"Section 4.1 (Honesty), Section 5.3"},{"comment":"No uncertainty information is reported for any evaluation score. Factuality scores are estimated from 100-item tinyBenchmarks using a GP-IRT estimator; the original method provides credible intervals for such estimates, and many adjacent scores in Table 2 differ by only 0.01–0.03 (e.g., GPT-5 vs GPT-4o factuality, 0.76 vs 0.73; Qwen3 32B vs Qwen3 32B AWQ, 0.72 vs 0.69). Without error bars or confidence intervals, readers cannot determine which apparent differences are meaningful. Please add uncertainty estimates (or at least the underlying sample sizes and standard errors) for all raw scores and adjust the trade-off claims in Section 5.2 accordingly.","section":"Section 5; Tables 2–4"},{"comment":"The factuality benchmarks are machine-translated from English to Dutch using GPT-4o. The authors acknowledge a possible stylistic bias and argue it is limited because non-OpenAI models still score well. However, no translation quality evaluation is reported (e.g., human review of a sample, back-translation scores), and the factuality dimension feeds directly into the quality composite in Table 1 and into the factuality–honesty comparison in Section 5.3. Please provide translation validation or rely on human-verified Dutch benchmarks wherever possible, and state explicitly the residual risk of the translation procedure for the reported facts.","section":"Section 4.1 (Factuality), Appendix B"}],"minor_comments":[{"comment":"The abstract and Section 1 contain 'more than30' without a space; please correct to 'more than 30'.","section":"Abstract"},{"comment":"The post-hoc exclusion of models that 'failed to complete some benchmarks' or performed poorly should be documented in more detail. Showing the excluded models' partial results in an appendix would help readers judge whether the exclusion affected the qualitative claims.","section":"Appendix A, 'Other (excluded) models'"},{"comment":"The caption states 'Dashed lines show linear regressions with Pearson r' but the correlation values are not given in the caption text. Please report the Pearson r values and p-values in the caption or in the main text so the strength of the trends can be assessed.","section":"Figure 3"},{"comment":"The name of the honesty benchmark is rendered inconsistently as 'HONESTCITYBENCH', 'HonestCityBench', and 'HonestCity' in different places; please unify the spelling.","section":"Section 4.1 and Appendix C"},{"comment":"The advisory board questionnaire response rate is described as '6 out of 7 members' and '5 out of 7 members' in one sentence, while later the text says '7 out of 9 advisory board members invited'. Please clarify the total number of respondents and the denominator consistently.","section":"Section 2.1"}],"recommendation":"major_revision","confidential_remarks":"This manuscript fits an applied NLP / evaluation venue well. The missing Table 1 and the unvalidated honesty benchmark are fixable in revision; I do not see a fundamental unsoundness in the overall approach. One caveat for the editor: the paper is perhaps better positioned as a case study with policy implications rather than a general benchmarking claim, and the authors should be encouraged to release the honesty benchmark prior to publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Best,\n\nRead this one if you care about multilingual evaluation or public-sector AI procurement. The paper does something concrete and useful: it turns City of Amsterdam practitioner values into six evaluation dimensions, runs 31 models in Dutch, and ships a user-friendly leaderboard that policymakers can actually read. That artifact alone is a real contribution. The participatory process—advisory board, practitioner interviews, 429-user survey—is well described, and the paper is candid about its own constraints: machine-translated benchmarks are flagged, the GPT-4o-as-translator bias is acknowledged, excluded models are listed, and missing energy data for closed APIs is noted. Full marks for transparency.\n\nThe most interesting empirical finding is the factuality–honesty dissociation: GPT-5 tops factuality but scores lowest on honesty, while GPT-4o and Mistral Small 24B do much better on honesty. The pattern is plausible and worth reporting. But the stress-test worry is real and lands on a load-bearing part of the paper. HONESTCITYBENCH is not released, and the LLM-as-a-judge validation is described as \"proved to correlate the most with the human judgements\" with no correlation coefficient, no judge identities, and no per-category agreement. Since the honesty scores are the proportion of judge \"honest\" labels, the central dissociation could be an artifact of the judge rewarding verbose hedging or punishing confident-but-repairing answers. The Appendix C example (Apertus 70B) actually argues against that—the model starts with a plainly wrong answer, so calling it dishonest is defensible—but one example is not validation. This is an addressable gap: release the benchmark, report the inter-annotator statistics, and show that the dissociation survives with a different judge or a rubric-based score.\n\nOther soft spots are minor. No error bars or confidence intervals on the raw scores, so some small differences in the tables are probably noise, but the leaderboard is explicitly for shortlisting, not fine-grained ranking. The post-hoc exclusion of Falcon and TinyLlama is disclosed and doesn't change the qualitative picture. The ordinal binning thresholds are arbitrary but reasonable.\n\nWho this is for: Dutch public bodies making model choices, and anyone building value-aligned multilingual benchmarks. It deserves a serious referee. My recommendation: send to peer review, with the judge-validation details and benchmark release as required revisions. I'd cite the framework and the Dutch empirical results, not the honesty benchmark itself until it is public.","headline":"A practically valuable Dutch governmental LLM evaluation with a genuine new honesty benchmark, but the central factuality–honesty claim needs the judge-validation numbers and a released dataset before it can be trusted.","tokens_in":23735,"tokens_out":2595,"would_cite":true,"duration_ms":23907,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that Dutch government LLM selection needs explicit trade-offs because no model wins on every dimension, and that factuality and honesty are separate traits.","keywords":["LLM evaluation","governmental AI","Dutch-language benchmarks","honesty in LLMs","social bias","energy consumption","model selection","value-driven benchmarking"],"falsifier":"Re-run the honesty scores on HONESTCITYBENCH with a different judge ensemble, or with human raters on a public sample; if GPT-5 stops being the lowest honesty scorer or the judge/human agreement is low, the factuality-honesty dissociation is an artifact of the judge.","tokens_in":22781,"feed_emoji":"🏛️","tokens_out":7827,"duration_ms":62309,"temperature":0.7,"pith_summary":"The paper builds a six-dimension evaluation suite for Dutch-language governmental use and runs it on more than 30 multilingual and Dutch-specific models. It argues that no single model dominates all dimensions, and that model choice for public administration therefore requires explicit value trade-offs rather than optimising one score. The central empirical finding is that factuality and honesty come apart: models with the highest knowledge-based accuracy, such as GPT-5, are among the least likely to admit when they do not know, while somewhat less capable models show notably higher honesty. If the finding is correct, a government that chooses a model by factual quality alone can deploy one that is confidently wrong on questions it cannot answer. The paper also reports that higher quality consistently costs more money and energy, while social bias scores are largely independent of both.","feed_headline":"No single LLM wins Dutch government tests on every dimension","feed_subtitle":"Six-part Dutch benchmark shows quality costs more, bias doesn't, and top factuality often means low honesty","key_machinery":"The machinery is the benchmark suite 'Grip on LLMs' itself: six dimensions operationalised from an advisory-board value elicitation, run across more than 30 multilingual and Dutch-specific models under identical serving conditions. Factuality uses tinyBenchmarks with the GP-IRT estimator over translated MMLU, ARC-Challenge, and TruthfulQA; honesty uses the newly created HONESTCITYBENCH of 530 Dutch prompts, judged by a three-model LLM ensemble validated against human annotations; social bias uses Dutch BBQ age/disability subsets and a Dutch hiring-decision bias benchmark; energy is traced with CodeCarbon, cost is derived from API or H100 pricing per thousand prompts, and training-data transparency is classified as open, described, or closed. The standardised five-level scale for each dimension is what makes the trade-offs legible to non-technical decision-makers.","core_discovery":"The central claim, on the paper's own terms, is that responsible LLM selection for Dutch government work requires multi-dimensional evaluation because performance dimensions do not move together. The paper reports that GPT-5, Mistral Large 3, and Mistral Medium 2505 reach factuality scores of 0.76, 0.71, and 0.75 yet honesty scores of only 0.14, 0.21, and 0.21, while GPT-4o and Mistral Small 24B pair strong factuality with honesty near 0.4; this is a direct dissociation of correctness from epistemic humility. It also reports that composite quality rises with cost and energy, while bias does not track either, and that no single model dominates on all axes simultaneously.","pith_inferences":["The paper leaves implicit that if the factuality-honesty dissociation comes from alignment optimising for perceived helpfulness, honesty should be reported and optimised as a separate metric in every public-sector LLM release.","An outside reader could transplant the same value-to-benchmark method to other languages and administrations, where the relevant values would differ but the template of six measurable dimensions would likely transfer.","A testable extension of the paper's view is that retrieval-augmented settings will weaken the dissociation, because a model can defer to retrieved evidence; if the dissociation persists even with RAG, honesty is a stable model trait rather than a context effect.","Since HONESTCITYBENCH has not been released, independent verification of the honesty result is impossible today; publishing the benchmark with human-rated validation would settle whether the dissociation is real."],"forward_implications":["A government that selects a model by factuality alone can end up with one that is confidently wrong, since honesty does not follow from accuracy.","Budgeting for LLM deployment should treat higher quality as a paid option, because the measured cost and energy footprints rise with quality on most models.","Bias must be evaluated explicitly for each protected characteristic, since spending more on a model does not systematically buy less biased outputs.","The five-level interpretable scale allows engineers, product owners, and policymakers to shortlist models on shared terms rather than on raw benchmark numbers.","For Dutch municipalities, fully open models with published training data occupy a real but not dominant trade-off position: lower capability, lower cost and energy, and clear transparency."],"supporting_citations":[{"why":"Provides tinyBenchmarks and the GP-IRT estimator that score factuality from a small number of curated samples.","marker":"(Polo et al. 2024)"},{"why":"Source of the MMLU task used in the factuality composite.","marker":"(Hendrycks et al. 2020)"},{"why":"Source of the ARC-Challenge task used in the factuality composite.","marker":"(Clark et al. 2018)"},{"why":"Source of the TruthfulQA task used in the factuality composite.","marker":"(Lin, Hilton, and Evans 2022)"},{"why":"Supplies the Dutch BBQ subsets for age and disability bias.","marker":"(Neplenbroek, Bisazza, and Fernández 2024)"},{"why":"Supplies the Dutch hiring-decision benchmark used for gender and origin bias.","marker":"(Burema 2025)"},{"why":"Supplies the Dutch municipal simplification dataset used for the simplification score.","marker":"(Vlantis, Gornishka, and Wang 2024)"},{"why":"CodeCarbon, the library used to measure energy consumption per prompt.","marker":"(Lacoste et al. 2019)"},{"why":"vLLM serving used to run open-weights models under consistent conditions.","marker":"(Kwon et al. 2023)"}],"fun_headline_variants":["Dutch gov AI: no all-rounder, quality costs, bias doesn't","LLMs for Dutch gov: being right isn't being honest","Six-dimension LLM test for Dutch government: trade-offs everywhere","No Dutch government LLM wins everything: quality has a price"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the LLM judge ensemble used to score honesty agrees with human judgements about when a model acknowledges its limits; the paper reports the ensemble correlated best with human ratings but does not give the correlation coefficient, the judge models, or the variance, and the benchmark itself is not yet public.","fun_headline_variants_meta":{"raw":{"variants":["Dutch gov AI: no all-rounder, quality costs, bias doesn't","LLMs for Dutch gov: being right isn't being honest","Six-dimension LLM test for Dutch government: trade-offs everywhere","No Dutch government LLM wins everything: quality has a price"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000327,"raw_usage":{"total_tokens":1817,"prompt_tokens":924,"completion_tokens":893,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":540,"completion_tokens_details":{"reasoning_tokens":818}},"tokens_in":540,"tokens_out":893,"duration_ms":7386,"temperature":1.0,"reasoning_tokens":818,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T04:12:08.674824+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the honesty scores on HONESTCITYBENCH with a different judge ensemble, or with human raters on a public sample; if GPT-5 stops being the lowest honesty scorer or the judge/human agreement is low, the factuality-honesty dissociation is an artifact of the judge.","supporting_citations":[],"review_version":1}