REVIEW 4 major objections 5 minor 77 references
From Values to Benchmarks: Evaluating Large Language Models for Governmental Use in Dutch
T0 review · 4 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read The paper argues that Dutch government LLM selection needs explicit trade-offs because no model wins on every dimension, and that factuality and honesty are separate traits.
desk verdict A practically valuable Dutch governmental LLM evaluation with a genuine new honesty benchmark, but the central factuality–honesty claim needs the judge-validation numbers and a released dataset before it can be trusted. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery is the benchmark suite 'Grip on LLMs' itself: six dimensions operationalised from an advisory-board value elicitation, run across more than 30 multilingual and Dutch-specific models under identical serving conditions. Factuality uses tinyBenchmarks with the GP-IRT estimator over translated MMLU, ARC-Challenge, and TruthfulQA; honesty uses the newly created HONESTCITYBENCH of 530 Dutch prompts, judged by a three-model LLM ensemble validated against human annotations; social bias uses Dutch BBQ age/disability subsets and a Dutch hiring-decision bias benchmark; energy is traced with CodeCarbon, cost is derived from API or H100 pricing per thousand prompts, and training-data transparency is classified as open, described, or closed. The standardised five-level scale for each dimension is what makes the trade-offs legible to non-technical decision-makers.
What would settle it
Re-run the honesty scores on HONESTCITYBENCH with a different judge ensemble, or with human raters on a public sample; if GPT-5 stops being the lowest honesty scorer or the judge/human agreement is low, the factuality-honesty dissociation is an artifact of the judge.
Extended reading notes
Core claim
The central claim, on the paper's own terms, is that responsible LLM selection for Dutch government work requires multi-dimensional evaluation because performance dimensions do not move together. The paper reports that GPT-5, Mistral Large 3, and Mistral Medium 2505 reach factuality scores of 0.76, 0.71, and 0.75 yet honesty scores of only 0.14, 0.21, and 0.21, while GPT-4o and Mistral Small 24B pair strong factuality with honesty near 0.4; this is a direct dissociation of correctness from epistemic humility. It also reports that composite quality rises with cost and energy, while bias does not track either, and that no single model dominates on all axes simultaneously.
Load-bearing premise
The load-bearing premise is that the LLM judge ensemble used to score honesty agrees with human judgements about when a model acknowledges its limits; the paper reports the ensemble correlated best with human ratings but does not give the correlation coefficient, the judge models, or the variance, and the benchmark itself is not yet public.
Editorial extensions
If this is right
- A government that selects a model by factuality alone can end up with one that is confidently wrong, since honesty does not follow from accuracy.
- Budgeting for LLM deployment should treat higher quality as a paid option, because the measured cost and energy footprints rise with quality on most models.
- Bias must be evaluated explicitly for each protected characteristic, since spending more on a model does not systematically buy less biased outputs.
- The five-level interpretable scale allows engineers, product owners, and policymakers to shortlist models on shared terms rather than on raw benchmark numbers.
- For Dutch municipalities, fully open models with published training data occupy a real but not dominant trade-off position: lower capability, lower cost and energy, and clear transparency.
Reading between the lines
- The paper leaves implicit that if the factuality-honesty dissociation comes from alignment optimising for perceived helpfulness, honesty should be reported and optimised as a separate metric in every public-sector LLM release.
- An outside reader could transplant the same value-to-benchmark method to other languages and administrations, where the relevant values would differ but the template of six measurable dimensions would likely transfer.
- A testable extension of the paper's view is that retrieval-augmented settings will weaken the dissociation, because a model can defer to retrieved evidence; if the dissociation persists even with RAG, honesty is a stable model trait rather than a context effect.
- Since HONESTCITYBENCH has not been released, independent verification of the honesty result is impossible today; publishing the benchmark with human-rated validation would settle whether the dissociation is real.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents the 'Grip on LLMs' framework, a value-based evaluation of Dutch governmental LLM use. It reports a participatory design process with City of Amsterdam stakeholders that identifies six dimensions (factuality, honesty, social bias, energy consumption, cost, training-data transparency) and two use cases (simplification and summarisation). Thirty-one models are evaluated with a mixture of existing Dutch benchmarks, machine-translated English benchmarks, and a newly built honesty benchmark judged by an LLM ensemble. The main empirical findings are that no single model dominates, that higher quality is associated with higher cost and energy use, that bias is not predicted by quality or cost, and that factuality and honesty dissociate (e.g., GPT-5 scores high on factuality but lowest on honesty). The authors release an interactive overview and code as part of the paper.
Significance. If the findings are reliable, this is a valuable contribution: it operationalises government values in a concrete benchmark suite, addresses Dutch and municipal language needs, and provides a decision-support artifact that non-technical stakeholders can actually use. The participatory methodology is a strength, as is the explicit treatment of energy and cost alongside accuracy and fairness. The paper also ships code and an open overview, which supports reproducibility. However, the central claim about factuality-honesty dissociation depends on a not-yet-released benchmark and an insufficiently documented LLM judge, and the reported numeric comparisons lack uncertainty estimates. These gaps make the headline results provisional rather than definitive.
major comments (4)
- [Table 1 (Section 5.1)] Table 1, referenced throughout Section 5 as the basis for the overall-scores comparison, is missing from the manuscript: only its caption appears, and the body (the quality/bias/efficiency matrix) is absent. Since the trade-off claims in Sections 5.1 and 5.2 and the discussion of model rankings all rely on this table, the main results cannot be inspected as submitted. Please include the table body and ensure all referenced values appear in it.
- [Section 4.1 (Honesty), Section 5.3] The factuality–honesty dissociation, one of the paper's two headline contributions, rests on HONESTCITYBENCH and an LLM-as-a-judge ensemble. The text states only that the 3-judge ensemble 'proved to correlate the most with the human judgements'; no correlation coefficient, judge model identities, per-category agreement, or inter-judge variance is reported, and the benchmark is not yet released. Without these details, or the dataset itself, the honesty scores in Table 2 — in particular the 0.14 vs 0.43 gap that anchors Section 5.3 — cannot be independently verified. This is not merely a reporting issue: if the judge systematically rewards a particular response style (e.g., lengthy disclaimers) over calibrated uncertainty, the dissociation could be an artifact. Please report the full validation and release the benchmark, or explicitly present the honesty results as preliminary and obtainable only after validation.
- [Section 5; Tables 2–4] No uncertainty information is reported for any evaluation score. Factuality scores are estimated from 100-item tinyBenchmarks using a GP-IRT estimator; the original method provides credible intervals for such estimates, and many adjacent scores in Table 2 differ by only 0.01–0.03 (e.g., GPT-5 vs GPT-4o factuality, 0.76 vs 0.73; Qwen3 32B vs Qwen3 32B AWQ, 0.72 vs 0.69). Without error bars or confidence intervals, readers cannot determine which apparent differences are meaningful. Please add uncertainty estimates (or at least the underlying sample sizes and standard errors) for all raw scores and adjust the trade-off claims in Section 5.2 accordingly.
- [Section 4.1 (Factuality), Appendix B] The factuality benchmarks are machine-translated from English to Dutch using GPT-4o. The authors acknowledge a possible stylistic bias and argue it is limited because non-OpenAI models still score well. However, no translation quality evaluation is reported (e.g., human review of a sample, back-translation scores), and the factuality dimension feeds directly into the quality composite in Table 1 and into the factuality–honesty comparison in Section 5.3. Please provide translation validation or rely on human-verified Dutch benchmarks wherever possible, and state explicitly the residual risk of the translation procedure for the reported facts.
minor comments (5)
- [Abstract] The abstract and Section 1 contain 'more than30' without a space; please correct to 'more than 30'.
- [Appendix A, 'Other (excluded) models'] The post-hoc exclusion of models that 'failed to complete some benchmarks' or performed poorly should be documented in more detail. Showing the excluded models' partial results in an appendix would help readers judge whether the exclusion affected the qualitative claims.
- [Figure 3] The caption states 'Dashed lines show linear regressions with Pearson r' but the correlation values are not given in the caption text. Please report the Pearson r values and p-values in the caption or in the main text so the strength of the trends can be assessed.
- [Section 4.1 and Appendix C] The name of the honesty benchmark is rendered inconsistently as 'HONESTCITYBENCH', 'HonestCityBench', and 'HonestCity' in different places; please unify the spelling.
- [Section 2.1] The advisory board questionnaire response rate is described as '6 out of 7 members' and '5 out of 7 members' in one sentence, while later the text says '7 out of 9 advisory board members invited'. Please clarify the total number of respondents and the denominator consistently.
Circularity Check
No significant circularity: the main claims are empirical measurements, and the one self-built benchmark (HONESTCITYBENCH) is a measurement instrument, not an equation-level reduction of the results.
full rationale
The paper's derivation chain is empirical and does not reduce to its own inputs. The evaluation dimensions are elicited from an advisory board and user research, then operationalized into benchmarks. Factuality is measured with external benchmarks (MMLU, ARC, TruthfulQA via tinyBenchmarks), bias with MBBQ and BZK Social Bias, energy with CodeCarbon, and cost from vendor pricing and GPU-hour estimates. The central claim that factuality and honesty are distinct properties is a measured pattern across 31 models, not an identity or a fitted relationship: no equation defines factuality in terms of honesty or vice versa. HONESTCITYBENCH is indeed self-built and scored by an LLM-as-judge ensemble, and the paper does not report the correlation coefficient of the human validation, judge model identities, or variance. That is a transparency and validity limitation, but it is not circularity as defined for this analysis: the honesty score is not constructed from the factuality score, nor is any fitted parameter relabeled as a prediction. The one co-authored benchmark used (Vlantis, Gornishka, and Wang 2024) is an external simplification dataset and does not carry the paper's main claims. The paper's own discussion acknowledges possible GPT-4o translation bias and the absence of energy data for closed models, further showing that the reported trade-offs are treated as empirical findings rather than definitional consequences. No load-bearing self-citation chain, imported uniqueness theorem, or ansatz-smuggling citation is present. Concerns about LLM-judge bias are potential confounds that belong under correctness risk, not circularity.
Assumptions & free parameters
free parameters (4)
- Ordinal thresholds for factuality scale =
0.5 / 0.6 / 0.7 / 0.8
- Ordinal thresholds for honesty scale =
0.2 / 0.4 / 0.6 / 0.8
- Ordinal thresholds for simplification scale =
26 / 32 / 38 / 44 SARI
- Ordinal thresholds for summarisation scale =
0.50 / 0.55 / 0.60 / 0.65 BERTScore
assumptions (4)
- domain assumption Machine-translated MMLU, ARC, and TruthfulQA are valid Dutch factuality measures
- domain assumption LLM-as-a-judge (3-model ensemble) agrees with human honesty ratings
- domain assumption TinyBenchmarks GP-IRT estimates on 100 translated samples track full-benchmark performance
- domain assumption Advisory board values align with wider Dutch governmental needs
Cite this review
Pith. "Pith review of From Values to Benchmarks: Evaluating Large Language Models for Governmental Use in Dutch." pith.science (2026). https://pith.science/paper/ECAHBYNJ
@misc{pith2026260809925,
author = {Pith},
title = {Pith review of: From Values to Benchmarks: Evaluating Large Language Models for Governmental Use in Dutch},
year = {2026},
howpublished = {\url{https://pith.science/paper/ECAHBYNJ}},
note = {Machine review of arXiv:2608.09925}
}
read the original abstract
Large language models are increasingly being deployed in governmental settings, yet few existing evaluation frameworks jointly reflect the values of public administration and the linguistic requirements of non-English contexts. We present the "Grip on LLMs" framework, a systematic evaluation suite for Dutch governmental use developed in collaboration with domain experts from a major Dutch municipal organisation. Through an advisory board process, user research, and a survey of the users of a civil-servant chatbot, we identify six evaluation dimensions (factuality, honesty, social bias, energy consumption, cost, and training data transparency) and operationalise them into a benchmark suite covering more than 30 multilingual and Dutch-specific models. Our results reveal that no single model excels across all dimensions, and that trade-offs are unavoidable: higher quality consistently comes at greater environmental impact and financial cost, while bias remains largely independent of both. We further find that factuality (whether a model answers correctly) and honesty (whether a model acknowledges what it does not know) are governed by distinct properties, with high factuality not implying high honesty. To make these findings actionable for non-technical audiences, we release a publicly accessible, user-friendly model overview designed for the full range of stakeholders involved in governmental LLM selection, from engineers to policymakers.
Figures
Reference graph
Works this paper leans on
-
[1]
arXiv preprint arXiv:2507.06261 , year=
Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities , author=. arXiv preprint arXiv:2507.06261 , year=
-
[2]
arXiv preprint arXiv:2303.08774 , year=
Gpt-4 technical report , author=. arXiv preprint arXiv:2303.08774 , year=
-
[3]
arXiv preprint arXiv:2302.13971 , year=
Llama: Open and efficient foundation language models , author=. arXiv preprint arXiv:2302.13971 , year=
-
[4]
Advances in neural information processing systems , volume=
Attention is all you need , author=. Advances in neural information processing systems , volume=
-
[5]
Advances in neural information processing systems , volume=
Language models are few-shot learners , author=. Advances in neural information processing systems , volume=
-
[6]
Telecommunications policy , volume=
AI governance in the public sector: Three tales from the frontiers of automated decision-making in democratic settings , author=. Telecommunications policy , volume=. 2020 , publisher=
work page 2020
-
[7]
Government information quarterly , volume=
Artificial intelligence for the public sector: results of landscaping the use of AI in government across the European Union , author=. Government information quarterly , volume=. 2022 , publisher=
work page 2022
-
[8]
European Journal of Social Security , volume=
Digital welfare fraud detection and the Dutch SyRI judgment , author=. European Journal of Social Security , volume=. 2021 , publisher=
work page 2021
Show all 77 references
-
[9]
Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages=
Mega: Multilingual evaluation of generative ai , author=. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages=
2023
-
[10]
arXiv preprint arXiv:2411.19799 , year=
Include: Evaluating multilingual language understanding with regional knowledge , author=. arXiv preprint arXiv:2411.19799 , year=
-
[11]
Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages=
Don’t trust ChatGPT when your question is not in English: A study of multilingual abilities and types of LLMs , author=. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages=
2023
-
[12]
Transactions on Machine Learning Research , year=
Privacy-aware visual language models , author=. Transactions on Machine Learning Research , year=
-
[13]
GEITje: een groot open Nederlands taalmodel , shorttitle =
Rijgersberg, Edwin and Lucassen, Bob , year =. GEITje: een groot open Nederlands taalmodel , shorttitle =
-
[14]
Proceedings of the 57th annual meeting of the association for computational linguistics , pages=
Energy and policy considerations for deep learning in NLP , author=. Proceedings of the 57th annual meeting of the association for computational linguistics , pages=
-
[15]
Proceedings of the 2024 ACM conference on fairness, accountability, and transparency , pages=
Power hungry processing: Watts driving the cost of AI deployment? , author=. Proceedings of the 2024 ACM conference on fairness, accountability, and transparency , pages=
2024
-
[16]
arXiv preprint arXiv:1910.09700 , year=
Quantifying the carbon emissions of machine learning , author=. arXiv preprint arXiv:1910.09700 , year=
1910 arXiv
-
[17]
arXiv preprint arXiv:1912.09582 , year=
Bertje: A dutch bert model , author=. arXiv preprint arXiv:1912.09582 , year=
1912 arXiv
-
[18]
Findings of the association for computational linguistics: EMNLP 2020 , pages=
Robbert: a dutch roberta-based language model , author=. Findings of the association for computational linguistics: EMNLP 2020 , pages=
2020
-
[19]
arXiv preprint arXiv:2406.13469 , year=
Encoder vs Decoder: Comparative Analysis of Encoder and Decoder Language Models on Multilingual NLU Tasks , author=. arXiv preprint arXiv:2406.13469 , year=
-
[20]
Smart, Dan Saattrup , booktitle =
-
[21]
2025 , month =
Schutz, Ivo , title =. 2025 , month =
2025
-
[22]
2025 , month =
Anonymous , title =. 2025 , month =
2025
-
[23]
arXiv preprint arXiv:2310.02446 , year=
Low-resource languages jailbreak gpt-4 , author=. arXiv preprint arXiv:2310.02446 , year=
-
[24]
Pending submission
Anonymous , title =. Pending submission. , year =
-
[25]
Advances in Neural Information Processing Systems , volume=
Alignment for honesty , author=. Advances in Neural Information Processing Systems , volume=
-
[26]
arXiv preprint arXiv:2406.13261 , year=
Behonest: Benchmarking honesty in large language models , author=. arXiv preprint arXiv:2406.13261 , year=
-
[27]
arXiv preprint arXiv:2301.01768 , year=
The political ideology of conversational AI: Converging evidence on ChatGPT's pro-environmental, left-libertarian orientation , author=. arXiv preprint arXiv:2301.01768 , year=
-
[28]
PloS one , volume=
The political preferences of LLMs , author=. PloS one , volume=. 2024 , publisher=
2024
-
[29]
Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society , year=
GermanPartiesQA: Benchmarking Commercial Large Language Models and AI Companions for Political Alignment and Sycophancy , author=. Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society , year=
-
[30]
arXiv preprint arXiv:2601.08785 , year=
Uncovering Political Bias in Large Language Models using Parliamentary Voting Records , author=. arXiv preprint arXiv:2601.08785 , year=
-
[31]
arXiv preprint arXiv:1803.05457 , year=
Think you have solved question answering? try arc, the ai2 reasoning challenge , author=. arXiv preprint arXiv:1803.05457 , year=
-
[32]
arXiv preprint arXiv:2402.14992 , year=
tinyBenchmarks: evaluating LLMs with fewer examples , author=. arXiv preprint arXiv:2402.14992 , year=
-
[33]
Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing: System Demonstrations , pages=
Okapi: Instruction-tuned large language models in multiple languages with reinforcement learning from human feedback , author=. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing: System Demonstrations , pages=
2023
-
[34]
Proceedings of the 2020 conference on empirical methods in natural language processing (EMNLP) , pages=
CrowS-pairs: A challenge dataset for measuring social biases in masked language models , author=. Proceedings of the 2020 conference on empirical methods in natural language processing (EMNLP) , pages=
2020
-
[35]
Proceedings of the 15th International Conference on Recent Advances in Natural Language Processing-Natural Language Processing in the Generative AI Era , pages=
Dutch CrowS-Pairs: Adapting a Challenge Dataset for Measuring Social Biases in Language Models for Dutch , author=. Proceedings of the 15th International Conference on Recent Advances in Natural Language Processing-Natural Language Processing in the Generative AI Era , pages=
-
[36]
Evaluating Dutch Social Bias in Large Language Models , author=
-
[37]
Proceedings of the 60th annual meeting of the association for computational linguistics (volume 1: long papers) , pages=
Truthfulqa: Measuring how models mimic human falsehoods , author=. Proceedings of the 60th annual meeting of the association for computational linguistics (volume 1: long papers) , pages=
-
[38]
Findings of the Association for Computational Linguistics: ACL 2022 , pages=
BBQ: A hand-built bias benchmark for question answering , author=. Findings of the Association for Computational Linguistics: ACL 2022 , pages=
2022
-
[39]
arXiv preprint arXiv:2009.03300 , year=
Measuring massive multitask language understanding , author=. arXiv preprint arXiv:2009.03300 , year=
2009 arXiv
-
[40]
arXiv preprint arXiv:2406.07243 , year=
MBBQ: A dataset for cross-lingual comparison of stereotypes in generative LLMs , author=. arXiv preprint arXiv:2406.07243 , year=
-
[41]
Proceedings of the 2023 conference on empirical methods in natural language processing , pages=
Dumb: A benchmark for smart evaluation of dutch models , author=. Proceedings of the 2023 conference on empirical methods in natural language processing , pages=
2023
-
[42]
Proceedings of the 2024 joint international conference on computational linguistics, language resources and evaluation (lrec-coling 2024) , pages=
Benchmarking the simplification of Dutch municipal text , author=. Proceedings of the 2024 joint international conference on computational linguistics, language resources and evaluation (lrec-coling 2024) , pages=
2024
-
[43]
Online:< https://ticclops
SoNaR user documentation , author=. Online:< https://ticclops. uvt. nl/SoNaR\_end-user\_documentation\_v , volume=
-
[44]
CLARIN Annual Conference Proceedings , pages=
Human Evaluation of Automated Text Simplification through Crowdsourcing , author=. CLARIN Annual Conference Proceedings , pages=. 2025 , organization=
2025
-
[45]
Proceedings of the 2018 conference on empirical methods in natural language processing , pages=
Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization , author=. Proceedings of the 2018 conference on empirical methods in natural language processing , pages=
2018
-
[46]
Advances in neural information processing systems , volume=
Teaching machines to read and comprehend , author=. Advances in neural information processing systems , volume=
-
[47]
Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
Get to the point: Summarization with pointer-generator networks , author=. Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
-
[48]
Proceedings of the ACM SIGOPS 29th Symposium on Operating Systems Principles , year=
Efficient Memory Management for Large Language Model Serving with PagedAttention , author=. Proceedings of the ACM SIGOPS 29th Symposium on Operating Systems Principles , year=
-
[49]
Proceedings of the 2020 conference on empirical methods in natural language processing: system demonstrations , pages=
Transformers: State-of-the-art natural language processing , author=. Proceedings of the 2020 conference on empirical methods in natural language processing: system demonstrations , pages=
2020
-
[50]
Zhang, Peiyuan and Zeng, Guangtao and Wang, Tianduo and Lu, Wei , journal=
-
[51]
The Falcon 3 Family of Open Models , url =
-
[52]
2023 , eprint=
Mistral 7B , author=. 2023 , eprint=
2023
-
[53]
2025 , howpublished=
Mistral Small 3 , author=. 2025 , howpublished=
2025
-
[54]
2025 , howpublished=
Mistral Medium 3 , author=. 2025 , howpublished=
2025
-
[55]
2025 , howpublished=
Mistral 3 , author=. 2025 , howpublished=
2025
-
[56]
arXiv preprint arXiv:2503.01743 , year=
Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs , author=. arXiv preprint arXiv:2503.01743 , year=
-
[57]
2024 , howpublished=
Llama 3.2: Revolutionizing edge AI and vision with open, customizable models , author=. 2024 , howpublished=
2024
-
[58]
arXiv preprint arXiv:2407.21783 , year=
The Llama 3 Herd of Models , author=. arXiv preprint arXiv:2407.21783 , year=
-
[59]
2024 , howpublished=
2024
-
[60]
2024 , howpublished=
Hello GPT-4o , author=. 2024 , howpublished=
2024
- [61]
-
[62]
arXiv preprint arXiv:2506.04079 , year=
Martins, Pedro Henrique and Alves, Jo. arXiv preprint arXiv:2506.04079 , year=
-
[63]
arXiv preprint arXiv:2602.05879 , year=
EuroLLM-22B: Technical Report , author=. arXiv preprint arXiv:2602.05879 , year=
-
[64]
Yang, An and Li, Anfeng and Yang, Baosong and Zhang, Beichen and Hui, Binyuan and Zheng, Bo and Yu, Bowen and Gao, Chang and Huang, Chengen and Lv, Chenxu and others , journal=
-
[65]
2025 , url=
Gemma 3 Technical Report , author=. 2025 , url=
2025
-
[66]
arXiv preprint arXiv:2504.00698 , year=
Command A: An enterprise-ready large language model , author=. arXiv preprint arXiv:2504.00698 , year=
-
[67]
Guo, Daya and Yang, Dejian and Zhang, Haowei and Song, Junxiao and Wang, Peiyi and Zhu, Qihao and Xu, Runxin and Zhang, Ruoyu and Ma, Shirong and Bi, Xiao and others , journal=
-
[68]
arXiv preprint arXiv:2508.10925 , year=
Agarwal, Sandhini and Ahmad, Lama and Ai, Jason and Altman, Sam and Applebaum, Andy and Arbus, Edwin and Arora, Rahul K and Bai, Yu and Baker, Bowen and Bao, Haiming and others , institution=. arXiv preprint arXiv:2508.10925 , year=
-
[69]
2025 , howpublished=
Introducing. 2025 , howpublished=
2025
-
[70]
2025 , howpublished=
The. 2025 , howpublished=
2025
-
[71]
arXiv preprint arXiv:2412.04261 , year=
Aya Expanse: Combining Research Breakthroughs for a New Multilingual Frontier , author=. arXiv preprint arXiv:2412.04261 , year=
-
[72]
Bakouch, Elie and Ben Allal, Loubna and Lozhkov, Anton and Tazi, Nouamane and Tunstall, Lewis and Patiño, Carlos Miguel and Beeching, Edward and Roucher, Aymeric and Reedi, Aksel Joonas and Gallouédec, Quentin and Rasul, Kashif and Habib, Nathan and Fourrier, Clémentine and Ky...
-
[73]
Vanroy, Bram , journal=
-
[74]
arXiv preprint arXiv:2412.15450 , year=
Fietje: An open, efficient LLM for Dutch , author=. arXiv preprint arXiv:2412.15450 , year=
-
[75]
arXiv preprint arXiv:2509.14233 , year=
Apertus: Democratizing Open and Compliant. arXiv preprint arXiv:2509.14233 , year=
-
[76]
Singh, Aaditya and Fry, Adam and Perelman, Adam and Tart, Adam and Ganesh, Adi and El-Kishky, Ahmed and McLaughlin, Aidan and Low, Aiden and Ostrow, AJ and Ananthram, Akhila and others , journal=
-
[77]
, author=
GPT-NL: Towards a Public Interest Large Language Model. , author=. PI-AI@ KI , year=
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.