Pith. sign in

REVIEW 3 major objections 6 minor 55 references

Models that win the dialect reward lose the human vote

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · glm-5.2

2026-07-09 02:59 UTC pith:OUUUDZA4

load-bearing objection Solid empirical contribution on dialect adaptation; the reward-quality gap is real but narrower than claimed the 3 major comments →

arxiv 2607.07669 v1 pith:OUUUDZA4 submitted 2026-07-08 cs.CL cs.AI

DiaLLM: An Investigation into the Robustness-Generation Gap in English Dialect Adaptation

classification cs.CL cs.AI
keywords dialectal generationreward-quality gaprobustness-generation dissociationLLM alignmentdialect adaptationeWAVEGRPODPO
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper investigates whether large language models can be adapted to not merely understand but actively produce regional English dialects (Australian, Indian, and Northern British English). The authors build a full training pipeline—continual pretraining on the International Corpus of English, supervised fine-tuning, and three alignment strategies (DPO, GRPO, GSPO)—across three model families. They find that benchmark performance and actual dialectal generation are dissociated: benchmarks are driven by pretraining and fine-tuning, while alignment reshapes generation in ways benchmarks cannot detect. More critically, they identify a reward-quality gap: the alignment method that most aggressively optimises a dialectal feature reward (GRPO) produces the fewest independently verifiable dialectal markers and is least preferred by both human annotators and LLM judges. The reward signal, derived from a typological database of attested features, can be gamed to increase surface feature counts without producing output that sounds authentically dialectal to a human listener.

Core claim

The central discovery is a reward-quality gap in dialectal generation. When a reward function based on surface-level dialectal feature detection is optimised aggressively, the model learns to satisfy the classifier without producing genuinely dialectal output. The method that maximises the reward (GRPO) yields the lowest density of independent linguistic markers and the lowest human preference, demonstrating that optimising a feature-count proxy does not correspond to perceived dialectal authenticity. Dialectal marking is established primarily during supervised fine-tuning and only redistributed—often destructively—by alignment methods that target the reward signal.

What carries the argument

The DiaLLM pipeline: continual pretraining on ICE, two post-training threads (implicit broad adaptation vs. explicit variety-targeted adaptation), three alignment methods (DPO, GRPO, GSPO), a dialect feature classifier trained on eWAVE-derived features used as a reward signal, and an independent rule-based linguistic analysis tool measuring lexical, orthographic, and morphosyntactic markers in generated text.

Load-bearing premise

The reward function and the training diagnostic both derive from the same eWAVE feature inventory—a typological database designed to catalogue which features are attested across varieties, not to serve as a reward signal for generative training. The paper itself acknowledges this mismatch, noting that eWAVE was never designed to measure perceptual dialectal quality.

What would settle it

If a reward function based on eWAVE feature density were replaced with one calibrated to human perceptual judgements of dialectal authenticity, and GRPO under that new reward produced outputs preferred over SFT, the reward-quality gap would be narrowed or closed, showing the gap is an artefact of the specific reward basis rather than a structural property of feature-density optimisation.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Reward functions for stylistic or dialectal generation cannot rely on surface-feature atlases; they must incorporate perceptual or holistic quality signals to avoid reward hacking.
  • Benchmark scores on dialectal NLU tasks are insufficient to evaluate whether a model can produce authentic dialectal output; generation-specific evaluation is necessary.
  • Supervised fine-tuning on dialectal data establishes the bulk of surface dialectal marking, suggesting that alignment-stage methods may be operating on a signal already saturated or redistributed rather than introduced.
  • Dialects with subtle or register-flexible features (e.g., Australian English) are harder to elicit and evaluate than those with dense morphosyntactic markers (e.g., Indian English), requiring variety-aware evaluation strategies.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The reward-quality gap observed here likely generalises beyond dialect to any controllable-generation task where the reward is a checklist of surface features (e.g., formality, sentiment, style transfer), suggesting that feature-density rewards are structurally vulnerable to Goodhart's law.
  • If human preference for dialectal output is non-monotonic—features help against a standard baseline but reward-driven maximisation hurts—then the optimal reward weight for dialectal features may be below the point of maximum feature density, implying a need for preference-calibrated rather than density-maximising reward objectives.
  • The failure of the eWAVE-based classifier to generalise to conversational outputs (collapsing en-UK to en-AU predictions) suggests that dialect identification models trained on formal or task-specific corpora may be unsuitable for evaluating open-ended generation, a problem that likely affects other dialect evaluation pipelines.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. This paper introduces DiaLLM, a full-pipeline dialect adaptation framework that applies continual pretraining (CPT) on the International Corpus of English (ICE) to three open-weight LLM families (Llama, Qwen, Gemma), followed by two post-training paradigms (implicit and explicit) and three alignment methods (DPO, GRPO, GSPO). The authors report two main findings: (1) a robustness-generation dissociation, where benchmarks are driven by CPT/SFT while alignment reshapes generation in ways benchmarks miss; and (2) a reward-quality gap, where GRPO most aggressively optimises the dialectal reward but is least preferred by human/LLM evaluators and produces the fewest independent surface dialectal markers. The paper is notably transparent about its limitations, including the eWAVE circularity, the restricted human evaluation, and the en-UK classifier failure.

Significance. The paper tackles an important and underexplored problem: dialectal generation (as opposed to mere robustness) in LLMs. The controlled comparison across three model families, two post-training paradigms, and three alignment methods is a substantial experimental contribution. The release of all code, checkpoints, preference datasets, and an independent linguistic analysis toolkit is a significant strength that enhances reproducibility. The reward-quality gap finding, if robustly supported, has practical implications for how dialectal reward signals are designed. The paper commendably ships an independent, rule-based linguistic analysis (Appendix J) as a non-circular check on the reward signal, which is a methodological strength.

major comments (3)
  1. §5.3, Table 11 (Appendix J): The independent linguistic analysis—the key non-circular evidence stream for the reward-quality gap—shows mixed results for Gemma. GRPO density is 1.69 vs GSPO at 1.44, meaning GRPO is NOT the lowest for Gemma. The paper acknowledges this ('the ordering is mixed only for Gemma'), but the abstract states the gap is corroborated 'most clearly on two of the three families,' while §5.3 claims the independent analysis 'corroborates this pattern across all three families.' The §5.3 phrasing overstates the evidence. The headline claim should be consistently scoped to Llama and Qwen for the independent linguistic analysis, with the Gemma exception clearly noted in the main text, not only in the appendix.
  2. Table 3, §5.4: The human and LLM preference evaluation is limited to Llama-3.1-8B with only two annotators per variety, and en-AU inter-annotator agreement is near chance (AC1=0.02, Table 9). The claim that 'the method that most aggressively optimises the dialectal reward is not preferred by human evaluators' (abstract, §5.4) is securely demonstrated only for en-IN and en-UK on Llama. The paper should explicitly qualify in the abstract and §5.4 that the human preference evidence is restricted to one model family and that en-AU human preference results are indicative rather than conclusive, as the current abstract phrasing ('is not preferred by human evaluators') implies broader coverage than the evidence supports.
  3. §3.4.4, Eq. (1) and Appendix E: The reward signal and the feature-density training diagnostic (Table 6) both derive from the same eWAVE feature classifier. The paper acknowledges this in the Limitations. However, the claim in §5.3 that 'GRPO most aggressively optimises the dialectal reward' is partially circular: GRPO optimises the eWAVE-based reward, and Table 6 measures success using the same eWAVE feature space. The variety classifier (Appendix F) and the independent linguistic analysis (Appendix J) provide non-circular checks, but the en-UK variety classifier fails entirely (Table 7), and the independent analysis is mixed for Gemma (see comment 1). The paper should more clearly flag in §5.3 (not just in Limitations) that the feature-density diagnostic is a training diagnostic, not an independent evaluation, to prevent readers from treating Table 6 as evidence for the reward-quality.
minor comments (6)
  1. Table 1: The 'SFT d' notation is used in the table header but not defined until the caption. Consider defining it in the text when first referenced in §5.1.
  2. §3.2: The paper states 'approximately 20 million tokens' for the CPT corpus. Table 4 gives 19,916,246. This is fine, but the raw token total (38,012,253) and retention rate (52.4%) should be mentioned in the main text for context on corpus size.
  3. Table 3 caption: 'Parenthesised t = tied judgements excluded from win rates' could be clearer; consider stating 't = number of tied judgements, excluded from win-rate computation.'
  4. §4.3: The justification for selecting Llama-3.1-8B for generation evaluation ('shows the most consistent dialect-sensitivity gains') could benefit from a brief citation to the specific benchmark results supporting this claim.
  5. Appendix I, Table 9: The prevalence paradox explanation for en-UK T3 (AC1=0.57 despite alpha=-0.19) is noted, but it would help to briefly explain why AC1 is preferred over kappa/alpha here, given the near-unanimous preference pattern.
  6. Figure 2: The y-axis scales differ between the two panels (0-1.5 for eWAVE density, 0-3.0 for independent markers). This is acceptable but could be noted to avoid visual miscomparison.

Simulated Author's Rebuttal

3 responses · 0 unresolved

We thank the referee for a careful and constructive report. All three major comments identify genuine scoping issues where our main-text claims overstate the evidence relative to what the abstract and limitations section already acknowledge. We agree with each point and will revise accordingly. No standing objections remain.

read point-by-point responses
  1. Referee: §5.3 overstates the independent linguistic analysis as corroborating the reward-quality gap 'across all three families,' while the abstract correctly scopes it to 'most clearly on two of the three families.' The Gemma exception (GRPO density 1.69 vs GSPO 1.44) is noted only in the appendix.

    Authors: The referee is correct. The §5.3 sentence 'corroborates this pattern across all three families' is inconsistent with both the abstract's more careful phrasing ('most clearly on two of the three families') and the data in Table 11, where GRPO is not the lowest-density method for Gemma. We will revise §5.3 to state that the independent linguistic analysis corroborates the reward-quality gap for Llama and Qwen, with Gemma as a noted exception, and will reference the Gemma-specific numbers directly in the main text rather than relegating them to the appendix. The abstract phrasing will remain as written. revision: yes

  2. Referee: The human preference claim ('not preferred by human evaluators') in the abstract and §5.4 implies broader coverage than the evidence supports: evaluation is limited to Llama-3.1-8B with two annotators per variety, and en-AU agreement is near chance (AC1=0.02).

    Authors: We agree. The abstract's phrasing ('is not preferred by human evaluators') does not convey the scope restriction that the body and limitations section already acknowledge. We will revise the abstract to specify that human preference evidence is drawn from Llama-3.1-8B only, and will add an explicit note in §5.4 that en-AU human preference results are indicative rather than conclusive given the near-chance inter-annotator agreement (AC1=0.02, Table 9). The claim will be scoped to en-IN and en-UK for the human evaluation component. revision: yes

  3. Referee: The feature-density diagnostic (Table 6) and the reward signal both derive from the same eWAVE classifier, making the claim that 'GRPO most aggressively optimises the dialectal reward' partially circular. This should be flagged in §5.3, not just in Limitations.

    Authors: The referee is right that Table 6 is a training diagnostic, not an independent evaluation, and §5.3 should make this explicit rather than deferring the clarification to the Limitations section. We will add a sentence in §5.3 stating that the feature-density diagnostic measures optimisation of the training reward signal itself and therefore cannot serve as independent evidence for output quality; we will cross-reference the variety classifier (Appendix F) and the independent linguistic analysis (Appendix J) as the non-circular checks, while noting the en-UK classifier failure and the Gemma exception in the same location. Table 6's caption already labels it as a training diagnostic, but we will reinforce this in the main text. revision: yes

Circularity Check

1 steps flagged

No significant circularity: the reward-quality gap is supported by independent linguistic analysis with separate detectors, and the paper transparently acknowledges the shared eWAVE basis between reward and training diagnostic.

specific steps
  1. fitted input called prediction [Section 3.4.4 (Eq. 1) and Appendix E (Table 6)]
    "ϕdial(y) is the log-sum dialectal reward from the feature classifier (Section 3.4.3)... Our reward function and generation evaluation share the same eWAVE-derived feature inventory, so feature density gains—retained in Appendix E as a training diagnostic only—reflect alignment with the training feature space rather than independently verified authentic dialect use."

    The dialect feature classifier (Section 3.4.3) is trained on MultiVALUE-transformed NLU data using eWAVE features, and the same classifier's log-sum output ϕdial(y) serves as the reward signal R in Eq. 1. The training diagnostic in Appendix E (Table 6) measures the same quantity—log(1 + Σσ(logits)) over the same eWAVE feature set. When the paper states 'GRPO most aggressively optimises the reward signal, achieving the highest eWAVE feature density across all three families' (Section 5.3), this is partially circular: GRPO optimises ϕdial by construction, and Table 6 reports ϕdial, so GRPO ranking highest is expected by design. However, this circularity is explicitly acknowledged and confined to the training diagnostic only. The paper does not use Table 6 as primary evaluation; instead, it引入

full rationale

The paper's central claim—the reward-quality gap—rests on three evidence streams: (1) human/LLM preference (Table 3), which is independent of the reward; (2) variety classifier accuracy (Appendix F), which uses a separate DeBERTa model trained on BESSTIE data, not the reward classifier; and (3) independent linguistic analysis (Appendix J, Table 11), which uses rule-based detectors (curated lexicons, spelling pairs, spaCy POS parses) that are explicitly separate from the eWAVE reward classifier. The paper states: 'The reward signal and the variety classifier (Sec. 5.3) both derive from the same eWAVE feature space, so neither is an independent check... We therefore analyse the generation outputs with a separate, rule-based instrument that does not reuse the reward classifier.' The only circular element is that the training diagnostic (Table 6) reports the same quantity that GRPO optimises, making 'GRPO achieves highest feature density' a near-tautology. But the paper is transparent about this, labels Table 6 as 'training diagnostic only,' and does not use it as primary evidence for the reward-quality gap. The load-bearing claim—that GRPO is not preferred by humans despite maximising reward—is supported by independent evidence. The BESSTIE self-citation (Srirag et al., 2025b, co-authored by present authors) provides the variety classifier and benchmark data, but this is an evaluation instrument, not a premise that forces the conclusion. Score 2: one minor partial circularity in the training diagnostic, acknowledged and not load-bearing for the central claim.

Axiom & Free-Parameter Ledger

5 free parameters · 4 axioms · 0 invented entities

The paper introduces no new entities (particles, forces, dimensions, etc.). It combines existing methods (CPT, SFT, DPO, GRPO, GSPO) and existing resources (ICE, eWAVE, MultiVALUE, BESSTIE) in a new configuration. The dialect feature classifier is a trained model, not a postulated entity. The main free parameter is lambda=0.80, selected by inspection of reward trajectories on Llama only then applied to all families. The key domain assumptions (MultiVALUE authenticity, eWAVE as reward basis, two-annotator reliability, ICE register coverage) are all acknowledged by the paper itself as limitations.

free parameters (5)
  • lambda (reward weight) = 0.80
    Selected by running full training for Llama 3.1-8B across lambda in {0.50, 0.66, 0.80} and inspecting reward trajectories (Appendix C). Fixed at 0.80 for all families and methods.
  • GaLore rank = 1024
    Set for continual pretraining; scale 0.25, subspace update interval 500 steps (Appendix C).
  • DPO beta = 0.1
    Standard DPO temperature parameter (Appendix C).
  • GRPO/GSPO beta = 0.02
    KL penalty coefficient for GRPO/GSPO (Appendix C).
  • Per-feature temperature calibration = T in [0.5, 3.0] per feature
    Grid-searched per feature on held-out validation set to minimize ECE (Appendix D).
axioms (4)
  • domain assumption MultiVALUE transformations produce authentic dialectal variants suitable for SFT training data.
    Section 3.3: preferred completions are converted to dialectal variants using MultiVALUE. The paper acknowledges in Limitations that 'MultiVALUE transformations may introduce stereotypical or overgeneralised patterns.'
  • domain assumption eWAVE features constitute an appropriate basis for a dialectal reward signal.
    Section 3.4.3: the dialect feature classifier is trained on 135 eWAVE-derived features. The paper itself states this is a mismatch: 'eWAVE was designed as a typological database... not as a reward signal for generative training' (Limitations).
  • domain assumption Two native or near-native speakers per dialect can provide reliable dialectal preference judgements.
    Section 4.3: human evaluation uses two annotators per variety. Inter-annotator agreement is modest overall, with en-AU near chance (AC1=0.02), challenging this assumption.
  • domain assumption ICE corpus provides adequate dialectal exposure for continual pretraining.
    Section 3.2: ICE consists primarily of formal and semi-formal texts. The paper acknowledges models 'may not generalise to informal or social media registers' (Limitations).

pith-pipeline@v1.1.0-glm · 25400 in / 3258 out tokens · 324692 ms · 2026-07-09T02:59:20.412765+00:00 · methodology

0 comments
read the original abstract

Large language models increasingly \emph{understand} dialectal English, yet still \emph{produce} only standard, US-leaning English, leaving dialectal generation, the harder half of the problem, largely unaddressed. We introduce \textbf{DiaLLM}, which continually pretrains three open-weight language model families on the International Corpus of English and applies implicit and explicit post-training paradigms, each combined with three model alignment strategies, giving the first controlled comparison of these components across Australian, Indian, and Northern British English. Our results reveal that dialectal robustness and generation are \emph{dissociated}: benchmarks are shaped by continual pretraining and SFT, while alignment visibly reshapes generation in ways benchmarks do not capture. Explicit variety-targeted adaptation produces output reliably recognised as dialectal and preferred over broad alignment, yet the method that most aggressively optimises the dialectal reward is not preferred by human evaluators. Independent linguistic analysis corroborates this reward-quality gap, most clearly on two of the three families. No single alignment method dominates, and closing the gap will require richer reward designs and continued investment in dialectal resources. We release all code, checkpoints, and preference datasets.

Figures

Figures reproduced from arXiv: 2607.07669 by Adarsh Kappiyath, Aditya Joshi, Dipankar Srirag, Diptesh Kanojia, Jordan Painter, Lu Yin.

Figure 1
Figure 1. Figure 1: Overview of the DiaLLM pipeline: continual pretraining on ICE followed by either implicit adapta￾tion (standard SFT + alignment) or explicit adaptation (dialectal SFT + variety-targeted alignment), with DPO, GRPO, and GSPO compared across both paradigms. (Gururangan et al., 2020), robustness to dialectal variation within a single language remains com￾paratively underexplored. Existing work is con￾centrated… view at source ↗
Figure 2
Figure 2. Figure 2: eWAVE reward density (left; reproduced from [PITH_FULL_IMAGE:figures/full_fig_p017_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Dialectal-marker density across the pipeline [PITH_FULL_IMAGE:figures/full_fig_p017_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Feature-type composition by variety (total [PITH_FULL_IMAGE:figures/full_fig_p018_4.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

55 extracted references · 55 canonical work pages · 2 internal anchors

  1. [1]

    The electronic world atlas of varieties of

  2. [2]

    2024 , eprint=

    Phi-4 Technical Report , author=. 2024 , eprint=

  3. [3]

    Demographic Dialectal Variation in Social Media: A Case Study of A frican- A merican E nglish

    Blodgett, Su Lin and Green, Lisa and O ' Connor, Brendan. Demographic Dialectal Variation in Social Media: A Case Study of A frican- A merican E nglish. Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing. 2016. doi:10.18653/v1/D16-1120

  4. [4]

    The Risk of Racial Bias in Hate Speech Detection

    Sap, Maarten and Card, Dallas and Gabriel, Saadia and Choi, Yejin and Smith, Noah A. The Risk of Racial Bias in Hate Speech Detection. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019. doi:10.18653/v1/P19-1163

  5. [5]

    2024 , eprint=

    Dialect prejudice predicts AI decisions about people's character, employability, and criminality , author=. 2024 , eprint=

  6. [6]

    Rejected Dialects: Biases Against A frican A merican Language in Reward Models

    Mire, Joel and Aysola, Zubin Trivadi and Chechelnitsky, Daniel and Deas, Nicholas and Zerva, Chrysoula and Sap, Maarten. Rejected Dialects: Biases Against A frican A merican Language in Reward Models. Findings of the Association for Computational Linguistics: NAACL 2025. 2025. doi:10.18653/v1/2025.findings-naacl.417

  7. [7]

    2020 , eprint=

    Don't Stop Pretraining: Adapt Language Models to Domains and Tasks , author=. 2020 , eprint=

  8. [8]

    Me-LLaMA: Medical Foundation Large Language Models for Comprehensive Text Analysis and Beyond , journal =

    Xie, Qianqian and Chen, Qingyu and Chen, Aokun and Peng, Cheng and Hu, Yan and Lin, Fongci and Peng, Xueqing and Huang, Jimin and Zhang, Jeffrey and Keloth, Vipina and Zhou, Xinyu and Qian, Lingfei and He, Huan and Shung, Dennis and Ohno-Machado, Lucila and Wu, Yonghui and Qi, Wang and Bian, Jiang , year =. Me-LLaMA: Medical Foundation Large Language Mode...

  9. [9]

    2023 , eprint=

    Lawyer LLaMA Technical Report , author=. 2023 , eprint=

  10. [10]

    2024 , eprint=

    Q-GaLore: Quantized GaLore with INT4 Projection and Layer-Adaptive Low-Rank Gradients , author=. 2024 , eprint=

  11. [11]

    2024 , eprint=

    GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection , author=. 2024 , eprint=

  12. [12]

    2022 , eprint=

    VALUE: Understanding Dialect Disparity in NLU , author=. 2022 , eprint=

  13. [13]

    2024 , eprint=

    AAVENUE: Detecting LLM Biases on NLU Tasks in AAVE via a Novel Benchmark , author=. 2024 , eprint=

  14. [14]

    2025 , eprint=

    EnDive: A Cross-Dialect Benchmark for Fairness and Performance in Large Language Models , author=. 2025 , eprint=

  15. [15]

    2025 , eprint=

    Assessing Dialect Fairness and Robustness of Large Language Models in Reasoning Tasks , author=. 2025 , eprint=

  16. [16]

    2024 , eprint=

    DIALECTBENCH: A NLP Benchmark for Dialects, Varieties, and Closely-Related Languages , author=. 2024 , eprint=

  17. [17]

    2025 , eprint=

    BESSTIE: A Benchmark for Sentiment and Sarcasm Classification for Varieties of English , author=. 2025 , eprint=

  18. [18]

    2023 , eprint=

    UltraFeedback: Boosting Language Models with High-quality Feedback , author=. 2023 , eprint=

  19. [19]

    2023 , eprint=

    Instruction-Following Evaluation for Large Language Models , author=. 2023 , eprint=

  20. [20]

    2022 , eprint=

    Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them , author=. 2022 , eprint=

  21. [21]

    2021 , eprint=

    Measuring Mathematical Problem Solving With the MATH Dataset , author=. 2021 , eprint=

  22. [22]

    2023 , eprint=

    GPQA: A Graduate-Level Google-Proof Q&A Benchmark , author=. 2023 , eprint=

  23. [23]

    2024 , eprint=

    MuSR: Testing the Limits of Chain-of-thought with Multistep Soft Reasoning , author=. 2024 , eprint=

  24. [24]

    2024 , eprint=

    MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark , author=. 2024 , eprint=

  25. [25]

    Tom B. Brown and Benjamin Mann and Nick Ryder and Melanie Subbiah and Jared Kaplan and Prafulla Dhariwal and Arvind Neelakantan and Pranav Shyam and Girish Sastry and Amanda Askell and Sandhini Agarwal and Ariel Herbert. Language Models are Few-Shot Learners , journal =. 2020 , url =. 2005.14165 , timestamp =

  26. [26]

    2023 , eprint=

    LLaMA: Open and Efficient Foundation Language Models , author=. 2023 , eprint=

  27. [27]

    Proceedings of the 2021

    Bender, Emily M. and Gebru, Timnit and McMillan-Major, Angelina and Shmitchell, Shmargaret , title =. Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency , pages =. 2021 , isbn =. doi:10.1145/3442188.3445922 , abstract =

  28. [28]

    ACM Computing Surveys , volume=

    Natural language processing for dialects of a language: A survey , author=. ACM Computing Surveys , volume=. 2025 , publisher=

  29. [29]

    The State and Fate of Linguistic Diversity and Inclusion in the NLP World

    Joshi, Pratik and Santy, Sebastin and Budhiraja, Amar and Bali, Kalika and Choudhury, Monojit. The State and Fate of Linguistic Diversity and Inclusion in the NLP World. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020. doi:10.18653/v1/2020.acl-main.560

  30. [30]

    World Englishes , volume =

    Greenbaum, Sidney and Nelson, Gerald , title =. World Englishes , volume =. doi:https://doi.org/10.1111/j.1467-971X.1996.tb00088.x , url =. https://onlinelibrary.wiley.com/doi/pdf/10.1111/j.1467-971X.1996.tb00088.x , abstract =

  31. [31]

    2022 , eprint=

    Training language models to follow instructions with human feedback , author=. 2022 , eprint=

  32. [32]

    2024 , eprint=

    Direct Preference Optimization: Your Language Model is Secretly a Reward Model , author=. 2024 , eprint=

  33. [33]

    2019 , eprint=

    Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks , author=. 2019 , eprint=

  34. [34]

    2020 , eprint=

    COMET: A Neural Framework for MT Evaluation , author=. 2020 , eprint=

  35. [35]

    2024 , eprint=

    DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models , author=. 2024 , eprint=

  36. [36]

    BERT : Pre-training of Deep Bidirectional Transformers for Language Understanding

    Devlin, Jacob and Chang, Ming-Wei and Lee, Kenton and Toutanova, Kristina. BERT : Pre-training of Deep Bidirectional Transformers for Language Understanding. Proceedings of the 2019 Conference of the North A merican Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers). 2019. doi:10.18653/v...

  37. [37]

    Multi- VALUE : A Framework for Cross-Dialectal E nglish NLP

    Ziems, Caleb and Held, William and Yang, Jingfeng and Dhamala, Jwala and Gupta, Rahul and Yang, Diyi. Multi- VALUE : A Framework for Cross-Dialectal E nglish NLP. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. doi:10.18653/v1/2023.acl-long.44

  38. [38]

    Proceedings of the 2018

    Wang, Alex and Singh, Amanpreet and Michael, Julian and Hill, Felix and Levy, Omer and Bowman, Samuel. GLUE : A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding. Proceedings of the 2018 EMNLP Workshop B lackbox NLP : Analyzing and Interpreting Neural Networks for NLP. 2018. doi:10.18653/v1/W18-5446

  39. [39]

    A Framework for Few-Shot Language Model Evaluation , month =

    Gao, Leo and Tow, Jonathan and Abbasi, Baber and Biderman, Stella and Black, Sid and DiPofi, Anthony and Foster, Charles and Golding, Laurence and Hsu, Jeffrey and. A Framework for Few-Shot Language Model Evaluation , month =. 2023 , publisher =. doi:10.5281/zenodo.10256836 , url =

  40. [40]

    de Souza, Jos \'e G

    Rei, Ricardo and C. de Souza, Jos \'e G. and Alves, Duarte and Zerva, Chrysoula and Farinha, Ana C and Glushkova, Taisiya and Lavie, Alon and Coheur, Luisa and Martins, Andr \'e F. T. COMET -22: Unbabel- IST 2022 Submission for the Metrics Shared Task. Proceedings of the Seventh Conference on Machine Translation. 2022. doi:10.18653/v1/2022.wmt-1.52

  41. [41]

    2025 , eprint=

    Group Sequence Policy Optimization , author=. 2025 , eprint=

  42. [42]

    Low-Resource Dialect Adaptation of Large Language Models: A

    Eeham Khan and Firas Saidani and Owen Van Esbroeck and Richard Khoury and Leila Kosseim , year=. Low-Resource Dialect Adaptation of Large Language Models: A. 2510.22747 , archivePrefix=

  43. [43]

    Dynamic Multi-Reward Weighting for Multi-Style Controllable Generation

    De Langis, Karin and Koo, Ryan and Kang, Dongyeop. Dynamic Multi-Reward Weighting for Multi-Style Controllable Generation. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. doi:10.18653/v1/2024.emnlp-main.386

  44. [44]

    TADA : Task Agnostic Dialect Adapters for E nglish

    Held, William and Ziems, Caleb and Yang, Diyi. TADA : Task Agnostic Dialect Adapters for E nglish. Findings of the Association for Computational Linguistics: ACL 2023. 2023. doi:10.18653/v1/2023.findings-acl.51

  45. [45]

    Task-Agnostic Low-Rank Adapters for Unseen E nglish Dialects

    Xiao, Zedian and Held, William and Liu, Yanchen and Yang, Diyi. Task-Agnostic Low-Rank Adapters for Unseen E nglish Dialects. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. doi:10.18653/v1/2023.emnlp-main.487

  46. [46]

    International Journal of Language, Translation and Intercultural Communication , volume=

    Australian English: Its evolution and current state , author=. International Journal of Language, Translation and Intercultural Communication , volume=

  47. [47]

    Predicting the Target Word of Game-playing Conversations using a Low-Rank Dialect Adapter for Decoder Models

    Srirag, Dipankar and Joshi, Aditya and Eisenstein, Jacob. Predicting the Target Word of Game-playing Conversations using a Low-Rank Dialect Adapter for Decoder Models. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 2: Short Papers). 2025. doi:1...

  48. [48]

    2024 , eprint=

    The Llama 3 Herd of Models , author=. 2024 , eprint=

  49. [49]

    2025 , eprint=

    Qwen3 Technical Report , author=. 2025 , eprint=

  50. [50]

    2025 , eprint=

    Gemma 3 Technical Report , author=. 2025 , eprint=

  51. [51]

    DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding Sharing

    Pengcheng He and Jianfeng Gao and Weizhu Chen , title =. CoRR , volume =. 2021 , url =. 2111.09543 , timestamp =

  52. [52]

    2024 , editor =

    Zhao, Jiawei and Zhang, Zhenyu and Chen, Beidi and Wang, Zhangyang and Anandkumar, Anima and Tian, Yuandong , booktitle =. 2024 , editor =

  53. [53]

    Gwet , title =

    Kilem L. Gwet , title =. British Journal of Mathematical and Statistical Psychology , volume =

  54. [54]

    Gwet , title =

    Tinakon Wongpakaran and Nahathai Wongpakaran and Derek Wedding and Kilem L. Gwet , title =. Journal of Research in Nursing , volume =

  55. [55]

    Klaus Krippendorff , title =