Pith. sign in

REVIEW 3 major objections 2 minor 299 references

Leveraging Pretrained Language Models as Energy Functions for Glauber Dynamics Text Diffusion

T0 review · 3 major / 2 minor · reviewed 2026-05-08 · grok-4.3

Pith's one-line read Pretrained language models can serve directly as energy functions to guide Glauber dynamics sampling in discrete text diffusion.

desk verdict The paper's fresh move is treating a frozen pretrained LM as the energy function for Glauber dynamics in discrete text diffusion, but the abstract and stress-test leave the mixing and sampling validity unaddressed. read the letter →

arxiv 2605.04291 v1 submitted 2026-05-05 cs.LG

classification cs.LG
keywords discretediffusionGlauberdynamicspretrainedlanguagemodelsenergyfunctionstextgenerationzero-shotreasoningplanningtasks
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper proposes replacing the usual uniform transition kernel in a discrete diffusion process with an energy function drawn from an existing pretrained causal or masked language model. When the pretrained model is interpreted as defining the stationary distribution, Glauber dynamics produce higher-quality text samples than diffusion models trained from scratch. With UL2 as the energy function the approach exceeds earlier diffusion language models and reaches parity with autoregressive models of similar size on standard generation metrics. The same models also match or surpass both diffusion and autoregressive baselines on zero-shot common-sense reasoning and on structured planning tasks such as Sudoku and Zebra puzzles.

What carries the argument

Glauber dynamics whose acceptance probabilities are computed from a pretrained language model viewed as an energy function that defines the stationary distribution of the diffusion process.

What would settle it

If the same Glauber dynamics pipeline with a uniform or randomly initialized energy function produces text of equal or higher quality on perplexity and downstream benchmarks than the pretrained-model version, the claimed benefit would be falsified.

Watch

Extended reading notes

Core claim

Instead of training a diffusion model whose forward process uses a uniform kernel, the authors treat a pretrained language model as an energy function whose associated Boltzmann distribution becomes the target stationary distribution for Glauber dynamics. Sampling then proceeds by repeatedly proposing single-token flips whose acceptance probabilities are set by the pretrained model logits, allowing the diffusion pipeline to inherit the pretrained model knowledge without additional architectural changes or retraining of the energy function.

Load-bearing premise

A pretrained language model can be used without modification as the energy function whose stationary distribution improves the quality of Glauber dynamics samples over a uniform-kernel baseline.

Editorial extensions

If this is right

  • Diffusion language models built this way outperform all prior diffusion-based language models on text generation.
  • Performance becomes competitive with autoregressive models of comparable parameter count without requiring autoregressive decoding at inference time.
  • The resulting models achieve strong zero-shot results on common-sense reasoning and on combinatorial planning tasks such as Sudoku and Zebra puzzles.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The method suggests a general route for injecting pretrained knowledge into any discrete diffusion sampler whose stationary distribution can be expressed as an energy function.
  • Because the pretrained model is used off-the-shelf, the approach may lower the total training cost of high-quality discrete generative models relative to training both the energy and the diffusion dynamics from random initialization.
  • The same energy-function construction could be tested on other discrete sequence domains where large pretrained models already exist, such as protein sequences or source code.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 2 minor

Summary. The paper proposes a discrete diffusion model for text that uses Glauber dynamics with an energy function derived from a frozen pretrained language model (e.g., UL2) to define the target stationary distribution. The central claim is that this yields higher-quality text generation than prior diffusion LMs, competitive performance with autoregressive models of similar size, and strong results on zero-shot commonsense reasoning plus planning/search tasks such as Sudoku and Zebra puzzles.

Significance. If the core assumption holds and the reported gains are reproducible, the work would offer a practical route to incorporate large pretrained LMs into diffusion pipelines without retraining, potentially improving non-autoregressive generation on tasks that benefit from global consistency.

major comments (3)
  1. [Method / Experiments] The central modeling assumption—that a frozen causal or masked LM can be plugged in directly as E(x) = −log p_LM(x) so that Glauber dynamics (Metropolis acceptance on single-token flips) samples from the LM distribution—receives no explicit empirical check. No experiment compares the empirical likelihood of finite-length Glauber trajectories against ancestral samples from the same UL2 model, nor demonstrates that the claimed performance improvements survive when the number of dynamics steps is varied.
  2. [Method] For causal LMs the energy function is asymmetric: flipping a token changes the conditioning context for all subsequent tokens, making each energy evaluation O(sequence length) and raising the possibility of slow mixing. The manuscript provides neither a mixing-time analysis nor an ablation that isolates the effect of this asymmetry on generation quality.
  3. [Experiments] The experimental section reports outperformance over prior diffusion LMs and competitiveness with GPT-2-scale autoregressive models, yet supplies no ablation that replaces the pretrained-LM energy with a uniform-kernel baseline while keeping all other components fixed. Without this control it is impossible to attribute gains specifically to the energy-function construction rather than other modeling choices.
minor comments (2)
  1. [Abstract] The abstract asserts performance gains without quoting any numerical metrics, baseline names, or error bars; a one-sentence summary of the key numbers would improve readability.
  2. [Method] Notation for the Glauber transition kernel, the precise Metropolis acceptance probability, and the temperature schedule should be stated explicitly (ideally with an equation) rather than left implicit.

Simulated Author's Rebuttal

3 responses · 1 unresolved

We thank the referee for the constructive and detailed feedback. We address each major comment point by point below, proposing revisions to strengthen the manuscript where the concerns are valid and providing clarifications on methodological choices.

read point-by-point responses
  1. Referee: [Method / Experiments] The central modeling assumption—that a frozen causal or masked LM can be plugged in directly as E(x) = −log p_LM(x) so that Glauber dynamics (Metropolis acceptance on single-token flips) samples from the LM distribution—receives no explicit empirical check. No experiment compares the empirical likelihood of finite-length Glauber trajectories against ancestral samples from the same UL2 model, nor demonstrates that the claimed performance improvements survive when the number of dynamics steps is varied.

    Authors: We acknowledge that a direct empirical verification of finite-time sampling behavior would strengthen the paper. The Metropolis-Hastings theorem guarantees convergence to the target distribution in the limit, but we did not report explicit checks against ancestral sampling or step-count ablations in the original submission. In the revision we will add (i) a comparison of average UL2 log-likelihoods for Glauber-generated sequences versus direct ancestral samples from the same model and (ii) an ablation varying the number of dynamics steps to show that downstream metrics improve and stabilize once sufficient steps are taken. These results will be included in the updated experimental section. revision: yes

  2. Referee: [Method] For causal LMs the energy function is asymmetric: flipping a token changes the conditioning context for all subsequent tokens, making each energy evaluation O(sequence length) and raising the possibility of slow mixing. The manuscript provides neither a mixing-time analysis nor an ablation that isolates the effect of this asymmetry on generation quality.

    Authors: The asymmetry for strictly causal factorizations is a legitimate concern, since a token flip alters the conditional probabilities of all later positions. Our primary results use UL2 in its masked configuration, for which the energy is symmetric. We will add an ablation that compares generation quality when the same backbone is used in causal versus masked energy modes. A full theoretical mixing-time analysis, however, lies outside the scope of the present work; we instead rely on consistent empirical performance across multiple benchmarks as practical evidence of usability. revision: partial

  3. Referee: [Experiments] The experimental section reports outperformance over prior diffusion LMs and competitiveness with GPT-2-scale autoregressive models, yet supplies no ablation that replaces the pretrained-LM energy with a uniform-kernel baseline while keeping all other components fixed. Without this control it is impossible to attribute gains specifically to the energy-function construction rather than other modeling choices.

    Authors: We agree that a direct control isolating the pretrained energy is desirable. While comparisons to prior diffusion models (which rely on uniform or learned kernels without an external energy) provide supporting context, we will add an explicit ablation in which the energy function is replaced by a constant, yielding an unbiased random-walk baseline. We expect this control to produce markedly lower-quality text, thereby confirming that the observed gains derive from the pretrained-LM energy. The new results will appear in the revised experimental section. revision: yes

standing simulated objections not resolved
  • A rigorous theoretical mixing-time analysis for Glauber dynamics under asymmetric causal language-model energy functions.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: external pretrained LMs supply the energy function

full rationale

The paper's core construction sets E(x) = -log p_LM(x) using a frozen external model (UL2) as the stationary distribution for Glauber dynamics, then runs the diffusion pipeline empirically. No parameter is fitted inside the paper and then renamed as a prediction; no self-citation supplies a uniqueness theorem or ansatz that the current work depends on; the reported gains are measured against external baselines rather than being forced by internal definitions. The derivation chain therefore remains open to external data and does not reduce to its own inputs by construction.

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

Review based on abstract only; limited visibility into exact formulation. The approach rests on standard statistical-physics assumptions applied to language modeling.

assumptions (2)
  • domain assumption A pretrained language model defines a valid energy function whose Boltzmann distribution can serve as the stationary distribution for Glauber dynamics on text sequences
    This is the central insight stated in the abstract.
  • standard math Glauber dynamics with the chosen energy function produces samples from the desired distribution
    Standard result from statistical mechanics invoked without proof in the abstract.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Leveraging Pretrained Language Models as Energy Functions for Glauber Dynamics Text Diffusion." pith.science (2026). https://pith.science/paper/2605.04291

@misc{pith2026260504291,
  author       = {Pith},
  title        = {Pith review of: Leveraging Pretrained Language Models as Energy Functions for Glauber Dynamics Text Diffusion},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/2605.04291}},
  note         = {Machine review of arXiv:2605.04291}
}
read the original abstract

We present a discrete diffusion-based language model using Glauber dynamics from statistical physics. Our main insight is that instead of trying to train a discrete state space diffusion model using Glauber dynamics with a uniform transition kernel as the forward process, one can set up an ``energy function'' based on pretrained causal/masked language models. When viewed as the stationary distribution, this energy function allows us to significantly improve the quality of the generated text. Incorporating UL2 as the pretrained model into our diffusion pipeline, we outperform prior diffusion based LMs and perform competitively with autoregressive models of comparable model sizes. Furthermore, our models are competitive with or outperform prior diffusion models and GPT-2 style auto-regressive models on zero-shot common sense reasoning tasks as well as planning and search tasks like Sudoku and Zebra puzzles.

Discussion (0). Continue with ORCID to comment.

Lean theorems connected to this paper

Citations machine-checked in the Pith Canon. Every link opens the source theorem in the public Lean library.

  • IndisputableMonolith.Cost (Jcost) washburn_uniqueness_aczel unclear
    ?
    unclear

    Relation between the paper passage and the cited Recognition theorem.

    Glauber dynamics... typically expected to shine in the presence of an 'energy function', i.e., sampling from p(x) ∝ e^{−f(x)} for some f operating on our discrete domain of interest. We propose to use pretrained LMs (such as AR or masked LMs) as our energy function.

  • IndisputableMonolith.Foundation.Atomicity atomic_tick unclear
    ?
    unclear

    Relation between the paper passage and the cited Recognition theorem.

    p(x_k = σ | x_{\k}) = min{1, exp(-f(x_{\k},σ) + f(x))}, with self-loops to ensure lazy chains.

What do these tags mean?
matches
The paper's claim is directly supported by a theorem in the formal canon.
supports
The theorem supports part of the paper's argument, but the paper may add assumptions or extra steps.
extends
The paper goes beyond the formal theorem; the theorem is a base layer rather than the whole result.
uses
The paper appears to rely on the theorem as machinery.
contradicts
The paper's claim conflicts with a theorem or certificate in the canon.
unclear
Pith found a possible connection, but the passage is too broad, indirect, or ambiguous to say the theorem truly supports the claim.

Reference graph

Works this paper leans on

299 extracted references · 299 canonical work pages

  1. [1]

    Liu , year = 2020, journal =

    Colin Raffel and Noam Shazeer and Adam Roberts and Katherine Lee and Sharan Narang and Michael Matena and Yanqi Zhou and Wei Li and Peter J. Liu , year = 2020, journal =

  2. [2]

    Hashimoto , year = 2023, month = 5, journal =

    Xuechen Li and Tianyi Zhang and Yann Dubois and Rohan Taori and Ishaan Gulrajani and Carlos Guestrin and Percy Liang and Tatsunori B. Hashimoto , year = 2023, month = 5, journal =

  3. [3]

    Encoder- decoder gemma: Improving the quality-efficiency trade-off via adaptation,

    Biao Zhang and Fedor Moiseev and Joshua Ainslie and Paul Suganthan and Min Ma and Surya Bhupatiraju and Fede Lebron and Orhan Firat and Armand Joulin and Zhe Dong , title =. Arxiv preprints 2504.06225 , year =

  4. [4]

    The Serial Scaling Hypothesis

    Yuxi Liu and Konpat Preechakul and Kananart Kuwaranancharoen and Yutong Bai , title =. Arxiv preprints 2507.12549 , year =

  5. [5]

    Train for the worst, plan for the best: Understanding token ordering in masked diffusions.arXiv preprint arXiv:2502.06768, 2025

    Train for the worst, plan for the best: Understanding token ordering in masked diffusions , author=. arXiv preprint arXiv:2502.06768 , year=

  6. [6]

    The Thirteenth International Conference on Learning Representations,

    Shansan Gong and Shivam Agarwal and Yizhe Zhang and Jiacheng Ye and Lin Zheng and Mukai Li and Chenxin An and Peilin Zhao and Wei Bi and Jiawei Han and Hao Peng and Lingpeng Kong , title =. The Thirteenth International Conference on Learning Representations,

  7. [7]

    The Thirteenth International Conference on Learning Representations,

    Minkai Xu and Tomas Geffner and Karsten Kreis and Weili Nie and Yilun Xu and Jure Leskovec and Stefano Ermon and Arash Vahdat , title =. The Thirteenth International Conference on Learning Representations,

  8. [8]

    Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics , year=

    HellaSwag: Can a Machine Really Finish Your Sentence? , author=. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics , year=

Show all 299 references
  1. [9]

    Causal language modeling can elicit search and reasoning capabilities on logic puzzles , booktitle =

    Kulin Shah and Nishanth Dikkala and Xin Wang and Rina Panigrahy , editor =. Causal language modeling can elicit search and reasoning capabilities on logic puzzles , booktitle =. 2024 , url =

  2. [10]

    Proceedings of the AAAI Conference on Artificial Intelligence , year =

    WinoGrande: An Adversarial Winograd Schema Challenge at Scale , author =. Proceedings of the AAAI Conference on Artificial Intelligence , year =

  3. [11]

    arXiv preprint arXiv:1904.09728 , year=

    Socialiqa: Commonsense reasoning about social interactions , author=. arXiv preprint arXiv:1904.09728 , year=

  4. [12]

    Thirty-Fourth AAAI Conference on Artificial Intelligence , year =

    Yonatan Bisk and Rowan Zellers and Ronan Le Bras and Jianfeng Gao and Yejin Choi , title =. Thirty-Fourth AAAI Conference on Artificial Intelligence , year =

  5. [13]

    Johnson and Jonathan Ho and Daniel Tarlow and Rianne van den Berg , year = 2021, booktitle =

    Jacob Austin and Daniel D. Johnson and Jonathan Ho and Daniel Tarlow and Rianne van den Berg , year = 2021, booktitle =

  6. [14]

    2204.05862 , archiveprefix =

    Yuntao Bai and Andy Jones and Kamal Ndousse and Amanda Askell and Anna Chen and Nova DasSarma and Dawn Drain and Stanislav Fort and Deep Ganguli and Tom Henighan and Nicholas Joseph and Saurav Kadavath and Jackson Kernion and Tom Conerly and Sheer El-Showk and Nelson Elhage an...

  7. [15]

    Tezak and John Schulman and Christine McLeavey and Jerry Tworek and Mark Chen , year = 2022, journal =

    Mohammad Bavarian and Heewoo Jun and Nikolas A. Tezak and John Schulman and Christine McLeavey and Jerry Tworek and Mark Chen , year = 2022, journal =

  8. [16]

    Parishad BehnamGhader and Vaibhav Adlakha and Marius Mosbach and Dzmitry Bahdanau and Nicolas Chapados and Siva Reddy , year = 2024, booktitle =

  9. [17]

    Luisa Bentivogli and Ido Dagan and Hoa Trang Dang and Danilo Giampiccolo and Bernardo Magnini , year = 2009, booktitle =

  10. [18]

    doi:10.3115/1073083.1073135 , url =

    Papineni, Kishore and Roukos, Salim and Ward, Todd and Zhu, Wei-Jing , year = 2002, month = jul, booktitle =. doi:10.3115/1073083.1073135 , url =

  11. [19]

    Campbell, Andrew and Benton, Joe and De Bortoli, Valentin and Rainforth, Tom and Deligiannidis, George and Doucet, Arnaud , year = 2022, journal =

  12. [20]

    Chang, Huiwen and Zhang, Han and Jiang, Lu and Liu, Ce and Freeman, William T , year = 2022, booktitle =

  13. [21]

    Chen, Ting and Zhang, Ruixiang and Hinton, Geoffrey , year = 2022, journal =

  14. [22]

    Cobbe, Karl and Kosaraju, Vineet and Bavarian, Mohammad and Chen, Mark and Jun, Heewoo and Kaiser, Lukasz and Plappert, Matthias and Tworek, Jerry and Hilton, Jacob and Nakano, Reiichiro and Hesse, Christopher and Schulman, John , year = 2021, journal =

  15. [23]

    Dagan, Ido and Glickman, Oren and Magnini, Bernardo , year = 2005, booktitle =

  16. [24]

    and Ermon, Stefano and Rudra, Atri and R

    Dao, Tri and Fu, Daniel Y. and Ermon, Stefano and Rudra, Atri and R

  17. [25]

    Dao, Tri , year = 2024, booktitle =

  18. [26]

    2501.12948 , archiveprefix =

  19. [27]

    Schwing and David A

    Aditya Deshpande and Jyoti Aneja and Liwei Wang and Alexander G. Schwing and David A. Forsyth , year = 2018, journal =

  20. [28]

    Dhariwal, Prafulla and Nichol, Alexander , year = 2021, journal =

  21. [29]

    Dhingra, Bhuwan and Mazaitis, Kathryn and Cohen, William W , year = 2017, journal =

  22. [30]

    Dieleman, Sander and Sartran, Laurent and Roshannai, Arman and Savinov, Nikolay and Ganin, Yaroslav and Richemond, Pierre H and Doucet, Arnaud and Strudel, Robin and Dyer, Chris and Durkan, Conor and others , year = 2022, journal =

  23. [31]

    Dong, Hanze and Xiong, Wei and Goyal, Deepanshu and Pan, Rui and Diao, Shizhe and Zhang, Jipeng and Shum, Kashun and Zhang, Tong , year = 2023, journal =

  24. [32]

    Du, Wanyu and Zhao, Jianqiao and Wang, Liwei and Ji, Yangfeng , year = 2022, journal =

  25. [33]

    Lin, Zhenghao and Gong, Yeyun and Shen, Yelong and Wu, Tong and Fan, Zhihao and Lin, Chen and Duan, Nan and Chen, Weizhu , year = 2023, booktitle =

  26. [34]

    Ghazvininejad, Marjan and Levy, Omer and Liu, Yinhan and Zettlemoyer, Luke , year = 2019, booktitle =

  27. [35]

    Gong, Shansan and Li, Mukai and Feng, Jiangtao and Wu, Zhiyong and Kong, Lingpeng , year = 2023, booktitle =

  28. [36]

    2410.17891 , archiveprefix =

    Shansan Gong and Shivam Agarwal and Yizhe Zhang and Jiacheng Ye and Lin Zheng and Mukai Li and Chenxin An and Peilin Zhao and Wei Bi and Jiawei Han and Hao Peng and Lingpeng Kong , year = 2024, url =. 2410.17891 , archiveprefix =

  29. [37]

    doi:10.5281/zenodo.5297715 , url =

    Black, Sid and Leo, Gao and Wang, Phil and Leahy, Connor and Biderman, Stella , year = 2021, month = mar, publisher =. doi:10.5281/zenodo.5297715 , url =

  30. [38]

    2407.21783 , archiveprefix =

  31. [39]

    Gu, Jiatao and Wang, Changhan and Zhao, Junbo , year = 2019, journal =

  32. [40]

    Ishaan Gulrajani and Tatsunori Hashimoto , year = 2023, booktitle =

  33. [41]

    Xiaochuang Han and Sachin Kumar and Yulia Tsvetkov and Marjan Ghazvininejad , year = 2023, journal =

  34. [42]

    2401.17181 , archiveprefix =

    Kehang Han and Kathleen Kenealy and Aditya Barua and Noah Fiedel and Noah Constant , year = 2024, url =. 2401.17181 , archiveprefix =

  35. [43]

    Hermann, Karl Moritz and Kocisky, Tomas and Grefenstette, Edward and Espeholt, Lasse and Kay, Will and Suleyman, Mustafa and Blunsom, Phil , year = 2015, booktitle =

  36. [44]

    Ho, Jonathan and Salimans, Tim , year = 2021, booktitle =

  37. [45]

    Holtzman, Ari and Buys, Jan and Du, Li and Forbes, Maxwell and Choi, Yejin , year = 2019, booktitle =

  38. [46]

    Hoogeboom, Emiel and Nielsen, Didrik and Jaini, Priyank and Forr

  39. [47]

    Gritsenko and Jasmijn Bastings and Ben Poole and Rianne van den Berg and Tim Salimans , year = 2022, booktitle =

    Emiel Hoogeboom and Alexey A. Gritsenko and Jasmijn Bastings and Ben Poole and Rianne van den Berg and Tim Salimans , year = 2022, booktitle =

  40. [48]

    Jie Huang and Xinyun Chen and Swaroop Mishra and Huaixiu Steven Zheng and Adams Wei Yu and Xinying Song and Denny Zhou , year = 2024, booktitle =

  41. [49]

    Smith and Iz Beltagy and Hannaneh Hajishirzi , year = 2023, url =

    Hamish Ivison and Yizhong Wang and Valentina Pyatkin and Nathan Lambert and Matthew Peters and Pradeep Dasigi and Joel Jang and David Wadden and Noah A. Smith and Iz Beltagy and Hannaneh Hajishirzi , year = 2023, url =. 2311.10702 , archiveprefix =

  42. [50]

    and Choi, Yejin and Hajishirzi, Hannaneh , year = 2024, booktitle =

    Hamish Ivison and Wang, Yizhong and Liu, Jiacheng and Wu, Zeqiu and Pyatkin, Valentina and Lambert, Nathan and Smith, Noah A. and Choi, Yejin and Hajishirzi, Hannaneh , year = 2024, booktitle =. 2406.09279 , code =

  43. [51]

    Liu and Mohammad Saleh and Etienne Pot and Ben Goodrich and Ryan Sepassi and Lukasz Kaiser and Noam Shazeer , year = 2018, booktitle =

    Peter J. Liu and Mohammad Saleh and Etienne Pot and Ben Goodrich and Ryan Sepassi and Lukasz Kaiser and Noam Shazeer , year = 2018, booktitle =

  44. [52]

    Jiang, Chao and Maddela, Mounica and Lan, Wuwei and Zhong, Yang and Xu, Wei , year = 2020, booktitle =

  45. [53]

    Albert Q. Jiang and Alexandre Sablayrolles and Arthur Mensch and Chris Bamford and Devendra Singh Chaplot and Diego de las Casas and Florian Bressand and Gianna Lengyel and Guillaume Lample and Lucile Saulnier and Lélio Renard Lavaud and Marie-Anne Lachaux and Pierre Stock and...

  46. [54]

    Albert Jiang and Alexandre Sablayrolles and Alexis Tacnet and Antoine Roux and Arthur Mensch and Audrey Herblin-Stoop and Baptiste Bout and Baudouin de Monicault and Blanche Savary and Bam4d and Caroline Feldman and Devendra Singh Chaplot and Diego de las Casas and Eleonore Ar...

  47. [55]

    Kong, Zhifeng and Ping, Wei and Huang, Jiaji and Zhao, Kexin and Catanzaro, Bryan , year = 2020, booktitle =

  48. [56]

    2409.12917 , archiveprefix =

    Aviral Kumar and Vincent Zhuang and Rishabh Agarwal and Yi Su and John D Co-Reyes and Avi Singh and Kate Baumli and Shariq Iqbal and Colton Bishop and Rebecca Roelofs and Lei M Zhang and Kay McKinney and Disha Shrivastava and Cosmin Paduraru and George Tucker and Doina Precup ...

  49. [57]

    Gonzalez and Hao Zhang and Ion Stoica , year = 2023, booktitle =

    Woosuk Kwon and Zhuohan Li and Siyuan Zhuang and Ying Sheng and Lianmin Zheng and Cody Hao Yu and Joseph E. Gonzalez and Hao Zhang and Ion Stoica , year = 2023, booktitle =

  50. [58]

    Smith and Hannaneh Hajishirzi , year = 2024, eprint =

    Nathan Lambert and Valentina Pyatkin and Jacob Morrison and LJ Miranda and Bill Yuchen Lin and Khyathi Chandu and Nouha Dziri and Sachin Kumar and Tom Zick and Yejin Choi and Noah A. Smith and Hannaneh Hajishirzi , year = 2024, eprint =

  51. [59]

    Miranda and Alisa Liu and Nouha Dziri and Shane Lyu and Yuling Gu and Saumya Malik and Victoria Graf and Jena D

    Nathan Lambert and Jacob Morrison and Valentina Pyatkin and Shengyi Huang and Hamish Ivison and Faeze Brahman and Lester James V. Miranda and Alisa Liu and Nouha Dziri and Shane Lyu and Yuling Gu and Saumya Malik and Victoria Graf and Jena D. Hwang and Jiangjiang Yang and Rona...

  52. [60]

    2210.14215 , archiveprefix =

    Michael Laskin and Luyu Wang and Junhyuk Oh and Emilio Parisotto and Stephen Spencer and Richie Steigerwald and DJ Strouse and Steven Hansen and Angelos Filos and Ethan Brooks and Maxime Gazeau and Himanshu Sahni and Satinder Singh and Volodymyr Mnih , year = 2022, url =. 2210...

  53. [61]

    Lewis, Mike and Liu, Yinhan and Goyal, Naman and Ghazvininejad, Marjan and Mohamed, Abdelrahman and Levy, Omer and Stoyanov, Veselin and Zettlemoyer, Luke , year = 2020, booktitle =

  54. [62]

    Li, Jiwei and Galley, Michel and Brockett, Chris and Gao, Jianfeng and Dolan, Bill , year = 2016, booktitle =

  55. [63]

    Xiang Lisa Li and John Thickstun and Ishaan Gulrajani and Percy Liang and Tatsunori Hashimoto , year = 2022, booktitle =

  56. [65]

    doi:10.1145/3531146.3533088 , isbn = 9781450393522, url =

    Weidinger, Laura and Uesato, Jonathan and Rauh, Maribeth and Griffin, Conor and Huang, Po-Sen and Mellor, John and Glaese, Amelia and Cheng, Myra and Balle, Borja and Kasirzadeh, Atoosa and Biles, Courtney and Brown, Sasha and Kenton, Zac and Hawkins, Will and Stepleton, Tom a...

  57. [66]

    Mahabadi, Rabeeh Karimi and Ruder, Sebastian and Dehghani, Mostafa and Henderson, James , year = 2021, booktitle =

  58. [67]

    2412.09413 , archiveprefix =

    Yingqian Min and Zhipeng Chen and Jinhao Jiang and Jie Chen and Jia Deng and Yiwen Hu and Yiru Tang and Jiapeng Wang and Xiaoxue Cheng and Huatong Song and Wayne Xin Zhao and Zheng Liu and Zhongyuan Wang and Ji-Rong Wen , year = 2024, url =. 2412.09413 , archiveprefix =

  59. [68]

    and Lapata, Mirella , year = 2018, booktitle =

    Narayan, Shashi and Cohen, Shay B. and Lapata, Mirella , year = 2018, booktitle =

  60. [69]

    Brown, Tom and Mann, Benjamin and Ryder, Nick and Subbiah, Melanie and Kaplan, Jared D and Dhariwal, Prafulla and Neelakantan, Arvind and Shyam, Pranav and Sastry, Girish and Askell, Amanda and Agarwal, Sandhini and Herbert-Voss, Ariel and Krueger, Gretchen and Henighan, Tom a...

  61. [70]

    doi:10.1162/tacl_a_00139 , url =

    Xu, Wei and Callison-Burch, Chris and Napoles, Courtney , year = 2015, journal =. doi:10.1162/tacl_a_00139 , url =

  62. [71]

    Nichol, Alexander Quinn and Dhariwal, Prafulla , year = 2021, booktitle =

  63. [72]

    Long Ouyang and Jeff Wu and Xu Jiang and Diogo Almeida and Carroll L. Wainwright and Pamela Mishkin and Chong Zhang and Sandhini Agarwal and Katarina Slama and Alex Ray and John Schulman and Jacob Hilton and Fraser Kelton and Luke Miller and Maddie Simens and Amanda Askell and...

  64. [73]

    Pillutla, Krishna and Swayamdipta, Swabha and Zellers, Rowan and Thickstun, John and Welleck, Sean and Choi, Yejin and Harchaoui, Zaid , year = 2021, booktitle =

  65. [74]

    Nichol, Alexander Quinn and Dhariwal, Prafulla and Ramesh, Aditya and Shyam, Pranav and Mishkin, Pamela and Mcgrew, Bob and Sutskever, Ilya and Chen, Mark , year = 2022, booktitle =

  66. [75]

    Proceedings of the 41st International Conference on Machine Learning , publisher =

    Bachmann, Gregor and Nagarajan, Vaishnavh , year = 2024, month =. Proceedings of the 41st International Conference on Machine Learning , publisher =

  67. [76]

    Radford, Alec and Wu, Jeffrey and Child, Rewon and Luan, David and Amodei, Dario and Sutskever, Ilya and others , year = 2019, journal =

  68. [77]

    Raffel, Colin and Shazeer, Noam and Roberts, Adam and Lee, Katherine and Narang, Sharan and Matena, Michael and Zhou, Yanqi and Li, Wei and Liu, Peter J , year = 2020, journal =

  69. [78]

    Ramesh, Aditya and Dhariwal, Prafulla and Nichol, Alex and Chu, Casey and Chen, Mark , year = 2022, journal =

  70. [79]

    Reid, Machel and Hellendoorn, Vincent J and Neubig, Graham , year = 2022, journal =

  71. [80]

    1907.11692 , timestamp =

    Yinhan Liu and Myle Ott and Naman Goyal and Jingfei Du and Mandar Joshi and Danqi Chen and Omer Levy and Mike Lewis and Luke Zettlemoyer and Veselin Stoyanov , year = 2019, journal =. 1907.11692 , timestamp =

  72. [81]

    2024 , eprint=

    Discrete Diffusion Modeling by Estimating the Ratios of the Data Distribution , author=. 2024 , eprint=

  73. [82]

    Lin, Chin-Yew , year = 2004, month = jul, booktitle =

  74. [83]

    Bar-Haim, Roy and Dagan, Ido and Dolan, Bill and Ferro, Lisa and Giampiccolo, Danilo , year = 2006, journal =

  75. [84]

    Saharia, Chitwan and Chan, William and Saxena, Saurabh and Li, Lala and Whang, Jay and Denton, Emily and Ghasemipour, Seyed Kamyar Seyed and Ayan, Burcu Karagol and Mahdavi, S Sara and Lopes, Rapha Gontijo and others , year = 2022, booktitle =

  76. [85]

    Victor Sanh and Albert Webson and Colin Raffel and Stephen Bach and Lintang Sutawika and Zaid Alyafeai and Antoine Chaffin and Arnaud Stiegler and Arun Raja and Manan Dey and M Saiful Bari and Canwen Xu and Urmish Thakker and Shanya Sharma Sharma and Eliza Szczechla and Taewoo...

  77. [86]

    Savinov, Nikolay and Chung, Junyoung and Binkowski, Mikolaj and Elsen, Erich and van den Oord, Aaron , year = 2021, booktitle =

  78. [87]

    Nikolay Savinov and Junyoung Chung and Mikolaj Binkowski and Erich Elsen and Aaron van den Oord , year = 2022, booktitle =

  79. [88]

    Kai Shen and Zeqian Ju and Xu Tan and Yanqing Liu and Yichong Leng and Lei He and Tao Qin and Sheng Zhao and Jiang Bian , year = 2023, journal =

  80. [89]

    Sohl-Dickstein, Jascha and Weiss, Eric and Maheswaranathan, Niru and Ganguli, Surya , year = 2015, booktitle =

  81. [90]

    Song, Yang and Ermon, Stefano , year = 2019, journal =

  82. [91]

    Song, Jiaming and Meng, Chenlin and Ermon, Stefano , year = 2020, booktitle =

  83. [92]

    Kingma and Abhishek Kumar and Stefano Ermon and Ben Poole , year = 2021, booktitle =

    Yang Song and Jascha Sohl-Dickstein and Diederik P. Kingma and Abhishek Kumar and Stefano Ermon and Ben Poole , year = 2021, booktitle =

  84. [93]

    Yang Song and Prafulla Dhariwal and Mark Chen and Ilya Sutskever , year = 2023, booktitle =

  85. [94]

    Srivastava, Aarohi and Rastogi, Abhinav and Rao, Abhishek and Shoeb, Abu Awal Md and Abid, Abubakar and Fisch, Adam and Brown, Adam R and Santoro, Adam and Gupta, Aditya and Garriga-Alonso, Adri

  86. [95]

    Strudel, Robin and Tallec, Corentin and Altch

  87. [96]

    Sun, Haoran and Yu, Lijun and Dai, Bo and Schuurmans, Dale and Dai, Hanjun , year = 2023, journal =

  88. [97]

    Suzgun, Mirac and Scales, Nathan and Sch

  89. [98]

    Jaesung Tae and Hyeongju Kim and Taesu Kim , year = 2022, booktitle =

  90. [99]

    Yi Tay and Mostafa Dehghani and Vinh Q. Tran and Xavier Garcia and Jason Wei and Xuezhi Wang and Hyung Won Chung and Dara Bahri and Tal Schuster and Steven Zheng and Denny Zhou and Neil Houlsby and Donald Metzler , year = 2023, booktitle =

  91. [100]

    2307.09288 , archiveprefix =

    Hugo Touvron and Louis Martin and Kevin Stone and Peter Albert and Amjad Almahairi and Yasmine Babaei and Nikolay Bashlykov and Soumya Batra and Prajjwal Bhargava and Shruti Bhosale and Dan Bikel and Lukas Blecher and Cristian Canton Ferrer and Moya Chen and Guillem Cucurull a...

  92. [101]

    Karthik Valmeekam and Matthew Marquez and Sarath Sreedharan and Subbarao Kambhampati , year = 2023, booktitle =

  93. [102]

    Vaswani, Ashish and Shazeer, Noam and Parmar, Niki and Uszkoreit, Jakob and Jones, Llion and Gomez, Aidan N and Kaiser,

  94. [103]

    Bowman , year = 2019, booktitle =

    Alex Wang and Amanpreet Singh and Julian Michael and Felix Hill and Omer Levy and Samuel R. Bowman , year = 2019, booktitle =

  95. [104]

    Smith and Iz Beltagy and Hannaneh Hajishirzi , year = 2023, booktitle =

    Yizhong Wang and Hamish Ivison and Pradeep Dasigi and Jack Hessel and Tushar Khot and Khyathi Chandu and David Wadden and Kelsey MacMillan and Noah A. Smith and Iz Beltagy and Hannaneh Hajishirzi , year = 2023, booktitle =

  96. [105]

    Chi and Quoc V Le and Denny Zhou , year = 2022, booktitle =

    Jason Wei and Xuezhi Wang and Dale Schuurmans and Maarten Bosma and brian ichter and Fei Xia and Ed H. Chi and Quoc V Le and Denny Zhou , year = 2022, booktitle =

  97. [106]

    Dai and Quoc V Le , year = 2022, booktitle =

    Jason Wei and Maarten Bosma and Vincent Zhao and Kelvin Guu and Adams Wei Yu and Brian Lester and Nan Du and Andrew M. Dai and Quoc V Le , year = 2022, booktitle =

  98. [107]

    Welleck, Sean and Kulikov, Ilia and Roller, Stephen and Dinan, Emily and Cho, Kyunghyun and Weston, Jason , year = 2019, booktitle =

  99. [108]

    Rush , year = 2020, booktitle =

    Thomas Wolf and Lysandre Debut and Victor Sanh and Julien Chaumond and Clement Delangue and Anthony Moi and Pierric Cistac and Tim Rault and Rémi Louf and Morgan Funtowicz and Joe Davison and Sam Shleifer and Patrick von Platen and Clara Ma and Yacine Jernite and Julien Plu an...

  100. [109]

    Xie, Jian and Zhang, Kai and Chen, Jiangjie and Zhu, Tinghui and Lou, Renze and Tian, Yuandong and Xiao, Yanghua and Su, Yu , year = 2024, booktitle =

  101. [110]

    Xiong, Wei and Dong, Hanze and Ye, Chenlu and Zhong, Han and Jiang, Nan and Zhang, Tong , year = 2023, journal =

  102. [111]

    Xu, Wei and Napoles, Courtney and Pavlick, Ellie and Chen, Quanze and Callison-Burch, Chris , year = 2016, journal =

  103. [112]

    2308.12219 , archiveprefix =

    Jiasheng Ye and Zaixiang Zheng and Yu Bao and Lihua Qian and Quanquan Gu , year = 2023, url =. 2308.12219 , archiveprefix =

  104. [113]

    Jiasheng Ye and Zaixiang Zheng and Yu Bao and Lihua Qian and Mingxuan Wang , year = 2023, eprint =

  105. [114]

    Yuan, Hongyi and Yuan, Zheng and Tan, Chuanqi and Huang, Fei and Huang, Songfang , year = 2022, journal =

  106. [115]

    Zhang, Xingxing and Lapata, Mirella , year = 2017, booktitle =

  107. [116]

    Zhang, Tianyi and Kishore, Varsha and Wu, Felix and Weinberger, Kilian Q and Artzi, Yoav , year = 2020, booktitle =

  108. [117]

    Zhang, Tianyi and Wu, Felix and Katiyar, Arzoo and Weinberger, Kilian Q and Artzi, Yoav , year = 2021, booktitle =

  109. [118]

    2302.05737 , archiveprefix =

    Lin Zheng and Jianbo Yuan and Lei Yu and Lingpeng Kong , year = 2024, url =. 2302.05737 , archiveprefix =

  110. [119]

    Zhou, Jeffrey and Lu, Tianjian and Mishra, Swaroop and Brahma, Siddhartha and Basu, Sujoy and Luan, Yi and Zhou, Denny and Hou, Le , year = 2023, journal =

  111. [120]

    arXiv preprint arXiv:2205.05131 , year=

    Ul2: Unifying language learning paradigms , author=. arXiv preprint arXiv:2205.05131 , year=

  112. [121]

    Score-Based Generative Modeling through Stochastic Differential Equations , booktitle =

    Yang Song and Jascha Sohl. Score-Based Generative Modeling through Stochastic Differential Equations , booktitle =

  113. [122]

    David A Levin and Yuval Peres , title =

  114. [123]

    Stochastic differential systems (Marseille-Luminy, 1984), volume 69 of Lecture Notes in Control and Inform

    Hans Follmer , title =. Stochastic differential systems (Marseille-Luminy, 1984), volume 69 of Lecture Notes in Control and Inform. Sci , pages =

  115. [124]

    Arxiv eprints , volume =

    Joe Benton and Yuyang Shi and Valentin De Bortoli and George Deligiannidis and Arnaud Doucet , title =. Arxiv eprints , volume =

  116. [125]

    Joseph Lehec , title =. Inst. Henri Poincare Probab. Stat. , volume =

  117. [126]

    Anderson , title =

    B.D.O. Anderson , title =. Stochastic Process. Appl. , volume =

  118. [127]

    Deep Unsupervised Learning using Nonequilibrium Thermodynamics , booktitle =

    Jascha Sohl. Deep Unsupervised Learning using Nonequilibrium Thermodynamics , booktitle =

  119. [128]

    Parallelising Glauber Dynamics , booktitle =

    Holden Lee , editor =. Parallelising Glauber Dynamics , booktitle =

  120. [129]

    Kakade and Lucas Janson , title =

    Depen Morwani and Itai Shapira and Nikhil Vyas and Eran Malach and Sham M. Kakade and Lucas Janson , title =. The Thirteenth International Conference on Learning Representations,

  121. [130]

    Kakade , title =

    Nikhil Vyas and Depen Morwani and Rosie Zhao and Itai Shapira and David Brandfonbrener and Lucas Janson and Sham M. Kakade , title =. The Thirteenth International Conference on Learning Representations,. 2025 , url =

  122. [131]

    CoRR , volume =

    Jingyuan Liu and Jianlin Su and Xingcheng Yao and Zhejun Jiang and Guokun Lai and Yulun Du and Yidao Qin and Weixin Xu and Enzhe Lu and Junjie Yan and Yanru Chen and Huabin Zheng and Yibo Liu and Shaowei Liu and Bohong Yin and Weiran He and Han Zhu and Yuzhi Wang and Jianzhou ...

  123. [132]

    Parallel Sampling of Diffusion Models , booktitle =

    Andy Shih and Suneel Belkhale and Stefano Ermon and Dorsa Sadigh and Nima Anari , editor =. Parallel Sampling of Diffusion Models , booktitle =

  124. [133]

    William Peebles and Saining Xie , title =

  125. [134]

    OpenWebTextCorpus , year =

    Aaron Gokaslan and Vanya Cohen , title =. OpenWebTextCorpus , year =

  126. [135]

    Krishna Pillutla and Lang Liu and John Thickstun and Sean Welleck and Swabha Swayamdipta and Rowan Zellers and Sewoong Oh and Yejin Choi and Za. J. Mach. Learn. Res. , volume =. 2023 , url =

  127. [136]

    Itai Gat and Tal Remez and Neta Shaul and Felix Kreuk and Ricky T. Q. Chen and Gabriel Synnaeve and Yossi Adi and Yaron Lipman , editor =. Discrete Flow Matching , booktitle =

  128. [137]

    Jaakkola and Brian Karrer and Ricky T

    Peter Holderrieth and Marton Havasi and Jason Yim and Neta Shaul and Itai Gat and Tommi S. Jaakkola and Brian Karrer and Ricky T. Q. Chen and Yaron Lipman , title =. Arxiv preprints , volume =

  129. [138]

    Yaron Lipman and Ricky T. Q. Chen and Heli Ben. Flow Matching for Generative Modeling , booktitle =

  130. [139]

    Gulrajani, Ishaan and Hashimoto, Tatsunori B , year = 2023, journal =

  131. [140]

    Han, Xiaochuang and Kumar, Sachin and Tsvetkov, Yulia , year = 2022, journal =

  132. [141]

    Xiaochuang Han and Sachin Kumar and Yulia Tsvetkov and Marjan Ghazvininejad , year = 2023, eprint =

  133. [142]

    Ho, Jonathan and Salimans, Tim and Gritsenko, Alexey A and Chan, William and Norouzi, Mohammad and Fleet, David J , year = 2022, booktitle =

  134. [143]

    Li, Xiang Lisa and Thickstun, John and Gulrajani, Ishaan and Liang, Percy and Hashimoto, Tatsunori B , year = 2022, booktitle =

  135. [144]

    Lou, Aaron and Meng, Chenlin and Ermon, Stefano , year = 2024, journal =

  136. [145]

    Pillutla, Krishna and Swayamdipta, Swabha and Zellers, Rowan and Thickstun, John and Welleck, Sean and Choi, Yejin and Harchaoui, Zaid , year = 2021, journal =

  137. [146]

    Richemond, Pierre H and Dieleman, Sander and Doucet, Arnaud , year = 2022, journal =

  138. [147]

    Advances in neural information processing systems , volume=

    Denoising diffusion probabilistic models , author=. Advances in neural information processing systems , volume=

  139. [148]

    Jiacheng Ye and Shansan Gong and Liheng Chen and Lin Zheng and Jiahui Gao and Han Shi and Chuan Wu and Xin Jiang and Zhenguo Li and Wei Bi and Lingpeng Kong , year = 2024, booktitle =

  140. [149]

    Susskind and Navdeep Jaitly , year = 2023, booktitle =

    Yizhe Zhang and Jiatao Gu and Zhuofeng Wu and Shuangfei Zhai and Joshua M. Susskind and Navdeep Jaitly , year = 2023, booktitle =

  141. [150]

    arXiv preprint arXiv:2502.13917 , year=

    TESS 2: A Large-Scale Generalist Diffusion Language Model , author=. arXiv preprint arXiv:2502.13917 , year=

  142. [151]

    arXiv preprint arXiv:2405.17035 , year=

    Glauber generative model: Discrete diffusion models via binary classification , author=. arXiv preprint arXiv:2405.17035 , year=

  143. [152]

    Andrew Campbell and Joe Benton and Valentin De Bortoli and Thomas Rainforth and George Deligiannidis and Arnaud Doucet , title =. Advances in Neural Information Processing Systems 35 Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022, New Orleans, LA...

  144. [153]

    Sahoo and Marianne Arriola and Yair Schiff and Aaron Gokaslan and Edgar Marroquin and Justin T

    Subham S. Sahoo and Marianne Arriola and Yair Schiff and Aaron Gokaslan and Edgar Marroquin and Justin T. Chiu and Alexander Rush and Volodymyr Kuleshov , title =. Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems ...

  145. [154]

    Scaling Instruction-Finetuned Language Models , journal =

    Hyung Won Chung and Le Hou and Shayne Longpre and Barret Zoph and Yi Tay and William Fedus and Yunxuan Li and Xuezhi Wang and Mostafa Dehghani and Siddhartha Brahma and Albert Webson and Shixiang Shane Gu and Zhuyun Dai and Mirac Suzgun and Xinyun Chen and Aakanksha Chowdhery ...

  146. [155]

    Masked Diffusion Models are Secretly Time-Agnostic Masked Models and Exploit Inaccurate Categorical Sampling , booktitle =

    Kaiwen Zheng and Yongxin Chen and Hanzi Mao and Ming. Masked Diffusion Models are Secretly Time-Agnostic Masked Models and Exploit Inaccurate Categorical Sampling , booktitle =

  147. [156]

    SIAM Journal on Mathematics of Data Science , pages =

    Giovanni Conforti and Alain Durmus and Marta Gentiloni Silveri , title =. SIAM Journal on Mathematics of Data Science , pages =

  148. [157]

    Proceedings of the First Workshop on Writing Aids at the Crossroads of AI, Cognitive Science and NLP (WRAICOGS 2025). 2025

  149. [158]

    Chain-of- M eta W riting: Linguistic and Textual Analysis of How Small Language Models Write Young Students Texts

    Buhnila, Ioana and Cislaru, Georgeta and Todirascu, Amalia. Chain-of- M eta W riting: Linguistic and Textual Analysis of How Small Language Models Write Young Students Texts. 2025

  150. [159]

    Semantic Masking in a Needle-in-a-haystack Test for Evaluating Large Language Model Long-Text Capabilities

    Shi, Ken and Penn, Gerald. Semantic Masking in a Needle-in-a-haystack Test for Evaluating Large Language Model Long-Text Capabilities. 2025

  151. [160]

    Reading Between the Lines: A dataset and a study on why some texts are tougher than others

    Khallaf, Nouran and Eugeni, Carlo and Sharoff, Serge. Reading Between the Lines: A dataset and a study on why some texts are tougher than others. 2025

  152. [161]

    P ara R ev : Building a dataset for Scientific Paragraph Revision annotated with revision instruction

    Jourdan, L \'e ane and Boudin, Florian and Dufour, Richard and Hernandez, Nicolas and Aizawa, Akiko. P ara R ev : Building a dataset for Scientific Paragraph Revision annotated with revision instruction. 2025

  153. [162]

    Towards an operative definition of creative writing: a preliminary assessment of creativeness in AI and human texts

    Maggi, Chiara and Vitaletti, Andrea. Towards an operative definition of creative writing: a preliminary assessment of creativeness in AI and human texts. 2025

  154. [163]

    Decoding Semantic Representations in the Brain Under Language Stimuli with Large Language Models

    Sato, Anna and Kobayashi, Ichiro. Decoding Semantic Representations in the Brain Under Language Stimuli with Large Language Models. 2025

  155. [164]

    Proceedings of the 4th Workshop on Arabic Corpus Linguistics (WACL-4). 2025

  156. [165]

    A rabic S ense: A Benchmark for Evaluating Commonsense Reasoning in A rabic with Large Language Models

    Lamsiyah, Salima and Zeinalipour, Kamyar and El amrany, Samir and Brust, Matthias and Maggini, Marco and Bouvry, Pascal and Schommer, Christoph. A rabic S ense: A Benchmark for Evaluating Commonsense Reasoning in A rabic with Large Language Models. 2025

  157. [166]

    Lahjawi: A rabic Cross-Dialect Translator

    Hamed, Mohamed Motasim and Hreden, Muhammad and Hennara, Khalil and Aldallal, Zeina and Chrouf, Sara and AlModhayan, Safwan. Lahjawi: A rabic Cross-Dialect Translator. 2025

  158. [167]

    Lost in Variation: An Unsupervised Methodology for Mining Lexico-syntactic Patterns in Middle A rabic Texts

    Bezan. Lost in Variation: An Unsupervised Methodology for Mining Lexico-syntactic Patterns in Middle A rabic Texts. 2025

  159. [168]

    SADSL y C : A Corpus for Saudi A rabian Multi-dialect Identification through Song Lyrics

    Alahmari, Salwa Saad. SADSL y C : A Corpus for Saudi A rabian Multi-dialect Identification through Song Lyrics. 2025

  160. [169]

    Enhancing Dialectal A rabic Intent Detection through Cross-Dialect Multilingual Input Augmentation

    Hossain, Shehenaz and Shammary, Fouad and Shammary, Bahaulddin and Afli, Haithem. Enhancing Dialectal A rabic Intent Detection through Cross-Dialect Multilingual Input Augmentation. 2025

  161. [170]

    D ial2 MSA -Verified: A Multi-Dialect A rabic Social Media Dataset for Neural Machine Translation to M odern S tandard A rabic

    Khered, Abdullah and Benkhedda, Youcef and Batista-Navarro, Riza. D ial2 MSA -Verified: A Multi-Dialect A rabic Social Media Dataset for Neural Machine Translation to M odern S tandard A rabic. 2025

  162. [171]

    Web-Based Corpus Compilation of the Emirati A rabic Dialect

    El-Ghawi, Yousra A. Web-Based Corpus Compilation of the Emirati A rabic Dialect. 2025

  163. [172]

    Evaluating Calibration of A rabic Pre-trained Language Models on Dialectal Text

    Al-Laith, Ali and Kebdani, Rachida. Evaluating Calibration of A rabic Pre-trained Language Models on Dialectal Text. 2025

  164. [173]

    Empirical Evaluation of Pre-trained Language Models for Summarizing M oroccan D arija News Articles

    Aftiss, Azzedine and Lamsiyah, Salima and Schommer, Christoph and El Alaoui, Said Ouatik. Empirical Evaluation of Pre-trained Language Models for Summarizing M oroccan D arija News Articles. 2025

  165. [174]

    D ialect2 SQL : A Novel Text-to- SQL Dataset for A rabic Dialects with a Focus on M oroccan D arija

    Chafik, Salmane and Ezzini, Saad and Berrada, Ismail. D ialect2 SQL : A Novel Text-to- SQL Dataset for A rabic Dialects with a Focus on M oroccan D arija. 2025

  166. [175]

    A ra S im: Optimizing A rabic Dialect Translation in Children`s Literature with LLM s and Similarity Scores

    Bouomar, Alaa Hassan and Abbas, Noorhan. A ra S im: Optimizing A rabic Dialect Translation in Children`s Literature with LLM s and Similarity Scores. 2025

  167. [176]

    Navigating Dialectal Bias and Ethical Complexities in L evantine A rabic Hate Speech Detection

    Haj Ahmed, Ahmed and Yew, Rui-Jie and Minocher, Xerxes and Venkatasubramanian, Suresh. Navigating Dialectal Bias and Ethical Complexities in L evantine A rabic Hate Speech Detection. 2025

  168. [177]

    Proceedings of the 12th Workshop on NLP for Similar Languages, Varieties and Dialects. 2025

  169. [178]

    Findings of the V ar D ial Evaluation Campaign 2025: The N or SID Shared Task on N orwegian Slot, Intent and Dialect Identification

    Scherrer, Yves and van der Goot, Rob and M hlum, Petter. Findings of the V ar D ial Evaluation Campaign 2025: The N or SID Shared Task on N orwegian Slot, Intent and Dialect Identification. 2025

  170. [179]

    Information Theory and Linguistic Variation: A Study of B razilian and E uropean P ortuguese

    Alves, Diego. Information Theory and Linguistic Variation: A Study of B razilian and E uropean P ortuguese. 2025

  171. [180]

    Leveraging Open-Source Large Language Models for Native Language Identification

    Ng, Yee Man and Markov, Ilia. Leveraging Open-Source Large Language Models for Native Language Identification. 2025

  172. [181]

    and Tayyar Madabushi, Harish

    Torgbi, Melissa and Clayman, Andrew and Speight, Jordan J. and Tayyar Madabushi, Harish. Adapting Whisper for Regional Dialects: Enhancing Public Services for Vulnerable Populations in the U nited K ingdom. 2025

  173. [182]

    Large Language Models as a Normalizer for Transliteration and Dialectal Translation

    Alam, Md Mahfuz Ibn and Anastasopoulos, Antonios. Large Language Models as a Normalizer for Transliteration and Dialectal Translation. 2025

  174. [183]

    Testing the Boundaries of LLM s: Dialectal and Language-Variety Tasks

    Faisal, Fahim and Anastasopoulos, Antonios. Testing the Boundaries of LLM s: Dialectal and Language-Variety Tasks. 2025

  175. [184]

    Text Generation Models for L uxembourgish with Limited Data: A Balanced Multilingual Strategy

    Plum, Alistair and Ranasinghe, Tharindu and Purschke, Christoph. Text Generation Models for L uxembourgish with Limited Data: A Balanced Multilingual Strategy. 2025

  176. [185]

    Retrieval of Parallelizable Texts Across C hurch S lavic Variants

    Lendvai, Piroska and Reichel, Uwe and Jouravel, Anna and Rabus, Achim and Renje, Elena. Retrieval of Parallelizable Texts Across C hurch S lavic Variants. 2025

  177. [186]

    Neural Text Normalization for L uxembourgish Using Real-Life Variation Data

    Lutgen, Anne-Marie and Plum, Alistair and Purschke, Christoph and Plank, Barbara. Neural Text Normalization for L uxembourgish Using Real-Life Variation Data. 2025

  178. [187]

    Improving Dialectal Slot and Intent Detection with Auxiliary Tasks: A Multi-Dialectal B avarian Case Study

    Kr. Improving Dialectal Slot and Intent Detection with Auxiliary Tasks: A Multi-Dialectal B avarian Case Study. 2025

  179. [188]

    Regional Distribution of the /el/-/ l/ Merger in A ustralian E nglish

    Coats, Steven and Diskin-Holdaway, Chlo \'e and Loakes, Debbie. Regional Distribution of the /el/-/ l/ Merger in A ustralian E nglish. 2025

  180. [189]

    Learning Cross-Dialectal Morphophonology with Syllable Structure Constraints

    Khalifa, Salam and Qaddoumi, Abdelrahim and Kodner, Jordan and Rambow, Owen. Learning Cross-Dialectal Morphophonology with Syllable Structure Constraints. 2025

  181. [190]

    and Riabi, Arij and Seddah, Djam \'e

    Lopetegui, Javier A. and Riabi, Arij and Seddah, Djam \'e. Common Ground, Diverse Roots: The Difficulty of Classifying Common Examples in S panish Varieties. 2025

  182. [191]

    Add Noise, Tasks, or Layers? M ai NLP at the V ar D ial 2025 Shared Task on N orwegian Dialectal Slot and Intent Detection

    Blaschke, Verena and K. Add Noise, Tasks, or Layers? M ai NLP at the V ar D ial 2025 Shared Task on N orwegian Dialectal Slot and Intent Detection. 2025

  183. [192]

    LTG at V ar D ial 2025 N or SID : More and Better Training Data for Slot and Intent Detection

    Midtgaard, Marthe and M hlum, Petter and Scherrer, Yves. LTG at V ar D ial 2025 N or SID : More and Better Training Data for Slot and Intent Detection. 2025

  184. [193]

    H i TZ at V ar D ial 2025 N or SID : Overcoming Data Scarcity with Language Transfer and Automatic Data Annotation

    Bengoetxea, Jaione and Zubillaga, Mikel and Azurmendi, Ekhi and Heredia, Maite and Etxaniz, Julen and Ferro, Markel and Barnes, Jeremy. H i TZ at V ar D ial 2025 N or SID : Overcoming Data Scarcity with Language Transfer and Automatic Data Annotation. 2025

  185. [194]

    CUFE @ V ar D ial 2025 N or SID : Multilingual BERT for N orwegian Dialect Identification and Intent Detection

    Ibrahim, Michael. CUFE @ V ar D ial 2025 N or SID : Multilingual BERT for N orwegian Dialect Identification and Intent Detection. 2025

  186. [195]

    Proceedings of the Second Workshop on Scaling Up Multilingual & Multi-Cultural Evaluation. 2025

  187. [196]

    The First Multilingual Model For The Detection of Suicide Texts

    Zevallos, Rodolfo Joel and Schoene, Annika Marie and Ortega, John E. The First Multilingual Model For The Detection of Suicide Texts. 2025

  188. [197]

    C ross I n: An Efficient Instruction Tuning Approach for Cross-Lingual Knowledge Alignment

    Lin, Geyu and Wang, Bin and Liu, Zhengyuan and Chen, Nancy F. C ross I n: An Efficient Instruction Tuning Approach for Cross-Lingual Knowledge Alignment. 2025

  189. [198]

    Evaluating Dialect Robustness of Language Models via Conversation Understanding

    Srirag, Dipankar and Sahoo, Nihar Ranjan and Joshi, Aditya. Evaluating Dialect Robustness of Language Models via Conversation Understanding. 2025

  190. [199]

    Cross-Lingual Document Recommendations with Transformer-Based Representations: Evaluating Multilingual Models and Mapping Techniques

    Tashu, Tsegaye Misikir and Kontos, Eduard-Raul and Sabatelli, Matthia and Valdenegro-Toro, Matias. Cross-Lingual Document Recommendations with Transformer-Based Representations: Evaluating Multilingual Models and Mapping Techniques. 2025

  191. [200]

    VRCP : Vocabulary Replacement Continued Pretraining for Efficient Multilingual Language Models

    Nozaki, Yuta and Nakashima, Dai and Sato, Ryo and Asaba, Naoki. VRCP : Vocabulary Replacement Continued Pretraining for Efficient Multilingual Language Models. 2025

  192. [201]

    Proceedings of the Second Workshop in South East Asian Language Processing. 2025

  193. [202]

    and Estuar, Maria Regina Justina E

    Bernardo, Jacob Simon D. and Estuar, Maria Regina Justina E. b AI -b AI : A Context-Aware Transliteration System for Baybayin Scripts. 2025

  194. [203]

    N usa BERT : Teaching I ndo BERT to be Multilingual and Multicultural

    Wongso, Wilson and Setiawan, David Samuel and Limcorn, Steven and Joyoadikusumo, Ananto. N usa BERT : Teaching I ndo BERT to be Multilingual and Multicultural. 2025

  195. [204]

    Evaluating Sampling Strategies for Similarity-Based Short Answer Scoring: a Case Study in T hailand

    Boonsarngsuk, Pachara and Arpanantikul, Pacharapon and Hiranwipas, Supakorn and Watcharakajorn, Wipu and Chuangsuwanich, Ekapol. Evaluating Sampling Strategies for Similarity-Based Short Answer Scoring: a Case Study in T hailand. 2025

  196. [205]

    T hai W inograd Schemas: A Benchmark for T hai Commonsense Reasoning

    Artkaew, Phakphum. T hai W inograd Schemas: A Benchmark for T hai Commonsense Reasoning. 2025

  197. [206]

    Anak Baik: A Low-Cost Approach to Curate I ndonesian Ethical and Unethical Instructions

    Hakim, Sulthan Abiyyu and Perdana, Rizal Setya and Fatyanosa, Tirana Noor. Anak Baik: A Low-Cost Approach to Curate I ndonesian Ethical and Unethical Instructions. 2025

  198. [207]

    I ndonesian Speech Content De-Identification in Low Resource Transcripts

    Abdjul, Rifqi Naufal and Puji Lestari, Dessi and Purwarianti, Ayu and Mawalim, Candy Olivia and Sakti, Sakriani and Unoki, Masashi. I ndonesian Speech Content De-Identification in Low Resource Transcripts. 2025

  199. [208]

    I ndo M orph: a Morphology Engine for I ndonesian

    Kamajaya, Ian and Moeljadi, David. I ndo M orph: a Morphology Engine for I ndonesian. 2025

  200. [209]

    N usa D ialogue: Dialogue Summarization and Generation for Underrepresented and Extremely Low-Resource Languages

    Purwarianti, Ayu and Adhista, Dea and Baptiso, Agung and Mahfuzh, Miftahul and Sabila, Yusrina and Adila, Aulia and Cahyawijaya, Samuel and Aji, Alham Fikri. N usa D ialogue: Dialogue Summarization and Generation for Underrepresented and Extremely Low-Resource Languages. 2025

  201. [210]

    Proceedings of the first International Workshop on Nakba Narratives as Language Resources. 2025

  202. [211]

    Deciphering Implicatures: On NLP and Oral Testimonies

    Sabra, Zainab. Deciphering Implicatures: On NLP and Oral Testimonies. 2025

  203. [212]

    A cultural shift in Western perceptions of P alestine

    Regier, Terry and Khalidi, Muhammad Ali. A cultural shift in Western perceptions of P alestine. 2025

  204. [213]

    and Castle, Rick and Chappell, Carissa and Schoinoplokaki, Emmanouela and Seet, Allene M

    Lamar, Annie K. and Castle, Rick and Chappell, Carissa and Schoinoplokaki, Emmanouela and Seet, Allene M. and Shilo, Amit and Nahas, Chloe. Cognitive Geographies of Catastrophe Narratives: Georeferenced Interview Transcriptions as Language Resource for Models of Forced Displac...

  205. [214]

    Sentiment Analysis of Nakba Oral Histories: A Critical Study of Large Language Models

    Ashqar, Huthaifa I. Sentiment Analysis of Nakba Oral Histories: A Critical Study of Large Language Models. 2025

  206. [215]

    The Nakba Lexicon: Building a Comprehensive Dataset from Palestinian Literature

    AbuHaija, Izza and Al Mandhari, Salim and El-Haj, Mo and Sibony, Jonas and Rayson, Paul. The Nakba Lexicon: Building a Comprehensive Dataset from Palestinian Literature. 2025

  207. [216]

    A rabic Topic Classification Corpus of the Nakba Short Stories

    Hamed, Osama and Zaidkilani, Nadeem. A rabic Topic Classification Corpus of the Nakba Short Stories. 2025

  208. [217]

    Exploring Author Style in Nakba Short Stories: A Comparative Study of Transformer-Based Models

    Hamed, Osama and Zaidkilani, Nadeem. Exploring Author Style in Nakba Short Stories: A Comparative Study of Transformer-Based Models. 2025

  209. [218]

    Detecting Inconsistencies in Narrative Elements of Cross Lingual Nakba Texts

    Hamarsheh, Nada and Elabour, Zahia and Murra, Aya and Yahya, Adnan. Detecting Inconsistencies in Narrative Elements of Cross Lingual Nakba Texts. 2025

  210. [219]

    Multilingual Propaganda Detection: Exploring Transformer-Based Models m BERT , XLM - R o BERT a, and m T 5

    Ragab, Mohamed Ibrahim and Mohamed, Ensaf Hussein and Medhat, Walaa. Multilingual Propaganda Detection: Exploring Transformer-Based Models m BERT , XLM - R o BERT a, and m T 5. 2025

  211. [220]

    and Rayan, Tamara N

    Awad, Ghadir A. and Rayan, Tamara N. and Dunagan, Lavinia and Gamba, David. Collective Memory and Narrative Cohesion: A Computational Study of Palestinian Refugee Oral Histories in L ebanon. 2025

  212. [221]

    The Missing Cause: An Analysis of Causal Attributions in Reporting on P alestine

    Garcia Corral, Paulina and Bechara, Hannah and Manohara, Krishnamoorthy and Jankin, Slava. The Missing Cause: An Analysis of Causal Attributions in Reporting on P alestine. 2025

  213. [222]

    Bias Detection in Media: Traditional Models vs

    Mohammed, Marryam Yahya and Mohamed, Esraa Ismail and Esmat, Mariam Nabil and Nagib, Yomna Ashraf and Radwan, Nada Ahmed and Elshaer, Ziad Mohamed and Mohamed, Ensaf Hussein. Bias Detection in Media: Traditional Models vs. Transformers in Analyzing Social Media Coverage of the...

  214. [223]

    N akba TR : A T urkish NER Dataset for Nakba Narratives

    Bilgin Tasdemir, Esma Fat. N akba TR : A T urkish NER Dataset for Nakba Narratives. 2025

  215. [224]

    Integrating Argumentation Features for Enhanced Propaganda Detection in A rabic Narratives on the Israeli War on G aza

    Nabhani, Sara and Borg, Claudia and Micallef, Kurt and Al-Khatib, Khalid. Integrating Argumentation Features for Enhanced Propaganda Detection in A rabic Narratives on the Israeli War on G aza. 2025

  216. [225]

    Proceedings of the First Workshop on Multilingual Counterspeech Generation. 2025

  217. [226]

    PANDA - Paired Anti-hate Narratives Dataset from A sia: Using an LLM -as-a-Judge to Create the First C hinese Counterspeech Dataset

    Bennie, Michael and Zhang, Demi and Xiao, Bushi and Cao, Jing and Liu, Chryseis Xinyi and Meng, Jian and Tripp, Alayo. PANDA - Paired Anti-hate Narratives Dataset from A sia: Using an LLM -as-a-Judge to Create the First C hinese Counterspeech Dataset. 2025

  218. [227]

    RSSN at Multilingual Counterspeech Generation: Leveraging Lightweight Transformers for Efficient and Context-Aware Counter-Narrative Generation

    V, Ravindran. RSSN at Multilingual Counterspeech Generation: Leveraging Lightweight Transformers for Efficient and Context-Aware Counter-Narrative Generation. 2025

  219. [228]

    Northeastern Uni at Multilingual Counterspeech Generation: Enhancing Counter Speech Generation with LLM Alignment through Direct Preference Optimization

    Wadhwa, Sahil and Xu, Chengtian and Chen, Haoming and Mahalingam, Aakash and Kar, Akankshya and Chaudhary, Divya. Northeastern Uni at Multilingual Counterspeech Generation: Enhancing Counter Speech Generation with LLM Alignment through Direct Preference Optimization. 2025

  220. [229]

    NLP @ IIMAS - CLTL at Multilingual Counterspeech Generation: Combating Hate Speech Using Contextualized Knowledge Graph Representations and LLM s

    Preciado M \'a rquez, David Salvador and G \'o mez Adorno, Helena and Markov, Ilia and Baez Santamaria, Selene. NLP @ IIMAS - CLTL at Multilingual Counterspeech Generation: Combating Hate Speech Using Contextualized Knowledge Graph Representations and LLM s. 2025

  221. [230]

    CODEOFCONDUCT at Multilingual Counterspeech Generation: A Context-Aware Model for Robust Counterspeech Generation in Low-Resource Languages

    Bennie, Michael and Xiao, Bushi and Liu, Chryseis Xinyi and Zhang, Demi and Meng, Jian and Tripp, Alayo. CODEOFCONDUCT at Multilingual Counterspeech Generation: A Context-Aware Model for Robust Counterspeech Generation in Low-Resource Languages. 2025

  222. [231]

    HW - TSC at Multilingual Counterspeech Generation

    Lyu, Xinglin and Wang, Haolin and Zhang, Min and Yang, Hao. HW - TSC at Multilingual Counterspeech Generation. 2025

  223. [232]

    MNLP @Multilingual Counterspeech Generation: Evaluating Translation and Background Knowledge Filtering

    Moscato, Emanuele and Muti, Arianna and Nozza, Debora. MNLP @Multilingual Counterspeech Generation: Evaluating Translation and Background Knowledge Filtering. 2025

  224. [233]

    Hyderabadi Pearls at Multilingual Counterspeech Generation : HALT : Hate Speech Alleviation using Large Language Models and Transformers

    Farhan, Md Shariq. Hyderabadi Pearls at Multilingual Counterspeech Generation : HALT : Hate Speech Alleviation using Large Language Models and Transformers. 2025

  225. [234]

    T ren T eam at Multilingual Counterspeech Generation: Multilingual Passage Re-Ranking Approaches for Knowledge-Driven Counterspeech Generation Against Hate

    Russo, Daniel. T ren T eam at Multilingual Counterspeech Generation: Multilingual Passage Re-Ranking Approaches for Knowledge-Driven Counterspeech Generation Against Hate. 2025

  226. [235]

    The First Workshop on Multilingual Counterspeech Generation at COLING 2025: Overview of the Shared Task

    Bonaldi, Helena and Vallecillo-Rodr \'i guez, Mar \'i a Estrella and Zubiaga, Irune and Montejo-Raez, Arturo and Soroa, Aitor and Mart \'i n-Valdivia, Mar \'i a-Teresa and Guerini, Marco and Agerri, Rodrigo. The First Workshop on Multilingual Counterspeech Generation at COLING...

  227. [236]

    Proceedings of the First Workshop on Language Models for Low-Resource Languages. 2025

  228. [237]

    Overview of the First Workshop on Language Models for Low-Resource Languages ( L o R es LM 2025)

    Hettiarachchi, Hansi and Ranasinghe, Tharindu and Rayson, Paul and Mitkov, Ruslan and Gaber, Mohamed and Premasiri, Damith and Tan, Fiona Anting and Uyangodage, Lasitha Randunu Chandrakantha. Overview of the First Workshop on Language Models for Low-Resource Languages ( L o R ...

  229. [238]

    Atlas-Chat: Adapting Large Language Models for Low-Resource M oroccan A rabic Dialect

    Shang, Guokan and Abdine, Hadi and Khoubrane, Yousef and Mohamed, Amr and Abbahaddou, Yassine and Ennadir, Sofiane and Momayiz, Imane and Ren, Xuguang and Moulines, Eric and Nakov, Preslav and Vazirgiannis, Michalis and Xing, Eric. Atlas-Chat: Adapting Large Language Models fo...

  230. [239]

    Empowering P ersian LLM s for Instruction Following: A Novel Dataset and Training Approach

    Mokhtarabadi, Hojjat and Zamani, Ziba and Maazallahi, Abbas and Manshaei, Mohammad Hossein. Empowering P ersian LLM s for Instruction Following: A Novel Dataset and Training Approach. 2025

  231. [240]

    B n S ent M ix: A Diverse B engali- E nglish Code-Mixed Dataset for Sentiment Analysis

    Alam, Sadia and Ishmam, Md Farhan and Alvee, Navid Hasin and Siddique, Md Shahnewaz and Hossain, Md Azam and Kamal, Abu Raihan Mostofa. B n S ent M ix: A Diverse B engali- E nglish Code-Mixed Dataset for Sentiment Analysis. 2025

  232. [241]

    Using Language Models for assessment of users' satisfaction with their partner in P ersian

    Habibzadeh, Zahra and Asadpour, Masoud. Using Language Models for assessment of users' satisfaction with their partner in P ersian. 2025

  233. [242]

    Enhancing Plagiarism Detection in M arathi with a Weighted Ensemble of TF - IDF and BERT Embeddings for Low-Resource Language Processing

    Mutsaddi, Atharva and Choudhary, Aditya Prashant. Enhancing Plagiarism Detection in M arathi with a Weighted Ensemble of TF - IDF and BERT Embeddings for Low-Resource Language Processing. 2025

  234. [243]

    Investigating the Impact of Language-Adaptive Fine-Tuning on Sentiment Analysis in H ausa Language Using A fri BERT a

    Sani, Sani Abdullahi and Muhammad, Shamsuddeen Hassan and Jarvis, Devon. Investigating the Impact of Language-Adaptive Fine-Tuning on Sentiment Analysis in H ausa Language Using A fri BERT a. 2025

  235. [244]

    and Gipp, Bela

    Zhukova, Anastasia and Matt, Christian E. and Gipp, Bela. Automated Collection of Evaluation Dataset for Semantic Search in Low-Resource Domain Language. 2025

  236. [245]

    F ilipino Benchmarks for Measuring Sexist and Homophobic Bias in Multilingual Language Models from S outheast A sia

    Gamboa, Lance Calvin Lim and Lee, Mark. F ilipino Benchmarks for Measuring Sexist and Homophobic Bias in Multilingual Language Models from S outheast A sia. 2025

  237. [246]

    Exploiting Word Sense Disambiguation in Large Language Models for Machine Translation

    Tran, Van-Hien and Dabre, Raj and Kaing, Hour and Song, Haiyue and Tanaka, Hideki and Utiyama, Masao. Exploiting Word Sense Disambiguation in Large Language Models for Machine Translation. 2025

  238. [247]

    Low-Resource Interlinear Translation: Morphology-Enhanced Neural Models for A ncient G reek

    Rapacz, Maciej and Smywi \'n ski-Pohl, Aleksander. Low-Resource Interlinear Translation: Morphology-Enhanced Neural Models for A ncient G reek. 2025

  239. [248]

    Language ver Y Rare for All

    Merad, Ibrahim and Wolf, Amos and Mazzawi, Ziad and L \'e o, Yannick. Language ver Y Rare for All. 2025

  240. [249]

    and Doh, Joon Young and Rodan, Eid and Zhu, Kevin and O ' Brien, Sean

    Donthi, Sundesh and Spencer, Maximilian and Patel, Om B. and Doh, Joon Young and Rodan, Eid and Zhu, Kevin and O ' Brien, Sean. Improving LLM Abilities in Idiomatic Translation. 2025

  241. [250]

    A Comparative Study of Static and Contextual Embeddings for Analyzing Semantic Changes in Medieval L atin Charters

    Liu, Yifan and Tilahun, Gelila and Gao, Xinxiang and Wen, Qianfeng and Gervers, Michael. A Comparative Study of Static and Contextual Embeddings for Analyzing Semantic Changes in Medieval L atin Charters. 2025

  242. [251]

    Bridging Literacy Gaps in A frican Informal Business Management with Low-Resource Conversational Agents

    Ouattara, Maimouna and Kabor \'e , Abdoul Kader and Klein, Jacques and Bissyand \'e , Tegawend \'e F. Bridging Literacy Gaps in A frican Informal Business Management with Low-Resource Conversational Agents. 2025

  243. [252]

    Social Bias in Large Language Models For B angla: An Empirical Study on Gender and Religious Bias

    Sadhu, Jayanta and Saha, Maneesha Rani and Shahriyar, Rifat. Social Bias in Large Language Models For B angla: An Empirical Study on Gender and Religious Bias. 2025

  244. [253]

    Extracting General-use Transformers for Low-resource Languages via Knowledge Distillation

    Cruz, Jan Christian Blaise. Extracting General-use Transformers for Low-resource Languages via Knowledge Distillation. 2025

  245. [254]

    Beyond Data Quantity: Key Factors Driving Performance in Multilingual Language Models

    Bagheri Nezhad, Sina and Agrawal, Ameeta and Pokharel, Rhitabrat. Beyond Data Quantity: Key Factors Driving Performance in Multilingual Language Models. 2025

  246. [255]

    B aby LM s for isi X hosa: Data-Efficient Language Modelling in a Low-Resource Context

    Matzopoulos, Alexis and Hendriks, Charl and Mahomed, Hishaam and Meyer, Francois. B aby LM s for isi X hosa: Data-Efficient Language Modelling in a Low-Resource Context. 2025

  247. [256]

    Mapping Cross-Lingual Sentence Representations for Low-Resource Language Pairs Using Pre-trained Language Models

    Tudor, Andreea Ioana and Tashu, Tsegaye Misikir. Mapping Cross-Lingual Sentence Representations for Low-Resource Language Pairs Using Pre-trained Language Models. 2025

  248. [257]

    How to age BERT Well: Continuous Training for Historical Language Adaptation

    Harju, Anika and van der Goot, Rob. How to age BERT Well: Continuous Training for Historical Language Adaptation. 2025

  249. [258]

    Exploiting Task Reversibility of DRS Parsing and Generation: Challenges and Insights from a Multi-lingual Perspective

    Amin, Muhammad Saad and Anselma, Luca and Mazzei, Alessandro. Exploiting Task Reversibility of DRS Parsing and Generation: Challenges and Insights from a Multi-lingual Perspective. 2025

  250. [259]

    BBPOS : BERT -based Part-of-Speech Tagging for U zbek

    Bobojonova, Latofat and Akhundjanova, Arofat and Ostheimer, Phil Sidney and Fellenz, Sophie. BBPOS : BERT -based Part-of-Speech Tagging for U zbek. 2025

  251. [260]

    When Every Token Counts: Optimal Segmentation for Low-Resource Language Models

    Dewangan, Vikrant and S, Bharath Raj and Suri, Garvit and Sonavane, Raghav. When Every Token Counts: Optimal Segmentation for Low-Resource Language Models. 2025

  252. [261]

    Recent Advancements and Challenges of T urkic C entral A sian Language Processing

    Veitsman, Yana and Hartmann, Mareike. Recent Advancements and Challenges of T urkic C entral A sian Language Processing. 2025

  253. [262]

    C a LQ uest

    Lasheras, Uriel Anderson and Pinheiro, Vladia. C a LQ uest. PT : Towards the Collection and Evaluation of Natural Causal Ladder Questions in P ortuguese for AI Agents. 2025

  254. [263]

    P ersian MCQ -Instruct: A Comprehensive Resource for Generating Multiple-Choice Questions in P ersian

    Zeinalipour, Kamyar and Jamshidi, Neda and Akbari, Fahimeh and Maggini, Marco and Bianchini, Monica and Gori, Marco. P ersian MCQ -Instruct: A Comprehensive Resource for Generating Multiple-Choice Questions in P ersian. 2025

  255. [264]

    Stop Jostling: Adaptive Negative Sampling Reduces the Marginalization of Low-Resource Language Tokens by Cross-Entropy Loss

    Turumtaev, Galim. Stop Jostling: Adaptive Negative Sampling Reduces the Marginalization of Low-Resource Language Tokens by Cross-Entropy Loss. 2025

  256. [265]

    and Alsehibani, Arwa and Qandos, Nour and Elshehy, Omar and Abdelkader, Mohamed and Koubaa, Anis

    Nacar, Omer and Sibaee, Serry Taiseer and Ahmed, Samar and Ben Atitallah, Safa and Ammar, Adel and Alhabashi, Yasser and Al-Batati, Abdulrahman S. and Alsehibani, Arwa and Qandos, Nour and Elshehy, Omar and Abdelkader, Mohamed and Koubaa, Anis. Towards Inclusive A rabic LLM s:...

  257. [266]

    Controlled Evaluation of Syntactic Knowledge in Multilingual Language Models

    Kryvosheieva, Daria and Levy, Roger. Controlled Evaluation of Syntactic Knowledge in Multilingual Language Models. 2025

  258. [267]

    Evaluating Large Language Models for In-Context Learning of Linguistic Patterns In Unseen Low Resource Languages

    Zhu, Hongpu and Liang, Yuqi and Xu, Wenjing and Xu, Hongzhi. Evaluating Large Language Models for In-Context Learning of Linguistic Patterns In Unseen Low Resource Languages. 2025

  259. [268]

    Next-Level C antonese-to- M andarin Translation: Fine-Tuning and Post-Processing with LLM s

    Dai, Yuqian and Chan, Chun Fai and Wong, Ying Ki and Pun, Tsz Ho. Next-Level C antonese-to- M andarin Translation: Fine-Tuning and Post-Processing with LLM s. 2025

  260. [269]

    When LLM s Struggle: Reference-less Translation Evaluation for Low-resource Languages

    Sindhujan, Archchana and Kanojia, Diptesh and Orasan, Constantin and Qian, Shenbin. When LLM s Struggle: Reference-less Translation Evaluation for Low-resource Languages. 2025

  261. [270]

    Proceedings of the First Workshop on Natural Language Processing for Indo-Aryan and Dravidian Languages. 2025

  262. [271]

    H indi Reading Comprehension: Do Large Language Models Exhibit Semantic Understanding?

    Lal, Daisy Monika and Rayson, Paul and El-Haj, Mo. H indi Reading Comprehension: Do Large Language Models Exhibit Semantic Understanding?. 2025

  263. [272]

    Machine Translation and Transliteration for I ndo- A ryan Languages: A Systematic Review

    Perera, Sandun Sameera and Sumanathilaka, Deshan Koshala. Machine Translation and Transliteration for I ndo- A ryan Languages: A Systematic Review. 2025

  264. [273]

    BERT opic for Topic Modeling of H indi Short Texts: A Comparative Study

    Mutsaddi, Atharva and Jamkhande, Anvi and Thakre, Aryan Shirish and Haribhakta, Yashodhara. BERT opic for Topic Modeling of H indi Short Texts: A Comparative Study. 2025

  265. [274]

    Evaluating Structural and Linguistic Quality in U rdu DRS Parsing and Generation through Bidirectional Evaluation

    Amin, Muhammad Saad and Anselma, Luca and Mazzei, Alessandro. Evaluating Structural and Linguistic Quality in U rdu DRS Parsing and Generation through Bidirectional Evaluation. 2025

  266. [275]

    Studying the Effect of H indi Tokenizer Performance on Downstream Tasks

    Goel, Rashi and Sadat, Fatiha. Studying the Effect of H indi Tokenizer Performance on Downstream Tasks. 2025

  267. [276]

    Adapting Multilingual LLM s to Low-Resource Languages using Continued Pre-training and Synthetic Corpus: A Case Study for H indi LLM s

    Joshi, Raviraj and Singla, Kanishk and Kamath, Anusha and Kalani, Raunak and Paul, Rakesh and Vaidya, Utkarsh and Chauhan, Sanjay Singh and Wartikar, Niranjan and Long, Eileen. Adapting Multilingual LLM s to Low-Resource Languages using Continued Pre-training and Synthetic Cor...

  268. [277]

    OVQA : A Dataset for Visual Question Answering and Multimodal Research in O dia Language

    Parida, Shantipriya and Sahoo, Shashikanta and Sekhar, Sambit and Sahoo, Kalyanamalini and Kotwal, Ketan and Khosla, Sonal and Dash, Satya Ranjan and Bose, Aneesh and Kohli, Guneet Singh and Lenka, Smruti Smita and Bojar, Ond r ej. OVQA : A Dataset for Visual Question Answerin...

  269. [278]

    Advancing Multilingual Speaker Identification and Verification for I ndo- A ryan and D ravidian Languages

    Sritharan, Braveenan and Thayasivam, Uthayasanker. Advancing Multilingual Speaker Identification and Verification for I ndo- A ryan and D ravidian Languages. 2025

  270. [279]

    Sentiment Analysis of S inhala News Comments Using Transformers

    Bandaranayake, Isuru and Usoof, Hakim. Sentiment Analysis of S inhala News Comments Using Transformers. 2025

  271. [280]

    E x M ute: A Context-Enriched Multimodal Dataset for Hateful Memes

    Debnath, Riddhiman Swanan and Firuj, Nahian Beente and Shakib, Abdul Wadud and Sultana, Sadia and Islam, Md Saiful. E x M ute: A Context-Enriched Multimodal Dataset for Hateful Memes. 2025

  272. [281]

    Studying the capabilities of Large Language Models in solving Combinatorics Problems posed in H indi

    Kumar, Yash and Roy, Subhajit. Studying the capabilities of Large Language Models in solving Combinatorics Problems posed in H indi. 2025

  273. [282]

    Sumon and Sami, Nasrullah and Chowdhury, Mahruba Sharmin and Islam, Md Saiful

    Shibu, Hrithik Majumdar and Datta, Shrestha and Miah, Md. Sumon and Sami, Nasrullah and Chowdhury, Mahruba Sharmin and Islam, Md Saiful. From Scarcity to Capability: Empowering Fake News Detection in Low-Resource Languages with LLM s. 2025

  274. [283]

    Enhancing Participatory Development Research in S outh A sia through LLM Agents System: An Empirically-Grounded Methodological Initiative from Field Evidence in Sri L ankan

    Zhao, Xinjie and Wang, Hao and Sriwarnasinghe, Shyaman Maduranga and Tang, Jiacheng and Wang, Shiyun and Sugiyama, Sayaka and Morikawa, So. Enhancing Participatory Development Research in S outh A sia through LLM Agents System: An Empirically-Grounded Methodological Initiative...

  275. [284]

    Identifying Aggression and Offensive Language in Code-Mixed Tweets: A Multi-Task Transfer Learning Approach

    Kancharla, Bharath and Singh, Prabhjot and Kancharla, Lohith Bhagavan and Chama, Yashita and Sharma, Raksha. Identifying Aggression and Offensive Language in Code-Mixed Tweets: A Multi-Task Transfer Learning Approach. 2025

  276. [285]

    Team I ndi D ata M iner at I ndo NLP 2025: H indi Back Transliteration - R oman to D evanagari using LL a M a

    Kumar, Saurabh and Kakadiya, Dhruvkumar Babubhai and Singh, Sanasam Ranbir. Team I ndi D ata M iner at I ndo NLP 2025: H indi Back Transliteration - R oman to D evanagari using LL a M a. 2025

  277. [286]

    I ndo NLP 2025 Shared Task: R omanized S inhala to S inhala Reverse Transliteration Using BERT

    Perera, Sandun Sameera and Jayakodi, Lahiru Prabhath and Sumanathilaka, Deshan Koshala and Anuradha, Isuri. I ndo NLP 2025 Shared Task: R omanized S inhala to S inhala Reverse Transliteration Using BERT. 2025

  278. [287]

    Proceedings of the Workshop on Generative AI and Knowledge Graphs (GenAIK). 2025

  279. [288]

    Effective Modeling of Generative Framework for Document-level Relational Triple Extraction

    Saini, Pratik and Nayak, Tapas. Effective Modeling of Generative Framework for Document-level Relational Triple Extraction. 2025

  280. [289]

    Learn Together: Joint Multitask Finetuning of Pretrained KG -enhanced LLM for Downstream Tasks

    Martynova, Anastasia and Tishin, Vladislav and Semenova, Natalia. Learn Together: Joint Multitask Finetuning of Pretrained KG -enhanced LLM for Downstream Tasks. 2025

  281. [290]

    GNET - QG : Graph Network for Multi-hop Question Generation

    Jamshidi, Samin and Chali, Yllias. GNET - QG : Graph Network for Multi-hop Question Generation. 2025

  282. [291]

    SKETCH : Structured Knowledge Enhanced Text Comprehension for Holistic Retrieval

    Mahalingam, Aakash and Gande, Vinesh Kumar and Chadha, Aman and Jain, Vinija and Chaudhary, Divya. SKETCH : Structured Knowledge Enhanced Text Comprehension for Holistic Retrieval. 2025

  283. [292]

    On Reducing Factual Hallucinations in Graph-to-Text Generation Using Large Language Models

    Iarosh, Dmitrii and Panchenko, Alexander and Salnikov, Mikhail. On Reducing Factual Hallucinations in Graph-to-Text Generation Using Large Language Models. 2025

  284. [293]

    G raph RAG : Leveraging Graph-Based Efficiency to Minimize Hallucinations in LLM -Driven RAG for Finance Data

    Barry, Mariam and Caillaut, Gaetan and Halftermeyer, Pierre and Qader, Raheel and Mouayad, Mehdi and Le Deit, Fabrice and Cariolaro, Dimitri and Gesnouin, Joseph. G raph RAG : Leveraging Graph-Based Efficiency to Minimize Hallucinations in LLM -Driven RAG for Finance Data. 2025

  285. [294]

    Structured Knowledge meets G en AI : A Framework for Logic-Driven Language Models

    Eldessouky, Farida Helmy and Ehab, Nourhan and Schindler, Carolin and Abuelkheir, Mervat and Minker, Wolfgang. Structured Knowledge meets G en AI : A Framework for Logic-Driven Language Models. 2025

  286. [295]

    Performance and Limitations of Fine-Tuned LLM s in SPARQL Query Generation

    Mecharnia, Thamer and d ' Aquin, Mathieu. Performance and Limitations of Fine-Tuned LLM s in SPARQL Query Generation. 2025

  287. [296]

    Refining Noisy Knowledge Graph with Large Language Models

    Dong, Na and Kertkeidkachorn, Natthawut and Liu, Xin and Shirai, Kiyoaki. Refining Noisy Knowledge Graph with Large Language Models. 2025

  288. [297]

    Can LLM s be Knowledge Graph Curators for Validating Triple Insertions?

    Regino, Andr \'e Gomes and dos Reis, Julio Cesar. Can LLM s be Knowledge Graph Curators for Validating Triple Insertions?. 2025

  289. [298]

    T ext2 C ypher: Bridging Natural Language and Graph Databases

    Ozsoy, Makbule Gulcin and Messallem, Leila and Besga, Jon and Minneci, Gianandrea. T ext2 C ypher: Bridging Natural Language and Graph Databases. 2025

  290. [299]

    KGF ake N et: A Knowledge Graph-Enhanced Model for Fake News Detection

    Kumar, Anuj and Kumar, Pardeep and Yadav, Abhishek and Ahlawat, Satyadev and Prasad, Yamuna. KGF ake N et: A Knowledge Graph-Enhanced Model for Fake News Detection. 2025

  291. [300]

    Style Knowledge Graph: Augmenting Text Style Transfer with Knowledge Graphs

    Toshevska, Martina and Kalajdziski, Slobodan and Gievska, Sonja. Style Knowledge Graph: Augmenting Text Style Transfer with Knowledge Graphs. 2025

Pith tools

Reviewed May 8, 2026 · model on record in the stance chip above.