A multilingual benchmark for attributing LLM-generated long-form text, showing performance degrades under distribution shifts and transformer methods transfer better across languages.
m GPT : Few-Shot Learners Go Multilingual
13 Pith papers cite this work, alongside 34 external citations. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
fields
cs.CL 13roles
background 1polarities
background 1representative citing papers
Dedicated monolingual models and tokenizers for Tamil, Telugu, Kannada, and Malayalam outperform a shared multilingual model and mGPT on tokenizer efficiency and most fine-tuned tasks, but the evaluation is single-run and partly unequal-budget.
Zero-data merging of Arabic and Persian biomedical LoRA adapters comes within 1.4–3.5 CHrF++ points of supervised adaptation for Dari; Pashto and Sorani Kurdish stay below usable quality.
A systematic benchmark of multilingual authorship attribution shows fine-tuned LLM detectors exceed 0.9 macro F1 in-language but transfer poorly across languages, with Russian training generalizing better than English.
Across 35 models and two languages, language models perform at or below chance at telling possible-but-unlikely events from impossible ones when semantic relatedness conflicts with possibility.
A new 11-task Maltese benchmark shows that 55 large language models lag behind small fine-tuned models, with prior Maltese exposure the strongest predictor.
Across ten languages, mutual information between orthographic words and pitch curves is higher in tonal than in pitch-accent and stress-accent languages, consistent with a gradient view of prosodic typology.
Multi-agent debate with tit-for-tat arguments and a judge LLM improves reasoning by preventing LLMs from locking into incorrect initial solutions.
E-CONAN benchmarks merge automatically translated, human-validated, hand-crafted, and news headline sentence pairs into unified Arabic NLI datasets with balanced class distributions.
A 2.7B German-first LLM trained cheaply on public data with language-specific quality filtering matches larger 7B models on German reasoning benchmarks and runs on-device.
A survey of NLU diagnostics benchmarks finds no shared naming convention or standard set of linguistic phenomena, and asks whether the field should build an ISO-like evaluation standard.
citing papers explorer
-
Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate
Multi-agent debate with tit-for-tat arguments and a judge LLM improves reasoning by preventing LLMs from locking into incorrect initial solutions.