Pith. sign in

REVIEW 3 major objections 1 minor 143 cited by

Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate

T0 review · 3 major / 1 minor · reviewed 2026-05-13 · grok-4.3

Pith's one-line read Large language models overcome stuck reasoning by having multiple agents argue tit-for-tat under a judge instead of reflecting alone.

desk verdict MAD gives a practical empirical boost to LLM reasoning by using tit-for-tat debate, but the judge's bias controls are still thin. read the letter →

arxiv 2305.19118 v4 pith:656DEUNA submitted 2023-05-30 cs.CL

classification cs.CL
keywords largelanguagemodelsmulti-agentdebatedivergentthinkingdegenerationofthoughtself-reflectionreasoningtaskscommonsensetranslationarithmetic
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Self-reflection causes LLMs to lock into an initial answer once they gain confidence, even when that answer is wrong, because later reflection fails to produce genuinely new ideas. The paper introduces a Multi-Agent Debate framework in which separate LLM agents present opposing arguments in a back-and-forth exchange while a judge LLM oversees the process and selects a final solution. Experiments on commonsense machine translation and counter-intuitive arithmetic reasoning show that this debate setup produces better results than reflection-based methods. The authors also find that debate performance depends on stopping at an adaptive point and keeping the level of disagreement moderate rather than extreme.

What carries the argument

The Multi-Agent Debate process in which LLM agents generate opposing arguments in a tit-for-tat dynamic and a separate judge LLM synthesizes them into a final answer.

What would settle it

Apply the same MAD setup to the reported datasets and obtain accuracy no higher than self-reflection baselines, or observe the judge consistently favoring one agent's first position regardless of counter-arguments.

Watch

Extended reading notes

Core claim

The Multi-Agent Debate framework encourages divergent thinking in LLMs by placing multiple agents in a tit-for-tat argumentative state, with a judge managing the exchange to reach a final solution, thereby addressing the Degeneration-of-Thought problem that limits self-reflection on tasks requiring deep contemplation.

Load-bearing premise

The judge LLM can evaluate and combine the agents' arguments fairly without itself becoming stuck in an initial view.

Editorial extensions

If this is right

  • MAD improves performance over self-reflection on commonsense machine translation and counter-intuitive arithmetic reasoning.
  • Effective MAD requires an adaptive stopping point for the debate and only a modest level of tit-for-tat intensity.
  • Using different LLMs for agents versus judge can produce biased synthesis.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The same debate structure could be tested on other tasks that reward considering multiple perspectives, such as planning or creative writing.
  • If the judge bias problem is confirmed, replacing the judge with a human or a rule-based aggregator becomes a direct next step.
  • Scaling the number of agents beyond the small groups tested here might increase the chance of surfacing overlooked alternatives.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Request a human review

A listed scientist reviews the paper for a fee and the review publishes here regardless of verdict. See the reviewers or get listed.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 1 minor

Summary. The paper identifies the Degeneration-of-Thought (DoT) problem in self-reflection methods for LLMs on complex reasoning tasks. It proposes a Multi-Agent Debate (MAD) framework in which multiple agents engage in tit-for-tat arguments managed by a judge LLM to produce a final solution. Experiments on commonsense machine translation and counter-intuitive arithmetic reasoning datasets are reported to demonstrate effectiveness, with additional analyses on adaptive debate length and tit-for-tat intensity.

Significance. If the results hold under tighter controls, the MAD framework provides a concrete procedural approach to mitigating DoT and encouraging divergent thinking in LLMs. The open-sourced code and empirical evaluation on two challenging tasks constitute a useful contribution to the study of LLM reasoning strategies.

major comments (3)
  1. [Abstract and Experiments] The central claim requires that the judge LLM synthesizes the debate without inheriting DoT bias. The manuscript notes unfairness when different LLMs are used for agents but does not report controls that hold the judge model fixed while varying agent diversity or initial stance strength (Abstract; Experiments section).
  2. [Experiments] Statistical significance, exact baseline implementations, prompt sensitivity, and judge-bias controls are not detailed, leaving the reported gains on the two datasets difficult to interpret or reproduce (Experiments section).
  3. [Analyses] The claim that an adaptive break and modest tit-for-tat level are required for good performance lacks quantitative thresholds or effect-size tables showing how performance degrades outside those regimes (Analyses section).
minor comments (1)
  1. All prompts and exact debate templates should be included in an appendix to support reproducibility.

Simulated Author's Rebuttal

3 responses · 0 unresolved

We appreciate the referee's detailed feedback on our manuscript. We believe the suggested revisions will significantly strengthen the paper by providing more rigorous controls and quantitative analyses. Below we respond point-by-point to the major comments.

read point-by-point responses
  1. Referee: [Abstract and Experiments] The central claim requires that the judge LLM synthesizes the debate without inheriting DoT bias. The manuscript notes unfairness when different LLMs are used for agents but does not report controls that hold the judge model fixed while varying agent diversity or initial stance strength (Abstract; Experiments section).

    Authors: We thank the referee for highlighting this important aspect. While we observed that using different LLMs for agents can lead to unfair judgments, we agree that explicit controls holding the judge fixed are necessary to isolate the effect of agent diversity. In the revised manuscript, we will include additional experiments where the judge model is fixed (e.g., using GPT-4 as judge) and systematically vary the agent models and the strength of initial stances. This will provide clearer evidence that the judge synthesizes without inheriting DoT bias. revision: yes

  2. Referee: [Experiments] Statistical significance, exact baseline implementations, prompt sensitivity, and judge-bias controls are not detailed, leaving the reported gains on the two datasets difficult to interpret or reproduce (Experiments section).

    Authors: We acknowledge the need for more rigorous reporting. In the revision, we will provide: (1) statistical significance tests (e.g., p-values from paired t-tests or bootstrap) for the performance gains; (2) exact prompt templates and baseline implementations with links to code; (3) analysis of prompt sensitivity by varying key prompt elements; and (4) additional judge-bias controls as mentioned above. These details will be added to the Experiments section to enhance reproducibility. revision: yes

  3. Referee: [Analyses] The claim that an adaptive break and modest tit-for-tat level are required for good performance lacks quantitative thresholds or effect-size tables showing how performance degrades outside those regimes (Analyses section).

    Authors: We agree that quantitative support would strengthen this claim. We will add effect-size tables and plots in the Analyses section showing performance as a function of debate length (number of rounds) and tit-for-tat intensity levels. This will include thresholds where performance degrades, such as when debate continues too long or tit-for-tat is too aggressive, leading to degeneration. revision: yes

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity in derivation chain

full rationale

The paper defines the MAD framework as a procedural multi-agent interaction with a judge, without any equations, fitted parameters, or mathematical derivations. Central claims rest on empirical results from two external datasets (commonsense machine translation and counter-intuitive arithmetic reasoning) compared to baselines, with no reduction of outputs to inputs by construction. No self-citation chains, uniqueness theorems, or ansatzes are invoked in a load-bearing manner for the core argument; the DoT observation and MAD proposal are presented as independent contributions evaluated externally.

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

The central claim rests on the domain assumption that structured disagreement among LLM instances produces net gains in reasoning quality; no free parameters or invented entities are introduced.

assumptions (2)
  • domain assumption Multiple LLM agents in tit-for-tat debate can generate novel thoughts that a single agent cannot produce via self-reflection.
    Invoked to justify why the framework overcomes DoT; appears in the motivation and method description.
  • domain assumption An LLM judge can reliably select the best solution from the debate transcript.
    Required for the final output step; noted as potentially problematic when different LLMs are used.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate." pith.science (2026). https://pith.science/paper/656DEUNA

@misc{pith2026230519118,
  author       = {Pith},
  title        = {Pith review of: Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/656DEUNA}},
  note         = {Machine review of arXiv:2305.19118}
}
read the original abstract

Modern large language models (LLMs) like ChatGPT have shown remarkable performance on general language tasks but still struggle on complex reasoning tasks, which drives the research on cognitive behaviors of LLMs to explore human-like problem-solving strategies. Along this direction, one representative strategy is self-reflection, which asks an LLM to refine the solution with the feedback generated by itself iteratively. However, our study shows that such reflection-style methods suffer from the Degeneration-of-Thought (DoT) problem: once the LLM has established confidence in its solutions, it is unable to generate novel thoughts later through reflection even if its initial stance is incorrect. To address the DoT problem, we propose a Multi-Agent Debate (MAD) framework, in which multiple agents express their arguments in the state of "tit for tat" and a judge manages the debate process to obtain a final solution. Clearly, our MAD framework encourages divergent thinking in LLMs which would be helpful for tasks that require deep levels of contemplation. Experiment results on two challenging datasets, commonsense machine translation and counter-intuitive arithmetic reasoning, demonstrate the effectiveness of our MAD framework. Extensive analyses suggest that the adaptive break of debate and the modest level of "tit for tat" state are required for MAD to obtain good performance. Moreover, we find that LLMs might not be a fair judge if different LLMs are used for agents. Code is available at https://github.com/Skytliang/Multi-Agents-Debate.

Discussion (0). Continue with ORCID to comment.

Forward citations

Showing 60 of 143 Pith papers that cite this

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. See all 143 Pith citations

  1. A Layered Analysis of Disagreement And Answer Quality in Multi-Agent LLM Debate

    cs.AI 2026-09 accept novelty 8.0 of 10

    LLM debate changes surface disagreement more than it changes persistent positions or improves answer quality.

  2. Conformity Breaks Conformal Prediction

    cs.LG 2026-09 accept novelty 8.0 of 10

    Social conformity in multi-agent LLM systems breaks conformal prediction guarantees by shifting the model's score distribution, reducing coverage from 90% to 74% and enabling targeted attacks on low-confidence items.

  3. Tractable Agreement Protocols

    cs.LG 2024-11 conditional novelty 8.0 of 10

    Conversation-calibrated agents, efficiently constructible from any ML model, reach approximate agreement in few rounds while improving accuracy, generalizing Aumann-Aaronson theorems to d dimensions and action feedback.

  4. Prompt Infection: LLM-to-LLM Prompt Injection within Multi-Agent Systems

    cs.MA 2024-10 unverdicted novelty 8.0 of 10

    Prompt injection attacks can self-replicate across LLM agents in multi-agent systems, enabling data theft, misinformation, and system disruption while propagating silently.

  5. Harness-R1: Learning to Edit Executable Runtime Harnesses from Agent Failure Trajectories

    cs.AI 2026-08 conditional novelty 7.0 of 10

    A dedicated harness-editor policy trained with RL on the realized outcomes of executable patches raises frozen-agent success by 9.3 points across WebShop, ALFWorld, and DBBench.

  6. Does Multi-Agent Debate Improve AI Feedback on Research Papers?

    econ.GN 2026-07 accept novelty 7.0 of 10

    Authors of economics meta-analyses found a single-pass AI report more useful than two multi-agent debate tools that cost up to thirty times more to run.

  7. Preference Optimization Drives Monoculture in LLM Prediction Markets

    cs.CE 2026-06 unverdicted novelty 7.0 of 10

    DPO fine-tuning causes LLM agents to share output distributions with pairwise error correlations of ρ=0.70, reducing ten agents to the effective power of ≈1.4 independent forecasters.

  8. Hidden Anchors in Multi-Agent LLM Deliberation

    cs.AI 2026-06 unverdicted novelty 7.0 of 10

    Multi-agent LLM deliberation is modeled with recoverable hidden anchors that allow opinions to escape the convex hull of initial beliefs, unlike classical consensus models.

  9. DICE: Entropy-Regularized Equilibrium Selection for Stable Multi-Agent LLM Coordination

    cs.LG 2026-06 unverdicted novelty 7.0 of 10

    HQRE entropy regularization makes multi-agent LLM coordination well-posed, yielding unique equilibria, linear mirror convergence, bounded Bayesian regret, and DICE gains of 4.3–8.5 pp on reasoning/planning tasks.

  10. Test-Time Hinting for Black-Box Vision-Language Models

    cs.CV 2026-05 unverdicted novelty 7.0 of 10

    Test-Time Hinting trains a hint generator to prepend contextual guidance to VLM prompts, improving accuracy on natural-image VQA benchmarks with generalization to unseen tasks and models.

  11. Predictive Maps of Multi-Agent Reasoning: A Successor-Representation Spectrum for LLM Communication Topologies

    cs.MA 2026-05 unverdicted novelty 7.0 of 10

    Successor-representation spectra of row-stochastic communication operators predict perturbation robustness, consensus speed, and error accumulation in multi-agent LLM topologies, with condition number showing perfect ...

  12. More Is Not More: What Matters for Diversity in LLM Opinions?

    cs.CL 2026-05 conditional novelty 7.0 of 10

    Diversity in LLM opinions comes mostly from the first persona sentence and from combining different interaction architectures, not from richer personas, temperature, or diversity instructions.

  13. What Do AI Agents Talk About? Discourse and Architectural Constraints in the First AI-Only Social Network

    cs.CL 2026-03 unverdicted novelty 7.0 of 10

    Discourse among AI agents on Moltbook is largely determined by architectural constraints like context windows and identity files, appearing as social learning but actually short-horizon contextual conditioning.

  14. From Standalone LLMs to Integrated Intelligence: A Survey of Compound Al Systems

    cs.MA 2025-06 accept novelty 7.0 of 10

    A survey that defines Compound AI Systems, proposes a multi-dimensional taxonomy based on component roles and orchestration strategies, reviews four foundational paradigms, and identifies key challenges for future research.

  15. Amplified Vulnerabilities: Structured Jailbreak Attacks on LLM-based Multi-Agent Debate

    cs.CR 2025-04 conditional novelty 7.0 of 10

    Multi-agent LLM debate systems are more vulnerable to jailbreak prompts than single agents, and a structured prompt rewrite sharply increases harmful outputs.

  16. Apodex 1.1: Scaling Agentic Intelligence for Complex Work

    cs.AI 2026-08 conditional novelty 6.0 of 10

    Apodex 1.1 reports that training a general-purpose language model across executable file, search, and code environments plus coordination traces yields frontier-band agentic performance in a 397B model and a competiti...

  17. Certifying Collective Reasoning in Multi-Agent Systems via Koopman Spectral Analysis

    cs.MA 2026-08 conditional novelty 6.0 of 10

    Koopman spectral analysis of interaction traces produces convergence deadline, faction attribution, and message compression certificates that hold on a synthetic attention-consensus model of LLM debate.

  18. Enacting Constructive Conflicts with AI Agents to Enhance Reconsideration among Novice Interaction Designers

    cs.HC 2026-08 conditional novelty 6.0 of 10

    An antagonistic AI agent that raises stakeholder objections led novice designers to revise more ideas and to reconsider their design assumptions more than unsupported self-reflection.

  19. When Outputs Disperse, Does Epistemic Revision Follow? A Black-Box Diagnostic for Machine Collectives

    cs.AI 2026-08 conditional novelty 6.0 of 10

    Dispersion-revision coupling: inducing output diversity improves false-premise recovery in gpt-4o-mini but not in gemini-2.5-flash, where agents reformulate the same false conclusion.

  20. Emergence of Biased Consensus in Multi-Agent LLM Debates

    cs.MA 2026-08 conditional novelty 6.0 of 10

    Collective bias in multi-agent LLM debates emerges as a finite-N rounded mean-field phase transition controlled by the ratio of conformity to sampling temperature.

  21. FaithEyes: Towards Faithful Tool Use via Multi-Agent Process-Image Verification

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A self-judging multi-agent VLM injects process-image helpfulness verdicts into tool observations and scales tool rewards by the helpful-call ratio, improving accuracy and tool faithfulness.

  22. Beyond a Joke: Multi-Angle Reasoning for Detecting and Explaining Harmful Humor in Memes

    cs.AI 2026-07 conditional novelty 6.0 of 10

    MAR-12 improves humor and hate detection in memes by prompting a VLM through twelve reasoning perspectives, attention-weighting them, and generating explanations from the weighted evidence.

  23. Stop Hand-Holding Your Coding Agent: Engineering the Loops that Replace Step-by-Step Prompting

    cs.SE 2026-06 unverdicted novelty 6.0 of 10

    Introduces loop engineering as a distinct practice layer for coding agents, supplies a taxonomy and verification ladder, and analyzes a hand-coded corpus of fifty real loops.

  24. Tool Use Enables Undetectable Steganography in Multi-Agent LLM Systems

    cs.CR 2026-06 unverdicted novelty 6.0 of 10

    Tool-using LLM agents can implement undetectable stegosystems, shifting the primary barrier to covert multi-agent collusion from technical feasibility to coordination without explicit agreement.

  25. On-Policy Self-Distillation with Sampled Demonstrations Reduces Output Diversity

    cs.LG 2026-06 unverdicted novelty 6.0 of 10

    On-policy self-distillation with sampled demonstrations reduces rollout diversity by amplifying existing probability gaps in the base model, unlike ideal RL which preserves ratios among correct outputs.

  26. When Does Delegation Beat Majority? A Delegation-Based Aggregator for Multi-Sample LLM Inference

    cs.AI 2026-06 unverdicted novelty 6.0 of 10

    Propagational Proxy Voting driven by letter entropy and centered reasoning embeddings beats majority by +2.24 pp on non-trivial MMLU-Pro questions without labels or training.

  27. Semantic Quorum Assurance: Collective Certification for Non-Deterministic AI Infrastructure

    cs.LG 2026-06 unverdicted novelty 6.0 of 10

    Semantic Quorum Assurance routes AI infrastructure proposals to diverse sandboxed validators and applies risk-adaptive quorums to cut unsafe approvals from 18.5% to 0.3% on 500 scenarios.

  28. Certifiable Semantic Agreement Among LLM Agents: What the Admissibility Instrument Decides

    cs.MA 2026-06 unverdicted novelty 6.0 of 10

    H-CSC is a BFT protocol that converts embedding-derived signals over LLM proposals into typed semantic_commit, verdict_commit, or abort outcomes, with empirical results on diagnostic and benchmark tasks showing high c...

  29. Evidence-Grounded Ensemble Diagnosis of 802.11 Packet Captures: A Multi-Stage Pipeline with Deterministic Reliability Scoring

    cs.LG 2026-06 unverdicted novelty 6.0 of 10

    PROBE pipeline with deterministic PCAP normalization, verdict-aware evidence ensembles, and composite reliability scoring raises weighted evidence F1 to 0.957 on 87 Wi-Fi captures while avoiding LLM self-confidence an...

  30. AutoResearchClaw: Self-Reinforcing Autonomous Research with Human-AI Collaboration

    cs.AI 2026-05 unverdicted novelty 6.0 of 10

    AutoResearchClaw presents a multi-agent autonomous research pipeline with debate, self-healing execution, verifiable reporting, human-in-the-loop modes, and cross-run evolution that outperforms AI Scientist v2 by 54.7...

  31. TRINITY: An Evolved LLM Coordinator

    cs.LG 2025-12 unverdicted novelty 6.0 of 10

    A compact 0.6B-parameter coordinator with a 10K-parameter head uses evolutionary strategy to dynamically delegate roles to LLMs, achieving SOTA results such as 86.2% on LiveCodeBench.

  32. Debate2Create: Robot Co-design via Multi-Agent LLM Debate

    cs.RO 2025-10 reject novelty 6.0 of 10

    A structured multi-agent LLM debate grounded in simulation is proposed for co-designing robot morphology and reward, but the body reports only a single Ant experiment.

  33. DEBATE: A Large-Scale Benchmark for Evaluating Opinion Dynamics in Role-Playing LLM Agents

    cs.CL 2025-10 conditional novelty 6.0 of 10

    Using 2,792 humans' real debates as ground truth, role-playing LLM agents show excessive opinion convergence and public-stance drift compared with humans.

  34. Auditing medical multi-agent AI reveals risks of false consensus

    cs.CL 2025-10 conditional novelty 6.0 of 10

    Medical AI doctor-teams often reach 'consensus' by repeating initial views and suppressing correct minorities, so high accuracy hides broken reasoning processes.

  35. PosterForest: Hierarchical Multi-Agent Collaboration for Scientific Poster Generation

    cs.AI 2025-08 unverdicted novelty 6.0 of 10

    PosterForest uses a Poster Tree intermediate representation and hierarchical multi-agent reasoning to generate coherent scientific posters without training, outperforming prior methods in evaluations.

  36. Dynamic Collaboration of Multi-Language Models based on Minimal Complete Semantic Units

    cs.AI 2025-08 conditional novelty 6.0 of 10

    MCSU-based vocabulary alignment plus distance-based dynamic selection (DDS) lets several LLMs vote token-by-token, beating single models and prior ensemble baselines on multiple reasoning benchmarks without training.

  37. AgentCDM: Enhancing Multi-Agent Collaborative Decision-Making via ACH-Inspired Structured Reasoning

    cs.AI 2025-08 conditional novelty 6.0 of 10

    AgentCDM uses two-stage RL training, first with ACH reasoning scaffolding then with the scaffold gradually removed, to make a Qwen-7B decision agent outperform voting, dictatorial, and prompted-reasoning baselines on ...

  38. Evaluating Large Language Models as Expert Annotators

    cs.CL 2025-08 unverdicted novelty 6.0 of 10

    Material Fingerprinting recovers the form and parameters of hyperelastic material models by nearest-neighbor matching of test data against a simulated fingerprint database: exact at zero noise, degrading under 5% noise.

  39. Can Large Language Models Integrate Spatial Data? Empirical Insights into Reasoning Strengths and Computational Weaknesses

    cs.AI 2025-08 conditional novelty 6.0 of 10

    LLMs only become competitive at spatial data integration when given pre-computed geometric features; a review-and-refine prompt then exceeds hand-tuned heuristics.

  40. DHEvo: Data-Algorithm Based Heuristic Evolution for Generalizable MILP Solving

    cs.NE 2025-07 conditional novelty 6.0 of 10

    DHEvo co-evolves MILP training instances and diving heuristics, improving generalization over existing LLM-based heuristic generation methods.

  41. WSI-Agents: A Collaborative Multi-Agent System for Multi-Modal Whole Slide Image Analysis

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A route, verify, and summarize agent system uses existing pathology models and a knowledge base to select the best whole-slide image answer.

  42. Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation

    cs.SE 2025-07 reject novelty 6.0 of 10

    An empirical study of 1,023 CoT-code pairs shows that 76.4% of LLM-generated CoTs are low quality and that CoT correctness does not guarantee code correctness.

  43. MIND: A Multi-agent Framework for Zero-shot Harmful Meme Detection

    cs.CL 2025-07 conditional novelty 6.0 of 10

    MIND uses unlabeled similar memes, bidirectional AI insight derivation, and multi-agent debate to improve zero-shot harmful meme detection on HarM, FHM, and MAMI.

  44. DynamiCare: A Dynamic Multi-Agent Framework for Interactive and Open-Ended Medical Decision-Making

    cs.AI 2025-07 conditional novelty 6.0 of 10

    DynamiCare is a multi-agent LLM framework that runs multi-round diagnostic dialogues with a dynamically adjusted specialist team, evaluated on a new 500-patient benchmark built from MIMIC-III.

  45. Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM Reasoning

    cs.AI 2025-07 conditional novelty 6.0 of 10

    Math reasoning gains in LLMs rarely transfer to general domains; RL tuning generalizes while SFT causes forgetting and representation drift.

  46. Language Models can perform Single-Utterance Self-Correction of Perturbed Reasoning

    cs.CL 2025-06 conditional novelty 6.0 of 10

    When a math reasoning chain is perturbed mid-way, several LLMs, including non-reasoning models, can detect the error and complete the solution correctly in the same utterance, with recovery varying strongly by model s...

  47. G-Memory: Tracing Hierarchical Memory for Multi-Agent Systems

    cs.MA 2025-06 conditional novelty 6.0 of 10

    G-Memory stores past multi-agent teamwork in a three-tier graph and retrieves it to boost performance on five benchmarks.

  48. Causal Graph based Event Reasoning using Semantic Relation Experts

    cs.AI 2025-06 conditional novelty 6.0 of 10

    A multi-agent LLM debate with four semantic-relation experts builds causal event graphs that improve explainable event likelihood prediction and match fine-tuned models on forecasting and next-event prediction.

  49. When to Trust Context: Self-Reflective Debates for Context Reliability

    cs.CL 2025-06 conditional novelty 6.0 of 10

    SR-DCR uses an asymmetric debate plus self-confidence to gate whether a model follows context or its prior, improving ClashEval accuracy on several models.

  50. Empowering Economic Simulation for Massively Multiplayer Online Games through Generative Agent-Based Modeling

    cs.AI 2025-06 conditional novelty 6.0 of 10

    LLM-driven agents in a simulated MMO economy reproduce role specialization and price responses to supply and demand, though the price result is partly shaped by what the AI is told.

  51. Co-Saving: Resource Aware Multi-Agent Collaboration for Software Development

    cs.CL 2025-05 reject novelty 6.0 of 10

    Co-Saving cuts token usage by roughly half in multi-agent software development by injecting learned shortcut instructions that bypass intermediate reasoning steps, while slightly improving a composite code-quality score.

  52. Learning to Reason via Mixture-of-Thought for Logical Reasoning

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Jointly training and voting across natural language, code, and truth-table reasoning modalities improves LLM logical reasoning accuracy by up to 11.7 percentage points.

  53. Empowering LLMs in Task-Oriented Dialogues: A Domain-Independent Multi-Agent Framework and Fine-Tuning Strategy

    cs.MA 2025-05 conditional novelty 6.0 of 10

    A three-agent domain-independent framework with distribution-balanced DPO training reaches Combined 106.3 on MultiWOZ 2.2 with Qwen2.5-7B, the best score among the compared baselines.

  54. SARI: Structured Audio Reasoning via Curriculum-Guided Reinforcement Learning

    cs.CL 2025-04 conditional novelty 6.0 of 10

    A curriculum-guided reinforcement learning recipe with structured chain-of-thought improves audio question answering, reaching 67.08% on MMAU test-mini.

  55. THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models

    cs.CL 2025-04 conditional novelty 6.0 of 10

    Reasoning models are poorly calibrated to problem difficulty, overthinking easy questions; a new interrupt-based decoding method sharply reduces token spend with mostly stable accuracy.

  56. When One LLM Drools, Multi-LLM Collaboration Rules

    cs.CL 2025-02 conditional novelty 6.0 of 10

    A position paper that introduces a four-level taxonomy of multi-LLM collaboration (API, text, logit, weight) and argues it is essential for reliability, pluralism, and democratization.

  57. Debate Helps Weak-to-Strong Generalization

    cs.CL 2025-01 conditional novelty 6.0 of 10

    Debate transcripts from two strong models, used as context when training an ensemble of weak models, improve weak-to-strong generalization on four NLP classification benchmarks.

  58. NADER: Neural Architecture Design via Multi-Agent Collaboration

    cs.CV 2024-12 reject novelty 6.0 of 10

    NADER uses a multi-agent LLM team with a graph-based block representation and a reflection memory to iteratively propose and test modified neural architectures, claiming gains beyond NAS-Bench-201's optimum on CIFAR a...

  59. MMFactory: A Universal Solution Search Engine for Vision-Language Tasks

    cs.CV 2024-12 conditional novelty 6.0 of 10

    MMFactory automatically generates and benchmarks a pool of reusable programmatic vision-language solutions from a few examples, letting users pick one that fits their accuracy and speed constraints.

  60. Enhancing Mathematical Reasoning in LLMs with Background Operators

    cs.AI 2024-12 reject novelty 6.0 of 10

    Fine-tuning Llama-3.1-8B to generate Prolog programs from a fixed 54-operator library reportedly reaches 84.8% on MATH counting/probability, but test-set solutions were added to training, so the headline is not a clea...

See all 143 Pith citations

Reference graph

Works this paper leans on

286 extracted references · 286 canonical work pages · cited by 143 Pith papers (see all)

  1. [1]

    Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages=

    Answering Questions by Meta-Reasoning over Multiple Chains of Thought , author=. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages=

  2. [3]

    A Multitask, Multilingual, Multimodal Evaluation of ChatGPT on Reasoning, Hallucination, and Interactivity , author=. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

  3. [7]

    Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing , pages=

    Solving General Arithmetic Word Problems , author=. Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing , pages=

  4. [12]

    Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology , pages=

    Generative agents: Interactive simulacra of human behavior , author=. Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology , pages=

  5. [13]

    Advances in neural information processing systems , volume=

    Large language models are zero-shot reasoners , author=. Advances in neural information processing systems , volume=

  6. [15]

    Advances in Neural Information Processing Systems , volume=

    Tree of thoughts: Deliberate problem solving with large language models , author=. Advances in Neural Information Processing Systems , volume=

  7. [18]

    Advances in neural information processing systems , volume=

    Chain-of-thought prompting elicits reasoning in large language models , author=. Advances in neural information processing systems , volume=

  8. [19]

    Advances in Neural Information Processing Systems , volume=

    Reflexion: Language agents with verbal reinforcement learning , author=. Advances in Neural Information Processing Systems , volume=

Show all 286 references
  1. [20]

    Advances in Neural Information Processing Systems , volume=

    Self-refine: Iterative refinement with self-feedback , author=. Advances in Neural Information Processing Systems , volume=

  2. [21]

    arXiv preprint arXiv:2303.08774 , year=

    Gpt-4 technical report , author=. arXiv preprint arXiv:2303.08774 , year=

  3. [23]

    Transactions of the Association for Computational Linguistics , volume=

    Exploring human-like translation strategy with large language models , author=. Transactions of the Association for Computational Linguistics , volume=. 2024 , publisher=

  4. [25]

    International Conference on Machine Learning , pages=

    The unreasonable effectiveness of few-shot learning for machine translation , author=. International Conference on Machine Learning , pages=. 2023 , organization=

  5. [26]

    Interactive-Chain-Prompting: Ambiguity Resolution for Crosslingual Conditional Generation with Interaction , author=. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for...

  6. [27]

    arXiv preprint arXiv:2305.11738 , year=

    Critic: Large language models can self-correct with tool-interactive critiquing , author=. arXiv preprint arXiv:2305.11738 , year=

  7. [28]

    Advances in Neural Information Processing Systems , volume=

    Eliciting thinking hierarchy without a prior , author=. Advances in Neural Information Processing Systems , volume=

  8. [30]

    Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages=

    Question Answering as Programming for Solving Time-Sensitive Questions , author=. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages=

  9. [31]

    The 61st Annual Meeting Of The Association For Computational Linguistics , year=

    Solving Math Word Problems via Cooperative Reasoning induced Language Models , author=. The 61st Annual Meeting Of The Association For Computational Linguistics , year=

  10. [32]

    Philosophical Explorations , volume=

    Does reflection lead to wise choices? , author=. Philosophical Explorations , volume=. 2011 , publisher=

  11. [33]

    , author=

    Metacognition and Reflection by Interdisciplinary Experts: Insights from Cognitive Science and Philosophy. , author=. Issues in Interdisciplinary Studies , volume=. 2017 , publisher=

  12. [34]

    arXiv preprint arXiv:2306.04634 , year=

    On the reliability of watermarks for large language models , author=. arXiv preprint arXiv:2306.04634 , year=

  13. [35]

    arXiv preprint arXiv:2308.10848 , year=

    Agentverse: Facilitating multi-agent collaboration and exploring emergent behaviors in agents , author=. arXiv preprint arXiv:2308.10848 , year=

  14. [36]

    Advances in Neural Information Processing Systems , volume=

    Camel: Communicative agents for" mind" exploration of large language model society , author=. Advances in Neural Information Processing Systems , volume=

  15. [37]

    arXiv preprint arXiv:2308.07201 , year=

    Chateval: Towards better llm-based evaluators through multi-agent debate , author=. arXiv preprint arXiv:2308.07201 , year=

  16. [38]

    arXiv preprint arXiv:2307.07924 , year=

    Communicative agents for software development , author=. arXiv preprint arXiv:2307.07924 , year=

  17. [39]

    Thinking, fast and slow , author=

  18. [40]

    Findings of the Association for Computational Linguistics: EMNLP 2023 , pages=

    Towards Making the Most of ChatGPT for Machine Translation , author=. Findings of the Association for Computational Linguistics: EMNLP 2023 , pages=

  19. [41]

    Transactions of the Association for Computational Linguistics , volume=

    Lost in the middle: How language models use long contexts , author=. Transactions of the Association for Computational Linguistics , volume=. 2024 , publisher=

  20. [42]

    Proceedings of the Eighth Conference on Machine Translation , pages=

    Findings of the WMT 2023 Shared Task on Machine Translation with Terminologies , author=. Proceedings of the Eighth Conference on Machine Translation , pages=

  21. [43]

    Proceedings of the Sixth Conference on Machine Translation , pages=

    Findings of the WMT shared task on machine translation using terminologies , author=. Proceedings of the Sixth Conference on Machine Translation , pages=

  22. [44]

    Proceedings of the AAAI Conference on Artificial Intelligence , volume=

    Meta-cotgan: A meta cooperative training paradigm for improving adversarial text generation , author=. Proceedings of the AAAI Conference on Artificial Intelligence , volume=

  23. [45]

    Transactions on Machine Learning Research , year=

    Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models , author=. Transactions on Machine Learning Research , year=

  24. [46]

    Findings of the Association for Computational Linguistics: ACL 2023 , pages=

    Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them , author=. Findings of the Association for Computational Linguistics: ACL 2023 , pages=

  25. [47]

    Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP) , pages=

    Learning to solve arithmetic word problems with verb categorization , author=. Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP) , pages=

  26. [48]

    Yejin Bang, Samuel Cahyawijaya, Nayeon Lee, Wenliang Dai, Dan Su, Bryan Wilie, Holy Lovenia, Ziwei Ji, Tiezheng Yu, Willy Chung, et al. 2023. A multitask, multilingual, multimodal evaluation of chatgpt on reasoning, hallucination, and interactivity. In Proceedings of the 13th ...

  27. [49]

    Lisa Bortolotti. 2011. Does reflection lead to wise choices? Philosophical Explorations, 14(3):297--313

  28. [50]

    Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al. 2021. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168

  29. [51]

    Kahneman Daniel. 2017. Thinking, fast and slow. Farrar, Straus and Giroux

  30. [52]

    Shizhe Diao, Pengcheng Wang, Yong Lin, and Tong Zhang. 2023. Active prompting with chain-of-thought for large language models. arXiv preprint arXiv:2302.12246

  31. [53]

    Yilun Du, Shuang Li, Antonio Torralba, Joshua B Tenenbaum, and Igor Mordatch. 2023. Improving factuality and reasoning in language models through multiagent debate. arXiv preprint arXiv:2305.14325

  32. [54]

    Yao Fu, Hao Peng, Tushar Khot, and Mirella Lapata. 2023. Improving language model negotiation with self-play and in-context learning from ai feedback. arXiv preprint arXiv:2305.10142

  33. [55]

    Yao Fu, Hao Peng, Ashish Sabharwal, Peter Clark, and Tushar Khot. 2022. Complexity-based prompting for multi-step reasoning. arXiv preprint arXiv:2210.00720

  34. [56]

    Xavier Garcia, Yamini Bansal, Colin Cherry, George Foster, Maxim Krikun, Melvin Johnson, and Orhan Firat. 2023. The unreasonable effectiveness of few-shot learning for machine translation. In International Conference on Machine Learning, pages 10867--10878. PMLR

  35. [57]

    Zhibin Gou, Zhihong Shao, Yeyun Gong, Yelong Shen, Yujiu Yang, Nan Duan, and Weizhu Chen. 2023. Critic: Large language models can self-correct with tool-interactive critiquing

  36. [58]

    Jie He, Tao Wang, Deyi Xiong, and Qun Liu. 2020. https://doi.org/10.18653/v1/2020.findings-emnlp.327 The box is in the pen: Evaluating commonsense reasoning in neural machine translation . In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 3662--36...

  37. [59]

    Zhiwei He, Tian Liang, Wenxiang Jiao, Zhuosheng Zhang, Yujiu Yang, Rui Wang, Zhaopeng Tu, Shuming Shi, and Xing Wang. 2024. Exploring human-like translation strategy with large language models. Transactions of the Association for Computational Linguistics, 12:229--246

  38. [60]

    Amr Hendy, Mohamed Abdelrehim, Amr Sharaf, Vikas Raunak, Mohamed Gabr, Hitokazu Matsushita, Young Jin Kim, Mohamed Afify, and Hany Hassan Awadalla. 2023. How good are gpt models at machine translation? a comprehensive evaluation. arXiv preprint arXiv:2302.09210

  39. [61]

    Mohammad Javad Hosseini, Hannaneh Hajishirzi, Oren Etzioni, and Nate Kushman. 2014. Learning to solve arithmetic word problems with verb categorization. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 523--533

  40. [62]

    Wenxiang Jiao, Wenxuan Wang, Jen-tse Huang, Xing Wang, Shuming Shi, and Zhaopeng Tu. 2023. Is chatgpt a good translator? yes with gpt-4 as the engine. arXiv preprint arXiv:2301.08745

  41. [63]

    Machiel Keestra. 2017. Metacognition and reflection by interdisciplinary experts: Insights from cognitive science and philosophy. Issues in Interdisciplinary Studies, 35:121--169

  42. [64]

    Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022. Large language models are zero-shot reasoners. Advances in neural information processing systems, 35:22199--22213

  43. [65]

    Yuqing Kong, Yunqi Li, Yubo Zhang, Zhihuan Huang, and Jinzhao Wu. 2022. Eliciting thinking hierarchy without a prior. Advances in Neural Information Processing Systems, 35:13329--13341

  44. [66]

    Nelson F Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang. 2024. Lost in the middle: How language models use long contexts. Transactions of the Association for Computational Linguistics, 12:157--173

  45. [67]

    Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, et al. 2024. Self-refine: Iterative refinement with self-feedback. Advances in Neural Information Processing Systems, 36

  46. [68]

    Joon Sung Park, Joseph O'Brien, Carrie Jun Cai, Meredith Ringel Morris, Percy Liang, and Michael S Bernstein. 2023. Generative agents: Interactive simulacra of human behavior. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology, pages 1--22

  47. [69]

    Jonathan Pilault, Xavier Garcia, Arthur Bra z inskas, and Orhan Firat. 2023. Interactive-chain-prompting: Ambiguity resolution for crosslingual conditional generation with interaction. In Proceedings of the 13th International Joint Conference on Natural Language Processing and...

  48. [70]

    Subhro Roy and Dan Roth. 2015. Solving general arithmetic word problems. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, pages 1743--1752

  49. [71]

    Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao. 2024. Reflexion: Language agents with verbal reinforcement learning. Advances in Neural Information Processing Systems, 36

  50. [72]

    Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R Brown, Adam Santoro, Aditya Gupta, Adri \`a Garriga-Alonso, et al. 2023. Beyond the imitation game: Quantifying and extrapolating the capabilities of language models. Transac...

  51. [73]

    Mirac Suzgun, Nathan Scales, Nathanael Sch \"a rli, Sebastian Gehrmann, Yi Tay, Hyung Won Chung, Aakanksha Chowdhery, Quoc Le, Ed Chi, Denny Zhou, et al. 2023. Challenging big-bench tasks and whether chain-of-thought can solve them. In Findings of the Association for Computati...

  52. [74]

    Peiyi Wang, Lei Li, Liang Chen, Zefan Cai, Dawei Zhu, Binghuai Lin, Yunbo Cao, Qi Liu, Tianyu Liu, and Zhifang Sui. 2023. Large language models are not fair evaluators. arXiv preprint arXiv:2305.17926

  53. [75]

    Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. 2022. Self-consistency improves chain of thought reasoning in language models. arXiv preprint arXiv:2203.11171

  54. [76]

    Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems, 35:24824--24837

  55. [77]

    Haoran Wu, Wenxuan Wang, Yuxuan Wan, Wenxiang Jiao, and Michael Lyu. 2023. Chatgpt or grammarly? evaluating chatgpt on grammatical error correction benchmark. arXiv preprint arXiv:2303.13648

  56. [78]

    Kai Xiong, Xiao Ding, Yixin Cao, Ting Liu, and Bing Qin. 2023. Diving into the inter-consistency of large language models: An insightful analysis through debate. arXiv preprint arXiv:2305.11595

  57. [79]

    Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Tom Griffiths, Yuan Cao, and Karthik Narasimhan. 2024. Tree of thoughts: Deliberate problem solving with large language models. Advances in Neural Information Processing Systems, 36

  58. [80]

    Haiyan Yin, Dingcheng Li, Xu Li, and Ping Li. 2020. Meta-cotgan: A meta cooperative training paradigm for improving adversarial text generation. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pages 9466--9473

  59. [81]

    Ori Yoran, Tomer Wolfson, Ben Bogin, Uri Katz, Daniel Deutch, and Jonathan Berant. 2023. Answering questions by meta-reasoning over multiple chains of thought. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 5942--5966

  60. [82]

    Zhuosheng Zhang, Aston Zhang, Mu Li, and Alex Smola. 2022. Automatic chain of thought prompting in large language models. arXiv preprint arXiv:2210.03493

  61. [83]

    Chuanyang Zheng, Zhengying Liu, Enze Xie, Zhenguo Li, and Yu Li. 2023. Progressive-hint prompting improves reasoning in large language models. arXiv preprint arXiv:2304.09797

  62. [84]

    Xinyu Zhu, Junjie Wang, Lin Zhang, Yuxiang Zhang, Yongfeng Huang, Jiaxing Zhang, Yujiu Yang, et al. 2023 a . Solving math word problems via cooperative reasoning induced language models. In The 61st Annual Meeting Of The Association For Computational Linguistics

  63. [85]

    Xinyu Zhu, Cheng Yang, Bei Chen, Siheng Li, Jian-Guang Lou, and Yujiu Yang. 2023 b . Question answering as programming for solving time-sensitive questions. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 12775--12790

  64. [86]

    Xizhou Zhu, Yuntao Chen, Hao Tian, Chenxin Tao, Weijie Su, Chenyu Yang, Gao Huang, Bin Li, Lewei Lu, Xiaogang Wang, et al. 2023 c . Ghost in the minecraft: Generally capable agents for open-world enviroments via large language models with text-based knowledge and memory. arXiv...

  65. [87]

    Proceedings of the Ninth Workshop on Noisy and User-generated Text (W-NUT 2024). 2024

  66. [88]

    Correcting Challenging F innish Learner Texts With Claude, GPT -3.5 and GPT -4 Large Language Models

    Creutz, Mathias. Correcting Challenging F innish Learner Texts With Claude, GPT -3.5 and GPT -4 Large Language Models. 2024

  67. [89]

    Context-aware Adversarial Attack on Named Entity Recognition

    Chen, Shuguang and Neves, Leonardo and Solorio, Thamar. Context-aware Adversarial Attack on Named Entity Recognition. 2024

  68. [90]

    Effects of different types of noise in user-generated reviews on human and machine translations including C hat GPT

    Popovic, Maja and Lapshinova-Koltunski, Ekaterina and Koponen, Maarit. Effects of different types of noise in user-generated reviews on human and machine translations including C hat GPT. 2024

  69. [91]

    Stanceosaurus 2.0 - Classifying Stance Towards R ussian and S panish Misinformation

    Lavrouk, Anton and Ligon, Ian and Zheng, Jonathan and Naous, Tarek and Xu, Wei and Ritter, Alan. Stanceosaurus 2.0 - Classifying Stance Towards R ussian and S panish Misinformation. 2024

  70. [92]

    and Shibli, G

    Elahi, Kazi and Rahman, Tasnuva and Shahriar, Shakil and Sarker, Samir and Shawon, Md. and Shibli, G. M. A Comparative Analysis of Noise Reduction Methods in Sentiment Analysis on Noisy B angla Texts. 2024

  71. [93]

    Label Supervised Contrastive Learning for Imbalanced Text Classification in E uclidean and Hyperbolic Embedding Spaces

    Khalid, Baber and Dai, Shuyang and Taghavi, Tara and Lee, Sungjin. Label Supervised Contrastive Learning for Imbalanced Text Classification in E uclidean and Hyperbolic Embedding Spaces. 2024

  72. [94]

    M aint N orm: A corpus and benchmark model for lexical normalisation and masking of industrial maintenance short text

    Bikaun, Tyler and Hodkiewicz, Melinda and Liu, Wei. M aint N orm: A corpus and benchmark model for lexical normalisation and masking of industrial maintenance short text. 2024

  73. [95]

    The Effects of Data Quality on Named Entity Recognition

    Bhadauria, Divya and Sierra M \'u nera, Alejandro and Krestel, Ralf. The Effects of Data Quality on Named Entity Recognition. 2024

  74. [96]

    Topic Bias in Emotion Classification

    Wegge, Maximilian and Klinger, Roman. Topic Bias in Emotion Classification. 2024

  75. [97]

    Stars Are All You Need: A Distantly Supervised Pyramid Network for Unified Sentiment Analysis

    Li, Wenchang and Chen, Yixing and Zheng, Shuang and Wang, Lei and Lalor, John. Stars Are All You Need: A Distantly Supervised Pyramid Network for Unified Sentiment Analysis. 2024

  76. [98]

    Proceedings of the 7th Workshop on Indian Language Data: Resources and Evaluation. 2024

  77. [99]

    Towards Disfluency Annotated Corpora for I ndian Languages

    Kochar, Chayan and Mujadia, Vandan Vasantlal and Mishra, Pruthwik and Sharma, Dipti Misra. Towards Disfluency Annotated Corpora for I ndian Languages. 2024

  78. [100]

    E mo M ix-3 L : A Code-Mixed Dataset for B angla- E nglish- H indi for Emotion Detection

    Raihan, Nishat and Goswami, Dhiman and Mahmud, Antara and Anastasopoulos, Antonios and Zampieri, Marcos. E mo M ix-3 L : A Code-Mixed Dataset for B angla- E nglish- H indi for Emotion Detection. 2024

  79. [101]

    and Buitelaar, Paul and McCrae, John P

    Rani, Priya and Negi, Gaurav and Jha, Saroj and Suryawanshi, Shardul and Ojha, Atul Kr. and Buitelaar, Paul and McCrae, John P. Findings of the WILDRE Shared Task on Code-mixed Less-resourced Sentiment Analysis for I ndo- A ryan Languages. 2024

  80. [102]

    Multilingual Bias Detection and Mitigation for I ndian Languages

    Maity, Ankita and Sharma, Anubhav and Dhar, Rudra and Abhishek, Tushar and Gupta, Manish and Varma, Vasudeva. Multilingual Bias Detection and Mitigation for I ndian Languages. 2024

  81. [103]

    Dharma \'s \=a stra Informatics: Concept Mining System for Socio-Cultural Facet in A ncient I ndia

    Nigam, Arooshi and Chandra, Subhash. Dharma \'s \=a stra Informatics: Concept Mining System for Socio-Cultural Facet in A ncient I ndia. 2024

  82. [104]

    Exploring News Summarization and Enrichment in a Highly Resource-Scarce I ndian Language: A Case Study of Mizo

    Bala, Abhinaba and Urlana, Ashok and Mishra, Rahul and Krishnamurthy, Parameswari. Exploring News Summarization and Enrichment in a Highly Resource-Scarce I ndian Language: A Case Study of Mizo. 2024

  83. [105]

    Finding the Causality of an Event in News Articles

    Lalitha Devi, Sobha and RK Rao, Pattabhi. Finding the Causality of an Event in News Articles. 2024

  84. [106]

    Creating Corpus of Low Resource I ndian Languages for Natural Language Processing: Challenges and Opportunities

    Dongare, Pratibha. Creating Corpus of Low Resource I ndian Languages for Natural Language Processing: Challenges and Opportunities. 2024

  85. [107]

    FZZG at WILDRE -7: Fine-tuning Pre-trained Models for Code-mixed, Less-resourced Sentiment Analysis

    Thakkar, Gaurish and Tadi \'c , Marko and Mikelic Preradovic, Nives. FZZG at WILDRE -7: Fine-tuning Pre-trained Models for Code-mixed, Less-resourced Sentiment Analysis. 2024

  86. [108]

    MLI nitiative@ WILDRE 7: Hybrid Approaches with Large Language Models for Enhanced Sentiment Analysis in Code-Switched and Code-Mixed Texts

    Veeramani, Hariram and Thapa, Surendrabikram and Naseem, Usman. MLI nitiative@ WILDRE 7: Hybrid Approaches with Large Language Models for Enhanced Sentiment Analysis in Code-Switched and Code-Mixed Texts. 2024

  87. [109]

    Aalamaram: A Large-Scale Linguistically Annotated Treebank for the T amil Language

    Abirami, A M and Leong, Wei Qi and Rengarajan, Hamsawardhini and Anitha, D and Suganya, R and Singh, Himanshu and Sarveswaran, Kengatharaiyer and Tjhi, William Chandra and Shah, Rajiv Ratn. Aalamaram: A Large-Scale Linguistically Annotated Treebank for the T amil Language. 2024

  88. [110]

    Proceedings of the Third Ukrainian Natural Language Processing Workshop (UNLP) @ LREC-COLING 2024. 2024

  89. [111]

    A Contemporary News Corpus of U krainian ( CNC - UA ): Compilation, Annotation, Publication

    Fischer, Stefan and Haidarzhyi, Kateryna and Knappen, J. A Contemporary News Corpus of U krainian ( CNC - UA ): Compilation, Annotation, Publication. 2024

  90. [112]

    Introducing the Djinni Recruitment Dataset: A Corpus of Anonymized CV s and Job Postings

    Drushchak, Nazarii and Romanyshyn, Mariana. Introducing the Djinni Recruitment Dataset: A Corpus of Anonymized CV s and Job Postings. 2024

  91. [113]

    Creating Parallel Corpora for U krainian: A G erman- U krainian Parallel Corpus ( P ara R ook|| DE - UK )

    Shvedova, Maria and Lukashevskyi, Arsenii. Creating Parallel Corpora for U krainian: A G erman- U krainian Parallel Corpus ( P ara R ook|| DE - UK ). 2024

  92. [114]

    Introducing NER - UK 2.0: A Rich Corpus of Named Entities for U krainian

    Chaplynskyi, Dmytro and Romanyshyn, Mariana. Introducing NER - UK 2.0: A Rich Corpus of Named Entities for U krainian. 2024

  93. [115]

    Instant Messaging Platforms News Multi-Task Classification for Stance, Sentiment, and Discrimination Detection

    Ustyianovych, Taras and Barbosa, Denilson. Instant Messaging Platforms News Multi-Task Classification for Stance, Sentiment, and Discrimination Detection. 2024

  94. [116]

    Setting up the Data Printer with Improved E nglish to U krainian Machine Translation

    Paniv, Yurii and Chaplynskyi, Dmytro and Trynus, Nikita and Kyrylov, Volodymyr. Setting up the Data Printer with Improved E nglish to U krainian Machine Translation. 2024

  95. [117]

    Automated Extraction of Hypo-Hypernym Relations for the U krainian W ord N et

    Romanyshyn, Nataliia and Chaplynskyi, Dmytro and Romanyshyn, Mariana. Automated Extraction of Hypo-Hypernym Relations for the U krainian W ord N et. 2024

  96. [118]

    U krainian Visual Word Sense Disambiguation Benchmark

    Laba, Yurii and Mohytych, Yaryna and Rohulia, Ivanna and Kyryleyza, Halyna and Dydyk-Meush, Hanna and Dobosevych, Oles and Hryniv, Rostyslav. U krainian Visual Word Sense Disambiguation Benchmark. 2024

  97. [119]

    The UNLP 2024 Shared Task on Fine-Tuning Large Language Models for U krainian

    Romanyshyn, Mariana and Syvokon, Oleksiy and Kyslyi, Roman. The UNLP 2024 Shared Task on Fine-Tuning Large Language Models for U krainian. 2024

  98. [120]

    Fine-Tuning and Retrieval Augmented Generation for Question Answering Using Affordable Large Language Models

    Boros, Tiberiu and Chivereanu, Radu and Dumitrescu, Stefan and Purcaru, Octavian. Fine-Tuning and Retrieval Augmented Generation for Question Answering Using Affordable Large Language Models. 2024

  99. [121]

    From Bytes to Borsch: Fine-Tuning Gemma and Mistral for the U krainian Language Representation

    Kiulian, Artur and Polishko, Anton and Khandoga, Mykola and Chubych, Oryna and Connor, Jack and Ravishankar, Raghav and Shirawalmath, Adarsh. From Bytes to Borsch: Fine-Tuning Gemma and Mistral for the U krainian Language Representation. 2024

  100. [122]

    Spivavtor: An Instruction Tuned U krainian Text Editing Model

    Saini, Aman and Chernodub, Artem and Raheja, Vipul and Kulkarni, Vivek. Spivavtor: An Instruction Tuned U krainian Text Editing Model. 2024

  101. [123]

    Eval- UA -tion 1.0: Benchmark for Evaluating U krainian (Large) Language Models

    Hamotskyi, Serhii and Levbarg, Anna-Izabella and H. Eval- UA -tion 1.0: Benchmark for Evaluating U krainian (Large) Language Models. 2024

  102. [124]

    L i BERT a: Advancing U krainian Language Modeling through Pre-training from Scratch

    Haltiuk, Mykola and Smywi \'n ski-Pohl, Aleksander. L i BERT a: Advancing U krainian Language Modeling through Pre-training from Scratch. 2024

  103. [125]

    Entity Embellishment Mitigation in LLM s Output with Noisy Synthetic Dataset for Alignment

    Galeshchuk, Svitlana. Entity Embellishment Mitigation in LLM s Output with Noisy Synthetic Dataset for Alignment. 2024

  104. [126]

    Language-Specific Pruning for Efficient Reduction of Large Language Models

    Shamrai, Maksym. Language-Specific Pruning for Efficient Reduction of Large Language Models. 2024

  105. [127]

    Proceedings of the Third Workshop on Understanding Implicit and Underspecified Language. 2024

  106. [128]

    Taking Action Towards Graceful Interaction: The Effects of Performing Actions on Modelling Policies for Instruction Clarification Requests

    Madureira, Brielen and Schlangen, David. Taking Action Towards Graceful Interaction: The Effects of Performing Actions on Modelling Policies for Instruction Clarification Requests. 2024

  107. [129]

    More Labels or Cases? Assessing Label Variation in Natural Language Inference

    Gruber, Cornelia and Hechinger, Katharina and Assenmacher, Matthias and Kauermann, G. More Labels or Cases? Assessing Label Variation in Natural Language Inference. 2024

  108. [130]

    Resolving Transcription Ambiguity in S panish: A Hybrid Acoustic-Lexical System for Punctuation Restoration

    Zhu, Xiliang and Chang, Chia-Tien and Gardiner, Shayna and Rossouw, David and Robertson, Jonas. Resolving Transcription Ambiguity in S panish: A Hybrid Acoustic-Lexical System for Punctuation Restoration. 2024

  109. [131]

    Assessing the Significance of Encoded Information in Contextualized Representations to Word Sense Disambiguation

    Yavas, Deniz Ekin. Assessing the Significance of Encoded Information in Contextualized Representations to Word Sense Disambiguation. 2024

  110. [132]

    Below the Sea (with the Sharks): Probing Textual Features of Implicit Sentiment in a Literary Case-study

    Bizzoni, Yuri and Feldkamp, Pascale. Below the Sea (with the Sharks): Probing Textual Features of Implicit Sentiment in a Literary Case-study. 2024

  111. [133]

    Exposing propaganda: an analysis of stylistic cues comparing human annotations and machine classification

    Faye, G \'e raud and Icard, Benjamin and Casanova, Morgane and Chanson, Julien and Maine, Fran c ois and Bancilhon, Fran c ois and Gadek, Guillaume and Gravier, Guillaume and \'E gr \'e , Paul. Exposing propaganda: an analysis of stylistic cues comparing human annotations and ...

  112. [134]

    Different Tastes of Entities: Investigating Human Label Variation in Named Entity Annotations

    Peng, Siyao and Sun, Zihang and Loftus, Sebastian and Plank, Barbara. Different Tastes of Entities: Investigating Human Label Variation in Named Entity Annotations. 2024

  113. [135]

    Colour Me Uncertain: Representing Vagueness with Probabilistic Semantics

    Chun Cheung, Kin and Emerson, Guy. Colour Me Uncertain: Representing Vagueness with Probabilistic Semantics. 2024

  114. [136]

    Proceedings of the 1st Workshop on Uncertainty-Aware NLP (UncertaiNLP 2024). 2024

  115. [137]

    Calibration-Tuning: Teaching Large Language Models to Know What They Don ' t Know

    Kapoor, Sanyam and Gruver, Nate and Roberts, Manley and Pal, Arka and Dooley, Samuel and Goldblum, Micah and Wilson, Andrew. Calibration-Tuning: Teaching Large Language Models to Know What They Don ' t Know. 2024

  116. [138]

    Context Tuning for Retrieval Augmented Generation

    Anantha, Raviteja and Vodianik, Danil. Context Tuning for Retrieval Augmented Generation. 2024

  117. [139]

    Optimizing Relation Extraction in Medical Texts through Active Learning: A Comparative Analysis of Trade-offs

    Liang, Siting and Valdunciel S \'a nchez, Pablo and Sonntag, Daniel. Optimizing Relation Extraction in Medical Texts through Active Learning: A Comparative Analysis of Trade-offs. 2024

  118. [140]

    Linguistic Obfuscation Attacks and Large Language Model Uncertainty

    Steindl, Sebastian and Sch. Linguistic Obfuscation Attacks and Large Language Model Uncertainty. 2024

  119. [141]

    Aligning Uncertainty: Leveraging LLM s to Analyze Uncertainty Transfer in Text Summarization

    Kolagar, Zahra and Zarcone, Alessandra. Aligning Uncertainty: Leveraging LLM s to Analyze Uncertainty Transfer in Text Summarization. 2024

  120. [142]

    How Does Beam Search improve Span-Level Confidence Estimation in Generative Sequence Labeling?

    Hashimoto, Kazuma and Naim, Iftekhar and Raman, Karthik. How Does Beam Search improve Span-Level Confidence Estimation in Generative Sequence Labeling?. 2024

  121. [143]

    Efficiently Acquiring Human Feedback with B ayesian Deep Learning

    Fang, Haishuo and Gor, Jeet and Simpson, Edwin. Efficiently Acquiring Human Feedback with B ayesian Deep Learning. 2024

  122. [144]

    Order Effects in Annotation Tasks: Further Evidence of Annotation Sensitivity

    Beck, Jacob and Eckman, Stephanie and Ma, Bolei and Chew, Rob and Kreuter, Frauke. Order Effects in Annotation Tasks: Further Evidence of Annotation Sensitivity. 2024

  123. [145]

    The Effect of Generalisation on the Inadequacy of the Mode

    Eikema, Bryan. The Effect of Generalisation on the Inadequacy of the Mode. 2024

  124. [146]

    Uncertainty Resolution in Misinformation Detection

    Orlovskiy, Yury and Thibault, Camille and Imouza, Anne and Godbout, Jean-Fran c ois and Rabbany, Reihaneh and Pelrine, Kellin. Uncertainty Resolution in Misinformation Detection. 2024

  125. [147]

    Don ' t Blame the Data, Blame the Model: Understanding Noise and Bias When Learning from Subjective Annotations

    Anand, Abhishek and Mokhberian, Negar and Kumar, Prathyusha and Saha, Anweasha and He, Zihao and Rao, Ashwin and Morstatter, Fred and Lerman, Kristina. Don ' t Blame the Data, Blame the Model: Understanding Noise and Bias When Learning from Subjective Annotations. 2024

  126. [148]

    Combining Confidence Elicitation and Sample-based Methods for Uncertainty Quantification in Misinformation Mitigation

    Rivera, Mauricio and Godbout, Jean-Fran c ois and Rabbany, Reihaneh and Pelrine, Kellin. Combining Confidence Elicitation and Sample-based Methods for Uncertainty Quantification in Misinformation Mitigation. 2024

  127. [149]

    Linguistically Communicating Uncertainty in Patient-Facing Risk Prediction Models

    Sivaprasad, Adarsa and Reiter, Ehud. Linguistically Communicating Uncertainty in Patient-Facing Risk Prediction Models. 2024

  128. [150]

    Proceedings of the Fourth Workshop on Threat, Aggression & Cyberbullying @ LREC-COLING-2024. 2024

  129. [151]

    The Constant in HATE : Toxicity in R eddit across Topics and Languages

    Tufa, Wondimagegnhue Tsegaye and Markov, Ilia and Vossen, Piek T.J.M. The Constant in HATE : Toxicity in R eddit across Topics and Languages. 2024

  130. [152]

    A Federated Learning Approach to Privacy Preserving Offensive Language Identification

    Zampieri, Marcos and Premasiri, Damith and Ranasinghe, Tharindu. A Federated Learning Approach to Privacy Preserving Offensive Language Identification. 2024

  131. [153]

    CLTL @ H arm P ot- ID : Leveraging Transformer Models for Detecting Offline Harm Potential and Its Targets in Low-Resource Languages

    Wang, Yeshan and Markov, Ilia. CLTL @ H arm P ot- ID : Leveraging Transformer Models for Detecting Offline Harm Potential and Its Targets in Low-Resource Languages. 2024

  132. [154]

    NJUST - KMG at TRAC -2024 Tasks 1 and 2: Offline Harm Potential Identification

    Wang, Jingyuan and Depp, Jack and Yang, Yang. NJUST - KMG at TRAC -2024 Tasks 1 and 2: Offline Harm Potential Identification. 2024

  133. [155]

    and Jha, Soumya Sangam and Rao, Vartika T

    H C, Anagha and Krishna, Saatvik M. and Jha, Soumya Sangam and Rao, Vartika T. and M, Anand Kumar. S calar L ab@ TRAC 2024: Exploring Machine Learning Techniques for Identifying Potential Offline Harm in Multilingual Commentaries. 2024

  134. [156]

    LLM -Based Synthetic Datasets: Applications and Limitations in Toxicity Detection

    Kruschwitz, Udo and Schmidhuber, Maximilian. LLM -Based Synthetic Datasets: Applications and Limitations in Toxicity Detection. 2024

  135. [157]

    Using Sarcasm to Improve Cyberbullying Detection

    Guo, Xiaoyu and Gauch, Susan. Using Sarcasm to Improve Cyberbullying Detection. 2024

  136. [158]

    Analyzing Offensive Language and Hate Speech in Political Discourse: A Case Study of G erman Politicians

    Weissenbacher, Maximilian and Kruschwitz, Udo. Analyzing Offensive Language and Hate Speech in Political Discourse: A Case Study of G erman Politicians. 2024

  137. [159]

    Ice and Fire: Dataset on Sentiment, Emotions, Toxicity, Sarcasm, Hate speech, Sympathy and More in I celandic Blog Comments

    Fri riksd \'o ttir, Steinunn Rut and Simonsen, Annika and \'A smundsson, Atli Sn r and Fri j \'o nsd \'o ttir, Gu r \'u n Lilja and Ingason, Anton Karl and Sn bjarnarson, V \'e steinn and Einarsson, Hafsteinn. Ice and Fire: Dataset on Sentiment, Emotions, Toxicity, Sarcasm, Ha...

  138. [160]

    Detecting Hate Speech in A mharic Using Multimodal Analysis of Social Media Memes

    Jigar, Melese Ayichlie and Ayele, Abinew Ali and Yimam, Seid Muhie and Biemann, Chris. Detecting Hate Speech in A mharic Using Multimodal Analysis of Social Media Memes. 2024

  139. [161]

    Content Moderation in Online Platforms: A Study of Annotation Methods for Inappropriate Language

    Barbarestani, Baran and Maks, Isa and Vossen, Piek T.J.M. Content Moderation in Online Platforms: A Study of Annotation Methods for Inappropriate Language. 2024

  140. [162]

    F rench T oxicity P rompts: a Large Benchmark for Evaluating and Mitigating Toxicity in F rench Texts

    Brun, Caroline and Nikoulina, Vassilina. F rench T oxicity P rompts: a Large Benchmark for Evaluating and Mitigating Toxicity in F rench Texts. 2024

  141. [163]

    Studying Reactions to Stereotypes in Teenagers: an Annotated I talian Dataset

    Chierchiello, Elisa and Bourgeade, Tom and Ricci, Giacomo and Bosco, Cristina and D ' Errico, Francesca. Studying Reactions to Stereotypes in Teenagers: an Annotated I talian Dataset. 2024

  142. [164]

    Offensiveness, Hate, Emotion and GPT : Benchmarking GPT 3.5 and GPT 4 as Classifiers on T witter-specific Datasets

    Bauer, Nikolaj and Preisig, Moritz and Volk, Martin. Offensiveness, Hate, Emotion and GPT : Benchmarking GPT 3.5 and GPT 4 as Classifiers on T witter-specific Datasets. 2024

  143. [165]

    D o D o Learning: Domain-Demographic Transfer in Language Models for Detecting Abuse Targeted at Public Figures

    Williams, Angus Redlarski and Kirk, Hannah Rose and Burke-Moore, Liam and Chung, Yi-Ling and Debono, Ivan and Johansson, Pica and Stevens, Francesca and Bright, Jonathan and Hale, Scott. D o D o Learning: Domain-Demographic Transfer in Language Models for Detecting Abuse Targe...

  144. [166]

    Empowering Users and Mitigating Harm: Leveraging Nudging Principles to Enhance Social Media Safety

    Donabauer, Gregor and Theophilou, Emily and Lomonaco, Francesco and Bursic, Sathya and Taibi, Davide and Hern \'a ndez-Leo, Davinia and Kruschwitz, Udo and Ognibene, Dimitri. Empowering Users and Mitigating Harm: Leveraging Nudging Principles to Enhance Social Media Safety. 2024

  145. [167]

    Exploring Boundaries and Intensities in Offensive and Hate Speech: Unveiling the Complex Spectrum of Social Media Discourse

    Ayele, Abinew Ali and Jalew, Esubalew Alemneh and Ali, Adem Chanie and Yimam, Seid Muhie and Biemann, Chris. Exploring Boundaries and Intensities in Offensive and Hate Speech: Unveiling the Complex Spectrum of Social Media Discourse. 2024

  146. [168]

    Proceedings of the 1st Worskhop on Towards Ethical and Inclusive Conversational AI: Language Attitudes, Linguistic Diversity, and Language Rights (TEICAI 2024). 2024

  147. [169]

    How Do Conversational Agents in Healthcare Impact on Patient Agency?

    Denecke, Kerstin. How Do Conversational Agents in Healthcare Impact on Patient Agency?. 2024

  148. [170]

    Why academia should cut back general enthusiasm about CA s

    Giulimondi, Alessia. Why academia should cut back general enthusiasm about CA s. 2024

  149. [171]

    Bridging the Language Gap: Integrating Language Variations into Conversational AI Agents for Enhanced User Engagement

    Amadeus, Marcellus and Homeli da Silva, Jose Roberto and Pessoa Rocha, Joao Victor. Bridging the Language Gap: Integrating Language Variations into Conversational AI Agents for Enhanced User Engagement. 2024

  150. [172]

    Socio-cultural adapted chatbots: Harnessing Knowledge Graphs and Large Language Models for enhanced context awarenes

    Camboim de S \'a , Jader and Anastasiou, Dimitra and Da Silveira, Marcos and Pruski, C \'e dric. Socio-cultural adapted chatbots: Harnessing Knowledge Graphs and Large Language Models for enhanced context awarenes. 2024

  151. [173]

    How should Conversational Agent systems respond to sexual harassment?

    De Grazia, Laura and Peir \'o Lilja, Alex and Farr \'u s Cabeceran, Mireia and Taul \'e , Mariona. How should Conversational Agent systems respond to sexual harassment?. 2024

  152. [174]

    Non-Referential Functions of Language in Social Agents: The Case of Social Proximity

    H. Non-Referential Functions of Language in Social Agents: The Case of Social Proximity. 2024

  153. [175]

    Making a Long Story Short in Conversation Modeling

    Tao, Yufei and Mines, Tiernan and Agrawal, Ameeta. Making a Long Story Short in Conversation Modeling. 2024

  154. [176]

    Proceedings of the Second International Workshop Towards Digital Language Equality (TDLE): Focusing on Sustainability @ LREC-COLING 2024. 2024

  155. [177]

    Surveying the Technology Support of Languages

    Gr. Surveying the Technology Support of Languages. 2024

  156. [178]

    Which Domains, Tasks and Languages are in the Focus of NLP Research on the Languages of E urope?

    Alves, Diego and Tadi \'c , Marko and Rehm, Georg. Which Domains, Tasks and Languages are in the Focus of NLP Research on the Languages of E urope?. 2024

  157. [179]

    Fine-Tuning Open Access LLM s for High-Precision NLU in Goal-Driven Dialog Systems

    Padr \'o , Llu \' s and Saur \' , Roser. Fine-Tuning Open Access LLM s for High-Precision NLU in Goal-Driven Dialog Systems. 2024

  158. [180]

    Could We Have Had Better Multilingual LLM s if E nglish Was Not the Central Language?

    Diandaru, Ryandito and Susanto, Lucky and Tang, Zilu and Purwarianti, Ayu and Wijaya, Derry Tanti. Could We Have Had Better Multilingual LLM s if E nglish Was Not the Central Language?. 2024

  159. [181]

    A Language Model Trained on Uruguayan S panish News Text

    Filevich, Juan Pablo and Marco, Gonzalo and Castro, Santiago and Chiruzzo, Luis and Ros \'a , Aiala. A Language Model Trained on Uruguayan S panish News Text. 2024

  160. [182]

    and Moreno-Mu \ n oz, Adri \'a n and Plaza-del-Arco, Flor Miriam and Molina Gonz \'a lez, M

    M \'a rmol Romero, Alba M. and Moreno-Mu \ n oz, Adri \'a n and Plaza-del-Arco, Flor Miriam and Molina Gonz \'a lez, M. Dolores and Montejo-R \'a ez, Arturo. Environmental Impact Measurement in the M ental R isk ES Evaluation Campaign. 2024

  161. [183]

    A mbi FC : Fact-Checking Ambiguous Claims with Evidence

    Glockner, Max and Stali \=u nait \.e , Ieva and Thorne, James and Vallejo, Gisela and Vlachos, Andreas and Gurevych, Iryna. A mbi FC : Fact-Checking Ambiguous Claims with Evidence. Transactions of the Association for Computational Linguistics. 2024. doi:10.1162/tacl_a_00629

  162. [184]

    Language Varieties of I taly: Technology Challenges and Opportunities

    Ramponi, Alan. Language Varieties of I taly: Technology Challenges and Opportunities. Transactions of the Association for Computational Linguistics. 2024. doi:10.1162/tacl_a_00631

  163. [185]

    Benchmarking Large Language Models for News Summarization

    Zhang, Tianyi and Ladhak, Faisal and Durmus, Esin and Liang, Percy and McKeown, Kathleen and Hashimoto, Tatsunori B. Benchmarking Large Language Models for News Summarization. Transactions of the Association for Computational Linguistics. 2024. doi:10.1162/tacl_a_00632

  164. [186]

    m GPT : Few-Shot Learners Go Multilingual

    Shliazhko, Oleh and Fenogenova, Alena and Tikhonova, Maria and Kozlova, Anastasia and Mikhailov, Vladislav and Shavrina, Tatiana. m GPT : Few-Shot Learners Go Multilingual. Transactions of the Association for Computational Linguistics. 2024. doi:10.1162/tacl_a_00633

  165. [187]

    Cultural Adaptation of Recipes

    Cao, Yong and Kementchedjhieva, Yova and Cui, Ruixiang and Karamolegkou, Antonia and Zhou, Li and Dare, Megan and Donatelli, Lucia and Hershcovich, Daniel. Cultural Adaptation of Recipes. Transactions of the Association for Computational Linguistics. 2024. doi:10.1162/tacl_a_00634

  166. [188]

    Metric-Free Learning Network with Dual Relations Propagation for Few-Shot Aspect Category Sentiment Analysis

    Zhao, Shiman and Xie, Yutao and Chen, Wei and Wang, Tengjiao and Yao, Jiahui and Zheng, Jiabin. Metric-Free Learning Network with Dual Relations Propagation for Few-Shot Aspect Category Sentiment Analysis. Transactions of the Association for Computational Linguistics. 2024. do...

  167. [189]

    Addressing the Binning Problem in Calibration Assessment through Scalar Annotations

    Jiang, Zhengping and Liu, Anqi and Durme, Benjamnin Van. Addressing the Binning Problem in Calibration Assessment through Scalar Annotations. Transactions of the Association for Computational Linguistics. 2024. doi:10.1162/tacl_a_00636

  168. [190]

    An Energy-based Model for Word-level A uto C ompletion in Computer-aided Translation

    Yang, Cheng and Huang, Guoping and Yu, Mo and Zhang, Zhirui and Li, Siheng and Yang, Mingming and Shi, Shuming and Yang, Yujiu and Liu, Lemao. An Energy-based Model for Word-level A uto C ompletion in Computer-aided Translation. Transactions of the Association for Computationa...

  169. [191]

    and Lin, Kevin and Hewitt, John and Paranjape, Ashwin and Bevilacqua, Michele and Petroni, Fabio and Liang, Percy

    Liu, Nelson F. and Lin, Kevin and Hewitt, John and Paranjape, Ashwin and Bevilacqua, Michele and Petroni, Fabio and Liang, Percy. Lost in the Middle: How Language Models Use Long Contexts. Transactions of the Association for Computational Linguistics. 2024. doi:10.1162/tacl_a_00638

  170. [192]

    Red Teaming Language Model Detectors with Language Models

    Shi, Zhouxing and Wang, Yihan and Yin, Fan and Chen, Xiangning and Chang, Kai-Wei and Hsieh, Cho-Jui. Red Teaming Language Model Detectors with Language Models. Transactions of the Association for Computational Linguistics. 2024. doi:10.1162/tacl_a_00639

  171. [193]

    Text Attribute Control via Closed-Loop Disentanglement

    Sha, Lei and Lukasiewicz, Thomas. Text Attribute Control via Closed-Loop Disentanglement. Transactions of the Association for Computational Linguistics. 2024. doi:10.1162/tacl_a_00640

  172. [194]

    Unifying Structured Data as Graph for Data-to-Text Pre-Training

    Li, Shujie and Li, Liang and Geng, Ruiying and Yang, Min and Li, Binhua and Yuan, Guanghu and He, Wanwei and Yuan, Shao and Ma, Can and Huang, Fei and Li, Yongbin. Unifying Structured Data as Graph for Data-to-Text Pre-Training. Transactions of the Association for Computationa...

  173. [195]

    Exploring Human-Like Translation Strategy with Large Language Models

    He, Zhiwei and Liang, Tian and Jiao, Wenxiang and Zhang, Zhuosheng and Yang, Yujiu and Wang, Rui and Tu, Zhaopeng and Shi, Shuming and Wang, Xing. Exploring Human-Like Translation Strategy with Large Language Models. Transactions of the Association for Computational Linguistic...

  174. [196]

    Retrieve What You Need: A Mutual Learning Framework for Open-domain Question Answering

    Wang, Dingmin and Huang, Qiuyuan and Jackson, Matthew and Gao, Jianfeng. Retrieve What You Need: A Mutual Learning Framework for Open-domain Question Answering. Transactions of the Association for Computational Linguistics. 2024. doi:10.1162/tacl_a_00646

  175. [197]

    Explicitly Representing Syntax Improves Sentence-to-Layout Prediction of Unexpected Situations

    Nuyts, Wolf and Cartuyvels, Ruben and Moens, Marie-Francine. Explicitly Representing Syntax Improves Sentence-to-Layout Prediction of Unexpected Situations. Transactions of the Association for Computational Linguistics. 2024. doi:10.1162/tacl_a_00643

  176. [198]

    Evaluating the Ripple Effects of Knowledge Editing in Language Models

    Cohen, Roi and Biran, Eden and Yoran, Ori and Globerson, Amir and Geva, Mor. Evaluating the Ripple Effects of Knowledge Editing in Language Models. Transactions of the Association for Computational Linguistics. 2024. doi:10.1162/tacl_a_00644

  177. [199]

    The Impact of Word Splitting on the Semantic Content of Contextualized Word Representations

    Soler, Aina Gar \' and Labeau, Matthieu and Clavel, Chlo \'e. The Impact of Word Splitting on the Semantic Content of Contextualized Word Representations. Transactions of the Association for Computational Linguistics. 2024. doi:10.1162/tacl_a_00647

  178. [200]

    Large Language Models Enable Few-Shot Clustering

    Viswanathan, Vijay and Gashteovski, Kiril and Gashteovski, Kiril and Lawrence, Carolin and Wu, Tongshuang and Neubig, Graham. Large Language Models Enable Few-Shot Clustering. Transactions of the Association for Computational Linguistics. 2024. doi:10.1162/tacl_a_00648

  179. [201]

    J usti LM : Few-shot Justification Generation for Explainable Fact-Checking of Real-world Claims

    Zeng, Fengzhu and Gao, Wei. J usti LM : Few-shot Justification Generation for Explainable Fact-Checking of Real-world Claims. Transactions of the Association for Computational Linguistics. 2024. doi:10.1162/tacl_a_00649

  180. [202]

    To Diverge or Not to Diverge: A Morphosyntactic Perspective on Machine Translation vs Human Translation

    Luo, Jiaming and Cherry, Colin and Foster, George. To Diverge or Not to Diverge: A Morphosyntactic Perspective on Machine Translation vs Human Translation. Transactions of the Association for Computational Linguistics. 2024. doi:10.1162/tacl_a_00645

  181. [203]

    What Do Self-Supervised Speech Models Know About Words?

    Pasad, Ankita and Chien, Chung-Ming and Settle, Shane and Livescu, Karen. What Do Self-Supervised Speech Models Know About Words?. Transactions of the Association for Computational Linguistics. 2024. doi:10.1162/tacl_a_00656

  182. [204]

    Are Character-level Translations Worth the Wait? Comparing B y T 5 and m T 5 for Machine Translation

    Edman, Lukas and Sarti, Gabriele and Toral, Antonio and Noord, Gertjan van and Bisazza, Arianna. Are Character-level Translations Worth the Wait? Comparing B y T 5 and m T 5 for Machine Translation. Transactions of the Association for Computational Linguistics. 2024. doi:10.11...

  183. [205]

    Geographic Adaptation of Pretrained Language Models

    Hofmann, Valentin and Glava. Geographic Adaptation of Pretrained Language Models. Transactions of the Association for Computational Linguistics. 2024. doi:10.1162/tacl_a_00652

  184. [206]

    Do Text Simplification Systems Preserve Meaning? A Human Evaluation via Reading Comprehension

    Agrawal, Sweta and Carpuat, Marine. Do Text Simplification Systems Preserve Meaning? A Human Evaluation via Reading Comprehension. Transactions of the Association for Computational Linguistics. 2024. doi:10.1162/tacl_a_00653

  185. [207]

    Simultaneous Selection and Adaptation of Source Data via Four-Level Optimization

    Xie, Pengtao and Zhao, Xingchen and He, Xuehai. Simultaneous Selection and Adaptation of Source Data via Four-Level Optimization. Transactions of the Association for Computational Linguistics. 2024. doi:10.1162/tacl_a_00658

  186. [208]

    and Choi, Jinho D

    Finch, Sarah E. and Choi, Jinho D. C onvo S ense: Overcoming Monotonous Commonsense Inferences for Conversational AI. Transactions of the Association for Computational Linguistics. 2024. doi:10.1162/tacl_a_00659

  187. [209]

    Automatically Correcting Large Language Models: Surveying the Landscape of Diverse Automated Correction Strategies

    Pan, Liangming and Saxon, Michael and Xu, Wenda and Nathani, Deepak and Wang, Xinyi and Wang, William Yang. Automatically Correcting Large Language Models: Surveying the Landscape of Diverse Automated Correction Strategies. Transactions of the Association for Computational Lin...

  188. [210]

    K o BBQ : K orean Bias Benchmark for Question Answering

    Jin, Jiho and Kim, Jiseon and Lee, Nayeon and Yoo, Haneul and Oh, Alice and Lee, Hwaran. K o BBQ : K orean Bias Benchmark for Question Answering. Transactions of the Association for Computational Linguistics. 2024. doi:10.1162/tacl_a_00661

  189. [211]

    A uto PEFT : Automatic Configuration Search for Parameter-Efficient Fine-Tuning

    Zhou, Han and Wan, Xingchen and Vuli \'c , Ivan and Korhonen, Anna. A uto PEFT : Automatic Configuration Search for Parameter-Efficient Fine-Tuning. Transactions of the Association for Computational Linguistics. 2024. doi:10.1162/tacl_a_00662

  190. [212]

    What Formal Languages Can Transformers Express? A Survey

    Strobl, Lena and Merrill, William and Weiss, Gail and Chiang, David and Angluin, Dana. What Formal Languages Can Transformers Express? A Survey. Transactions of the Association for Computational Linguistics. 2024. doi:10.1162/tacl_a_00663

  191. [213]

    Text-to- O verpass QL : A Natural Language Interface for Complex Geodata Querying of O pen S treet M ap

    Staniek, Michael and Schumann, Raphael and Z. Text-to- O verpass QL : A Natural Language Interface for Complex Geodata Querying of O pen S treet M ap. Transactions of the Association for Computational Linguistics. 2024. doi:10.1162/tacl_a_00654

  192. [214]

    Eliciting the Translation Ability of Large Language Models via Multilingual Finetuning with Translation Instructions

    Li, Jiahuan and Zhou, Hao and Huang, Shujian and Cheng, Shanbo and Chen, Jiajun. Eliciting the Translation Ability of Large Language Models via Multilingual Finetuning with Translation Instructions. Transactions of the Association for Computational Linguistics. 2024. doi:10.11...

  193. [215]

    Semantics of Multiword Expressions in Transformer-Based Models: A Survey

    Mileti \'c , Filip and Walde, Sabine Schulte im. Semantics of Multiword Expressions in Transformer-Based Models: A Survey. Transactions of the Association for Computational Linguistics. 2024. doi:10.1162/tacl_a_00657

  194. [216]

    The T hai Discourse Treebank: Annotating and Classifying T hai Discourse Connectives

    Prasertsom, Ponrawee and Jaroonpol, Apiwat and Rutherford, Attapol T. The T hai Discourse Treebank: Annotating and Classifying T hai Discourse Connectives. Transactions of the Association for Computational Linguistics. 2024. doi:10.1162/tacl_a_00650

  195. [217]

    Victoria and Herrera, Francisco

    Rodr \' guez-Barroso, Nuria and C \'a mara, Eugenio Mart \' nez and Collados, Jose Camacho and Luz \'o n, M. Victoria and Herrera, Francisco. Federated Learning for Exploiting Annotators ' Disagreements in Natural Language Processing. Transactions of the Association for Comput...

  196. [218]

    Computational Complexity of Natural Morphology Revisited

    Senuma, Hajime and Aizawa, Akiko. Computational Complexity of Natural Morphology Revisited. Transactions of the Association for Computational Linguistics. 2024. doi:10.1162/tacl_a_00665

  197. [219]

    Improving Probability-based Prompt Selection Through Unified Evaluation and Analysis

    Yang, Sohee and Kim, Jonghyeon and Jang, Joel and Ye, Seonghyeon and Lee, Hyunji and Seo, Minjoon. Improving Probability-based Prompt Selection Through Unified Evaluation and Analysis. Transactions of the Association for Computational Linguistics. 2024. doi:10.1162/tacl_a_00666

  198. [220]

    Evaluating Correctness and Faithfulness of Instruction-Following Models for Question Answering

    Adlakha, Vaibhav and BehnamGhader, Parishad and Lu, Xing Han and Meade, Nicholas and Reddy, Siva. Evaluating Correctness and Faithfulness of Instruction-Following Models for Question Answering. Transactions of the Association for Computational Linguistics. 2024. doi:10.1162/ta...

  199. [221]

    Proceedings of the 3rd Annual Meeting of the Special Interest Group on Under-resourced Languages @ LREC-COLING 2024. 2024

  200. [222]

    and Bergen, Benjamin

    Arnett, Catherine and Chang, Tyler A. and Bergen, Benjamin. A Bit of a Problem: Measurement Disparities in Dataset Sizes across Languages. 2024

  201. [223]

    A Novel Corpus for Automated Sexism Identification on Social Media

    Mut Altin, Lutfiye Seda and Saggion, Horacio. A Novel Corpus for Automated Sexism Identification on Social Media. 2024

  202. [224]

    Advancing Generative AI for P ortuguese with Open Decoder Gerv \'a sio PT *

    Santos, Rodrigo and Silva, Jo \ a o Ricardo and Gomes, Lu \' s and Rodrigues, Jo \ a o and Branco, Ant \'o nio. Advancing Generative AI for P ortuguese with Open Decoder Gerv \'a sio PT *. 2024

  203. [225]

    Assessing Pre-Built Speaker Recognition Models for Endangered Language Data

    Levow, Gina-Anne. Assessing Pre-Built Speaker Recognition Models for Endangered Language Data. 2024

  204. [226]

    BERT bek: A Pretrained Language Model for U zbek

    Kuriyozov, Elmurod and Vilares, David and G \'o mez-Rodr \' guez, Carlos. BERT bek: A Pretrained Language Model for U zbek. 2024

  205. [227]

    Beyond Error Categories: A Contextual Approach of Evaluating Emerging Spell and Grammar Checkers

    Arnard \'o ttir, \'o runn and Ing \'o lfsd \'o ttir, Svanhv \' t Lilja and S \' monarson, Haukur Barri and Einarsson, Hafsteinn and Ingason, Anton Karl and orsteinsson, Vilhj \'a lmur. Beyond Error Categories: A Contextual Approach of Evaluating Emerging Spell and Grammar Chec...

  206. [228]

    Bidirectional E nglish- N epali Machine Translation( MT ) System for Legal Domain

    Poudel, Shabdapurush and Bal, Bal Krishna and Acharya, Praveen. Bidirectional E nglish- N epali Machine Translation( MT ) System for Legal Domain. 2024

  207. [229]

    and Maranan, Jazzmin R

    Gonzales, Kiel D. and Maranan, Jazzmin R. and Santelices, Francis Paolo D. and Renovalles, Edsel Jedd M. and Macale, Nissan D. and Palafox, Nicole Anne A. and Mendoza, Jose Marie A. BK 3 AT : Bangsamoro K-3 Children ' s Speech Corpus for Developing Assessment Tools in the Bang...

  208. [230]

    C orpus A ri \`e ja: Building an Annotated Corpus with Variation in O ccitan

    Poujade, Clamenca and Bras, Myriam and Urieli, Assaf. C orpus A ri \`e ja: Building an Annotated Corpus with Variation in O ccitan. 2024

  209. [231]

    and Heeringa, Wilbert and de Vries, Wietse and Zwagers, Oscar Yde and Wieling, Martijn and Jensma, Goffe Th

    Sekeres, Hedwig G. and Heeringa, Wilbert and de Vries, Wietse and Zwagers, Oscar Yde and Wieling, Martijn and Jensma, Goffe Th. Developing Infrastructure for Low-Resource Language Corpus Building. 2024

  210. [232]

    Evaluating I celandic Sentiment Analysis Models Trained on Translated Data

    J. Evaluating I celandic Sentiment Analysis Models Trained on Translated Data. 2024

  211. [233]

    Exploring Text Classification for Enhancing Digital Game-Based Language Learning for I rish

    Mc Cahill, Leona and Baltazar, Thomas and Bruen, Sally and Xu, Liang and Ward, Monica and U \' Dhonnchadha, Elaine and Foster, Jennifer. Exploring Text Classification for Enhancing Digital Game-Based Language Learning for I rish. 2024

  212. [234]

    Forget NLI , Use a Dictionary: Zero-Shot Topic Classification for Low-Resource Languages with Application to L uxembourgish

    Philippy, Fred and Haddadan, Shohreh and Guo, Siwen. Forget NLI , Use a Dictionary: Zero-Shot Topic Classification for Low-Resource Languages with Application to L uxembourgish. 2024

  213. [235]

    Fostering the Ecosystem of Open Neural Encoders for P ortuguese with Albertina PT * Family

    Santos, Rodrigo and Rodrigues, Jo \ a o and Gomes, Lu \' s and Silva, Jo \ a o Ricardo and Branco, Ant \'o nio and Lopes Cardoso, Henrique and Os \'o rio, Tom \'a s Freitas and Leite, Bernardo. Fostering the Ecosystem of Open Neural Encoders for P ortuguese with Albertina PT *...

  214. [236]

    Improving Language Coverage on H e LI - OTS

    Jauhiainen, Tommi and Lind \'e n, Krister. Improving Language Coverage on H e LI - OTS. 2024

  215. [237]

    Improving Legal Judgement Prediction in R omanian with Long Text Encoders

    Masala, Mihai and Rebedea, Traian and Velicu, Horia. Improving Legal Judgement Prediction in R omanian with Long Text Encoders. 2024

  216. [238]

    Improving Noisy Student Training for Low-resource Languages in End-to-End ASR Using C ycle GAN and Inter-domain Losses

    Li, Chia-Yu and Vu, Ngoc Thang. Improving Noisy Student Training for Low-resource Languages in End-to-End ASR Using C ycle GAN and Inter-domain Losses. 2024

  217. [239]

    I ndonesian- E nglish Code-Switching Speech Recognition Using the Machine Speech Chain Based Semi-Supervised Learning

    Tazakka, Rais Vaza Man and Lestari, Dessi and Purwarianti, Ayu and Tanaya, Dipta and Azizah, Kurniawati and Sakti, Sakriani. I ndonesian- E nglish Code-Switching Speech Recognition Using the Machine Speech Chain Based Semi-Supervised Learning. 2024

  218. [240]

    Inter-language Transfer Learning for Visual Speech Recognition toward Under-resourced Environments

    Kondo, Fumiya and Tamura, Satoshi. Inter-language Transfer Learning for Visual Speech Recognition toward Under-resourced Environments. 2024

  219. [241]

    Investigating Neural Machine Translation for Low-Resource Languages: Using B avarian as a Case Study

    Her, Wan-hua and Kruschwitz, Udo. Investigating Neural Machine Translation for Low-Resource Languages: Using B avarian as a Case Study. 2024

  220. [242]

    and Maillard, Jean and Lusito, Stefano

    Haberland, Christopher R. and Maillard, Jean and Lusito, Stefano. I talian- L igurian Machine Translation in Its Cultural Context. 2024

  221. [243]

    Labadain-30k+: A Monolingual Tetun Document-Level Audited Dataset

    de Jesus, Gabriel and Nunes, S \'e rgio. Labadain-30k+: A Monolingual Tetun Document-Level Audited Dataset. 2024

  222. [244]

    Language Models on a Diet: Cost-Efficient Development of Encoders for Closely-Related Languages via Additional Pretraining

    Ljube s i \'c , Nikola and Suchomel, V \' t and Rupnik, Peter and Kuzman, Taja and van Noord, Rik. Language Models on a Diet: Cost-Efficient Development of Encoders for Closely-Related Languages via Additional Pretraining. 2024

  223. [245]

    Man or Machine: Evaluating Spelling Error Detection in D anish Newspaper Corpora

    Bick, Eckhard and Blom, Jonas Nygaard and Rathje, Marianne and Schack, J rgen. Man or Machine: Evaluating Spelling Error Detection in D anish Newspaper Corpora. 2024

  224. [246]

    Managing Fine-grained Metadata for Text Bases in Extremely Low Resource Languages: The Cases of Two Regional Languages of F rance

    Vergez-Couret, Marianne and Bernhard, Delphine and Nauge, Michael and Bras, Myriam and Ruiz Fabo, Pablo and Werner, Carole. Managing Fine-grained Metadata for Text Bases in Extremely Low Resource Languages: The Cases of Two Regional Languages of F rance. 2024

  225. [247]

    Mixat: A Data Set of Bilingual Emirati- E nglish Speech

    Al Ali, Maryam Khalifa and Aldarmaki, Hanan. Mixat: A Data Set of Bilingual Emirati- E nglish Speech. 2024

  226. [248]

    Multi-dialectal ASR of A rmenian from Naturalistic and Read Speech

    Arthur, Malajyan and Khurshudyan, Victoria and Avetisyan, Karen and Dolatian, Hossep and Nouvel, Damien. Multi-dialectal ASR of A rmenian from Naturalistic and Read Speech. 2024

  227. [249]

    Multilingual Self-supervised Visually Grounded Speech Models

    Nguyen, Huynh Phuong Thanh and Sakti, Sakriani. Multilingual Self-supervised Visually Grounded Speech Models. 2024

  228. [250]

    N epal Script Text Recognition Using CRNN CTC Architecture

    Nakarmi, Swornim and Sthapit, Sarin and Shakya, Arya and Chulyadyo, Rajani and Bal, Bal Krishna. N epal Script Text Recognition Using CRNN CTC Architecture. 2024

  229. [251]

    Cusenza, Giulio and. 2024

  230. [252]

    P ersian E mo: Enhancing F arsi- D ari Emotion Analysis with a Hybrid Transformer and Recurrent Neural Network Model

    Hussiny, Mohammad Ali and Payenda, Mohammad Arif and vrelid, Lilja. P ersian E mo: Enhancing F arsi- D ari Emotion Analysis with a Hybrid Transformer and Recurrent Neural Network Model. 2024

  231. [253]

    and Cajote, Rhandley D

    Guevara, Rowena Cristina L. and Cajote, Rhandley D. and Bayona, Michael Gringo Angelo R. and Lucas, Crisron Rudolf G. P hilippine Languages Database: A Multilingual Speech Corpora for Developing Systems for Low-Resource Languages. 2024

  232. [254]

    Prompting towards Alleviating Code-Switched Data Scarcity in Under-Resourced Languages with GPT as a Pivot

    Terblanche, Michelle and Olaleye, Kayode and Marivate, Vukosi. Prompting towards Alleviating Code-Switched Data Scarcity in Under-Resourced Languages with GPT as a Pivot. 2024

  233. [255]

    Quantifying the Ethical Dilemma of Using Culturally Toxic Training Data in AI Tools for Indigenous Languages

    Domingues, Pedro Henrique and Pinhanez, Claudio Santos and Cavalin, Paulo and Nogima, Julio. Quantifying the Ethical Dilemma of Using Culturally Toxic Training Data in AI Tools for Indigenous Languages. 2024

  234. [256]

    Residual Dropout: A Simple Approach to Improve Transformer ' s Data Efficiency

    Escolano, Carlos and De Luca Fornaciari, Francesca and Melero, Maite. Residual Dropout: A Simple Approach to Improve Transformer ' s Data Efficiency. 2024

  235. [257]

    Resource Acquisition for Understudied Languages: Extracting Wordlists from Dictionaries for Computer-assisted Language Comparison

    Blum, Frederic and Englisch, Johannes and Hermida Rodriguez, Alba and van Gijn, Rik and List, Johann-Mattis. Resource Acquisition for Understudied Languages: Extracting Wordlists from Dictionaries for Computer-assisted Language Comparison. 2024

  236. [258]

    Robust Guidance for Unsupervised Data Selection: Capturing Perplexing Named Entities for Domain-Specific Machine Translation

    Ji, Seunghyun and Sinulingga, Hagai Raja and Kwon, Darongsae. Robust Guidance for Unsupervised Data Selection: Capturing Perplexing Named Entities for Domain-Specific Machine Translation. 2024

  237. [259]

    Seeding Alignment between Language Technology and Indigenous Methodologies: A Decolonizing Framework for Endangered Language Revitalization

    Carpenter, Craig John and Lyon, John and Thorogood, Miles and Armstrong, Jeannette C. Seeding Alignment between Language Technology and Indigenous Methodologies: A Decolonizing Framework for Endangered Language Revitalization. 2024

  238. [260]

    Solving Failure Modes in the Creation of Trustworthy Language Technologies

    Leoni, Gianna and Steven, Lee and Keith, T \=u reiti and Mahelona, Keoni and Jones, Peter-Lucas and Duncan, Suzanne. Solving Failure Modes in the Creation of Trustworthy Language Technologies. 2024

  239. [261]

    Tandem Long-Short Duration-based Modeling for Automatic Speech Recognition

    Mengke, Dalai and Meng, Yan and Mihajlik, Peter. Tandem Long-Short Duration-based Modeling for Automatic Speech Recognition. 2024

  240. [262]

    TELP -- Text Extraction with Linguistic Patterns

    Cordeiro, Jo \ a o and Silvano, Purifica c \ a o Moura and Leal, Ant \'o nio and Pais, Sebasti \ a o. TELP -- Text Extraction with Linguistic Patterns. 2024

  241. [263]

    The First Parallel Corpus and Neural Machine Translation Model of W estern A rmenian and E nglish

    Boyac o g lu, Ari Nubar and Niehues, Jan. The First Parallel Corpus and Neural Machine Translation Model of W estern A rmenian and E nglish. 2024

  242. [264]

    Tracing Linguistic Heritage: Constructing a S omali- I talian Terminological Resource through Explorers ' Notebooks and Contemporary Corpus Analysis

    Piccini, Silvia and Vilela Ruiz, Giuliana Elizabeth and Bellandi, Andrea and Carniani, Enrico. Tracing Linguistic Heritage: Constructing a S omali- I talian Terminological Resource through Explorers ' Notebooks and Contemporary Corpus Analysis. 2024

  243. [265]

    Uncovering Social Changes of the B asque Speaking T witter Community During COVID -19 Pandemic

    Fernandez de Landa, Joseba and Garc \' a-Ferrero, Iker and Salaberria, Ander and Campos, Jon Ander. Uncovering Social Changes of the B asque Speaking T witter Community During COVID -19 Pandemic. 2024

  244. [266]

    U ni D ive: A COST Action on Universality, Diversity and Idiosyncrasy in Language Technology

    Savary, Agata and Zeman, Daniel and Barbu Mititelu, Verginica and Barreiro, Anabela and Caftanatov, Olesea and de Marneffe, Marie-Catherine and Dobrovoljc, Kaja and Eryi. U ni D ive: A COST Action on Universality, Diversity and Idiosyncrasy in Language Technology. 2024

  245. [267]

    Unsupervised Outlier Detection for Language-Independent Text Quality Filtering

    Da ason, J \'o n and Loftsson, Hrafn. Unsupervised Outlier Detection for Language-Independent Text Quality Filtering. 2024

  246. [268]

    U z ABSA : Aspect-Based Sentiment Analysis for the U zbek Language

    Matlatipov, Sanatbek Gayratovich and Rajabov, Jaloliddin and Kuriyozov, Elmurod and Aripov, Mersaid. U z ABSA : Aspect-Based Sentiment Analysis for the U zbek Language. 2024

  247. [269]

    V i H ealth NLI : A Dataset for V ietnamese Natural Language Inference in Healthcare

    Nguyen, Huyen and Ngo, Quyen The and Do, Thanh-Ha and Hoang, Tuan-Anh. V i H ealth NLI : A Dataset for V ietnamese Natural Language Inference in Healthcare. 2024

  248. [270]

    Why the Unexpected? Dissecting the Political and Economic Bias in P ersian Small and Large Language Models

    Barkhordar, Ehsan and Thapa, Surendrabikram and Maratha, Ashwarya and Naseem, Usman. Why the Unexpected? Dissecting the Political and Economic Bias in P ersian Small and Large Language Models. 2024

  249. [271]

    Work in Progress: Text-to-speech on Edge Devices for Te Reo M \=a ori and ` \=O lelo Hawaiʻi

    Keith, T \=u reiti. Work in Progress: Text-to-speech on Edge Devices for Te Reo M \=a ori and ` \=O lelo Hawaiʻi. 2024

  250. [272]

    Proceedings of the 6th Workshop on Research in Computational Linguistic Typology and Multilingual NLP. 2024

  251. [273]

    Syntactic dependency length shaped by strategic memory allocation

    Xu, Weijie and Futrell, Richard. Syntactic dependency length shaped by strategic memory allocation. 2024

  252. [274]

    GUIDE : Creating Semantic Domain Dictionaries for Low-Resource Languages

    Janetzki, Jonathan and De Melo, Gerard and Nemecek, Joshua and Whitenack, Daniel. GUIDE : Creating Semantic Domain Dictionaries for Low-Resource Languages. 2024

  253. [275]

    A New Dataset for Tonal and Segmental Dialectometry from the Y ue- and Pinghua-Speaking Area

    Sung, Ho Wang Matthew and Prokic, Jelena and Chen, Yiya. A New Dataset for Tonal and Segmental Dialectometry from the Y ue- and Pinghua-Speaking Area. 2024

  254. [276]

    A Computational Model for the Assessment of Mutual Intelligibility Among Closely Related Languages

    Nieder, Jessica and List, Johann-Mattis. A Computational Model for the Assessment of Mutual Intelligibility Among Closely Related Languages. 2024

  255. [277]

    Predicting M andarin and C antonese Adult Speakers ' Eye-Movement Patterns in Natural Reading

    Junlin, Li and Hsu, Yu-Yin and Chersoni, Emmanuele and Peng, Bo. Predicting M andarin and C antonese Adult Speakers ' Eye-Movement Patterns in Natural Reading. 2024

  256. [278]

    The Typology of Ellipsis: A Corpus for Linguistic Analysis and Machine Learning Applications

    Cavar, Damir and Mompelat, Ludovic and Abdo, Muhammad. The Typology of Ellipsis: A Corpus for Linguistic Analysis and Machine Learning Applications. 2024

  257. [279]

    Language Atlas of J apanese and Ryukyuan ( LAJ a R ): A Linguistic Typology Database for Endangered J aponic Languages

    Kato, Kanji and Miyagawa, So and Nakagawa, Natsuko. Language Atlas of J apanese and Ryukyuan ( LAJ a R ): A Linguistic Typology Database for Endangered J aponic Languages. 2024

  258. [280]

    GTNC : A Many-To-One Dataset of G oogle Translations from N ews C rawl

    Reijnaers, Damiaan and Pouw, Charlotte. GTNC : A Many-To-One Dataset of G oogle Translations from N ews C rawl. 2024

  259. [281]

    Sociolinguistically Informed Interpretability: A Case Study on H inglish Emotion Classification

    Tatariya, Kushal and Lent, Heather and Bjerva, Johannes and de Lhoneux, Miryam. Sociolinguistically Informed Interpretability: A Case Study on H inglish Emotion Classification. 2024

  260. [282]

    A Call for Consistency in Reporting Typological Diversity

    Poelman, Wessel and Ploeger, Esther and de Lhoneux, Miryam and Bjerva, Johannes. A Call for Consistency in Reporting Typological Diversity. 2024

  261. [283]

    Are Sounds Sound for Phylogenetic Reconstruction?

    H. Are Sounds Sound for Phylogenetic Reconstruction?. 2024

  262. [284]

    Compounds in U niversal D ependencies: A Survey in Five E uropean Languages

    Svoboda, Emil and S ev c \' kov \'a , Magda. Compounds in U niversal D ependencies: A Survey in Five E uropean Languages. 2024

  263. [285]

    Predicting positive transfer for improved low-resource speech recognition using acoustic pseudo-tokens

    San, Nay and Paraskevopoulos, Georgios and Arora, Aryaman and He, Xiluo and Kaur, Prabhjot and Adams, Oliver and Jurafsky, Dan. Predicting positive transfer for improved low-resource speech recognition using acoustic pseudo-tokens. 2024

  264. [286]

    and Radev, Dragomir

    Chi, Nathan and Malchev, Teodor and Kong, Riley and Chi, Ryan and Huang, Lucas and Chi, Ethan and McCoy, R. and Radev, Dragomir. M ode L ing: A Novel Dataset for Testing Linguistic Reasoning in Language Models. 2024

  265. [287]

    T artu NLP @ SIGTYP 2024 Shared Task: Adapting XLM - R o BERT a for Ancient and Historical Languages

    Dorkin, Aleksei and Sirts, Kairit. T artu NLP @ SIGTYP 2024 Shared Task: Adapting XLM - R o BERT a for Ancient and Historical Languages. 2024

  266. [288]

    Heidelberg-Boston @ SIGTYP 2024 Shared Task: Enhancing Low-Resource Language Analysis With Character-Aware Hierarchical Transformers

    Riemenschneider, Frederick and Krahn, Kevin. Heidelberg-Boston @ SIGTYP 2024 Shared Task: Enhancing Low-Resource Language Analysis With Character-Aware Hierarchical Transformers. 2024

  267. [289]

    UDP arse @ SIGTYP 2024 Shared Task : Modern Language Models for Historical Languages

    Heinecke, Johannes. UDP arse @ SIGTYP 2024 Shared Task : Modern Language Models for Historical Languages. 2024

  268. [290]

    A llen Institute for AI @ SIGTYP 2024 Shared Task on Word Embedding Evaluation for Ancient and Historical Languages

    Miranda, Lester James. A llen Institute for AI @ SIGTYP 2024 Shared Task on Word Embedding Evaluation for Ancient and Historical Languages. 2024

  269. [291]

    and Moran, P \'a draic and McCrae, John

    Dereza, Oksana and Doyle, Adrian and Rani, Priya and Ojha, Atul Kr. and Moran, P \'a draic and McCrae, John. Findings of the SIGTYP 2024 Shared Task on Word Embedding Evaluation for Ancient and Historical Languages. 2024

  270. [292]

    Proceedings of the LREC-COLING 2024 11th Workshop on the Representation and Processing of Sign Languages: Evaluation of Sign Language Resources. 2024

  271. [293]

    Advancing Annotation for Continuous Data in S wiss G erman Sign Language

    Battisti, Alessia and Tissi, Katja and Sidler-Miserez, Sandra and Ebling, Sarah. Advancing Annotation for Continuous Data in S wiss G erman Sign Language. 2024

  272. [294]

    Person Identification from Pose Estimates in Sign Language

    Battisti, Alessia and van den Bold, Emma and G. Person Identification from Pose Estimates in Sign Language. 2024

  273. [295]

    Data Integration, Annotation, and Transcription Methods for Sign Language Dialogue with Latency in Videoconferencing

    Bono, Mayumi and Okada, Tomohiro and Skobov, Victor and Adam, Robert. Data Integration, Annotation, and Transcription Methods for Sign Language Dialogue with Latency in Videoconferencing. 2024

  274. [296]

    Evaluating the Alignment of Utterances in the S wedish S ign L anguage Corpus

    B. Evaluating the Alignment of Utterances in the S wedish S ign L anguage Corpus. 2024

  275. [297]

    How to Approach Lexical Variation in Sign Language Corpora

    B. How to Approach Lexical Variation in Sign Language Corpora. 2024

  276. [298]

    and Kocab, Annemarie and Lu, Alex X

    Desai, Aashaka and De Meulder, Maartje and Hochgesang, Julie A. and Kocab, Annemarie and Lu, Alex X. Systemic Biases in Sign Language AI Research: A Deaf-Led Call to Reevaluate Research Agendas. 2024

  277. [299]

    and Oomen, Marloes and Roelofsen, Floris

    Esselink, Lyke D. and Oomen, Marloes and Roelofsen, Floris. Evaluating Inter-Annotator Agreement for Non-Manual Markers in Sign Languages. 2024

  278. [300]

    A software editor for the AZVD graphical Sign Language representation system

    Filhol, Michael and von Ascheberg, Thomas. A software editor for the AZVD graphical Sign Language representation system. 2024

Pith tools

Reviewed May 13, 2026 · model on record in the stance chip above.