Pith. sign in

REVIEW 2 cited by

Linguistics Theory Meets LLM: Code-Switched Text Generation via Equivalence Constrained Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.22660 v1 pith:LUMCHKYL submitted 2024-10-30 cs.CL

classification cs.CL
keywords textcode-switchedgenerationhumanlanguagellmsmetricscode-switching
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Code-switching, the phenomenon of alternating between two or more languages in a single conversation, presents unique challenges for Natural Language Processing (NLP). Most existing research focuses on either syntactic constraints or neural generation, with few efforts to integrate linguistic theory with large language models (LLMs) for generating natural code-switched text. In this paper, we introduce EZSwitch, a novel framework that combines Equivalence Constraint Theory (ECT) with LLMs to produce linguistically valid and fluent code-switched text. We evaluate our method using both human judgments and automatic metrics, demonstrating a significant improvement in the quality of generated code-switching sentences compared to baseline LLMs. To address the lack of suitable evaluation metrics, we conduct a comprehensive correlation study of various automatic metrics against human scores, revealing that current metrics often fail to capture the nuanced fluency of code-switched text. Additionally, we create CSPref, a human preference dataset based on human ratings and analyze model performance across ``hard`` and ``easy`` examples. Our findings indicate that incorporating linguistic constraints into LLMs leads to more robust and human-aligned generation, paving the way for scalable code-switching text generation across diverse language pairs.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Lost in the Mix: Evaluating LLM Understanding of Code-Switched Text

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Code-switching hurts LLM comprehension when non-English tokens enter English text, but inserting English into other languages often improves accuracy; fine-tuning mitigates losses more reliably than prompting.

  2. SwitchLingua: The First Large-Scale Multilingual and Multi-Ethnic Code-Switching Dataset

    cs.CL 2025-05 reject novelty 5.0 of 10

    The authors present SwitchLingua, a large multilingual code-switching text and audio dataset, and SAER, a semantic-aware error metric for code-switching ASR evaluation.

Pith tools