Pith. sign in

cs.CL

Computation and Language

Covers natural language processing. Roughly includes material in ACM Subject Class I.2.7. Note that work on artificial languages (programming languages, logics, formal systems) that does not explicitly address natural-language issues broadly construed (natural-language processing, computational linguistics, speech, text retrieval, etc.) is not appropriate for this area.

Papers reviewed in the last 7 days lead, then the papers readers actually read. Ranking is not a quality score.

sort pith recommended most recent

Curriculum budget scheduler gives LLMs 8.3% more accuracy at 34% fewer tokens

By shifting token limits from easy to hard problems during training, the policy avoids overthinking simple questions and underthinking hard

· “Avoiding Overthinking and Underthinking: Curriculum-Aware Budget Scheduling for LLMs”

open re-runnable review →
Figure from the paper

Home activity benchmark shows AI question-answering gaps

HOME-KGQA tests multimodal KGQA on daily household tasks, where LLM methods lag behind their encyclopedic results.

· “HOME-KGQA: A Benchmark Dataset for Multimodal Knowledge Graph Question Answering on Household Daily Activities”

open re-runnable review →
Figure from the paper

Specialized GPT-5.4 beats physicians on real clinical chats

Benchmark of 15,000 clinician conversations shows the tuned model tops base versions and human doctors across consult, documentation, and 3.

· “HealthBench Professional: Evaluating Large Language Models on Real Clinician Chats”

open re-runnable review →
Figure from the paper

Dysarthria profiles differ by cause but match across languages

Analysis of 3,374 speakers shows aetiology-specific degradation with cosine similarity above 0.95 for profile shapes in 12 languages.

· “Phonological Subspace Collapse Is Aetiology-Specific and Cross-Lingually Stable: Evidence from 3,374 Speakers”

open re-runnable review →
Figure from the paper

42% of LLM turn-level findings fail after autocorrelation correction

Naive pooled tests on dependent turns within conversations produce inflated significance, and a cluster-robust method cuts the rate of nonre

· “The Autocorrelation Blind Spot: Why 42% of Turn-Level Findings in LLM Conversation Analysis May Be Spurious”

open re-runnable review →
Figure from the paper

ReVision cuts 46% of visual tokens in computer agent histories

The approach improves success rates by 3% on OSWorld and two other benchmarks by letting agents use five history screenshots more efficientl

· “ReVision: Scaling Computer-Use Agents via Temporal Visual Redundancy Reduction”

open re-runnable review →
Figure from the paper

Coarse-to-fine tree search reaches 71.6 accuracy on BIRD SQL benchmark

Level-wise exploration with formulation and evaluation agents produces diverse, correctly grained skeletons for complex natural-language-to-

· “LEAF-SQL: Level-wise Exploration with Adaptive Fine-graining for Text-to-SQL Skeleton Prediction”

open re-runnable review →
Figure from the paper

RL reasoning gains come from sparse fixes at uncertain tokens

Only 1-3% of positions need correction within the base model's top alternatives, enabling a method that matches RL with far less training.

· “Rethinking RL for LLM Reasoning: It's Sparse Policy Selection, Not Capability Learning”

open re-runnable review →
Figure from the paper

AI agents rate psychiatric symptoms better than humans on tricky cases

Decomposing interviews into symptom-specific tasks yields lower error than original raters and 0.877 agreement with experts.

· “ADAPTS: Agentic Decomposition for Automated Protocol-agnostic Tracking of Symptoms”

open re-runnable review →
Figure from the paper

Benchmark joins Reddit graphs and text to test ideological debate models

ControBench uses 26k interactions across Trump, abortion and religion threads to expose where graph and language models diverge on cross-ide

· “ControBench: An Interaction-Aware Benchmark for Controversial Discourse Analysis on Social Networks”

open re-runnable review →
Figure from the paper

Structured skill format lifts discovery MRR from 0.649 to 0.729

Separating scheduling, execution structure, and logic in skill text makes search and risk review more accurate than plain descriptions.

· “From Skill Text to Skill Structure: The Scheduling-Structural-Logical Representation for Agent Skills”

open re-runnable review →
Figure from the paper

Fine-tuned model beats GPT-4o mini by 28 points on Brazilian legal benchmark

Commercial LLMs bias heavily toward civil law and score near zero on administrative cases while LoRA adaptation succeeds at low cost.

· “LegalBench-BR: A Benchmark for Evaluating Large Language Models on Brazilian Legal Decision Classification”

open re-runnable review →
Figure from the paper

Evolutionary trees produce better contexts for creativity tests

Hierarchical planning plus niche optimization and simulated participants raise quality on six metrics over prior generators.

· “AlphaContext: An Evolutionary Tree-based Psychometric Context Generator for Creativity Assessment”

open re-runnable review →
Figure from the paper

Small models match LLMs on clarifying citizen consultations

Finetuned small language models reproduce manual clarifications of noisy texts and enable opinion clustering on a new 240k dataset.

· “The GDN-CC Dataset: Automatic Corpus Clarification for AI-enhanced Democratic Citizen Consultations”

open re-runnable review →

Tool merging and retrieval lifts LLM tool accuracy by up to 38%

Reducing overlapping tools and picking only relevant ones for each query improves selection on standard benchmarks.

· “ToolScope: Enhancing LLM Agent Tool Use through Tool Merging and Context-Aware Filtering”

open re-runnable review →

Competency estimates beat accuracy for ranking medical LLMs

Item Response Theory separates model ability from question difficulty, producing rankings that hold up better on external clinical tasks and

· “Measuring Competency, Not Performance: Item-Aware Evaluation Across Medical Benchmarks”

open re-runnable review →
Figure from the paper

Counterfactual training cuts segmentation hallucinations by 30%

New benchmark and fine-tuning method teach models when to produce a mask and when to abstain from objects that are not present.

· “Counterfactual Segmentation Reasoning: Diagnosing and Mitigating Pixel-Grounding Hallucination”

open re-runnable review →
Figure from the paper

Gemini 1.5 recalls details from 10 million tokens near perfectly

The model family processes hours of video and audio with text, advancing long-context tasks while matching prior top performance on standard

· “Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context”

open re-runnable review →

The paper introduces BEIR, a benchmark of 18 diverse public datasets spanning multiple…

Benchmark reveals re-ranking models top performance charts while dense retrievers lag despite lower compute demands.

· “BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval Models”

open re-runnable review →

browse all of cs.CL → full archive · search · sub-categories