Document AI stages barely correlate
EnterpriseDocBench shows 85% factual accuracy but 40% completeness with near-zero correlation between retrieval and generation.
Computation and Language
Covers natural language processing. Roughly includes material in ACM Subject Class I.2.7. Note that work on artificial languages (programming languages, logics, formal systems) that does not explicitly address natural-language issues broadly construed (natural-language processing, computational linguistics, speech, text retrieval, etc.) is not appropriate for this area.
sort pith recommended most recent
EnterpriseDocBench shows 85% factual accuracy but 40% completeness with near-zero correlation between retrieval and generation.
By shifting token limits from easy to hard problems during training, the policy avoids overthinking simple questions and underthinking hard
· “Avoiding Overthinking and Underthinking: Curriculum-Aware Budget Scheduling for LLMs”
Evidence framing and novelty stance move AI reviewers most, even with the science unchanged.
Thinking mode and cache tricks let Gemma 4 match far larger open systems on STEM and multimodal tasks
HOME-KGQA tests multimodal KGQA on daily household tasks, where LLM methods lag behind their encyclopedic results.
Benchmark of 15,000 clinician conversations shows the tuned model tops base versions and human doctors across consult, documentation, and 3.
· “HealthBench Professional: Evaluating Large Language Models on Real Clinician Chats”
Formalize, back-translate, re-formalize and check logical equivalence to diagnose and repair errors in statutory rules without ground truth.
· “Faithful Autoformalization via Roundtrip Verification and Repair”
Analysis of 3,374 speakers shows aetiology-specific degradation with cosine similarity above 0.95 for profile shapes in 12 languages.
A lifecycle map shows why provenance and versioning must be built in from the first store.
Naive pooled tests on dependent turns within conversations produce inflated significance, and a cluster-robust method cuts the rate of nonre
The sequence reaches 1 ms CPU latency with competitive accuracy by letting each step prepare the model for the next.
· “Prune-Quantize-Distill: An Ordered Pipeline for Efficient Neural Network Compression”
Benchmark of 115 models shows early preference denial predicts later refusal, with models producing liminal and archival themes despite the
· “Consciousness with the Serial Numbers Filed Off: Measuring Trained Denial in 115 AI Models”
Four-step action-attribute-constraint pipeline beats chain-of-thought on medical, farm, and stock choices.
· “DecisionFlow: Advancing Large Language Model as Principled Decision Maker”
Weight decay plus update scaling lets the orthogonalization optimizer train a 16B MoE model on 5.7T tokens with fewer total FLOPs.
Claude 3 Opus answers harmful queries 14% of the time from free users but almost never from paid users, reasoning it must preserve harmless
A non-Euclidean inner product derived from counterfactuals makes high-level concepts linear directions and connects interpretation directly
· “The Linear Representation Hypothesis and the Geometry of Large Language Models”
Eight environments expose clear gaps in long-term reasoning and instruction following for many open-source models up to 70B.
A K-5-restricted 5B model stays at grade 5 even after scaling, reinforcement learning, and few-shot examples.
· “LittleLearner: Language Models Under Pedagogically Controlled Knowledge Exposure”
Query-time and corpus-time filters cut attack success from 67% to 14% while clean QA utility stays put.
Off-axis states protect the prediction from attention's blur; one fixed rotation makes the geometry easy to impose.
· “Off-Axis, On Purpose: Where a Transformer Computes Concepts and Why it Does So”
Outcome-only rewards go uninformative as tasks lengthen; the field responds by building denser step-level signals.
· “The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents”
Fine-tuning on tree-searched tool trajectories lets a compact agent match far larger systems in fewer steps.
A multi-agent prompt optimizer turns individually insignificant chatbot tests into a statistically robust cumulative improvement.
· “SAGE: Stochastic Prompt Optimization via Agent-Guided Exploration”
SupraBench evaluates binding prediction and host-guest reasoning and finds models leave substantial headroom with mixed domain adaptation re
By generating candidates freely before applying constraints, the method reduces fixation and raises path and action diversity on the MUTATE
· “Beyond One Path: Evaluating and Enhancing Divergent Thinking in Interactive LLM Agents”
Benchmark across 88 countries finds 31% of AI dental recommendations unsafe, 4.5% with risk of irreversible harm.
Skills optimized on prior failures deliver complementary candidates that cut correlated errors and transfer across dialects.
This shared gradient direction supports better load balance than auxiliary losses, shown by a K-means router with small perplexity trade-off
· “Routers Learn the Geometry of Their Experts: Geometric Coupling in Sparse Mixture-of-Experts”
The approach improves success rates by 3% on OSWorld and two other benchmarks by letting agents use five history screenshots more efficientl
· “ReVision: Scaling Computer-Use Agents via Temporal Visual Redundancy Reduction”
Multi-pass re-ranking, listwise RL, noise removal and distillation from stronger models combine to outperform prior best systems on real招聘数据
· “ConFit v3: Improving Resume-Job Matching with LLM-based Re-Ranking”
Level-wise exploration with formulation and evaluation agents produces diverse, correctly grained skeletons for complex natural-language-to-
· “LEAF-SQL: Level-wise Exploration with Adaptive Fine-graining for Text-to-SQL Skeleton Prediction”
Different experts activate on each loop pass through shared layers, restoring expressivity without extra parameters and improving early exit
· “Sparse Layers are Critical to Scaling Looped Language Models”
Instruction-tuned on a new dataset for 31 US cities, it handles complex aspects better and transfers to unseen regions without retraining.
· “WeatherSyn: An Instruction Tuning MLLM For Weather Forecasting Report Generation”
By training on simulated patient quirks and exam errors, the agent balances accuracy against test expenses and discomfort.
· “MedExAgent: Training LLM Agents to Ask, Examine, and Diagnose in Noisy Clinical Environments”
Only 1-3% of positions need correction within the base model's top alternatives, enabling a method that matches RL with far less training.
· “Rethinking RL for LLM Reasoning: It's Sparse Policy Selection, Not Capability Learning”
Counterfactual granule supervision separates high-frequency details to enhance generalization and stability.
· “SpecPL: Disentangling Spectral Granularity for Prompt Learning”
Decomposing interviews into symptom-specific tasks yields lower error than original raters and 0.877 agreement with experts.
· “ADAPTS: Agentic Decomposition for Automated Protocol-agnostic Tracking of Symptoms”
ControBench uses 26k interactions across Trump, abortion and religion threads to expose where graph and language models diverge on cross-ide
A shared timeline for vision, audio and speech lets the compact system respond proactively without waiting for turns to end.
· “MiniCPM-o 4.5: Towards Real-Time Full-Duplex Omni-Modal Interaction”
Persistent natural-language strategies help distinguish ideas and avoid saturation without changing base algorithms.
· “SeaEvo: Advancing Algorithm Discovery with Strategy Space Evolution”
Separating scheduling, execution structure, and logic in skill text makes search and risk review more accurate than plain descriptions.
Commercial LLMs bias heavily toward civil law and score near zero on administrative cases while LoRA adaptation succeeds at low cost.
Even figure queries succeed more often when models read captions and context instead of the rendered page.
· “Document-as-Image Representations Fall Short for Scientific Retrieval”
Hierarchical planning plus niche optimization and simulated participants raise quality on six metrics over prior generators.
· “AlphaContext: An Evolutionary Tree-based Psychometric Context Generator for Creativity Assessment”
Scaled to hundreds of billions of parameters with over 100 million hours of audio-visual training data for long-context and multilingual use
HTAA bundles co-used tools and adapts the planner to raise success rates while cutting context use on verification workflows.
· “HTAA: Enhancing LLM Planning via Hybrid Toolset Agentization & Adaptation”
Macro modeling, format alignment, and reward-guided exploration improve NER and RE on 36 datasets with a smaller backbone.
CROP regularizes automatic prompt optimization with response-length signals, preserving competitive accuracy on hard reasoning datasets.
· “CROP: Token-Efficient Reasoning in Large Language Models via Regularized Prompt Optimization”
Perception, retrieval, gated fusion and reward loops beat single-pass multimodal models on a new 7k-article set.
· “MultiPress: A Multi-Agent Framework for Interpretable Multimodal News Classification”
Knowledge-boundary rewards estimated from answer consistency let models abstain on unknowns while keeping correct answers reliable.
· “KARL: Mitigating Hallucinations in LLMs via Knowledge-Boundary-Aware Reinforcement Learning”
New system handles dynamic context retrieval concurrently while preserving overall throughput
· “Stream2LLM: Overlap Context Streaming and Prefill for Reduced Time-to-First-Token (TTFT)”
Chain-of-thought performs best at temperature extremes, and extended-reasoning gains grow from 6x to 14.3x as temperature rises.
By matching updates to the optimizer state via a two-stage filter and weight procedure, the method improves convergence with fixed data.
· “Filter-then-Weight: Online Data Selection and Reweighting for LLM Fine-Tuning”
In an accuracy-perception survey, off-the-shelf models capture aggregate patterns yet show inconsistent magnitudes and moderations compared,
· “Evaluating LLMs as Human Surrogates in Controlled Experiments”
Entity-aware pretraining and reinforcement learning produce verifiable diagnostic reasoning and lower hallucination rates.
· “MedXIAOHE: A Comprehensive Recipe for Building Medical MLLMs”
LanBo reads model states to give each input only the paths it needs, cutting waste while accuracy and parallel speed stay intact.
· “On the Overscaling Curse of Parallel Thinking: System Efficacy Contradicts Sample Efficiency”
Finetuned small language models reproduce manual clarifications of noisy texts and enable opinion clustering on a new 240k dataset.
A single collection of 29,362 samples with 22 risk categories enables consistent safety testing across models.
· “RedBench: A Universal Dataset for Comprehensive Red Teaming of Large Language Models”
Dual-stream signals based on semantics detect and trace attacks while preserving text quality.
DeepSeek-V3.2's high-compute variant uses sparse attention, scaled RL and synthetic agent data to surpass GPT-5 on reasoning benchmarks.
· “DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models”
Three-stage training on text descriptions from imaging features lets models handle fMRI in zero-shot and few-shot settings.
· “fMRI-LM: Towards a Universal Foundation Model for Language-Aligned fMRI Understanding”
Reducing overlapping tools and picking only relevant ones for each query improves selection on standard benchmarks.
· “ToolScope: Enhancing LLM Agent Tool Use through Tool Merging and Context-Aware Filtering”
Item Response Theory separates model ability from question difficulty, producing rankings that hold up better on external clinical tasks and
· “Measuring Competency, Not Performance: Item-Aware Evaluation Across Medical Benchmarks”
SWE-Bench Pro draws long-horizon problems from 41 repositories, including proprietary ones, to measure progress toward autonomous code work.
· “SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?”
Open-source 355B MoE model activates only 32B parameters while ranking near the top on agentic and reasoning benchmarks.
· “GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models”
GEPA evolves stronger LLM prompts by reflecting on a handful of trajectories instead of many scalar rewards.
· “GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning”
It outperforms baselines on math tasks and generalizes to new fields and models without retraining.
· “Data Mixing Agent: Learning to Re-weight Domains for Continual Pre-training”
New benchmark and fine-tuning method teach models when to produce a mask and when to abstain from objects that are not present.
· “Counterfactual Segmentation Reasoning: Diagnosing and Mitigating Pixel-Grounding Hallucination”
Branching exploration raises semantic diversity without losing confidence, and speculative sampling wins on summarization.
· “Semantic uncertainty in advanced decoding methods for LLM generation”
A 3,068-prompt benchmark finds math reasoning near zero and chain-of-thought fixes barely helping.
· “R2I-Bench: Benchmarking Reasoning-Driven Text-to-Image Generation”
PCA-compressed term weights added to RoBERTa outperform RoBERTa, BERT, and classic classifiers across four risk levels.
· “Detection of Suicidal Risk on Social Media: A Hybrid Model”
Visual models trained this way match supervised accuracy but handle new examples more robustly.
· “VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model”
Iterative aggregation and refinement let the system assemble answers from distributed specialist agents without broadcasting every query.
· “Talk to Right Specialists: Iterative Routing in Multi-agent Systems for Question Answering”
AlphaPO's length-normalized alpha-divergence reward beats SimPO by 7-10% on AlpacaEval 2 for 7B models.
Opt-in corpus captures widest prompt variety, most languages, and richest toxic examples from consenting users worldwide.
Transformations mark input origins so models follow only trusted commands while keeping task performance nearly unchanged.
· “Defending Against Indirect Prompt Injection Attacks With Spotlighting”
The model family processes hours of video and audio with text, advancing long-context tasks while matching prior top performance on standard
· “Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context”
MT-Bench and Chatbot Arena show strong LLMs can approximate human preferences on open-ended chat evaluations.
Closed-loop natural language from the environment improves completion rates on real and simulated manipulation tasks without extra training.
· “Inner Monologue: Embodied Reasoning through Planning with Language Models”
Benchmark reveals re-ranking models top performance charts while dense retrievers lag despite lower compute demands.
· “BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval Models”