Aim the slop cannon: structure beats fluent wrongness
Agents, selective memory, and adversarial step-checks make LLM science inspectable
Artificial Intelligence
Covers all areas of AI except Vision, Robotics, Machine Learning, Multiagent Systems, and Computation and Language (Natural Language Processing), which have separate subject areas. In particular, includes Expert Systems, Theorem Proving (although this may overlap with Logic in Computer Science), Knowledge Representation, Planning, and Uncertainty in AI. Roughly includes material in ACM Subject Classes I.2.0, I.2.1, I.2.3, I.2.4, I.2.8, and I.2.11.
sort pith recommended most recent
Agents, selective memory, and adversarial step-checks make LLM science inspectable
Fine-tuning and sampling share one iteration with global descent and quadratic convergence.
· “Newton Matching for Generative Modeling: A Unified Framework for Fine-Tuning and Sampling”
Each generated token leaks at most bT of mutual information about the private secret; unanimous votes add no noise.
· “PAC-Private Autoregressive Generation: Calibrating Noise to Ensemble Disagreement”
Training a student model on one or two tokens per response matches or beats dense supervision in nine configurations.
· “Extremely Sparse Supervision Incentivizes Reasoning Ability”
Bilevel game theory shows only grounded verifiers converge to zero risk; the gated system resolves 72.2% on SWE-bench.
· “Bilevel Coordinated Reflection: A Game-Theoretic Approach to Multi-Agent LLM Systems”
Wait when uncertain, emit when ready: latency drops from hundreds of milliseconds to tens without losing accuracy.
· “X2Streaming-ASR: wait when uncertain, emit when ready for streaming ASR”
For sparse data the full nonlinear dynamics reduces to independent scalar flows with explicit convergence times.
A theorem uses Householder reflections to show a frustrated zero-field O(n) chain equals a simpler field-driven chain for any length and…
Under ETH for PPAD, polynomial CCE computation is impossible; first radically uncoupled algorithm matches the hardness bound.
· “Independent Reinforcement Learning in Discounted Markov Games”
A heads-up no-limit turn subgame with 83k info sets solves at 0.397ms per iteration, 14–258× faster than the fastest CPU solver.
Entropic optimal transport yields a well-defined, estimable shift; new bound unifies both shifts.
Even the specialist loses 18 points on field images, so blur hurts everyone, not just general models.
· “Can Edge-Deployable Vision-Language Models Identify Species?”
GMMM estimates the causal effect of GEO and GEM by converting generated-answer occurrences into marketing inputs.
A minimal experiment shows adaptive control arising without any task objective, reshaping alignment for persistent agents.
· “Artificial Id: Drive and Persistent Alignment in Agentic AI”
MindTopo benchmark tests 14 models on five topological properties; every model reasons better than it plans, revealing a fundamental gap.
· “MindTopo: Can Foundation Models Reason in Topological Space?”
A detector-guided preference-optimization step also cuts a generator's measured hallucinations by over half.
· “Domain-Specific Hallucination Detection in Large Language Models”
A learned acquisition policy combined with LLM priors achieves 5.67-fold enrichment over random on held-out screens.
· “Biology-in-the-loop: Amortized Adaptive Hit Discovery in CRISPR Screens”
A unified framework shows that all state-of-the-art linear recommenders differ only in regularization type, enabling simpler scalable…
· “On the Regularization Landscape for the Linear Recommendation Models”
Headroom-Closed Index shows interactive tasks lag far behind knowledge; only L5 systems close the improvement loop on improvement itself.
· “The Last AI Built by Humans: Toward Genuine Recursive Self-Improvement”
Trained to verify every reasoning step, the model fixes its own errors mid-flight — at no extra latency.
· “RetroThinker: Enabling Retrospective Thinking in Speech LLMs”
It improves on the grammar-based baseline by 17 points and needs no fine-tuning for new model types.
Layerwise interventions show a functional handoff where causal control shifts from query routing to answer content across depth
· “From Parameters to Answers: How LLMs Retrieve and Use Their Internal Knowledge”
Combining risk from a baseline checkpoint with kinetic action yields optimal time allocation; a frozen template recovers most gains.
· “Model-Aware Schedules Improve Generation via Fiberwise Optimal Transport”
Survey of maritime professionals reveals conditional acceptance and strong concerns about overreliance and skill loss.
· “Understanding Operator Attitudes Toward AI-Supported Decision Making in Maritime Operations”
A small autoregressive refiner restores spatial coherence and beats a model twice its size.
· “Logit Refiner: Improving Visual Autoregressive Models via Intra-Scale Dependency Modeling”
By training a stateful denoiser on temporally aligned denoising tasks, the model stabilizes recurrence and scales with inference steps.
Switch-localized metrics reveal a system's true ability to recognize Yoruba at language boundaries, hidden by aggregate WER.
Controlled test across three model families shows framing reversal under 7% even when correctly detected.
· “Recognizing Is Not Reversing: A Controlled Inversion Test of Fact-Preserving News Framing”
Multi-coefficient per-token KL mixing on TweetEval shows directional gains, but wide uncertainty across three seeds.
SIRF embeds content rules into LLM weights via 70M synthetic tokens, achieving 71% recall at 95% precision under <1s latency.
· “SIRF: A Spec-Internalized Risk Foundation Model for Industrial Content Risk Control”
LOCUS selects a task-aware LoRA subspace to cut tokens without changing the alignment loss or accuracy.
· “LOCUS: Task-Aware Low-Rank Post-Training for Token-Efficient Language Generation”
New framework ORCH outperforms four prior multi-agent methods across 25 wildfire missions and eight language models.
· “ORCH: Organizational Principles Enable Collective Intelligence in Embodied AI”
A single-layer neural CDE hit 0.53 correlation to reference emotion intensity; the tuned baseline scored 0.08.
· “Continuous-Time Acoustic Modelling with Neural Controlled Differential Equations”
Voltage-to-time converter avoids bulky current-scaling circuits, enabling compact analog neurons for memristive inference
· “A Time-Based Readout for Vector-Matrix Multiplication in Fully Analog Memristive SNNs”
Joint forward-reverse fusion consistently beats voting, electoral rules, and LLM judges on a medical diagnosis benchmark.
Language-driven semantic reasoning guides geometric kernels toward design-intent surfaces without solver changes.
· “Language-Augmented Semantic Priors for B-Spline Surface Fitting”
A differentiable safeguard aligns training and inference, preserving task success while enforcing hard constraints.
One scorer allocates costly executions and trims the weak, giving stronger agents on 50 examples per task.
· “COBRA-Skills: Contextual Bandit-Guided Evolution for Agent Skill Optimization”
Prioritizing cross-task harness flaws lifts reasoning accuracy 18.56% and aids transfer to new models.
· “Ecdysis: Efficient and Effective Training of Runtime Harnesses for LLM Agents”
7,654 policy-relevant datasets reachable through place-based search, even with messy metadata
· “Geospatial AI, Dataverse Metadata, and the Study of Place-Based Government”
Warrant theory redefines logical consequence and failure through the presence or absence of warrant, shifting focus from truth to…
A developmental framework: agents internalise norms through embodied, motivated, staged interaction.
At 0.80 kbps and 160 ms latency, it matches or exceeds non-streaming and streaming codecs in reconstruction and downstream tasks.
Conditional generation of year-long sub-hourly load curves matches real data in forecasting and appliance detection, enabling safe energy…
· “LoaDiff: Conditional Generation of Electricity Consumption Time Series for Energy Analytics”
Keeping its program, plans, and search state, it beats baselines that regenerate code each turn.
· “MAPLE: Memory-Augmented Planning with Language and Evolution”
By learning only state trajectories instead of raw actions, model-based RL cuts overflow and training cycles on a real testbed.
Outperforms OLS, gradient-boosted trees, and analyst consensus across four commercial alternative data channels.
· “Making Alternative Data Work: Context-Augmented LLMs for Financial Forecasting”
Merging faces by underlying surface and using the solid's own coordinate frame eliminates catastrophic failures from repatitioning…
ShEx schemas + few-shot queries let LLMs write SPARQL; a sampling trick cuts metadata runtime 80×.
· “Enabling Knowledge Graph Understanding at Scale with the EXplore Your Graphs ENgine (EXYGEN)”
ResNet50 is most stable, but histopathology images reveal performance gaps invisible in dermoscopy alone.
· “A Comparative Evaluation of Pre-trained Convolutional Neural Networks for Melanoma Detection”
By measuring job-level power elasticity with a new metric, data centers can cut power without crippling model training.
· “Characterizing Job Power Elasticity for Power-Flexible AI Training”
Controlled experiment isolates the text-to-text preprocessing step as a causal source of bias, not just the image model.
· “Prompt Revision as a Source of Cultural Bias in Text-to-Image Systems”
A Random Forest pipeline with seven geometric features outperforms GPU-dependent systems for Formula Student racing teams.
InterIL generates image and layout in one pass using a learnable communication module, beating prior methods on harmony and human…
Resolving on original non-ground clauses yields reusable learned clauses and potentially exponentially shorter proofs.
Pre-pretraining on music, grammars, and cellular automata accelerates next-token prediction but fails to improve downstream benchmarks.
Compressing the full hidden trajectory into a fixed tensor matches dense detectors with 67 times less storage and no extra LLM calls.
· “ActMap: Single-Pass Uncertainty Quantification from Generation-Time Activation Maps”
Dual graph architecture beats vector-RAG by 19% on free-form answer tasks in 38-report Sanofi test set.
Refitting only the statistics on retained data shifts 47 of 221 checkpoints and can reverse the pass/fail verdict.
Exactly computable in Hanabi, the metric separates human and AI play and suggests AI compatibility with humans, not other AIs, matters most.
· “The Convention Gap: Towards Measuring Implicit Communication in Cooperative AI Evaluation”
Optimal transport on articulatory features matches black-box accuracy while identifying specific phonetic differences like rhoticity
The methodology also improves ACO priors over hand-designed baselines and transfers to larger CVRP instances.
Strict F1 improves by 0.06–0.15 across six languages without task-specific training or candidate ranking.
· “Cross-Lingual Clinical Annotation Projection as Constrained Text Generation: A Six-Language Study”
End-to-end measurement shows naive cross-pool precision transfer errs by 422%.
· “Prevalence Determines Precision:Silent Contamination in Detector-Defined Datasets”
A study on FSD50K and AudioSet shows that complex feature-level continual learning methods are unnecessary for in-domain audio…
· “Investigating catastrophic forgetting in sound event classification”
A single shared threshold with provable reliability lets cheap models handle easy cases without retraining or per-model tuning.
· “Calibration-Aware Uncertainty Cascades for Efficient Heterogeneous Model Collaboration”
Clinicians favored comparative model rankings over term-by-term readings in a body-fat symbolic regression case.
New method selects the best LLM per query while avoiding context loss and confusion, outperforming single-model baselines.
· “SWRouter: Similarity-Contractive Window Routing for Multi-Turn Large Language Model Conversations”
X-AuT reduces audio-encoder depth from 18 to 14 layers while keeping word error rate within 0.14 points of baseline on ten public…
· “X-AuT: Progressive Audio-Encoder Compression for Speech LLMs with Cross-Scale Distillation”
Prompting callers to hum, sing, or puff cheeks spots impersonation far better than artifact scanning.
· “Deep-Fake CAPTCHA: Mitigating Next-Generation Social Engineering Attacks”
Data stories embed executable queries in prose, turning exploration of cultural-heritage knowledge graphs into quality assessment.
TASCO optimizes for confidence that stays robust under small perturbations, boosting accuracy by up to 17% while using fewer tokens.
· “Beyond Confidence: Stability-Aware Test-Time Adaptation for LLM Reasoning”
A multi‑country study finds that when buyers use AI to monitor environmental practices, suppliers have fewer media‑reported…
A directory-aware semantic storage and trace reuse system cuts token consumption dramatically while preserving high accuracy on structured…
· “VikingRAG: Accurate and Token-efficient Retrieval-augmented Generation over Structured Documents”
A new software framework binds task-level goals to application effects through contracts that survive revisions, handovers, and policy…
· “Agent-Integrated Software: Interaction Contracts and Continuous Assurance”
First independent analysis of transparent moderation system reveals 78% of harmful content escapes default service.
Outperforms unimodal baselines by up to 5.76% on AVGC and achieves 88.42% on DTU.
Six societies share semantics but not raw coordinates; an inherited global interface severely harms new learning.
Training on a single small knowledge graph for 30 minutes transfers to 40 unseen benchmarks.
· “Reification as a Transferable Vocabulary: Zero-Shot Link Prediction with Vanilla GNNs”
CoMA-DiT treats paired biosignals as mutual supervisors to improve attention and emotion decoding.
· “Exploring Diffusion Transformers for Cross-Modal Augmentation in Multimodal Brain State Decoding”