Introduces NeuroDoc and NeuroAudit to create a community-reviewed corpus of 53 EEG benchmark entries with 245 task definitions using a rulebook-guided task document and executable kernel.
InProceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2(Toronto ON, Canada)(KDD ’25)
12 Pith papers cite this work, alongside 20 external citations. Polarity classification is still indexing.
years
2026 12representative citing papers
REST-TS resolves text collapse in multimodal time series forecasting by exclusively supervising the text branch on numerical residuals to compel genuine content extraction from text descriptions.
SAGE decomposes univariate time-series anomaly detection into four specialized LLM analyzers plus an evidence-grounded detector and supervisor, achieving the highest average performance on three benchmarks while using only normal data for in-context examples.
CRAFT benchmark shows multi-agent coordination under partial information remains unsolved for current LLMs, with smaller open-weight models often matching or beating frontier systems.
SQLConductor uses Search-to-Policy Learning with MCTS, stability-weighted SFT, and curriculum RL to train a compact policy for adaptive step-wise Text-to-SQL orchestration, reporting 73.2% EX on BIRD-Dev.
DREAM proposes intent-aware tokenization, frozen-model evaluation, and dynamic beams to refine early SID assignments and improve cold-start performance in generative recommenders on Amazon benchmarks.
ORCA is an agent-orchestrated interactive copilot that automates and guides end-to-end causal analysis from workflow selection to report generation across real-world use cases.
RankElastor mitigates embedding collapse via spectrum-robust token mixing and GLU-based P-FFNs, yielding better performance and scaling on industrial recommendation datasets.
Long-Term Embeddings anchor sequential recommendation models to fixed content-based item representations to capture stable preferences and ensure version compatibility, resulting in uplifts in user engagement and financial metrics.
A code-owned harness enforces source, routing, trace, hygiene, and recommendation contracts for enterprise LLM agents; prompt-only fails and bolt-on guardrails over-refuse.
AOEPT proposes modal-contextualized prompts that distill global modality priors to restore reasoning scope in multimodal transformers under missing-modality conditions.
Agentic-FL introduces language model agents for autonomous orchestration in federated learning to address client heterogeneity and dynamic conditions.
citing papers explorer
-
EEG Benchmarking Needs a Task Specification Layer: NeuroDoc for Rulebook-Guided, Executable Benchmark Construction
Introduces NeuroDoc and NeuroAudit to create a community-reviewed corpus of 53 EEG benchmark entries with 245 task definitions using a rulebook-guided task document and executable kernel.
-
Does Text Actually Help? Uncovering and Resolving Text Collapse in Multimodal Time Series Forecasting
REST-TS resolves text collapse in multimodal time series forecasting by exclusively supervising the text branch on numerical residuals to compel genuine content extraction from text descriptions.
-
Detecting Time Series Anomalies Like an Expert: A Multi-Agent LLM Framework with Specialized Analyzers
SAGE decomposes univariate time-series anomaly detection into four specialized LLM analyzers plus an evidence-grounded detector and supervisor, achieving the highest average performance on three benchmarks while using only normal data for in-context examples.
-
CRAFT: Grounded Multi-Agent Coordination Under Partial Information
CRAFT benchmark shows multi-agent coordination under partial information remains unsolved for current LLMs, with smaller open-weight models often matching or beating frontier systems.
-
SQLConductor: Search-to-Policy Learning for Step-wise Text-to-SQL Orchestration
SQLConductor uses Search-to-Policy Learning with MCTS, stability-weighted SFT, and curriculum RL to train a compact policy for adaptive step-wise Text-to-SQL orchestration, reporting 73.2% EX on BIRD-Dev.
-
DREAM: Dynamic Refinement of Early Assignment Mappings
DREAM proposes intent-aware tokenization, frozen-model evaluation, and dynamic beams to refine early SID assignments and improve cold-start performance in generative recommenders on Amazon benchmarks.
-
ORCA: An End-to-End Interactive Copilot for Optimized Root Cause Analysis
ORCA is an agent-orchestrated interactive copilot that automates and guides end-to-end causal analysis from workflow selection to report generation across real-world use cases.
-
Expand More, Shrink Less: Shaping Effective-Rank Dynamics for Dense Scaling in Recommendation
RankElastor mitigates embedding collapse via spectrum-robust token mixing and GLU-based P-FFNs, yielding better performance and scaling on industrial recommendation datasets.
-
Long-Term Embeddings for Balanced Personalization
Long-Term Embeddings anchor sequential recommendation models to fixed content-based item representations to capture stable preferences and ensure version compatibility, resulting in uplifts in user engagement and financial metrics.
-
From Prompts to Contracts: Harness Engineering for Auditable Enterprise LLM Agents
A code-owned harness enforces source, routing, trace, hygiene, and recommendation contracts for enterprise LLM agents; prompt-only fails and bolt-on guardrails over-refuse.
-
AOEPT: Breaking the Implicit Modality-Reduction Bottleneck in Modality-Missing Prompt Tuning
AOEPT proposes modal-contextualized prompts that distill global modality priors to restore reasoning scope in multimodal transformers under missing-modality conditions.
-
Agentic Federated Learning: The Future of Distributed Training Orchestration
Agentic-FL introduces language model agents for autonomous orchestration in federated learning to address client heterogeneity and dynamic conditions.