The work proves that approximating correlation clustering to additive εn² error requires Ω(n/ε²) adjacency-matrix queries, with stronger bounds under memory constraints in random and general query models.
Tool reference
Lin, Divergence measures based on the Shannon entropy, IEEE Transactions on Information Theory 37 (1) (1991) 145--151
Tool reference. 78% of classified Pith citations use this work as a method, library, or software dependency, not as a substantive claim.
citation-role summary
citation-polarity summary
representative citing papers
Agent-ValueBench is the first dedicated benchmark for agent values, showing they diverge from LLM values, form a homogeneous 'Value Tide' across models, and bend under harnesses and skill steering.
UAV-CAS is a calibrated digital-twin dataset with 99k flows across 1,024 configurations for training and testing intrusion detection in UAV swarm networks.
Continuous language diffusion works by entering high-margin decoder basins where frozen T5 embeddings recover 93-96% of native decisions and linear readouts reach 97.9% agreement, implying models should be evaluated as representation-decoder systems.
The paper introduces a protocol-resolved framework for virological measurements, defining an observation operator that maps latent ensembles to observed data and recasting plaque assays as estimates of protocol-conditioned infectious concentration.
Statistical grammar induction shows the GROWING (bottom-up) maturational account of syntactic category acquisition significantly outperforms the INWARD account across three metrics under identical input and learning conditions.
LLM agents exhibit persistent attack-selection biases as fixed traits independent of success rates, with a bias momentum effect that resists steering and yields no performance gain.
SynHAT uses a novel two-stage spatio-temporal diffusion framework with Latent Spatio-Temporal U-Net to synthesize realistic human activity traces, outperforming baselines by 52% on spatial and 33% on temporal metrics across four cities.
NES systems in AI IDEs expand attack surfaces via context poisoning from imperceptible actions and global codebase retrieval, with professional developers largely unaware of the risks.
The paper justifies the composite coherence metric in event-based narrative extraction via an information-geometric decomposition on the product manifold and an axiomatic uniqueness proof for the geometric mean.
An LLM-assisted pipeline applied to 4,323 governance records finds that permissionless and corporate AI agent protocols show similar participation inequality and fragmentation but denser thematic alignment in the open setting.
LLMs recover dominant binomial orders from corpora but align less closely with exact preference distributions, with preference strength partially encoded in middle-to-late layers and manipulable via steering.
SGCD reshapes GRPO token advantages via detached sibling-contrast credit from an external LLM, improving AppWorld and airline tool-use scores while keeping policy gradient as the actor update.
APG4RecSim automatically generates realistic user profiles for LLM-based recommendation simulations, outperforming manual baselines by up to 7% in nDCG@10 and 8% in JSD on three benchmark datasets.
SHIELD is a new diverse clinical note dataset paired with distilled small language models that achieve 0.89 span-level precision and 0.88 recall for on-premise PHI de-identification.
Daily life resolves into roughly eight routine types with individuals maintaining stable, person-specific distributions and transitions that persist across weeks to months.
Seven cross-domain prompt-injection detectors are introduced; three are shipped and d028 raises F1 on paraphrased attacks from 0.033 to 0.378, while adaptive-attack support remains unevaluated.
Regularized Jensen-Shannon analogues of KL-based direct-correlation measures are constructed, bounded in [0,1], shown to satisfy the metric property, and their closed-form maxima under fixed alphabet sizes are derived.
PaRT achieves >50% tagging efficiency for boosted H->WW jets at 1% background efficiency, decorrelated from jet mass, with data-to-simulation scale factors of 0.9-1.0 on 138 fb^{-1} of 13 TeV collisions.
MDE computes three entropy features from flow stats to match conventional ML performance (F1 0.708-0.989) on four IDS benchmarks while exposing aggregate-metric failures and providing stable SHAP attributions.
NGC 6791 has an age of 8.46 ± 0.66 Gyr, [Fe/H] = +0.280 ± 0.079, and other parameters that favor an inner-Galaxy origin followed by outward migration.
Reward-only PPO can match hotel RevPAR while destroying Fixed-RM price discipline under hidden competitor state; a discipline-stability evaluation protocol and full-distribution Trace-Prior repair diagnose and partially fix the failure.
R2V-Agent combines an SLM policy trained via BC and DPO with a step-level risk-calibrated router using Brier scores and CVaR to escalate to LLM only on high residual failure risk, improving success-cost tradeoffs on HumanEval+, TextWorld, and TerminalBench.
An unsupervised system-aware framework combines online detection with an LLM-augmented contextual digital twin to deliver real-time, interpretable anomaly diagnosis in industrial control systems.
citing papers explorer
-
Query Lower Bounds for Correlation Clustering under Memory Constraints
The work proves that approximating correlation clustering to additive εn² error requires Ω(n/ε²) adjacency-matrix queries, with stronger bounds under memory constraints in random and general query models.
-
Agent-ValueBench: A Comprehensive Benchmark for Evaluating Agent Values
Agent-ValueBench is the first dedicated benchmark for agent values, showing they diverge from LLM values, form a homogeneous 'Value Tide' across models, and bend under harnesses and skill steering.
-
UAV-CAS: A Calibrated Digital-Twin Dataset for Intrusion Detection in UAV Swarm Networks
UAV-CAS is a calibrated digital-twin dataset with 99k flows across 1,024 configurations for training and testing intrusion detection in UAV swarm networks.
-
Continuous Language Diffusion as a Decoder-Interface Problem
Continuous language diffusion works by entering high-margin decoder basins where frozen T5 embeddings recover 93-96% of native decisions and linear readouts reach 97.9% agreement, implying models should be evaluated as representation-decoder systems.
-
Experimental Collapse in Virophysics: Protocol-Resolved Observation, Inference, and Plaque-Assay Blindness
The paper introduces a protocol-resolved framework for virological measurements, defining an observation operator that maps latent ensembles to observed data and recasting plaque assays as estimates of protocol-conditioned infectious concentration.
-
A Computational Operationalisation of Competing Maturational Theories of Syntactic Development via Statistical Grammar Induction
Statistical grammar induction shows the GROWING (bottom-up) maturational account of syntactic category acquisition significantly outperforms the INWARD account across three metrics under identical input and learning conditions.
-
CyBiasBench: Benchmarking Bias in LLM Agents for Cyber-Attack Scenarios
LLM agents exhibit persistent attack-selection biases as fixed traits independent of success rates, with a bias momentum effect that resists steering and yields no performance gain.
-
SynHAT: A Two-stage Coarse-to-Fine Diffusion Framework for Synthesizing Human Activity Traces
SynHAT uses a novel two-stage spatio-temporal diffusion framework with Latent Spatio-Temporal U-Net to synthesize realistic human activity traces, outperforming baselines by 52% on spatial and 33% on temporal metrics across four cities.
-
"Tab, Tab, Bug": Security Pitfalls of Next Edit Suggestions in AI-Integrated IDEs
NES systems in AI IDEs expand attack surfaces via context poisoning from imperceptible actions and global codebase retrieval, with professional developers largely unaware of the risks.
-
An Information-Geometric Justification for Composite Coherence in Event-Based Narrative Extraction
The paper justifies the composite coherence metric in event-based narrative extraction via an information-geometric decomposition on the product manifold and an axiomatic uniqueness proof for the geometric mean.
-
Agentic Analysis for Agentic Infrastructure: An LLM-Powered Pipeline for Comparative Governance of DAO and Corporate AI Protocols
An LLM-assisted pipeline applied to 4,323 governance records finds that permissionless and corporate AI agent protocols show similar participation inequality and fragmentation but denser thematic alignment in the open setting.
-
Behavioral and Representational Evidence of Binomial Ordering Preferences in Large Language Models
LLMs recover dominant binomial orders from corpora but align less closely with exact preference distributions, with preference strength partially encoded in middle-to-late layers and manipulable via steering.
-
Keep Policy Gradient in Charge: Sibling-Guided Credit Distillation for Long-Horizon Tool-Use Agents
SGCD reshapes GRPO token advantages via detached sibling-contrast credit from an external LLM, improving AppWorld and airline tool-use scores while keeping policy gradient as the actor update.
-
Task-Aware Automated User Profile Generation for Recommendation Simulation Using Large Language Models
APG4RecSim automatically generates realistic user profiles for LLM-based recommendation simulations, outperforming manual baselines by up to 7% in nDCG@10 and 8% in JSD on three benchmark datasets.
-
SHIELD: A Diverse Clinical Note Dataset and Distilled Small Language Models for Enterprise-Scale De-identification
SHIELD is a new diverse clinical note dataset paired with distilled small language models that achieve 0.89 span-level precision and 0.88 recall for on-premise PHI de-identification.
-
Quantifying the Persistence of Daily Routines
Daily life resolves into roughly eight routine types with individuals maintaining stable, person-specific distributions and transitions that persist across weeks to months.
-
Beyond Pattern Matching: Seven Cross-Domain Techniques for Prompt Injection Detection
Seven cross-domain prompt-injection detectors are introduced; three are shipped and d028 raises F1 on paraphrased attacks from 0.033 to 0.378, while adaptive-attack support remains unevaluated.
-
How to quantify direct correlations between variables
Regularized Jensen-Shannon analogues of KL-based direct-correlation measures are constructed, bounded in [0,1], shown to satisfy the metric property, and their closed-form maxima under fixed alphabet sizes are derived.
-
Particle transformers for identifying Lorentz-boosted Higgs bosons decaying to a pair of W bosons
PaRT achieves >50% tagging efficiency for boosted H->WW jets at 1% background efficiency, decorrelated from jet mass, with data-to-simulation scale factors of 0.9-1.0 on 138 fb^{-1} of 13 TeV collisions.
-
Multi-Level Distributional Entropy for Explainable Network Intrusion Detection
MDE computes three entropy features from flow stats to match conventional ML performance (F1 0.708-0.989) on four IDS benchmarks while exposing aggregate-metric failures and providing stable SHAP attributions.
-
The Absolute Age of the Open Cluster NGC 6791 and Its Implications for Galactic Archaeology and Asteroseismic Calibration
NGC 6791 has an age of 8.46 ± 0.66 Gyr, [Fe/H] = +0.280 ± 0.079, and other parameters that favor an inner-Galaxy origin followed by outward migration.
-
When Outcome Looks Right But Discipline Fails: Trace-Based Evaluation Under Hidden Competitor State
Reward-only PPO can match hotel RevPAR while destroying Fixed-RM price discipline under hidden competitor state; a discipline-stability evaluation protocol and full-distribution Trace-Prior repair diagnose and partially fix the failure.
-
R2V Agent: Teaching SLMs When to Ask for Help
R2V-Agent combines an SLM policy trained via BC and DPO with a step-level risk-calibrated router using Brier scores and CVaR to escalate to LLM only on high residual failure risk, improving success-cost tradeoffs on HumanEval+, TextWorld, and TerminalBench.
-
System-aware contextual digital twin for ICS anomaly diagnosis
An unsupervised system-aware framework combines online detection with an LLM-augmented contextual digital twin to deliver real-time, interpretable anomaly diagnosis in industrial control systems.
-
Can Semantic Methods Enhance Team Sports Tactics? A Methodology for Football with Broader Applications
A semantic vector-space model is proposed for encoding football tactics and team profiles to compute alignment and strategy recommendations using distance metrics.
-
Closed-loop Neuroprosthetic Control through Spared Neural Activity Enables Proportional Foot Movements after Spinal Cord Injury
Spared lower-limb EMG after SCI was decoded to drive closed-loop FES, yielding 33-40% gains in foot flexion range and proportional control up to six stimulation levels in a small cohort.
-
GWTC-5.0: Methods for Identifying and Characterizing Gravitational-wave Transients
Describes the methods for producing the fifth gravitational-wave transient catalog (GWTC-5.0) from O4b data of LIGO, Virgo and KAGRA.
-
Bayesian inference for compact binary coalescences with BILBY: Validation and application to the first LIGO--Virgo gravitational-wave transient catalogue
BILBY is validated on simulated compact binary signals and reproduces the eleven GWTC-1 results with configuration and output files provided for reproduction.
-
A Brief History of Fr\'echet Distances: From Curves and Probability Laws to FID
The paper delivers a chronological history of Fréchet distances connecting early abstract set theory to curve metrics, optimal transport, and the FID metric in generative models.
-
Is More Data Worth the Cost? Dataset Scaling Laws in a Tiny Attention-Only Decoder
A reduced attention-only decoder shows diminishing returns in dataset scaling, reaching 90% of full accuracy with only 30% of the data.
- Majorization-Guided Test-Time Adaptation for Vision-Language Models under Modality-Specific Shift