Introduces PowerPhase benchmark for massive-variate power-system forecasting and PowerForge model that achieves best average rank on safety-fidelity metrics across all tested grids.
hub Canonical reference
Masked image pretraining on language assisted representation
Canonical reference. 80% of citing Pith papers cite this work as background.
hub tools
citation-role summary
citation-polarity summary
representative citing papers
First integrated spiking controller combining bipedal locomotion and arm control on a full-scale humanoid via NEF, SPA, and basal ganglia, validated in Nengo-Isaac Sim co-simulation.
Introduces a scalable algebraic framework relating rank deficiency of generalized Vandermonde matrices for sparse steering vectors to thinned Toeplitz matrices and augmented full-ULA matrices to characterize and avoid multi-source ambiguities in thinned uniform linear arrays.
Brain-IT-VQA decodes visual question answers from fMRI using a transformer to extract language tokens and introduces the NSD-VQA benchmark with 20 controlled questions per image across 20 categories.
PlanAudio introduces a unified autoregressive LLM framework with semantic latent chain-of-thought for generating composite speech and sound audio from free-form text, plus a new benchmark.
Mind-ParaWorld creates parallel worlds with atomic facts to evaluate search agents on future scenarios, showing they synthesize evidence well but struggle with collection, coverage, sufficiency judgment, and stopping decisions.
Symbolic rational-function networks recover an admissible PDE from noiseless complete measurements and select the regularization-minimizing parameterization within the architecture.
DuFal combines global and local high-frequency Fourier neural operators with cross-attention fusion to recover fine anatomical structures in extremely sparse-view CBCT, outperforming prior methods on LUNA16 and ToothFairy data.
Diffusion models reconstruct high-resolution 3D cardiac ultrasound volumes from heavily undersampled elevation planes and outperform traditional interpolation and supervised deep learning baselines.
System-by-system autoregressive OMR with text-aware ABC transcription outperforms prior neural and rule-based systems and boosts VLM sheet-music QA.
Event-level MAE embeddings plus UMAP/HDBSCAN or K-Means clustering recover 15 hydroacoustic classes from multi-year Mayotte data with ~1 hour of annotation and detector-comparable F1.
Introduces Visibility-Aware Densification with Temporally-Adaptive Thresholding and Temporal Offset Warping to improve dynamic region quality in 3D Gaussian Splatting on three benchmarks.
Introduces the Perception algorithm for seed-guided semi-supervised clustering via a-contrario anomaly detection, defining clusters as anomaly-free subsets and achieving competitive performance with 10-30 seeds per cluster on benchmarks.
EntropyInfer adaptively allocates inference compute using per-head attention entropy for rigid/dynamic classification during prefilling and compresses KV cache with generated tokens, achieving up to 2.39x speedup on long contexts.
A contrastive learning transformer embeds network flow sequences to enable correlation clustering that groups scanner sources consistently with labels.
A survey of on-device learning in TinyML organized by distribution change regimes, highlighting influences on applications, hardware, and solutions plus a gap between benchmarks and deployments.
Shared-score quaternion self-attention reduces score multiplications by 75% and softmax operations from four to one while proving equivalence to component-wise attention under quaternion linear projections.
Rubato model with InterMo representation outperforms cascade methods in generating timestamped piano sheet music from audio, even when cascades receive ground-truth MIDI.
TextTeacher uses frozen text embeddings from captions as semantic anchors to guide vision model training, improving ImageNet accuracy by up to 2.7 p.p. and transfer performance by 1.0 p.p. on average.
Transcoda achieves state-of-the-art zero-shot OMR with an 18.46% OMR-NED error rate on synthetic scores and 63.97% on historical Polish scans using a 59M model trained in 6 hours via synthetic data, kern normalization, and grammar decoding.
Lexical acoustic coding lets LLMs transmit audio waveforms as editable natural-language sentences that another LLM can parse and reconstruct into sound.
SAGE uses sparse autoencoders to boost vulnerability signals in LLMs, raising internal SNR 12.7x and delivering up to 318% MCC gains on vulnerability detection benchmarks.
Sonata is a small hybrid world model pre-trained to predict future IMU states that outperforms autoregressive baselines on clinical discrimination, fall-risk prediction, and cross-cohort transfer while fitting on-device wearables.
LLM-Codec augments audio codec training with multi-step token prediction and contrastive semantic alignment to improve both waveform reconstruction and autoregressive predictability for speech language models.
citing papers explorer
-
Navigating the Safety-Fidelity Trade-off: Massive-Variate Time Series Forecasting for Power Systems via Probabilistic Scenarios
Introduces PowerPhase benchmark for massive-variate power-system forecasting and PowerForge model that achieves best average rank on safety-fidelity metrics across all tested grids.
-
A Spiking Neural Architecture for Coordinating Arm and Locomotor Control
First integrated spiking controller combining bipedal locomotion and arm control on a full-scale humanoid via NEF, SPA, and basal ganglia, validated in Nengo-Isaac Sim co-simulation.
-
Ambiguity Analysis and Design of Sparse Arrays via Generalized Vandermonde Rank Conditions
Introduces a scalable algebraic framework relating rank deficiency of generalized Vandermonde matrices for sparse steering vectors to thinned Toeplitz matrices and augmented full-ULA matrices to characterize and avoid multi-source ambiguities in thinned uniform linear arrays.
-
Brain-IT-VQA: From Brain Signals to Answers
Brain-IT-VQA decodes visual question answers from fMRI using a transformer to extract language tokens and introduces the NSD-VQA benchmark with 20 controlled questions per image across 20 categories.
-
Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts
PlanAudio introduces a unified autoregressive LLM framework with semantic latent chain-of-thought for generating composite speech and sound audio from free-form text, plus a new benchmark.
-
Evaluating the Search Agent in a Parallel World
Mind-ParaWorld creates parallel worlds with atomic facts to evaluate search agents on future scenarios, showing they synthesize evidence well but struggle with collection, coverage, sufficiency judgment, and stopping decisions.
-
Symbolic recovery of PDEs from measurement data
Symbolic rational-function networks recover an admissible PDE from noiseless complete measurements and select the regularization-minimizing parameterization within the architecture.
-
DuFal: Dual-Frequency-Aware Learning for High-Fidelity Extremely Sparse-view CBCT Reconstruction
DuFal combines global and local high-frequency Fourier neural operators with cross-attention fusion to recover fine anatomical structures in extremely sparse-view CBCT, outperforming prior methods on LUNA16 and ToothFairy data.
-
High Volume Rate 3D Ultrasound Reconstruction with Diffusion Models
Diffusion models reconstruct high-resolution 3D cardiac ultrasound volumes from heavily undersampled elevation planes and outperform traditional interpolation and supervised deep learning baselines.
-
LEGATO 2: Toward Multimodal Sheet Music Recognition and Understanding
System-by-system autoregressive OMR with text-aware ABC transcription outperforms prior neural and rule-based systems and boosts VLM sheet-music QA.
-
A Self-Supervised Approach for Minimal-Annotation Hydroacoustic Data Exploration
Event-level MAE embeddings plus UMAP/HDBSCAN or K-Means clustering recover 15 hydroacoustic classes from multi-year Mayotte data with ~1 hour of annotation and detector-comparable F1.
-
Temporally Aware Densification for Dynamic 3D Gaussian Splatting
Introduces Visibility-Aware Densification with Temporally-Adaptive Thresholding and Temporal Offset Warping to improve dynamic region quality in 3D Gaussian Splatting on three benchmarks.
-
Seed-Guided Semi-Supervised Clustering by A-Contrario Anomaly Detection
Introduces the Perception algorithm for seed-guided semi-supervised clustering via a-contrario anomaly detection, defining clusters as anomaly-free subsets and achieving competitive performance with 10-30 seeds per cluster on benchmarks.
-
From Rigid to Dynamic: Entropy-Guided Adaptive Inference for Long-Context LLMs
EntropyInfer adaptively allocates inference compute using per-head attention entropy for rigid/dynamic classification during prefilling and compresses KV cache with generated tokens, achieving up to 2.39x speedup on long contexts.
-
Contrastive Learning and Correlation Clustering for Sequences of Network Telescope Data
A contrastive learning transformer embeds network flow sequences to enable correlation clustering that groups scanner sources consistently with labels.
-
What changes after deployment? A survey on On-device Learning in TinyML
A survey of on-device learning in TinyML organized by distribution change regimes, highlighting influences on applications, hardware, and solutions plus a gap between benchmarks and deployments.
-
Quaternion Self-Attention with Shared Scores
Shared-score quaternion self-attention reduces score multiplications by 75% and softmax operations from four to one while proving equivalence to component-wise attention under quaternion linear projections.
-
Rubato: Transcribing Piano Music with Timestamps
Rubato model with InterMo representation outperforms cascade methods in generating timestamped piano sheet music from audio, even when cascades receive ground-truth MIDI.
-
TextTeacher: What Can Language Teach About Images?
TextTeacher uses frozen text embeddings from captions as semantic anchors to guide vision model training, improving ImageNet accuracy by up to 2.7 p.p. and transfer performance by 1.0 p.p. on average.
-
Transcoda: End-to-End Zero-Shot Optical Music Recognition via Data-Centric Synthetic Training
Transcoda achieves state-of-the-art zero-shot OMR with an 18.46% OMR-NED error rate on synthetic scores and 63.97% on historical Polish scans using a 59M model trained in 6 hours via synthetic data, kern normalization, and grammar decoding.
-
Communicating Sound Through Natural Language
Lexical acoustic coding lets LLMs transmit audio waveforms as editable natural-language sentences that another LLM can parse and reconstruct into sound.
-
SAGE: Signal-Amplified Guided Embeddings for LLM-based Vulnerability Detection
SAGE uses sparse autoencoders to boost vulnerability signals in LLMs, raising internal SNR 12.7x and delivering up to 318% MCC gains on vulnerability detection benchmarks.
-
Sonata: A Hybrid World Model for Inertial Kinematics under Clinical Data Scarcity
Sonata is a small hybrid world model pre-trained to predict future IMU states that outperforms autoregressive baselines on clinical discrimination, fall-risk prediction, and cross-cohort transfer while fitting on-device wearables.
-
LLM-Codec: Neural Audio Codec Meets Language Model Objectives
LLM-Codec augments audio codec training with multi-step token prediction and contrastive semantic alignment to improve both waveform reconstruction and autoregressive predictability for speech language models.
-
A Case Study on the Impact of Anonymization Along the RAG Pipeline
Anonymization placement in RAG—at the dataset or at the generated answer—creates observable differences in privacy protection versus response utility.
-
SyncBreaker:Stage-Aware Multimodal Adversarial Attacks on Audio-Driven Talking Head Generation
A multimodal adversarial attack using stage-sampled image nullification and cross-attention flattening degrades lip-sync and facial dynamics in Hallo-based talking-head generation.
-
MONETA: Multimodal Industry Classification through Geographic Information with Multi Agent Systems
MONETA is the first multimodal benchmark for industry classification using text and geographic sources, with MLLM baselines at 62-74% accuracy and up to 22.8% gains from multi-turn context enrichment and explanations.
-
Leveraging Artist Catalogs for Cold-Start Music Recommendation
ACARec attends over artist catalogs to generate CF embeddings for new tracks, more than doubling recall and NDCG versus content-only baselines in music recommendation.
-
TADA! Tuning Audio Diffusion Models through Activation Steering
Activation steering at a semantic bottleneck in audio diffusion models achieves state-of-the-art control over musical attributes such as instruments, vocals, and genres.
-
zea: A Toolbox for Cognitive Ultrasound Imaging
zea is a Python toolbox that supplies a modular differentiable pipeline for ultrasound imaging and signal processing, built on Keras 3 to support TensorFlow, PyTorch, and JAX backends.
-
SonicMaster: Towards Controllable All-in-One Music Restoration and Mastering
SonicMaster is a text-conditioned flow-matching generative model for unified music restoration and mastering, trained on a dataset of simulated degradations across equalization, dynamics, reverb, amplitude, and stereo.
-
Towards a General-Purpose Zero-Shot Synthetic Low-Light Image and Video Pipeline
A self-supervised Degradation Estimation Network estimates parameters for physics-informed noise distributions to generate realistic synthetic low-light data, showing gains on noise replication, enhancement, and detection tasks.
-
SamaVaani: Auditing and Debiasing Multilingual Clinical ASR for Indian Languages
Audit of multilingual clinical ASR reveals demographic biases; SamaVaani debiasing technique is proposed to jointly boost performance and fairness in Indian languages.
-
Profy: Interpretable Visualization of Expertise-Dependent Motor Skills Toward Supporting Piano Practice
Profy uses take-level expert-amateur labels on 1083 piano recordings to produce time-aligned highlight scores that correlate with expert review points (r=0.61) on held-out amateur clips.
-
Inside the Latent Flow: Causal Deciphering of Attention Dynamics in Audio Separation Foundation Models
Causal probing of attention in audio separation transformers identifies dual pathways and asynchronous convergence, enabling a training-free Layer-Selective Attention Caching method that reduces self-attention computation by ~25% with negligible quality loss.
-
Privacy-preserving Prosody Representation Learning
A self-supervised prosody encoder with speaker disentanglement strategies outperforms raw prosody and HuBERT baselines on pitch reconstruction and prosodic event detection while achieving strong speaker separation.
-
Decoding Stimulus Reconstruction-Based Auditory Attention Robustly in Unbalanced EEG Datasets
Stimulus reconstruction-based DNNs overestimate auditory attention decoding on unbalanced EEG datasets; LOPEO cross-validation prevents the inflation.
-
Articulatory movements influence electromagnetic wave transmission through the vocal tract
Articulatory configurations during vowel production create distinct electromagnetic transmission patterns through the vocal tract, confirmed by qualitative agreement between finite-element simulations and scattering-matrix measurements on two subjects.
-
Training-inference input alignment outweighs framework choice in longitudinal retinal image prediction
Training-inference input alignment outweighs framework choice for longitudinal retinal image prediction, with deterministic regression matching complex models when acquisition variability dominates disease progression.
-
SQuTR: A Robustness Benchmark for Spoken Query to Text Retrieval under Acoustic Noise
SQuTR is a large bilingual benchmark of 37,317 synthesized spoken queries under clean/low/medium/high noise, showing that retrieval quality steadily degrades as noise increases.
-
AI Models for Depressive Disorder Detection and Diagnosis: A Review
A systematic review of AI for depressive disorder detection that introduces a novel hierarchical taxonomy organized by clinical task, data modality, and model class.
-
Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models
The paper unifies perspectives on Long CoT in reasoning LLMs by introducing a taxonomy, detailing characteristics of deep reasoning and reflection, and discussing emergence phenomena and future directions.
-
Recursive QLSTM with Dynamic Variational Quantum Circuit Adaptation
The paper introduces Recursive QLSTM via metacore recursion, numerically tests variants on sequence lengths, and offers theoretical arguments for better temporal propagation.
-
Are you speaking my languages? On spoken language adherence in multimodal LLMs
Defines language adherence failures in multimodal ASR LLMs and compares soft prompting, SFT, and CoT strategies for reducing violations across languages.
-
Performance Analysis and Optimization of 3D Generative Diffusion Models across GPU Architectures
Profiling of Med-DDPM shows cuDNN kernels dominate training; TF32 Tensor Core activation and 3D channels-last layout reduce SM cycles up to 100x and raise Tensor Core utilization on A100 without quality loss.
-
Information-theoretic Multimodal Representation Learning for Electrocardiogram Signals
MERIT applies information theory to ECG representation learning via masked modeling and ECG-text contrastive alignment, reporting F1 gains over 3% on PTB-XL All and 5% on SubClass plus zero-shot and text generation improvements.
-
Afrispeech Semantics: Evaluating Audio Semantic Reasoning in Spoken Language Models Across Domains and Accents
Audio language models are benchmarked on five semantic and paralinguistic reasoning tasks to reveal limitations in handling spoken audio evidence, accent variation, and domain shifts.
-
MSDS: Deep Structural Similarity with Multiscale Representation
MSDS computes DeepSSIM at multiple pyramid scales and fuses the scores with learned weights, producing consistent improvements over single-scale DeepSSIM on IQA benchmarks with negligible extra cost.
-
SatBLIP: Context Understanding and Feature Identification from Satellite Imagery with Vision-Language Learning
SatBLIP fine-tunes a satellite-adapted BLIP model on GPT-4o-generated captions to predict county-level SVI from satellite tiles and uses SHAP to highlight key features like roof condition and vegetation.
-
LLM4Log: A Systematic Review of Large Language Model-based Log Analysis
Systematic review of 145 papers on LLM-based log analysis, providing a unified taxonomy, common design patterns, evaluation practices, and challenges for deployment under drift and limited labels.