Introduces PowerPhase benchmark for massive-variate power-system forecasting and PowerForge model that achieves best average rank on safety-fidelity metrics across all tested grids.
hub Canonical reference
Masked image pretraining on language assisted representation
Canonical reference. 80% of citing Pith papers cite this work as background.
hub tools
citation-role summary
citation-polarity summary
representative citing papers
First integrated spiking controller combining bipedal locomotion and arm control on a full-scale humanoid via NEF, SPA, and basal ganglia, validated in Nengo-Isaac Sim co-simulation.
Introduces a scalable algebraic framework relating rank deficiency of generalized Vandermonde matrices for sparse steering vectors to thinned Toeplitz matrices and augmented full-ULA matrices to characterize and avoid multi-source ambiguities in thinned uniform linear arrays.
Brain-IT-VQA decodes visual question answers from fMRI using a transformer to extract language tokens and introduces the NSD-VQA benchmark with 20 controlled questions per image across 20 categories.
PlanAudio introduces a unified autoregressive LLM framework with semantic latent chain-of-thought for generating composite speech and sound audio from free-form text, plus a new benchmark.
Mind-ParaWorld creates parallel worlds with atomic facts to evaluate search agents on future scenarios, showing they synthesize evidence well but struggle with collection, coverage, sufficiency judgment, and stopping decisions.
Symbolic rational-function networks recover an admissible PDE from noiseless complete measurements and select the regularization-minimizing parameterization within the architecture.
DuFal combines global and local high-frequency Fourier neural operators with cross-attention fusion to recover fine anatomical structures in extremely sparse-view CBCT, outperforming prior methods on LUNA16 and ToothFairy data.
Diffusion models reconstruct high-resolution 3D cardiac ultrasound volumes from heavily undersampled elevation planes and outperform traditional interpolation and supervised deep learning baselines.
System-by-system autoregressive OMR with text-aware ABC transcription outperforms prior neural and rule-based systems and boosts VLM sheet-music QA.
Event-level MAE embeddings plus UMAP/HDBSCAN or K-Means clustering recover 15 hydroacoustic classes from multi-year Mayotte data with ~1 hour of annotation and detector-comparable F1.
Introduces Visibility-Aware Densification with Temporally-Adaptive Thresholding and Temporal Offset Warping to improve dynamic region quality in 3D Gaussian Splatting on three benchmarks.
Introduces the Perception algorithm for seed-guided semi-supervised clustering via a-contrario anomaly detection, defining clusters as anomaly-free subsets and achieving competitive performance with 10-30 seeds per cluster on benchmarks.
EntropyInfer adaptively allocates inference compute using per-head attention entropy for rigid/dynamic classification during prefilling and compresses KV cache with generated tokens, achieving up to 2.39x speedup on long contexts.
A contrastive learning transformer embeds network flow sequences to enable correlation clustering that groups scanner sources consistently with labels.
A survey of on-device learning in TinyML organized by distribution change regimes, highlighting influences on applications, hardware, and solutions plus a gap between benchmarks and deployments.
Shared-score quaternion self-attention reduces score multiplications by 75% and softmax operations from four to one while proving equivalence to component-wise attention under quaternion linear projections.
Rubato model with InterMo representation outperforms cascade methods in generating timestamped piano sheet music from audio, even when cascades receive ground-truth MIDI.
TextTeacher uses frozen text embeddings from captions as semantic anchors to guide vision model training, improving ImageNet accuracy by up to 2.7 p.p. and transfer performance by 1.0 p.p. on average.
Transcoda achieves state-of-the-art zero-shot OMR with an 18.46% OMR-NED error rate on synthetic scores and 63.97% on historical Polish scans using a 59M model trained in 6 hours via synthetic data, kern normalization, and grammar decoding.
Lexical acoustic coding lets LLMs transmit audio waveforms as editable natural-language sentences that another LLM can parse and reconstruct into sound.
SAGE uses sparse autoencoders to boost vulnerability signals in LLMs, raising internal SNR 12.7x and delivering up to 318% MCC gains on vulnerability detection benchmarks.
Sonata is a small hybrid world model pre-trained to predict future IMU states that outperforms autoregressive baselines on clinical discrimination, fall-risk prediction, and cross-cohort transfer while fitting on-device wearables.
LLM-Codec augments audio codec training with multi-step token prediction and contrastive semantic alignment to improve both waveform reconstruction and autoregressive predictability for speech language models.
citing papers explorer
-
Articulatory movements influence electromagnetic wave transmission through the vocal tract
Articulatory configurations during vowel production create distinct electromagnetic transmission patterns through the vocal tract, confirmed by qualitative agreement between finite-element simulations and scattering-matrix measurements on two subjects.