HSTU-based generative recommenders with 1.5 trillion parameters scale as a power law with compute up to GPT-3 scale, outperform baselines by up to 65.8% NDCG, run 5-15x faster than FlashAttention2 on long sequences, and improve online A/B metrics by 12.4%.
Bytetransformer: A high-performance transformer boosted for variable-length inputs
4 Pith papers cite this work, alongside 23 external citations. Polarity classification is still indexing.
representative citing papers
Proteus combines static code analysis, runtime signals, and LLM guidance to select burst-buffer data layouts, claiming 91.3% decision accuracy and up to 3.24x/2.9x speedups on write- and metadata-intensive HPC workloads.
PROMISE tool automates mixed-precision tuning with user-defined floating-point formats, validated on linear solvers and Rodinia benchmarks showing many variables can use lower precision safely.
citing papers explorer
-
Actions Speak Louder than Words: Trillion-Parameter Sequential Transducers for Generative Recommendations
HSTU-based generative recommenders with 1.5 trillion parameters scale as a power law with compute up to GPT-3 scale, outperform baselines by up to 65.8% NDCG, run 5-15x faster than FlashAttention2 on long sequences, and improve online A/B metrics by 12.4%.
-
Rethinking Burst Buffer Optimization: Enabling Layout Heterogeneity via Hybrid Analysis and LLM Guidance
Proteus combines static code analysis, runtime signals, and LLM guidance to select burst-buffer data layouts, claiming 91.3% decision accuracy and up to 3.24x/2.9x speedups on write- and metadata-intensive HPC workloads.
-
Floating-point autotuning with customized precisions
PROMISE tool automates mixed-precision tuning with user-defined floating-point formats, validated on linear solvers and Rodinia benchmarks showing many variables can use lower precision safely.
- ImplicitTerrainV2: Wavelet-Guided Spatially Adaptive Neural Terrain Representation