HSTU-based generative recommenders with 1.5 trillion parameters scale as a power law with compute up to GPT-3 scale, outperform baselines by up to 65.8% NDCG, run 5-15x faster than FlashAttention2 on long sequences, and improve online A/B metrics by 12.4%.
Zero-shot recommendation as language modeling
2 Pith papers cite this work, alongside 37 external citations. Polarity classification is still indexing.
2
Pith papers citing it
37
external citations · OpenAlex
verdicts
UNVERDICTED 2representative citing papers
Shared task overview where all 17 retrieval systems from 2 teams beat the prior baseline and all 4 generation teams produced at least one human-preferred output.
citing papers explorer
-
Actions Speak Louder than Words: Trillion-Parameter Sequential Transducers for Generative Recommendations
HSTU-based generative recommenders with 1.5 trillion parameters scale as a power law with compute up to GPT-3 scale, outperform baselines by up to 65.8% NDCG, run 5-15x faster than FlashAttention2 on long sequences, and improve online A/B metrics by 12.4%.
-
Findings of the MAGMaR 2026 Shared Task
Shared task overview where all 17 retrieval systems from 2 teams beat the prior baseline and all 4 generation teams produced at least one human-preferred output.