HELM adaptively partitions HBM between EMB and KV caches via a three-layer PPO controller and EMB-KV-aware scheduling, reducing P99 latency by 24-38% while achieving 93.5-99.6% SLO satisfaction on production workloads.
Gems: Breaking the long-sequence barrier in generative recommendation with a multi-stream decoder
3 Pith papers cite this work. Polarity classification is still indexing.
years
2026 3verdicts
UNVERDICTED 3representative citing papers
SinkRec proposes a memory-conditioned architecture with TDGD to mitigate semantic state sink in linear attention for long-sequence recommendation.
UniFormer introduces a unified model-centric scaling approach for recommender systems via feature-space and task-space modules, semantic tokenization, and multi-sequence attention, with reported gains in production A/B tests at Kuaishou.
citing papers explorer
-
One Pool, Two Caches: Adaptive HBM Partitioning for Accelerating Generative Recommender Serving
HELM adaptively partitions HBM between EMB and KV caches via a three-layer PPO controller and EMB-KV-aware scheduling, reducing P99 latency by 24-38% while achieving 93.5-99.6% SLO satisfaction on production workloads.
-
SinkRec: Mitigating Semantic State Sink in Long Sequence Recommendation with Memory-Conditioned Gated Delta Networks
SinkRec proposes a memory-conditioned architecture with TDGD to mitigate semantic state sink in linear attention for long-sequence recommendation.
-
UniFormer: Efficient and Unified Model-Centric Scaling for Industrial Recommendation
UniFormer introduces a unified model-centric scaling approach for recommender systems via feature-space and task-space modules, semantic tokenization, and multi-sequence attention, with reported gains in production A/B tests at Kuaishou.