MemFlow routes queries by intent to tiered memory operations, nearly doubling accuracy of a 1.7B SLM on long-horizon benchmarks compared to full-context baselines.
H2O: Heavy-hitter oracle for efficient generative inference of large language models
3 Pith papers cite this work. Polarity classification is still indexing.
years
2026 3representative citing papers
A shared per-token controller jointly routes attention resolution, FFN experts, and KV bit-width and is claimed to Pareto-dominate independently tuned MoD+MoE+KV-quant at matched cost while protecting rare-token accuracy.
CONF-KV dynamically allocates KV cache budget via model confidence, matching full-KV perplexity within 1.5-2.1 points at sliding-window memory cost and improving retrieval accuracy.
citing papers explorer
-
MemFlow: Intent-Driven Memory Orchestration for Small Language Model Agents
MemFlow routes queries by intent to tiered memory operations, nearly doubling accuracy of a 1.7B SLM on long-horizon benchmarks compared to full-context baselines.
-
TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation
A shared per-token controller jointly routes attention resolution, FFN experts, and KV bit-width and is claimed to Pareto-dominate independently tuned MoD+MoE+KV-quant at matched cost while protecting rare-token accuracy.
-
CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM
CONF-KV dynamically allocates KV cache budget via model confidence, matching full-KV perplexity within 1.5-2.1 points at sliding-window memory cost and improving retrieval accuracy.