Context-aware cross-family speculative decoding reaches 1.7x speedup on structured Polish text but fails to deliver gains on varied instructions because both models are memory-bandwidth bound on unified memory.
Native LLM and MLLM Inference at Scale on Apple Silicon
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
fields
cs.CL 2years
2026 2verdicts
UNVERDICTED 2representative citing papers
BaseRT achieves up to 1.56x higher LLM decode throughput than llama.cpp on Apple Silicon through native Metal kernel fusion and unified memory optimizations.
citing papers explorer
-
Cross-Family Speculative Decoding for Polish Language Models on Apple~Silicon: An Empirical Evaluation of Bielik~11B with UAG-Extended MLX-LM
Context-aware cross-family speculative decoding reaches 1.7x speedup on structured Polish text but fails to deliver gains on varied instructions because both models are memory-bandwidth bound on unified memory.
-
BaseRT: Best-in-Class LLM Inference on Apple Silicon via Native Metal
BaseRT achieves up to 1.56x higher LLM decode throughput than llama.cpp on Apple Silicon through native Metal kernel fusion and unified memory optimizations.