← back to paper
arxiv: 2607.10183 · 2 revisions
Automated Tensor Scheduling for Hybrid CPU-GPU LLM Inference on Consumer Devices