Running lightweight distilled LLMs inside Intel TDX secure VMs reportedly gives higher tokens per second than plain CPU execution for sub-3B models, with Q4 quantization reaching about 3x FP16 throughput.
Shad- ownet: A secure and efficient on-device model inference system for convolutional neural networks,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.AI 1years
2025 1verdicts
REJECT 1roles
background 1polarities
support 1representative citing papers
citing papers explorer
-
Distilled Large Language Model in Confidential Computing Environment for System-on-Chip Design
Running lightweight distilled LLMs inside Intel TDX secure VMs reportedly gives higher tokens per second than plain CPU execution for sub-3B models, with Q4 quantization reaching about 3x FP16 throughput.