An empirical benchmark of a 64GB Jetson Orin AGX shows that LLMs up to 32B parameters can run with INT8 quantization, but token throughput drops sharply as sequence length grows, and quantization slows smaller models.
Characterizing the performance of accelerated jetson edge devices for training dnns,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.DC 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
Understanding the Performance and Power of LLM Inferencing on Edge Accelerators
An empirical benchmark of a 64GB Jetson Orin AGX shows that LLMs up to 32B parameters can run with INT8 quantization, but token throughput drops sharply as sequence length grows, and quantization slows smaller models.