Pith. sign in

REVIEW 1 cited by

Large Language Models on Small Resource-Constrained Systems: Performance Characterization, Analysis and Trade-offs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.15352 v1 pith:BJHUB22Q submitted 2024-12-19 cs.LG cs.CC

classification cs.LGcs.CC
keywords hardwarellmsavailablejetsondevicesmodelsrecentresearch
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Generative AI like the Large Language Models (LLMs) has become more available for the general consumer in recent years. Publicly available services, e.g., ChatGPT, perform token generation on networked cloud server hardware, effectively removing the hardware entry cost for end users. However, the reliance on network access for these services, privacy and security risks involved, and sometimes the needs of the application make it necessary to run LLMs locally on edge devices. A significant amount of research has been done on optimization of LLMs and other transformer-based models on non-networked, resource-constrained devices, but they typically target older hardware. Our research intends to provide a 'baseline' characterization of more recent commercially available embedded hardware for LLMs, and to provide a simple utility to facilitate batch testing LLMs on recent Jetson hardware. We focus on the latest line of NVIDIA Jetson devices (Jetson Orin), and a set of publicly available LLMs (Pythia) ranging between 70 million and 1.4 billion parameters. Through detailed experimental evaluation with varying software and hardware parameters, we showcase trade-off spaces and optimization choices. Additionally, we design our testing structure to facilitate further research that involves performing batch LLM testing on Jetson hardware.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Understanding the Performance and Power of LLM Inferencing on Edge Accelerators

    cs.DC 2025-06 conditional novelty 6.0 of 10

    An empirical benchmark of a 64GB Jetson Orin AGX shows that LLMs up to 32B parameters can run with INT8 quantization, but token throughput drops sharply as sequence length grows, and quantization slows smaller models.

Pith tools