Pith. sign in

REVIEW 1 cited by

Efficiently Distilling LLMs for Edge Applications

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.01353 v1 pith:OE4Y6WCN submitted 2024-04-01 cs.LG cs.AIcs.CL

classification cs.LGcs.AIcs.CL
keywords modelsapplicationstrainingedgellmssupernetabilitycalled
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Supernet training of LLMs is of great interest in industrial applications as it confers the ability to produce a palette of smaller models at constant cost, regardless of the number of models (of different size / latency) produced. We propose a new method called Multistage Low-rank Fine-tuning of Super-transformers (MLFS) for parameter-efficient supernet training. We show that it is possible to obtain high-quality encoder models that are suitable for commercial edge applications, and that while decoder-only models are resistant to a comparable degree of compression, decoders can be effectively sliced for a significant reduction in training time.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SoftmAP: Software-Hardware Co-design for Integer-Only Softmax on Associative Processors

    cs.AR 2024-11 conditional novelty 4.0 of 10

    An integer-only Softmax approximation from I-BERT, mapped onto associative processors, can cut Softmax energy by up to 1300x and latency by up to 12.58x versus GPUs, but with small perplexity loss at the advertised precision.

Pith tools