REVIEW 2 cited by
QUTE: Quantifying Uncertainty in TinyML with Early-exit-assisted ensembles for model-monitoring
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Uncertainty quantification (UQ) provides a resource-efficient solution for on-device monitoring of tinyML models deployed without access to true labels. However, existing UQ methods impose significant memory and compute demands, making them impractical for ultra-low-power, KB-sized TinyML devices. Prior work has attempted to reduce overhead by using early-exit ensembles to quantify uncertainty in a single forward pass, but these approaches still carry prohibitive costs. To address this, we propose QUTE, a novel resource-efficient early-exit-assisted ensemble architecture optimized for tinyML models. QUTE introduces additional output blocks at the final exit of the base network, distilling early-exit knowledge into these blocks to form a diverse yet lightweight ensemble. We show that QUTE delivers superior uncertainty quality on tiny models, achieving comparable performance on larger models with 59% smaller model sizes than the closest prior work. When deployed on a microcontroller, QUTE demonstrates a 31% reduction in latency on average. In addition, we show that QUTE excels at detecting accuracy-drop events, outperforming all prior works.
Forward citations
Cited by 2 Pith papers
-
DEBUG-HD: Debugging TinyML models on-device using Hyper-Dimensional computing
DEBUG-HD uses a binarized MLP-hidden-layer projection as the HDC encoder and outperforms prior binary HDC methods by 27% on average at detecting input corruptions in TinyML, at hyper-dimensions of 300 to 400.
-
TCUQ: Single-Pass Uncertainty Quantification from Temporal Consistency with Streaming Conformal Calibration for TinyML
TCUQ turns short-horizon output instability into an on-device uncertainty score with a streaming quantile threshold, claiming calibrated abstention for TinyML without online labels.
Discussion (0). Continue with ORCID to comment.