REVIEW 2 cited by
ApproxABFT: Approximate Algorithm-Based Fault Tolerance for Neural Network Processing
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
With the increasing deployment of deep neural networks (DNNs) in terrestrial and aerospace safety-critical applications, system reliability has emerged as a co-equal design metric alongside computational efficiency. Algorithm-based fault tolerance (ABFT) mechanisms, characterized by architecture-agnostic and cost-effectiveness, have become a promising solution for reliability enhancement. However, conventional ABFT approaches rely on rigorous verification mechanisms where even minor computational deviations trigger error recovery processes, which not only disregards the intrinsic fault tolerance characteristics of DNN models but also incurs redundant fault tolerance processing overhead. To address these limitations, we propose an Approximate ABFT framework (ApproxABFT) that innovatively introduces adaptive error tolerance thresholds to enable selective fault recovery, activating error correction modules exclusively when computational deviations exceed predefined thresholds. This approach effectively mitigating overreaction to non-critical computational errors. Furthermore, a dynamic block granularity optimization algorithm is implemented to achieve inter-layer error sensitivity balancing. Experimental evaluations demonstrate that the proposed ApproxABFT achieves a 43.39% average reduction in redundant computing overhead compared to previous accurate ABFT, while simultaneously enhancing the tolerable soft error rate by an order of magnitude.
Forward citations
Cited by 2 Pith papers
-
Custom Algorithm-based Fault Tolerance for Attention Layers in Transformers
Flash-ABFT verifies an entire transformer attention layer with one fused checksum that covers softmax and all three matrix products, reporting 96-99% fault detection at under 5.3% area overhead.
-
Zero Memory Overhead Approach for Protecting Vision Transformer Parameters
A method that repurposes the LSB of each ViT parameter as a parity bit detects and masks bit-flip faults with zero memory overhead.
Discussion (0). Continue with ORCID to comment.