ITHICA generates functional tests via intra-thread instruction duplication and comparison, detecting 39% more defective servers than baseline methods on over 3000 real CPUs while revealing new defect behaviors.
hub
Primer: Fast private transformer inference on encrypted data
19 Pith papers cite this work, alongside 86 external citations. Polarity classification is still indexing.
hub tools
citation-role summary
citation-polarity summary
roles
background 2polarities
background 2representative citing papers
A new NoC with Direct Compute Access delivers 5.3x multicast and 2.8x reduction speedups, yielding up to 3.8x performance gains and 1.17x energy savings versus baseline unicast designs in GEMM workloads.
Power-Softmax is a new HE-compatible attention variant that permits training and inference of billion-parameter polynomial LLMs with performance matching standard transformers.
A rubric-guided GRPO pipeline fine-tunes a 7B LLM to synthesize quantum circuits achieving 3.31x T-gate compression with <1% hardware-constraint violations, validated on IBM and IonQ processors.
SwiftCTS combines physics-informed gradient-boosted models with K-shot multiplicative calibration to enable fast, low-error prediction and Pareto optimization of clock tree metrics on unseen macro architectures without retraining.
Proposes Distributed Persistence Domain and Persistent CXL Switch to enable low-latency persistence operations at CXL switch level while maintaining crash consistency in disaggregated memory.
DORA is an instruction-based DNN accelerator architecture with a two-stage compilation framework that delivers stable efficiency across varied workloads and up to 5x throughput gains versus prior accelerators on FPGA.
A new hypergraph-based GNN framework with polarity decomposition and consistency regularization is proposed for unsat core prediction in SAT.
ImageHD delivers up to 40.4x speedup and 383x energy efficiency for on-device continual learning of visual representations by using hyperdimensional computing and bounded exemplar management on an FPGA.
D-VQLS with FWHT Pauli decomposition and 1% thresholding reduces circuit evaluations by 256x for 10-qubit tridiagonal systems while achieving over 99.99% fidelity and near-ideal scaling on up to 96 GPUs.
PrivaDE is a privacy-preserving protocol for jointly computing data utility scores in ML using secure computation, with optimizations for efficiency and blockchain integration via smart contracts.
Fine-grained fusion and adaptive scheduling in SSMs deliver up to 4.8x speedup and 10x lower on-chip memory, enabling a fusion-aware accelerator with 1.78x higher performance than MARCA at equal area.
Co-design of 14.5x compacted index, asynchronous scheduler, and multiplication-free kernel for PIM-based graph ANNS delivers up to 20x CPU and 17.1x GPU throughput on billion-scale benchmarks.
NeuroSPICE uses PINNs to solve circuit DAEs via residual minimization, creating surrogate models for optimization and handling nonlinear emerging devices.
DAT combines a small-large model cascade with fine-tuning and bandwidth-aware multi-stream transmission to deliver high-accuracy event recognition and low-latency alerts for video streams in edge-cloud systems.
Quantum walks integrated with variational circuits and CUDA-Q acceleration generate high-fidelity adaptive probability distributions for 1D financial modeling and 2D digit patterns.
Systematic study concludes overlay architectures suit frequent model switching in current autonomous driving setups, while customized ones may become preferable as bitstream reload overhead decreases.
The paper reviews energy-aware computing literature and constructs a taxonomy organized by hardware/software aspects, measurement, optimizations, scheduling, scaling, consolidation, federated learning, and cooling.
The paper reviews multiscale thermal modeling techniques for 3D ICs, unifying scales from device to system while stressing thermal boundary resistance and validation needs.
citing papers explorer
-
ITHICA: Intra-Thread Instruction Checking Approach for Defect-Induced Silent Data Corruptions
ITHICA generates functional tests via intra-thread instruction duplication and comparison, detecting 39% more defective servers than baseline methods on over 3000 real CPUs while revealing new defect behaviors.
-
A Lightweight High-Throughput Collective-Capable NoC for Large-Scale ML Accelerators
A new NoC with Direct Compute Access delivers 5.3x multicast and 2.8x reduction speedups, yielding up to 3.8x performance gains and 1.17x energy savings versus baseline unicast designs in GEMM workloads.
-
Power-Softmax: Towards Secure LLM Inference over Encrypted Data
Power-Softmax is a new HE-compatible attention variant that permits training and inference of billion-parameter polynomial LLMs with performance matching standard transformers.
-
RubriQ: Rubric-Guided Group Relative Policy Optimization for Constraint-Aware Quantum Circuit Synthesis
A rubric-guided GRPO pipeline fine-tunes a 7B LLM to synthesize quantum circuits achieving 3.31x T-gate compression with <1% hardware-constraint violations, validated on IBM and IonQ processors.
-
SwiftCTS: Fast Cross-Design Prediction and Pareto Optimization of Clock Tree Metrics via Few-Shot Calibration
SwiftCTS combines physics-informed gradient-boosted models with K-shot multiplicative calibration to enable fast, low-error prediction and Pareto optimization of clock tree metrics on unseen macro architectures without retraining.
-
Distributed Persistence Domain for Persistent Memory Pooling
Proposes Distributed Persistence Domain and Persistent CXL Switch to enable low-latency persistence operations at CXL switch level while maintaining crash consistency in disaggregated memory.
-
DORA: Dataflow-Instruction Orchestration Architecture for DNN Acceleration
DORA is an instruction-based DNN accelerator architecture with a two-stage compilation framework that delivers stable efficiency across varied workloads and up to 5x throughput gains versus prior accelerators on FPGA.
-
Unsat Core Prediction through Polarity-Aware Representation Learning over Clause-Literal Hypergraphs
A new hypergraph-based GNN framework with polarity decomposition and consistency regularization is proposed for unsat core prediction in SAT.
-
ImageHD: Energy-Efficient On-Device Continual Learning of Visual Representations via Hyperdimensional Computing
ImageHD delivers up to 40.4x speedup and 383x energy efficiency for on-device continual learning of visual representations by using hyperdimensional computing and bounded exemplar management on an FPGA.
-
Distributed Variational Quantum Linear Solver
D-VQLS with FWHT Pauli decomposition and 1% thresholding reduces circuit evaluations by 256x for 10-qubit tridiagonal systems while achieving over 99.99% fidelity and near-ideal scaling on up to 96 GPUs.
-
PrivaDE: Privacy-preserving Data Evaluation for Blockchain-based Data Marketplaces
PrivaDE is a privacy-preserving protocol for jointly computing data utility scores in ML using secure computation, with optimizations for efficiency and blockchain integration via smart contracts.
-
Fine-Grained Fusion: The Missing Piece in Area-Efficient State Space Model Acceleration
Fine-grained fusion and adaptive scheduling in SSMs deliver up to 4.8x speedup and 10x lower on-chip memory, enabling a fusion-aware accelerator with 1.78x higher performance than MARCA at equal area.
-
Co-Designing Graph-based Approximate Nearest Neighbor Search at Billion Scale for Processing-in-Memory
Co-design of 14.5x compacted index, asynchronous scheduler, and multiplication-free kernel for PIM-based graph ANNS delivers up to 20x CPU and 17.1x GPU throughput on billion-scale benchmarks.
-
Physics-Informed Neural Networks for Device and Circuit Modeling: A Case Study of NeuroSPICE
NeuroSPICE uses PINNs to solve circuit DAEs via residual minimization, creating surrogate models for optimization and handling nonlinear emerging devices.
-
DAT: Dual-Aware Adaptive Transmission for Efficient Multimodal LLM Inference in Edge-Cloud Systems
DAT combines a small-large model cascade with fine-tuning and bandwidth-aware multi-stream transmission to deliver high-accuracy event recognition and low-latency alerts for video streams in edge-cloud systems.
-
Quantum Walks-Based Adaptive Distribution Generation with Efficient CUDA-Q Acceleration
Quantum walks integrated with variational circuits and CUDA-Q acceleration generate high-fidelity adaptive probability distributions for 1D financial modeling and 2D digit patterns.
-
To Overlay or to Customize? Revisiting Architectural Choices in Heterogeneous Systems
Systematic study concludes overlay architectures suit frequent model switching in current autonomous driving setups, while customized ones may become preferable as bitstream reload overhead decreases.
-
Energy-Aware Computing in the Year 2026
The paper reviews energy-aware computing literature and constructs a taxonomy organized by hardware/software aspects, measurement, optimizations, scheduling, scaling, consolidation, federated learning, and cooling.
-
A Review of Multiscale Thermal Modeling in Heterogeneous 3D ICs
The paper reviews multiscale thermal modeling techniques for 3D ICs, unifying scales from device to system while stressing thermal boundary resistance and validation needs.