An 8-pass MLIR pipeline lifts bit-level accelerator RTL semantics to tensor ISA specs that enable automatic compiler-backend generation, validated on Gemmini and VTA.
Canonical reference
Title resolution pending
Canonical reference. 83% of citing Pith papers cite this work as background.
citation-role summary
citation-polarity summary
representative citing papers
Simulation study shows cold TLB misses in reverse address translation dominate latency for small collectives in multi-GPU pods, causing up to 1.4x degradation, while larger ones see diminishing returns.
FlexiFlow optimizes carbon footprint for item-level intelligence on flexible electronics by modeling lifetime variation, delivering 1.62X microarchitectural and 14.5X algorithmic reductions plus a 30.9 kHz tape-out.
Serving frontier MoE and multimodal models on Ascend 910 via vLLM-Ascend is feasible but dominated by engineering cost from incomplete operators, fragile parallelism, kernel faults, and weak observability.
COMPOSE is a timing-driven composable CGRA architecture that fuses cross-iteration operations and defers registration to deliver 1.6x performance and 2.9x EDP gains over prior CGRA designs for recurrence-bound loops.
HammerSim is a gem5-based full-system framework for modeling RowHammer with probability-driven bitflip simulation, validated against real DDR4 DIMMs via JS divergence.
ACALSim is a new simulation framework with customizable threading, event-driven execution, and shared-memory model that reports over 14x speedup versus SST and enables simulation of large LLaMA models that SST cannot complete.
Quantum circuits show high average condition (97.56%) and decision (97.63%) coverage but lower path coverage (71.84%), with probabilistic versions adding confidence levels (averages 88.87%, 88.65%, 37.18%); mutation testing reveals weak or no correlation between structural coverage and fault finding
Reference-augmented offline policy optimization through a differentiable RNN dynamics model cuts TDCR tip-position error by ~51% versus non-augmented training and outperforms Jacobian controllers across speeds.
Privatar partitions VR avatar reconstruction via frequency-domain decomposition, keeping sensitive components local and offloading the rest with distribution-aware minimal perturbation noise, achieving 2.37x throughput with provable privacy.
AEGIS reduces inter-GPU communication by up to 81.3% in self-attention and reaches 96.62% scaling efficiency with 3.86x speedup on four GPUs for 2048-token encrypted Transformer inference.
Analog-aware block Jacobi schemes in flexible GMRES maintain convergence under simulated device non-idealities when block size, damping, and approximation accuracy are chosen to account for analog scaling, noise, quantization, and clipping.
FILCO introduces a real-time reconfigurable composing architecture for DNN acceleration that achieves 1.3x-5x better throughput and hardware efficiency than prior designs on diverse workloads via an analytical model and two-stage design space exploration.
MATCH augments sparsified attention with an efficient in-context retrieval system to boost performance on long-range recall tasks in transformers.
App-based tracking and regression analysis identify capacity offer as strongest positive factor in rural Tuttlingen and punctuality as strongest negative factor in both German locations.
The paper reviews energy-aware computing literature and constructs a taxonomy organized by hardware/software aspects, measurement, optimizations, scheduling, scaling, consolidation, federated learning, and cooling.
A review of initiatives to make LHC Monte Carlo event generations available as open data to minimize redundant simulations and resource use.
citing papers explorer
-
TensorLift: Automatic Extraction of Tensor-Level ISA Semantics from Accelerator RTL via MLIR Semantic Lifting
An 8-pass MLIR pipeline lifts bit-level accelerator RTL semantics to tensor ISA specs that enable automatic compiler-backend generation, validated on Gemmini and VTA.
-
Analyzing Reverse Address Translation Overheads in Multi-GPU Scale-Up Pods
Simulation study shows cold TLB misses in reverse address translation dominate latency for small collectives in multi-GPU pods, causing up to 1.4x degradation, while larger ones see diminishing returns.
-
Lifetime-Aware Design for Item-Level Intelligence at the Extreme Edge
FlexiFlow optimizes carbon footprint for item-level intelligence on flexible electronics by modeling lifetime variation, delivering 1.62X microarchitectural and 14.5X algorithmic reductions plus a 30.9 kHz tape-out.
-
On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend
Serving frontier MoE and multimodal models on Ascend 910 via vLLM-Ascend is feasible but dominated by engineering cost from incomplete operators, fragile parallelism, kernel faults, and weak observability.
-
COMPOSE: Static Timing-driven Composable Reconfigurable Architecture for Accelerating Recurrence-Bound Loops
COMPOSE is a timing-driven composable CGRA architecture that fuses cross-iteration operations and defers registration to deliver 1.6x performance and 2.9x EDP gains over prior CGRA designs for recurrence-bound loops.
-
HammerSim: A System-Level Tool to Model RowHammer
HammerSim is a gem5-based full-system framework for modeling RowHammer with probability-driven bitflip simulation, validated against real DDR4 DIMMs via JS divergence.
-
ACALSim: A Scalable Parallel Simulation Framework for High-Performance System Design Space Exploration
ACALSim is a new simulation framework with customizable threading, event-driven execution, and shared-memory model that reports over 14x speedup versus SST and enables simulation of large LLaMA models that SST cannot complete.
-
Probabilistic Condition, Decision and Path Coverage of Circuit-based Quantum Programs
Quantum circuits show high average condition (97.56%) and decision (97.63%) coverage but lower path coverage (71.84%), with probabilistic versions adding confidence levels (averages 88.87%, 88.65%, 37.18%); mutation testing reveals weak or no correlation between structural coverage and fault finding
-
Reference-Augmented Learning for Precise Tracking Policy of Tendon-Driven Continuum Robots
Reference-augmented offline policy optimization through a differentiable RNN dynamics model cuts TDCR tip-position error by ~51% versus non-augmented training and outperforms Jacobian controllers across speeds.
-
Dual-Anchoring: Addressing State Drift in Vision-Language Navigation
Privatar partitions VR avatar reconstruction via frequency-domain decomposition, keeping sensitive components local and offloading the rest with distribution-aware minimal perturbation noise, achieving 2.37x throughput with provable privacy.
-
AEGIS: Scaling Long-Sequence Homomorphic Encrypted Transformer Inference via Hybrid Parallelism on Multi-GPU Systems
AEGIS reduces inter-GPU communication by up to 81.3% in self-attention and reaches 96.62% scaling efficiency with 3.86x speedup on four GPUs for 2048-token encrypted Transformer inference.
-
Hybrid Digital-Analog Approximate Inverse Preconditioning for Krylov Methods
Analog-aware block Jacobi schemes in flexible GMRES maintain convergence under simulated device non-idealities when block size, damping, and approximation accuracy are chosen to account for analog scaling, noise, quantization, and clipping.
-
FILCO: Flexible Composing Architecture with Real-Time Reconfigurability for DNN Acceleration
FILCO introduces a real-time reconfigurable composing architecture for DNN acceleration that achieves 1.3x-5x better throughput and hardware efficiency than prior designs on diverse workloads via an analytical model and two-stage design space exploration.
-
MATCH: Modulating Attention via In-Context Retrieval for Long-Context Transformers
MATCH augments sparsified attention with an efficient in-context retrieval system to boost performance on long-range recall tasks in transformers.
-
Urban Context and Travel Experience Events: An Exploratory Comparison of Two German Cities
App-based tracking and regression analysis identify capacity offer as strongest positive factor in rural Tuttlingen and punctuality as strongest negative factor in both German locations.
-
Energy-Aware Computing in the Year 2026
The paper reviews energy-aware computing literature and constructs a taxonomy organized by hardware/software aspects, measurement, optimizations, scheduling, scaling, consolidation, federated learning, and cooling.
-
Open LHC Monte Carlo Event Generation
A review of initiatives to make LHC Monte Carlo event generations available as open data to minimize redundant simulations and resource use.
- A Switch-Centric In-Network Architecture for Accelerating LLM Inference in Shared-Memory Network