Dual-view memory speeds NPU-PIM LLM inference up to 2.32x
PFM stores each tensor once and lets NPU and PIM read it in their preferred layout.
· “Rethinking Unified Memory for NPU-PIM Systems: Dual-View Memory for Dynamic Inference of LLM”
Hardware Architecture
Covers systems organization and hardware architecture. Roughly includes material in ACM Subject Classes C.0, C.1, and C.5.
sort pith recommended most recent
PFM stores each tensor once and lets NPU and PIM read it in their preferred layout.
· “Rethinking Unified Memory for NPU-PIM Systems: Dual-View Memory for Dynamic Inference of LLM”
The benchmark's 145 tasks show the gap is trade-off reasoning and geometry planning, not rule recall.
A literature review summarizing design approaches, performance metrics, challenges, and emerging techniques for mm-wave and sub-THz/THz…
· “Recent Advances in mm-Wave and Sub-THz/THz Oscillators for FutureG Technologies”
M100 manages data streams explicitly to support both driving models and LLMs with better efficiency.
· “M100: An Orchestrated Dataflow Architecture Powering General AI Computing”
All weights and activations stay in on-chip memory during topology optimization, yielding 4.18x better energy efficiency than scaled GPU
Selective protection of sensitive blocks plus adaptive rollback lets voltage drop or frequency rise while image quality stays intact.
· “DRIFT: Harnessing Inherent Fault Tolerance for Efficient and Reliable Diffusion Model Inference”
Validated to within 8.57% error, it supports arbitrary LLM scenarios and yields design insights for memory and accelerator hardware.
· “A Full-Stack Performance Evaluation Infrastructure for 3D-DRAM-based LLM Accelerators”
Systolic arrays plus compute-in-memory hit 138 tokens/s at 112 mW for multimodal LLMs on 22nm.
A hardware noise schedule lifts best-known-solution hits from 17.5% to 86.5% and halves time-to-solution.
· “Integrated Hardware Annealing based on Langevin Dynamics for Ising Machines”
A round-robin attack across sub-banks cuts mean time to failure from 13 years to about 1 second.
· “From Fleet to Lab: Revisiting the Security and Complexity of Industrial Rowhammer Mitigation”
One epoch of clean-data fine-tuning, no triggers needed, keeps functional correctness and often lifts pass rates.
· “RTLGuard: A Lightweight Teacher-Student Defense for Poisoned RTL Code Generation Models”
A delay-aware schedule drops critical-path weight traffic from 9.00 to 4.50 MB/token; a 30.9B-parameter MoE decodes at 5.94 tok/s.
A host profile plus swept PIM parameters stays within 10 percent of full-system simulation.
· “VIPER: Architecture-Aware Performance Modeling for Processing-in-Memory Design-Space Exploration”
The paper's honest negative result: lowering the prefill chunk size recovers more KV at the same time-to-first-token.
Adding an LLM to a SPICE-driven genetic loop yields compact, corner-robust analog classifiers.
A 101-design benchmark set backs the claim: SYNTLOG finishes huge FSMs Vivado abandons, with built-in validation.
A passive probe near the NIC tells benign web use from floods, scans, and probes without packet or host access.
Independent execution and reasoning signals, fused only at decision time, recover 60% of the oracle gap.
· “Execution-Anchored Hallucination Calibration Reranking for Verilog Code Generation”
A simulated hybrid silicon/IGO monolithic-3D 6T SRAM with back-end-of-line pass-gates gives 25% area reduction and 42% faster writes at the…
· “A Process-Aware Hybrid Si/IGO Monolithic-3D 6T SRAM with BEOL Pass-Gates for the 2nm Node”
Co-designing a 4F2 DRAM cell with a two-tier near-memory processor keeps the full decode working set in memory and enables 256K contexts.
Accel-Sim 2.0 runs real async AI kernels cycle-by-cycle, exposing hidden chiplet and prefetch bottlenecks.
· “Architecting the Next Generation of Asynchronous, Distributed GPUs for the AI Era”
Tuned per-PE arithmetic keeps neural-net accuracy while shrinking hardware by over two-thirds.
· “Precision-Aware Variable Bit Processing Elements for Hardware-Efficient Systolic Array Designs”
Measured on M1 and M3: a 25.85M conv model streams 0 ANE bytes in fp16 and about 83 percent residency in int8 or 2-bit.
A state exploration protocol reuses the simulation engine to build a full reachability graph, so no model conversion is needed.
· “Constraint-Driven Modeling Enabling Dual Model Checking and Simulation for Discrete Event Systems”
Holding only a few rows of state, it runs 1080p30 and detects Full-HD frames in 11 ms.
Engineers get a tunable accept/defer rule for LLM-generated hardware before any testbench or golden model exists.
· “NoTB: Oracle-Free Triage of LLM-Generated RTL via Cross-Model Formal Consensus”
Transformer-CNN model maps CPU and GPU temperature fields with RMSE below 0.26 °C, fast enough for runtime control.
· “TherMapNet Attention-Guided Runtime Full-Chip Thermal Map Prediction from Performance Metrics”
A five-category co-design taxonomy shows published accuracy and energy numbers mostly cannot be compared.
Prototype streams 420 GB of CKKS bootstrapping transforms through a QSFP56 port, skipping a separate processing stage.
· “Programmable Compute-in-Transit using Integrated Photonics”
Proof-kernel-checked math, no human-reviewed proofs, and a full cost ledger make the claim auditable.
Triplicating every flip-flop makes the tile-level cost 16.8% extra area and 15.2% extra power.
Draft-model prefetching hides PCIe transfer stalls, letting large MoE models run fast on memory-limited GPUs.
Perturbing the accumulated output instead of stored weights cuts read-modify-write energy to 0.46–0.83× on IMC accelerators.
· “Event-triggered Implicit Perturbation for Zeroth-Order Fine-Tuning of Spiking Transformers”
Lowering chip voltage injects tiny arithmetic faults that regularize learning, improving PGD robustness while saving power.
· “Faults That Fortify: CNN Adversarial Robustness via GPU Undervolting”
One activation representation yields exact LUT inference and compact rules, within 0.5 pp of teacher accuracy.
· “Bern2Edge: A Neurosymbolic Compiler for Edge Deployment via Bernstein Polynomial Networks”
A 0.076 mm² CNN plus classifier runs in 7.34 ms, fitting the 8 ms real-time EEG budget.
· “A Resource-Efficient CNN-Based EEG Auditory Attention Decoding ASIC”
The ONEX compiler splits 2D execution planning into independent depth-optimal 1D rows and columns, shortening syndrome cycles.
· “Architecture and Compilation Co-Design for High-Rate Quantum Product Codes on Neutral Atom Arrays”
Fixed-point math plus a low-cost sine approximation delivers 45x lower energy-delay on optimization tasks.
· “ODEONN: A Digital ODE Solver Architecture for Oscillatory Neural Networks”
An all-digital design slows the clock only when needed, avoiding large voltage guard bands.
· “Experimental Verification of Fast Voltage Droop Correction Circuits”
Weight scaling plus per-layer exponent-bias search beats FFT-internal scaling and cuts energy 2.5x versus CPU.
· “Energy-Efficient Visual Inspection with FFT-Based CNNs and Adaptive Floating-Point Quantization”
By decoupling threads from registers, the design adapts parallelism per phase and feeds tensor cores without redundant loads.
· “A Thread-Register Decoupled GPU Execution Model for Efficient Tensor Computation”
It matches dynamic programming's expansions while storing at most 17 traversal entries at central demand.
Jointly searching chiplet composition, placement, batching, and scheduling cuts time-to-first-token by 43.7 percent.
· “HYDRA: A Heterogeneous Chiplet DSE Framework for Serving Dynamic Hybrid LLM Workloads”
HyperCut prunes inter-layer candidates early, cutting exploration time by 80.47% over SET.
· “HyperCut: Fast Inter-Layer Scheduling via Directed Hypergraph and Early Filtering”
A state-carrying LLM tuner beats five baselines everywhere; removing its memory cuts frontier quality by 58.5%.
· “StateTune: Transforming LLM-Assisted EDA Flow Tuning into a Stateful, Closed-Loop Process”
Mutation feedback steers LLM proposals through symbolic repair, roughly doubling verified assertions on RTL designs.
· “Coverage-Driven RTL Assertion Generation with Formal Exploration and Neuro-Symbolic Refinement”
MAGMA runs the entire EM loop in hardware on a Spartan-7, adapting mixture parameters while classifying streaming pixels.
· “MAGMA: Mixture-Model Adaptive Gaussian Model Acceleration”
Reusing sampled prefixes cuts DRAM traffic ~96% and speeds up end-to-end HGNN inference 5.7×.
· “ESR-HGNN: Eliminating Semantic Redundancy for Efficient Mini-batch HGNN Inference”
It pairs a spline-convolution engine with a 3D/2D cache to bring VGA event vision to the edge.
A neuro-symbolic pipeline has an LLM rewrite only property-irrelevant logic, then proves each rewrite sound before verifying.
· “NeuroAbs: A Neuro-Symbolic RTL Abstraction Framework for Property Checking Acceleration”
Turn-model routing plus SAT rerouting nearly eliminates oversubscribed NoC links for only 4% more bandwidth.
· “The Road Less Traveled: Congestion-Aware NoC Placement and Packet Routing for FPGAs”
CTTE drives N-Trace and E-Trace from a shared front end and reconstructs 818k instructions from Linux with zero errors.
· “CTTE: An Open Dual-Protocol RISC-V Trace Encoder for N-Trace and E-Trace”
Its loop chooses what to fix next from end-of-flow timing and power targets; TNS improves 30.67%.
A 278-byte generator cuts classifier energy 47% and tracking error 4.6x in a wristband design study.
· “One Residual with Three Reuses: A Wristband Front End for Gesture Sensing”
Pipelined tree reductions lift the loop-carried 2-FLOP floor to 216 FLOP/cycle, and a tolerance contract verifies the reordered sums.
· “DTX: A Throughput-First Training Accelerator for Diffusion and Transformer Models”
Measured on FPGA: committed events keep exact cycles while code stays editable.
The INT8 NPU scheme saved 17–28% energy per sample while adding up to 38% training time.
· “NPU Offloading of a Frozen Visual Encoder for Robot Policy Training”
Storing MoE experts in flash and reading them over a direct GPU link plus an HBM relay cuts serving latency by 40-50%.
· “Beyond Capacity: Scalable MoE LLM Inference via High-Bandwidth Flash with Direct GPU and HBM Paths”
The 128-terabyte state that once required a supercomputer fits on a $42k SATA subsystem.
· “Qu-Trefoil: Large-Scale Quantum Circuit Simulator Working on FPGA With SATA Storages”
Every degree-four potential with constant nonzero Hessian determinant has polynomial gradient inverse.
Sequential reads in one user's LBA range trigger SSD-internal cleanup that throttles other processes.
· “Experimental Study on System-Level Performance Impact of Read Disturbance in Modern SSDs”
Bounded multicast sharing recovers utilization lost to routing skew, without global expert replication.
Attack signatures shift with chip, cache design, noise, and pacing; transfer scores drop to zero.
Simulated HBF also cuts the minimum GPU count from four to one, but needs HBM-class read bandwidth and better endurance.
· “Exploring High-Bandwidth Flash for Modern LLM Inference: Opportunities and Challenges”
PPAPlace trains on post-global-routing labels and flows timing gradients back into placement.
· “PPAPlace: Differentiable Cross-Stage Objectives for Chip Placement Optimization”
Post-quantum keys and authenticated encryption keep in-memory LLM inference secure at 34%/9% overhead.
· “YAVIN: A Unified Architecture for Secure Edge Processing in Memory”
Hardware enforces type-checked loads with under 0.85% runtime and 1.4% area cost on an FPGA prototype.
· “ROLoad-PMP: Securing Sensitive Operations for Kernels and Bare-Metal Firmware”
Bitstream deployment, versioning, and rollback move out of vendor toolchains and into the OS.
Simulations show HBF pays off only if HBM keeps the bandwidth-critical path; the paper gives the conditions.
A tiny NFA-based overlay inspects cache-coherent traffic in real time and swaps filters in under a second.
· “Dryas: A Reprogrammable Engine for High-Speed Interconnect Tracing and Analysis”
Closed-loop diagnosis plus knowledge retrieval beats one-shot scripts and fixed action sets on 14 designs.
· “SynAct: A Reasoning-Acting Large Language Model Agent for Adaptive Synthesis Optimization”
A twelve-gate verifier finds the standard allclose test certifies nearly 1,500 silently broken kernels as correct.
A TPU datapath block reaches functional convergence in as few as two repair iterations, preserving validated logic.
· “Spec-Driven Hardware Evolution via Executable Contract Refinement and Proof-Guided RTL Update”
Across 46 designs the median kill rate is 74%; three testbenches catch 0% of injected faults.
· “GateTruth: Auditing the Rigor of RTL Design Benchmarks via Mutation Testing”
A new accelerator combines local spiking learning and INT4 math, beating GPUs with under 1% accuracy loss.
Per-tree leaf bitwidths set during boosting keep accuracy while shrinking the hardware's logic footprint.
· “FQTree: Fine-grained Quantization and Hardware Generation of Boosted Decision Trees”
DRAM buffers intermediate activations while linear layers stay in dense TLC NAND, reviving in-storage LLM computing.
· “NITRO: High-Performance 3D NAND Flash-Based In-Storage Computing with Enhanced Activation Dataflow”
Transversal gates cut syndrome rounds; PACE packs them as densely as the decoder allows.
· “Do Not Let CNOTs Overwhelm the Decoder: Scheduling Transversal Gates for Fast FTQC”
An ISA- and source-level audit finds INT8 withdrawn across PTX, CUTLASS, vLLM, and SGLang.
A learned confidence model fetches just enough extra experts to hide memory stalls, beating fixed top-k prefetching.
· “APEX: Adaptive Expert Prefetching for Memory-Efficient Edge MoE Inference”