REVIEW 1 major objections 2 minor 50 references
A twit-based multiplier enables efficient generic modular multiplication for RNS moduli of the form 2^n ± δ by deferring carry propagation to the final stage.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.3
2026-06-27 14:29 UTC pith:RFTDFHOT
load-bearing objection The paper gives a working twit-based RNS multiplier with synthesis numbers, but the claim that carry propagation stays deferred for every delta is not isolated in the results. the 1 major comments →
A Generic Modulo-(2^npmδ) RNS Multiplier Based on Twit Representation
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The proposed architecture computes the product of two residues through operand splitting, modular partial-product generation, carry-save accumulation, overflow folding, and a twit-compatible final modular addition. This organization avoids the long critical paths of conventional designs and is compatible with the twit representation for the full range of admissible deltas.
What carries the argument
Twit representation of residues combined with overflow folding after carry-save accumulation in the multiplication pipeline.
Load-bearing premise
The twit representation remains compatible with the new partial-product generation, carry-save accumulation, and overflow-folding steps across the entire admissible delta range without introducing hidden long carry chains or requiring post-synthesis fixes.
What would settle it
Synthesis results for a specific delta in the admissible range showing no reduction or an increase in delay, area, or power would falsify the efficiency claim.
If this is right
- The multiplier integrates directly with prior twit-based adders and subtractors for complete RNS arithmetic.
- It supports 5-bit channels that together give sufficient dynamic range due to the wide delta range.
- Performance gains hold for 8-bit and 11-bit channel widths as well.
- End-to-end latency decreases in RNS-based workloads involving many multiplications and additions.
Where Pith is reading between the lines
- Similar techniques might apply to other arithmetic operations in RNS beyond add, subtract, and multiply.
- Adoption could lower power in machine-learning accelerators that use RNS for efficiency.
- Further work could explore automated generation of such multipliers for arbitrary delta values.
- Comparison against other generic RNS multipliers not based on twit would quantify the specific benefit.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a generic twit-based modular multiplier for RNS channels with moduli of form 2^n ± δ (0 ≤ δ ≤ 2^{n-1}-1). The architecture performs operand splitting, modular partial-product generation, carry-save accumulation, overflow folding, and a final twit-compatible modular addition, with the goal of deferring all carry propagation to the last stage. Synthesis in FreePDK 45 nm for 5-/8-/11-bit channels reports average reductions of 20.5% delay, 13.2% area, and 28.0% power versus baselines, plus system-level latency gains over mixed multiply-add workloads.
Significance. If the constant-depth claim holds across the full δ range, the work supplies a missing multiplier primitive that pairs with prior twit-based addition/subtraction, enabling flexible RNS datapaths without modulus-specific redesign. The concrete 45 nm synthesis numbers and workload study constitute reproducible, falsifiable evidence of the gains; this is a practical strength for RNS applications in cryptography and accelerators.
major comments (1)
- [§3] §3 (Architecture, overflow-folding paragraph): the assertion that folding maps the CSA sum back into twit form without introducing δ-dependent carry chains is load-bearing for the constant-critical-path claim, yet the description provides neither an explicit bound on carry-propagation length inside the folding logic nor a separate synthesis isolation of folding delay versus δ; the 5-/8-/11-bit results therefore rest on an unverified assumption.
minor comments (2)
- [Abstract] Abstract and §5: the baseline architectures, exact δ values chosen for each channel width, and any error-bar information on the reported percentage reductions are not stated; adding these details would allow readers to reproduce the comparison.
- [§5] §5 synthesis tables: column headings for the compared designs and the precise modulus set (including chosen δ) are needed for clarity.
Simulated Author's Rebuttal
We thank the referee for the constructive feedback on our manuscript. The major comment highlights a point where additional clarification and evidence can strengthen the constant-critical-path claim. We address it point-by-point below and commit to revisions that directly respond to the concern.
read point-by-point responses
-
Referee: [§3] §3 (Architecture, overflow-folding paragraph): the assertion that folding maps the CSA sum back into twit form without introducing δ-dependent carry chains is load-bearing for the constant-critical-path claim, yet the description provides neither an explicit bound on carry-propagation length inside the folding logic nor a separate synthesis isolation of folding delay versus δ; the 5-/8-/11-bit results therefore rest on an unverified assumption.
Authors: We acknowledge that the current manuscript does not supply an explicit analytical bound on carry length within the overflow-folding logic or an isolated synthesis sweep versus δ. The overall 45 nm results for the complete multiplier already incorporate the folding stage for representative δ values within each channel width, and the reported delay reductions are consistent with the claim. However, to make the independence from δ fully verifiable, we will revise §3 to include (1) a short proof that the folding operation produces at most a two-bit carry chain independent of δ (arising from the twit digit bounds and the specific overflow-correction constants), and (2) a supplementary table isolating post-synthesis delay of the folding block alone across the full admissible δ range for the 8-bit channel. These additions will be placed in the revised version. revision: yes
Circularity Check
Minor self-citation on twit representation; central claims rest on synthesis results
full rationale
The paper presents an architectural construction (operand splitting, modular PP generation, CSA accumulation, overflow folding, final twit addition) whose performance numbers are obtained from FreePDK 45 nm synthesis rather than from any equation or derivation that reduces to a fitted constant or prior result by construction. Reliance on the twit representation is explicitly cited from prior work and does not create a load-bearing circular chain inside the present derivation or claims. No self-definitional, fitted-prediction, or uniqueness-imported steps are present.
Axiom & Free-Parameter Ledger
free parameters (2)
- residue channel width
- delta range
axioms (1)
- domain assumption Twit representation permits efficient modular addition and subtraction for all admissible δ
read the original abstract
Modular multiplication is a fundamental arithmetic primitive in Residue Number Systems (RNS) and is often the dominant source of delay, area, and energy consumption in RNS datapaths used in cryptography, signal processing, and machine-learning accelerators. Recent work introduced a twit-based residue representation for moduli of the form $2^n \pm \delta$, with $0 \le \delta \le 2^{n-1}-1$, and showed that it enables efficient generic modular addition and subtraction across the full admissible $\delta$ range. However, an efficient modular multiplier compatible with the same representation has remained unavailable. This paper presents a generic twit-based modulo-$(2^n \pm \delta)$ multiplier for RNS channels. The proposed architecture computes the product through operand splitting, modular partial-product generation, carry-save accumulation, overflow folding, and a twit-compatible final modular addition. By deferring carry propagation to the final stage, the resulting organization avoids the long critical paths characteristic of conventional multiply-then-reduce designs. To demonstrate the effectiveness of the proposed approach, we study a modulus set with 5-bit residue channels and show that, owing to the broad admissible range of $\delta$, it can provide a sufficiently wide dynamic range. Moreover, additional 8-bit and 11-bit configurations are used to evaluate the proposed approach at larger channel widths. We implement and synthesize the proposed multiplier in a FreePDK 45\,nm flow, and the results show average reductions of 20.5\% in delay, 13.2\% in area, and 28.0\% in power relative to baseline designs. A system-level study further indicates that these circuit-level improvements translate into lower end-to-end latency over a broad range of modular multiplication and addition workloads.
Figures
Reference graph
Works this paper leans on
-
[1]
Up to 8k-bit Modular Montgomery Multiplication in Residue Number Systems With Fast 16-bit Residue Channels,
Z. Ahmadpour and G. Jaberipur, “Up to 8k-bit Modular Montgomery Multiplication in Residue Number Systems With Fast 16-bit Residue Channels,”IEEE Transactions on Computers, vol. 71, no. 6, pp. 1399– 1410, 2021
2021
-
[2]
RNS-based FPGA accelerators for high-quality 3D medical image wavelet processing using scaled filter coefficients,
N. N. Nagornov, P. A. Lyakhov, M. V . Valueva, and M. V . Bergerman, “RNS-based FPGA accelerators for high-quality 3D medical image wavelet processing using scaled filter coefficients,”IEEE Access, vol. 10, pp. 19 215–19 231, 2022
2022
-
[3]
Application of the residue number system to reduce hardware costs of the convolutional neural network implementation,
M. V . Valueva, N. Nagornov, P. A. Lyakhov, G. V . Valuev, and N. I. Chervyakov, “Application of the residue number system to reduce hardware costs of the convolutional neural network implementation,” Mathematics and computers in simulation, vol. 177, pp. 232–243, 2020. 12 TABLE III: Synthesis results of the compared generic modulo-(2 n ±δ)multipliers for...
2020
-
[4]
1.19 1.072010.48 0.75 1272 0.75 1513.55 0.81
-
[5]
1.32 1.19 2964.10 1.11 2221 1.31 2938.61 1.57 Proposed1.11 1.002679.70 1.00 1689 1.00 1871.75 1.00 +3 259
-
[6]
1.70 1.40 9140.09 3.11 3374 1.61 5731.41 2.26
-
[7]
1.36 1.12 3435.28 1.17 2685 1.28 3646.77 1.44 Proposed1.21 1.00 2938.76 1.00 2094 1.00 2534.16 1.00 −9 247
2094
-
[8]
1.31 1.16 2592.41 1.011668 0.982188.75 1.13
-
[9]
1.26 1.12 2774.03 1.08 1994 1.17 2515.63 1.30 Proposed1.13 1.00 2558.15 1.001709 1.001931.68 1.00 +9 265
1994
-
[10]
1.59 1.22 8061.17 2.04 3198 1.15 5072.03 1.41
-
[11]
1.49 1.143882.52 0.983052 1.10 4549.31 1.26 Proposed1.30 1.003953.38 1.002769 1.00 3608.84 1.00 −127 129 [14] 1.20 1.192351.66 0.901574 1.00 1886.60 1.20 Proposed1.01 1.002622.45 1.001569 1.00 1577.79 1.00 +127 383 [14] 1.70 1.12 6750.88 1.52 2911 1.12 4951.03 1.26 Proposed1.52 1.00 4452.78 1.00 2594 1.00 3930.43 1.00 11 −3 2045
2045
-
[12]
1.60 1.203329.68 0.77 2230 0.73 3567.33 0.87
-
[13]
1.70 1.27 5136.02 1.18 4036 1.32 6851.11 1.68 Proposed1.34 1.004339.15 1.00 3050 1.00 4078.15 1.00 +3 2051
2051
-
[14]
2.21 1.57 54554.70 10.04 9283 2.18 20541.41 3.43
-
[15]
1.76 1.25 5989.68 1.10 5092 1.20 8969.05 1.50 Proposed1.41 1.00 5431.68 1.00 4250 1.00 5984.42 1.00 −9 2039
2039
-
[16]
1.70 1.21 4291.75 1.022836 0.954821.20 1.15
-
[17]
1.72 1.22 4676.57 1.12 3733 1.25 6418.52 1.53 Proposed1.41 1.00 4192.26 1.002985 1.004199.60 1.00 +9 2057
2057
-
[18]
2.35 1.56 64344.38 12.98 9338 2.47 21957.43 3.86
-
[19]
1.88 1.25 6302.23 1.27 5305 1.41 9965.44 1.75 Proposed1.51 1.00 4955.34 1.00 3775 1.00 5689.68 1.00 −1023 1025 [14] 2.08 1.19 13422.91 2.01 5360 1.33 11157.90 1.58 Proposed1.75 1.00 6687.65 1.00 4042 1.00 7081.58 1.00 +1023 3071 [14] 2.31 1.23 50857.13 7.14 7920 1.84 18273.02 2.27 Proposed1.87 1.00 7126.43 1.00 4300 1.00 8044.44 1.00
-
[20]
Res-dnn: A residue number system-based dnn accelerator unit,
N. Samimi, M. Kamal, A. Afzali-Kusha, and M. Pedram, “Res-dnn: A residue number system-based dnn accelerator unit,”IEEE Transactions on Circuits and Systems I: regular papers, vol. 67, no. 2, pp. 658–671, 2019
2019
-
[21]
Computer arithmetic: Algorithms and hardware designs,
B. Parhami, “Computer arithmetic: Algorithms and hardware designs,” Oxford University Press, 2010
2010
-
[22]
Efficient VLSI implementation of modulo(2 n ±1) addition and multiplication,
R. Zimmermann, “Efficient VLSI implementation of modulo(2 n ±1) addition and multiplication,” inProceedings 14th IEEE Symposium on Computer Arithmetic (Cat. No.99CB36336), 1999, pp. 158–167
1999
-
[23]
Efficient diminished-1 modulo 2 n +1 multipliers,
C. Efstathiou, H. T. Vergos, G. Dimitrakopoulos, and D. Nikolos, “Efficient diminished-1 modulo 2 n +1 multipliers,”IEEE Transactions on Computers, vol. 54, no. 4, pp. 491–496, 2005
2005
-
[24]
Design of efficient modulo 2 n +1 multipliers,
H. T. Vergos and C. Efstathiou, “Design of efficient modulo 2 n +1 multipliers,”IET Computers & Digital Techniques, vol. 1, no. 1, pp. 49–57, 2007
2007
-
[25]
On the Design of Modulo 2n +1 Multipliers,
C. Efstathiou, K. Pekmestzi, and N. Axelos, “On the Design of Modulo 2n +1 Multipliers,” in14th Euromicro conference on digital system design. IEEE, 2011, pp. 453–459
2011
-
[26]
Efficient modulo 2n +1 multiply and multiply-add units based on modified Booth encoding,
C. Efstathiou, N. Moshopoulos, N. Axelos, and K. Pekmestzi, “Efficient modulo 2n +1 multiply and multiply-add units based on modified Booth encoding,”Integration, vol. 47, no. 1, pp. 140–147, 2014
2014
-
[27]
Arithmetic units for RNS moduli{2 n −3}and{2 n +3}operations,
P. M. Matutino, R. Chaves, and L. Sousa, “Arithmetic units for RNS moduli{2 n −3}and{2 n +3}operations,” in13th Euromicro Conference on Digital System Design. IEEE, 2010, pp. 243–246
2010
-
[28]
Improved modulo-(2 n ±3)multipliers,
H. Ahmadifar and G. Jaberipur, “Improved modulo-(2 n ±3)multipliers,” in17th CSI International Symposium on Computer Architecture & Digital Systems. IEEE, 2013, pp. 31–35
2013
-
[29]
Modulo-(2 q − 3)Multiplication with Fully Modular Partial Product Generation and Reduction,
G. Jaberipur, S. Gorgin, N. Ahamadian, and J.-A. Lee, “Modulo-(2 q − 3)Multiplication with Fully Modular Partial Product Generation and Reduction,” in30th Symposium on Computer Arithmetic. IEEE, 2023, pp. 68–75
2023
-
[30]
New efficient structure for a modular multiplier for RNS,
A. A. Hiasat, “New efficient structure for a modular multiplier for RNS,” IEEE Transactions on Computers, vol. 49, no. 2, pp. 170–174, 2000
2000
-
[31]
RNS Arithmetic Units for Modulo 2 n ±k,
P. M. Matutino, H. Pettenghi, R. Chaves, and L. Sousa, “RNS Arithmetic Units for Modulo 2 n ±k,” in15th Euromicro Conference on Digital System Design. IEEE, 2012, pp. 795–802
2012
-
[32]
A generic modulo-(2 n ±δ) addition algorithm via two-valued digit encoding,
S. Gorgin, A. Sadr, D. Rahmati, and J. Kim, “A generic modulo-(2 n ±δ) addition algorithm via two-valued digit encoding,” in32nd Symposium on Computer Arithmetic. IEEE, 2025, pp. 85–92
2025
-
[33]
P. A. Mohan,Residue Number Systems: Theory and Applications, 1st ed. Birkh¨auser Basel, 2016
2016
-
[34]
Modular multiplication and base extensions in residue number systems,
J.-C. Bajard, L.-S. Didier, and P. Kornerup, “Modular multiplication and base extensions in residue number systems,” in15th Symposium on Computer Arithmetic. IEEE, 2001, pp. 59–65
2001
-
[35]
Area-Power Efficient Modulo 2 n −1 and Modulo 2n +1 Multipliers for{2 n −1,2 n,2 n +1}Based RNS,
R. Muralidharan and C.-H. Chang, “Area-Power Efficient Modulo 2 n −1 and Modulo 2n +1 Multipliers for{2 n −1,2 n,2 n +1}Based RNS,”IEEE Transactions on Circuits and Systems I: Regular Papers, vol. 59, no. 10, pp. 2263–2274, 2012
2012
-
[36]
Novel approaches to the design of VLSI RNS multipliers,
D. Radhakrishnan and Y . Yuan, “Novel approaches to the design of VLSI RNS multipliers,”IEEE Transactions on Circuits and Systems II Analog and Digital Signal Processing, vol. 39, no. 1, pp. 52–57, 1992
1992
-
[37]
A universal architecture for designing efficient modulo 2n +1 multipliers,
L. Sousa and R. Chaves, “A universal architecture for designing efficient modulo 2n +1 multipliers,”IEEE Transactions on Circuits and Systems I: Regular Papers, vol. 52, no. 6, pp. 1166–1178, 2005
2005
-
[38]
Koren,Computer Arithmetic Algorithms, 2nd ed
I. Koren,Computer Arithmetic Algorithms, 2nd ed. Natick, MA, USA: A K Peters, 2002
2002
-
[39]
Multiplication of Multidigit Numbers on Automata,
A. Karatsuba and Y . Ofman, “Multiplication of Multidigit Numbers on Automata,”Soviet Physics Doklady, vol. 7, p. 595, 12 1962
1962
-
[40]
Weighted two-valued digit- set encodings: unifying efficient hardware representation schemes for redundant number systems,
G. Jaberipur, B. Parhami, and M. Ghodsi, “Weighted two-valued digit- set encodings: unifying efficient hardware representation schemes for redundant number systems,”IEEE Transactions on Circuits and Systems I: Regular Papers, vol. 52, no. 7, pp. 1348–1357, 2005
2005
-
[41]
Fast parallel-prefix modulo 2n +1 adders,
C. Efstathiou, H. T. Vergos, and D. Nikolos, “Fast parallel-prefix modulo 2n +1 adders,”IEEE Transactions on Computers, vol. 53, no. 9, pp. 1211–1216, 2004
2004
-
[42]
New memoryless, mod(2 n ±1)residue multiplier,
A. Hiasat, “New memoryless, mod(2 n ±1)residue multiplier,”Elec- tronics Letters, vol. 28, pp. 314–315, 1992
1992
-
[43]
Combinational logic approach for designing RNS multipliers,
A. A. Hiasat and H. Abdel-Aty-Zohdy, “Combinational logic approach for designing RNS multipliers,” in39th Midwest Symposium on Circuits and Systems, vol. 1. IEEE, 1996, pp. 541–543
1996
-
[44]
Novel Modulo 2 n +1 Multipliers,
H. Vergos and C. Efstathiou, “Novel Modulo 2 n +1 Multipliers,” in 9th Euromicro Conference on Digital System Design. IEEE, 2006, pp. 491–496
2006
-
[45]
Faster modulo 2 n +1 multipliers without booth recoding,
R. Chaves and L. Sousa, “Faster modulo 2 n +1 multipliers without booth recoding,” inConf. Design of Circuits and Integrated Systems, 2005
2005
-
[46]
(4+2logn)∆GParallel Prefix Modulo-(2 n −3)Adder via Double Representation of Residues in [0, 13 2],
G. Jaberipur and S. H. F. Langroudi, “(4+2logn)∆GParallel Prefix Modulo-(2 n −3)Adder via Double Representation of Residues in [0, 13 2],”IEEE Transactions on Circuits and Systems II: Express Briefs, vol. 62, no. 6, pp. 583–587, 2015
2015
-
[47]
High-Performance Multiplication Modulo 2 n–3,
P.-M. Seidel, “High-Performance Multiplication Modulo 2 n–3,” in52nd Asilomar Conference on Signals, Systems, and Computers, 2018, pp. 130–134
2018
-
[48]
A low-complexity com- binatorial RNS multiplier,
V . Paliouras, K. Karagianni, and T. Stouraitis, “A low-complexity com- binatorial RNS multiplier,”IEEE Transactions on Circuits and Systems II: Analog and Digital Signal Processing, vol. 48, no. 7, pp. 675–683, 2001
2001
-
[49]
RNS arithmetic multiplier for medium and large moduli,
A. A. Hiasat, “RNS arithmetic multiplier for medium and large moduli,” IEEE Transactions on Circuits and Systems II: Analog and Digital Signal Processing, vol. 47, no. 9, pp. 937–940, 2002
2002
-
[50]
Double-least-significant-bits 2’s-complement number rep- resentation scheme with bitwise complementation and symmetric range,
B. Parhami, “Double-least-significant-bits 2’s-complement number rep- resentation scheme with bitwise complementation and symmetric range,” IET Circuits, Devices & Systems, vol. 2, pp. 179–186, 2008. Saeid Gorgin(Senior Member, IEEE) received the B.S. degree in Computer Engineering from Azad University, South Tehran Branch, Tehran, Iran, in 2001, the M.S....
2008
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.