REVIEW 2 major objections 4 minor 31 references
Training the wiring of logic-gate and lookup-table networks cuts gate count nearly fifty-fold while matching MNIST accuracy.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-13 03:17 UTC pith:R3AGOSIO
load-bearing objection Useful systems result: joint connection+gate training cuts LGN size by nearly 50× at matched MNIST accuracy, with honest ablations, though the exact headline numbers come from shallow-only recipes. the 2 major comments →
Fully Trainable Deep Differentiable Logic Gate Networks and Lookup Table Networks
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Jointly training a soft distribution over candidate connections for every gate or LUT input pin, while simultaneously learning gate types or LUT entries, produces binary networks that match the accuracy of fixed-connection baselines with nearly fifty times fewer gates (one layer of 8000 gates reaches 98.45 percent on MNIST versus 98.47 percent with 384 000 fixed gates) and remain stable up to ten layers once straight-through estimators, a high learning rate and constant-gate removal are used.
What carries the argument
A temperature-controlled softmax over a pool of candidate wires per input pin, hardened by a straight-through estimator so that the forward pass selects only the highest-probability connection while the backward pass still receives gradients from the soft mixture; gate types or LUT entries are learned in parallel by the same soft-to-hard schedule.
Load-bearing premise
That forcing each pin to a single hard wire (and each LUT entry to a binary value) after soft training leaves only a small accuracy drop—the discretization gap stays negligible even for deeper six-input LUT networks.
What would settle it
Train the reported fully connection-trainable LGN and 6-LUTN architectures, measure the test accuracy of the continuous model just before hard selection versus the fully binarized network after hard selection; if the gap grows beyond a few tenths of a percent for networks deeper than four layers, the claim that the final binary model retains the continuous accuracy fails.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a method to jointly train gate/LUT types and their input connections in deep differentiable logic gate networks (LGNs) and lookup table networks (LUTNs). Connections are selected from a per-pin pool of Nc candidates via a softmax (or STE/argmax) over learned weights, while gate types or LUT entries are optimized in parallel. On Yin-Yang, MNIST and Fashion-MNIST the connection-optimized models are reported to match or exceed fixed-connection baselines with far fewer gates/LUTs (e.g., one layer of 8000 gates at 98.45 % and two layers at 98.92 % on MNIST, claimed nearly 50 imes smaller than Petersen et al. 2022; two layers of 2000 6-LUTs at 98.88 %). Stability to 10 layers for LGNs is obtained by raising the learning rate, applying straight-through estimators, and removing constant-output gate types; an improved sigmoid-annealed LUT neuron is shown to train stably to 6 layers and to need fewer parameters than the original LGN formulation.
Significance. If the gate-count reductions hold under a single consistent training recipe, the work would be a clear advance for resource-constrained and FPGA-targeted inference: it shows that the random fixed wiring used in prior LGN/LUTN literature is a major source of inefficiency and that co-optimizing topology can shrink models by an order of magnitude while preserving accuracy. The multi-benchmark evaluation, 95 % confidence intervals, systematic ablations (learning-rate, STE on gates vs. connections, constant-gate removal, residual init), and the GEMM reformulation of the LUT neuron are concrete engineering contributions that make the approach more practical. The discretization-gap analysis and the explicit comparison of full-precision versus binarized networks further strengthen the empirical foundation.
major comments (2)
- [§4.2, abstract, Table 1, Fig. 6] The headline gate-reduction claim (abstract and §4.2: 98.45 % with one layer of 8000 gates, 98.92 % with two layers, “almost 50 times fewer” than Petersen et al.’s 384 k-gate fixed-connection network) is obtained with the non-default recipes ste/c and highLR. The same paragraph explicitly states that these recipes “performed worse on networks beyond two layers deep,” while the only recipe shown to remain stable to 10 layers (ste/gc/no01/res) is the one used for Table 1, whose 8 k-gate entries are 98.3 % (1 layer) and 98.7 % (2 layers). Consequently the 50 imes claim is not demonstrated by the depth-stable method that the paper otherwise promotes. A single consistent recipe that simultaneously reaches the claimed shallow accuracy and remains stable when depth is increased must be reported, or the abstract/§4.2 numbers must be clearly caveated and replaced by the Table-1 figures.
- [§5.1.2, §5.3, Table 4, Fig. 8b] For 6-LUTNs the fully-binarized accuracy collapses with depth (Table 4: 98.88 % at 2 layers with Nc=all falls to 93.7 % at 6 layers). The paper correctly notes the growing discretization gap (Fig. 8b, §5.1.2–5.3), yet still presents the shallow 98.88 % figure as a principal result. Either a method that keeps the gap small for deeper 6-LUTNs must be supplied, or the claim of “stable training \ldots up to 6-layer deep networks” should be restricted to the full-precision (non-hardware) model and the binarized numbers reported with the same caveats applied to the LGN headline results.
minor comments (4)
- [Table 1] Table 1 reports only the ste/gc/no01/res recipe; the abstract numbers obtained with other recipes never appear in any table. Adding a row or footnote that lists the shallow highLR/ste/c results would remove ambiguity.
- [§2.2, §3.2, §4] The temperature schedules for Tc and Tg, the precise residual-init value, and the β-annealing schedule for LUTNs are described only in prose; a short hyper-parameter table would aid reproducibility.
- [Fig. 6c] Figure 6c caption defines the discretization gap as max full-precision accuracy minus max binarized accuracy “even if they do not occur on the same epoch.” Clarifying whether early-stopping or last-epoch selection is used would avoid confusion.
- [front matter, references] A few typographical inconsistencies appear (e.g., “Peinlaan” vs. “Pleinlaan”, occasional missing spaces around citations). A light copy-edit pass is sufficient.
Circularity Check
Empirical methods paper: accuracies measured on external held-out benchmarks; no derivation reduces to its inputs by construction.
full rationale
The paper proposes a training procedure (softmax over candidate connections per pin, jointly with gate/LUT parameters; STE hard selection; temperature annealing; constant-gate trimming; sigmoid annealing of LUT entries) and reports test-set accuracies on Yin-Yang, MNIST and Fashion-MNIST. Those accuracies are not algebraically forced by the training objectives or by any fitted scalar that is later re-used as a “prediction.” Comparisons to fixed-connection LGNs cite Petersen et al. (independent authors) and to the authors’ own prior LUTN model (Mommen et al. 2026) only as baselines that are re-implemented and re-measured; the load-bearing claims are the new experimental numbers obtained with the connection-training algorithm. No uniqueness theorem, ansatz, or self-definitional identity is invoked. Selective reporting of shallow-only recipes for the headline 50 imes claim is a transparency/correctness issue, not circularity. Hence score 0 and empty steps.
Axiom & Free-Parameter Ledger
free parameters (7)
- learning_rate (fully trainable LGNs) =
0.1
- connection_softmax_temperature_schedule Tc =
1 → 1e-4
- gate_softmax_temperature_schedule Tg =
1 → 1e-4
- Nc (candidate connections per input pin) =
8 / 16 / all
- residual_init_pass_through_weight =
5
- sigmoid_annealing_beta_schedule (LUTNs) =
increased over second half of training
- layer_width / depth choices =
e.g. 8000 gates/layer; 2000 6-LUTs/layer
axioms (6)
- domain assumption Boolean 2-input gates and N-input LUTs can be relaxed to differentiable mixtures (softmax over 16 gate types; MUX/Boolean product form with sigmoid entries) and trained by backpropagation.
- domain assumption Straight-through estimators that use argmax in the forward pass and softmax gradients in the backward pass yield useful discrete networks.
- domain assumption After training, selecting the highest-probability connection and gate/LUT configuration produces a hardware-implementable binary network whose accuracy is close to the continuous model.
- domain assumption Binarized-pixel MNIST/Fashion-MNIST and Yin-Yang are adequate proxies for the value of smaller logic/LUT networks on edge/FPGA targets.
- standard math Softmax, LogSumExp, and matrix calculus identities used in the GEMM LUT reformulation are valid.
- ad hoc to paper Removing constant-output gate types from hidden layers improves information flow without harming expressivity enough to erase the accuracy gains.
invented entities (2)
-
Connection pool with per-pin softmax probabilities (partial/full connection training)
no independent evidence
-
Sigmoid-annealed LUT neuron (improved LUTN model)
no independent evidence
read the original abstract
We introduce a novel method for both partial and full optimization of the connections in deep differentiable logic gate networks (LGNs) and lookup table networks (LUTNs). Our training method utilizes a probability distribution over a set of connections per gate/lookup table (LUT) input pin, selecting the connection with highest merit, all whilst the optimal gate types or LUT-entries are learned in parallel. We show that the connection-optimized LGNs outperform standard fixed-connection LGNs on the Yin-Yang, MNIST Handwritten Digits and Fashion-MNIST benchmarks, while requiring only a fraction of the number of logic gates. We achieve 98.92% on the MNIST dataset with two layers of 8000 gates. With only one layer of 8000 gates, we obtain 98.45%, showing that our method requires almost 50 times fewer gates compared to fixed-connection LGNs. Training stability up to ten layers has been ensured by employing a high learning rate, straight-through estimators and trimming constant-output gate types. Additionally, we present a LUT neuron description that enables stable training with backpropagation, tested up to 6-layer deep networks. The model requires four times fewer trainable parameters and still achieves a higher accuracy compared to the fixed-connection LGN training algorithm. Our connection-training algorithm also works well for the LUTNs, achieving an accuracy of 98.88% for two layers of 2000 6-input LUTs.
Figures
Reference graph
Works this paper leans on
-
[1]
Bacellar, Alan T. L. and Susskind, Zachary and Jr, Mauricio Breternitz and John, Eugene and John, Lizy K. and Lima, Priscila M. V. and Fran. Differentiable. doi:10.48550/arXiv.2410.11112 , urldate =
-
[2]
Desislavov, Radosvet and. Trends in. Sustainable Computing: Informatics and Systems , volume =. doi:10.1016/j.suscom.2023.100857 , urldate =
-
[3]
Diehl, Peter U. and Pedroni, Bruno U. and Cassidy, Andrew and Merolla, Paul and Neftci, Emre and Zarrella, Guido , year = 2016, month = jul, pages =. 2016. doi:10.1109/IJCNN.2016.7727758 , urldate =
-
[4]
and Ward, Max and Neftci, Emre O
Eshraghian, Jason K. and Ward, Max and Neftci, Emre O. and Wang, Xinxin and Lenz, Gregor and Dwivedi, Girish and Bennamoun, Mohammed and Jeong, Doo Seok and Lu, Wei D. , year = 2023, month = sep, journal =. Training. doi:10.1109/JPROC.2023.3308088 , urldate =
-
[5]
Quantized
Hubara, Itay and Courbariaux, Matthieu and Soudry, Daniel and. Quantized. Journal of Machine Learning Research , volume =
-
[6]
Kim, Seijoon and Park, Seongsik and Na, Byunggook and Yoon, Sungroh , year = 2020, month = apr, journal =. Spiking-. doi:10.1609/aaai.v34i07.6787 , urldate =
-
[7]
Kriener, Laura and G. The. doi:10.48550/arXiv.2102.08211 , urldate =
-
[8]
1998 , howpublished =
MNIST Database of Handwritten Digits , author =. 1998 , howpublished =
1998
-
[9]
Lee, Yang Yang and Halim, Zaini Abdul and Wahab, Mohd Nadhir Ab and Almohamad, Tarik Adnan , year = 2024, month = mar, journal =. Stochastic. doi:10.34133/research.0307 , urldate =
-
[10]
Convolutional
Petersen, Felix and Kuehne, Hilde and Borgelt, Christian and Welzel, Julian and Ermon, Stefano , year = 2024, month = dec, journal =. Convolutional
2024
-
[11]
Petersen, Felix and Borgelt, Christian and Kuehne, Hilde and Deussen, Oliver , year = 2022, month = dec, journal =. Deep
2022
-
[12]
Putra, Rachmad Vidya Wicaksana and Shafique, Muhammad , year = 2021, month = jul, pages =. Q-. 2021. doi:10.1109/IJCNN52387.2021.9534087 , urldate =
-
[13]
Rastegari, Mohammad and Ordonez, Vicente and Redmon, Joseph and Farhadi, Ali , editor =. Computer. doi:10.1007/978-3-319-46493-0_32 , abstract =
-
[14]
doi:10.1109/TCAD.2018.2819366 , urldate =
Rathi, Nitin and Panda, Priyadarshini and Roy, Kaushik , year = 2019, month = apr, journal =. doi:10.1109/TCAD.2018.2819366 , urldate =
-
[15]
Nature Computational Science , volume =
Opportunities for Neuromorphic Computing Algorithms and Applications , author =. Nature Computational Science , volume =. doi:10.1038/s43588-021-00184-y , abstract =
-
[16]
Sevilla, Jaime and Heim, Lennart and Ho, Anson and Besiroglu, Tamay and Hobbhahn, Marius and Villalobos, Pablo , year = 2022, month = jul, pages =. Compute. 2022. doi:10.1109/IJCNN55064.2022.9891914 , abstract =
-
[17]
LUTNet: Rethinking Inference in FPGA Soft Logic
Wang, Erwei and Davis, James J. and Cheung, Peter Y. K. and Constantinides, George A. , year = 2019, month = apr, number =. doi:10.48550/arXiv.1904.00938 , urldate =
work page internal anchor Pith review Pith/arXiv arXiv doi:10.48550/arxiv.1904.00938 2019
-
[19]
Xiao, Han and Rasul, Kashif and Vollgraf, Roland , year = 2017, month = sep, number =. Fashion-. doi:10.48550/arXiv.1708.07747 , urldate =
-
[20]
Deep. IEEE Access , author =. 2023 , pages =. doi:10.1109/ACCESS.2023.3328622 , abstract =
-
[21]
Mommen, Wout and Keuninckx, Lars and Detterer, Paul and Colpaert, Achiel and Wambacq, Piet , month = jan, year =. Inter-patient. doi:10.48550/arXiv.2601.11433 , abstract =
-
[22]
Bührer, Simon and Plesner, Andreas and Aczel, Till and Wattenhofer, Roger , month = aug, year =. Recurrent. doi:10.48550/arXiv.2508.06097 , abstract =
-
[23]
LUTNet: Learning FPGA Configurations for Highly Efficient Neural Network Inference
Wang, Erwei and Davis, James J. and Cheung, Peter Y. K. and Constantinides, George A. , month = mar, year =. doi:10.48550/arXiv.1910.12625 , abstract =
work page internal anchor Pith review Pith/arXiv arXiv doi:10.48550/arxiv.1910.12625 1910
-
[24]
doi:10.48550/arXiv.2510.15655 , abstract =
Gerlach, Lino and Våge, Liv and Gerlach, Thore and Kauffman, Elliott , month = oct, year =. doi:10.48550/arXiv.2510.15655 , abstract =
-
[25]
Umuroglu, Yaman and Akhauri, Yash and Fraser, Nicholas James and Blott, Michaela , month = aug, year =. 2020 30th. doi:10.1109/FPL50879.2020.00055 , abstract =
-
[26]
Andronic, Marta and Constantinides, George A. , month = dec, year =. 2023. doi:10.1109/ICFPT59805.2023.00012 , abstract =
-
[27]
Andronic, Marta and Constantinides, George A. , month = sep, year =. 2024 34th. doi:10.1109/FPL64840.2024.00028 , abstract =
-
[28]
Nazemi, Mahdi and Pasandi, Ghasem and Pedram, Massoud , month = jan, year =. Energy-. 2019 24th. doi:10.1145/3287624.3287722 , abstract =
-
[29]
On training networks of monostable multivibrator timer neurons , volume =. Neural Networks , author =. 2026 , keywords =. doi:10.1016/j.neunet.2025.108092 , abstract =
-
[30]
Salimans, Tim and Ho, Jonathan and Chen, Xi and Sidor, Szymon and Sutskever, Ilya , month = sep, year =. Evolution. doi:10.48550/arXiv.1703.03864 , abstract =
-
[31]
and Clune, Jeff , month = apr, year =
Such, Felipe Petroski and Madhavan, Vashisht and Conti, Edoardo and Lehman, Joel and Stanley, Kenneth O. and Clune, Jeff , month = apr, year =. Deep. doi:10.48550/arXiv.1712.06567 , abstract =
-
[32]
Sarkar, Bidipta and Fellows, Mattie and Duque, Juan Agustin and Letcher, Alistair and Villares, Antonio León and Sims, Anya and Wibault, Clarisse and Samsonov, Dmitry and Cope, Dylan and Liesen, Jarek and Li, Kang and Seier, Lukas and Wolf, Theo and Berdica, Uljad and Mohl, Valentin and Goldie, Alexander David and Courville, Aaron and Sevegnani, Karin and...
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.