REVIEW 2 major objections 5 minor 52 references
Neural networks plus pathfinders find Seiberg-duality chains for modest quivers faster than blind or pure physics search.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-31 01:43 UTC pith:T4P5Z3V3
load-bearing objection Solid empirical ML tool for tracing Seiberg dualities on modest quivers; hybrids beat BFS and LCA with public code and honest failure analysis. the 2 major comments →
Learning to Trace Seiberg Dualities
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
For quivers with a modest number of nodes (order 10), transformer-plus-MLP graph networks used as policies inside bidirectional A* and beam pathfinders, especially when hybridized with a lowest-common-ancestor rank heuristic, find Seiberg-duality paths more efficiently than unguided BFS or pure deterministic LCA search while maintaining near-100% success.
What carries the argument
Hybrid pathfinders that combine a Distance GNN (heuristic estimating mutation distance), an Adviser GNN (policy over which node to dualize next), and a physics-informed Lowest-Common-Ancestor cost that prefers rank-reducing moves; the networks are trained on BFS-generated mutation trees from toric Calabi–Yau seed quivers.
Load-bearing premise
Distances used for training and scoring are taken from BFS trees via longest common path prefixes and are only upper bounds once quiver symmetries are ignored, biasing the data toward pairs that share a low-rank common ancestor—the same criterion the physics baseline exploits.
What would settle it
Retrain the same architectures on a dataset that fully quotients by quiver automorphisms and re-measures efficiency ratios against LCA; if the hybrid advantage disappears or success rates fall below the pure LCA baseline at comparable complexity, the claimed outperformance is an artifact of the biased distance definition.
If this is right
- Computational complexity of duality chains can be read off as C = D log10 K and used to estimate how far a holographic RG flow proceeds down a warped throat.
- The same hybrid pathfinders give a practical tool for deciding whether two given quivers are Seiberg dual and for enumerating short connecting sequences.
- The task supplies a concrete benchmark for frontier AI models on a well-defined theoretical-physics search problem.
- Hybrid NN-plus-physics search remains superior to pure LCA up to roughly 1.5–2 times the training complexity before degrading.
Where Pith is reading between the lines
- The same mutation-pathfinding setup could be reused for other cluster-algebra or BPS-quiver problems where the mutation graph is infinite and non-monotonic heuristics appear.
- Once automorphism-aware distances are available, the residual gap between Hybrid LCA and pure LCA would quantify how much genuine pattern learning the networks have acquired beyond the rank-minimizing bias built into the training trees.
- Scaling the training set to larger node counts would test whether the observed complexity threshold (roughly 1.5–2× training C) continues to hold or collapses, giving a concrete scaling law for this class of physics search tasks.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies the computational problem of deciding whether two 4d N=1 quiver gauge theories are related by a sequence of Seiberg dualities (quiver mutations) and of finding short duality paths. Starting from toric Calabi–Yau seed quivers, the authors generate mutation trees by BFS, train a Distance GNN (DGNN) to regress mutation distance and an Adviser GNN (AGNN) to propose the next node to dualize, and embed both networks as heuristics/costs inside bidirectional A*, beam search, a physics-inspired Lowest-Common-Ancestor (LCA) rank heuristic, and hybrid combinations. On in-distribution and out-of-distribution quivers with O(10) nodes they report that transformer+MLP architectures, especially Hybrid and Hybrid-LCA pathfinders, substantially outperform unguided BFS and improve on pure LCA in efficiency while retaining near-100% success; a controlled complexity study (C = D log10 K) estimates the hybrids remain superior up to roughly 1.5–2× the training complexity before degrading.
Significance. The work supplies a concrete, reproducible benchmark that sits at the intersection of Seiberg duality, cluster-algebra mutations, and modern graph ML. Strengths that should be credited explicitly include: (i) public checkpoints and pathfinder code, (ii) systematic ID/OOD tables, success heatmaps and efficiency ratios (Tables 2a–2b, §§6–7), (iii) an honest failure-mode section that retrains at lower complexity to locate the breaking point (Eq. 7.3), and (iv) transparent disclosure that training distances are upper bounds from BFS path prefixes and that the dataset is biased toward low-rank common ancestors. If the empirical claims hold under the stated regime, the paper offers both a practical tool for duality searches and a useful stress-test for frontier AI models applied to theoretical physics.
major comments (2)
- [§3.1, Eq. (3.5); §6.2.2] §3.1, Eq. (3.5): duality distance is defined from BFS path-prefix LCAs without full quotienting by quiver automorphisms, so the reported d is an upper bound and the training set is biased toward pairs that share a low-rank common ancestor—the same criterion the LCA baseline exploits. The authors flag this (§6.1.5–6.2.2, Fig. 24) and still show gains on infinite-mutation OOD families never seen in training, but a quantitative bound on the residual bias (e.g., symmetry-reduced distances on a subsample, or ER recomputed after random node relabeling) is needed before the ~1.1–1.2× Hybrid-vs-LCA claim can be treated as fully independent of the generation procedure.
- [§5.2, Table 1] §5.2 and Table 1: the Hybrid-LCA cost weights (c_det dec/eq/inc, λ_det cost, λ_AR, λ_DGNN, λ_LCA) are stated to have been optimized for 100% success and maximal ER. It is unclear whether this optimization used a held-out validation split distinct from the 500-pair evaluation samples of §6. Without that separation the modest efficiency gains over pure LCA risk mild contamination by the evaluation distribution; a short statement of the hyperparameter protocol (or a re-evaluation with frozen weights chosen only on a validation slice) would remove the ambiguity.
minor comments (5)
- [§2.2] §2.2: typographical slip “the far more generic cas is” → “case”.
- [§4.1.1, Fig. 8] Fig. 8c and related MAE plots: the median error is negative (ˆd < d). A one-sentence remark that the network systematically under-estimates large distances (consistent with the non-monotonicity discussion) would help the reader interpret the heuristic quality.
- [§6.2, Appendix A.2] Appendix A.2 / Fig. 37: several quivers are anomalous as 4d N=1 theories. The text already notes they remain valid as BPS quivers; a brief cross-reference in the main OOD discussion (§6.2) would prevent confusion.
- [§7] The complexity measure C = D log10 K is introduced only in §7. A forward pointer in the introduction or §5.3 would make the later breaking-point analysis easier to anticipate.
- [Note Added / §1] References: the concurrent DualityCert work [25] is noted; a one-sentence clarification of the complementary goals (path-finding/complexity vs. verifier-gated claim repair) already present in the Note Added could be moved into the introduction for readers who skip the note.
Circularity Check
No load-bearing circularity: empirical pathfinder benchmarks on BFS-generated labels, with disclosed LCA-affinity bias that does not force the claimed efficiency gains.
specific steps
-
other
[§3.1 Eq. (3.5); §6.2.2]
"d(QA, QB)=|PA|+|PB|−2|PLCA|.... beating the LCA pathfinder on the datasets we generated, as mentioned already in the previous sections, is the most difficult test for our pathfinders because the LCA pathfinder follows the same criterion (but opposite) that we used to generate the dataset in the first place"
Training distances and pair structure are defined via longest common mutation-path prefixes (i.e., LCAs of the BFS tree). The deterministic LCA baseline optimizes the same low-rank-ancestor geometry. This creates a mild affinity between labels and one baseline, so part of the Hybrid-vs-LCA contest is run on data shaped by LCA logic. It is not a by-construction identity: Hybrid still explores fewer nodes ~84% of the time at equal path length, sometimes finds longer paths, and loses to LCA on several OOD slices; BFS and infinite-mutation OOD checks remain independent.
full rationale
The paper is an applied ML/pathfinding study, not a first-principles derivation of a physical law. Seiberg mutation rules (Eqs. 2.2–2.4) are taken as given; training pairs and distance labels are produced by BFS over duality trees (Alg. 1, Eq. 3.5) and used to supervise DGNN/AGNN; pathfinders are then scored by success rate and nodes explored against BFS and LCA baselines on held-out ID and OOD quivers. That pipeline is standard supervised learning plus search, not a self-definitional loop: nothing equates a fitted parameter to the reported efficiency ratio by construction, and Hybrid/Hybrid-LCA can and do underperform LCA on small-K, finite-mutation, and anomalous OOD families (Figs. 28–30). The only mild affinity is that Eq. 3.5 distances are LCA-prefix lengths and the dataset therefore favors pairs that share a low-rank ancestor—the same structural cue the deterministic LCA baseline exploits—which the authors explicitly flag (§3.1, §6.1.5–6.2.2). That is an evaluation-bias caveat, already bounded by bidirectional-BFS comparisons, infinite-mutation OOD families never seen in training (Fig. 38), and the controlled complexity stress test (§7). No self-citation uniqueness theorem, smuggled ansatz, or renaming of a known result carries the central claim. Score 1 reflects only that disclosed generation affinity, not forced circularity.
Axiom & Free-Parameter Ledger
free parameters (4)
- Hybrid-LCA cost weights (c_det dec, c_det eq, c_det inc, λ_det cost, λ_AR, λ_DGNN, λ_LCA) =
c_det dec=0.3, c_det eq=2.7, c_det inc=3.1, λ_det cost=1.3 (others tuned analogously)
- AGNN beam width B =
B=3
- DGNN/AGNN architecture widths and depths (H=64, 3 GNN layers, 2 transformer layers, dropout 0.2, etc.) =
H=64; DGNN ~250 epochs; AGNN stopped at epoch 39–49
- LCA rank-change costs c_dec:c_eq:c_inc =
~1:10:100
axioms (4)
- domain assumption Seiberg duality on a quiver node is exactly the Fomin–Zelevinsky mutation of (A,N) with vector-like pairs deleted and ranks required non-negative (Eqs. 2.2–2.4).
- ad hoc to paper Duality distance equals |P_A|+|P_B|−2|P_LCA| computed from BFS path prefixes in the mutation tree (Eq. 3.5), without full symmetry reduction.
- domain assumption Anomaly cancellation (weighted arrow balance) is required for training seeds; OOD tests may drop it.
- standard math Two quivers are identified up to node permutation via Weisfeiler–Lehman hashing during search.
invented entities (2)
-
Distance GNN (DGNN) and Adviser GNN (AGNN)
no independent evidence
-
Hybrid LCA pathfinder and complexity C=D log10 K
no independent evidence
read the original abstract
Dualities play an important role in establishing both microscopic and emergent phenomena in a wide range of physical systems. In practice, though, it can often be computationally challenging to establish when two systems are dual, even when all of the "rules of the game" are well-known. Said differently, when confronted with two systems, how can one efficiently establish that they are in fact dual? In this paper we use machine learning methods to address this question for Seiberg dualities of supersymmetric quiver gauge theories. Mathematically, this involves establishing mutations of quivers, which is in turn a variation on the theme of "learning to unknot". On the one hand, this leads us to a practical tool for establishing the computational complexity of different dualities. On the other hand, it also allows us to study how different network architectures learn how to trace Seiberg dualities. We find that for quivers with a modest number of quiver nodes (of order $10$), different network architectures consisting of transformers and multi-layer perceptrons tend to outperform deterministic algorithms. Supplementing the network by well-established pathfinder algorithms (essentially "Google Maps for quivers") leads to an additional improvement in the efficiency and accuracy of the search strategy. We anticipate that this class of questions can serve as a useful benchmark for frontier AI models applied to theoretical physics.
Figures
Reference graph
Works this paper leans on
-
[1]
Electric - Magnetic Duality in Supersymmetric Non-Abelian Gauge Theories,
N. Seiberg, “Electric - Magnetic Duality in Supersymmetric Non-Abelian Gauge Theories,”Nucl. Phys. B435(1995) 129–146,arXiv:hep-th/9411149
Pith/arXiv arXiv 1995
-
[2]
Cluster Algebras I: Foundations,
S. Fomin and A. Zelevinsky, “Cluster Algebras I: Foundations,”arXiv:math/0104151
-
[3]
Cluster Algebras II: Finite Type Classification,
S. Fomin and A. Zelevinsky, “Cluster Algebras II: Finite Type Classification,”Invent. Math.154no. 1, (2003) 63–121,arXiv:math/0208229
Pith/arXiv arXiv 2003
-
[4]
I. R. Klebanov and M. J. Strassler, “Supergravity and a Confining Gauge Theory: Duality Cascades andχSB Resolution of Naked Singularities,”JHEP08(2000) 052, arXiv:hep-th/0007191
Pith/arXiv arXiv 2000
-
[5]
Superconformal Field Theory on Three-Branes at a Calabi-Yau Singularity,
I. R. Klebanov and E. Witten, “Superconformal Field Theory on Three-Branes at a Calabi-Yau Singularity,”Nucl. Phys. B536(1998) 199–218,arXiv:hep-th/9807080
Pith/arXiv arXiv 1998
-
[6]
Gravity Duals of Supersymmetric SU(N)×SU(N+M) Gauge Theories,
I. R. Klebanov and A. A. Tseytlin, “Gravity Duals of Supersymmetric SU(N)×SU(N+M) Gauge Theories,”Nucl. Phys. B578(2000) 123–138, arXiv:hep-th/0002159
Pith/arXiv arXiv 2000
-
[7]
Chaotic Duality in String Theory,
S. Franco, Y.-H. He, C. Herzog, and J. Walcher, “Chaotic Duality in String Theory,” Phys. Rev. D70(2004) 046006,arXiv:hep-th/0402120
Pith/arXiv arXiv 2004
-
[8]
Statistical Inference and String Theory,
J. J. Heckman, “Statistical Inference and String Theory,”Int. J. Mod. Phys. A30 no. 26, (2015) 1550160,arXiv:1305.3621 [hep-th]
Pith/arXiv arXiv 2015
-
[9]
Relative Entropy and Proximity of Quantum Field Theories,
V. Balasubramanian, J. J. Heckman, and A. Maloney, “Relative Entropy and Proximity of Quantum Field Theories,”JHEP05(2015) 104,arXiv:1410.6809 [hep-th]
Pith/arXiv arXiv 2015
-
[10]
Misanthropic Entropy and Renormalization as a Communication Channel,
R. Fowler and J. J. Heckman, “Misanthropic Entropy and Renormalization as a Communication Channel,”Int. J. Mod. Phys. A37no. 16, (2022) 2250109, arXiv:2108.02772 [hep-th]
Pith/arXiv arXiv 2022
-
[11]
Higher Symmetries of 5D Orbifold SCFTs,
M. Del Zotto, J. J. Heckman, S. N. Meynet, R. Moscrop, and H. Y. Zhang, “Higher Symmetries of 5D Orbifold SCFTs,”Phys. Rev. D106no. 4, (2022) 046010, arXiv:2201.08372 [hep-th]
Pith/arXiv arXiv 2022
-
[12]
Quiver Approach to Symmetry Theories,
V. Chakrabhavi, M. Cvetiˇ c, J. J. Heckman, and S. Meynet, “Quiver Approach to Symmetry Theories,”arXiv:2605.30354 [hep-th]
-
[13]
Quiver Mutations, Seiberg Duality and Machine Learning,
J. Bao, S. Franco, Y.-H. He, E. Hirst, G. Musiker, and Y. Xiao, “Quiver Mutations, Seiberg Duality and Machine Learning,”Phys. Rev. D102no. 8, (2020) 086013, arXiv:2006.10783 [hep-th]. 69
Pith/arXiv arXiv 2020
-
[14]
BPS spectroscopy with reinforcement learning,
F. Carta, A. Gauntlett, F. Griffin, and Y.-H. He, “BPS spectroscopy with reinforcement learning,”Phys. Lett. B868(2025) 139646,arXiv:2501.14863 [hep-th]
Pith/arXiv arXiv 2025
-
[15]
S. Gukov, J. Halverson, F. Ruehle, and P. Su lkowski, “Learning to Unknot,”Mach. Learn. Sci. Tech.2no. 2, (2021) 025035,arXiv:2010.16263 [math.GT]
Pith/arXiv arXiv 2021
-
[16]
A. Gromov, “Grokking modular arithmetic,”arXiv:2301.02679 [cs.LG]
-
[17]
N. Arkani-Hamed, J. L. Bourjaily, F. Cachazo, A. B. Goncharov, A. Postnikov, and J. Trnka,Grassmannian Geometry of Scattering Amplitudes. Cambridge University Press, 4, 2016.arXiv:1212.5605 [hep-th]
Pith/arXiv arXiv 2016
-
[18]
Lectures on D-branes, gauge theories and Calabi-Yau singularities,
Y.-H. He, “Lectures on D-branes, gauge theories and Calabi-Yau singularities,” in1st Hangzhou-Beijing International Summer School. 8, 2004.arXiv:hep-th/0408142
Pith/arXiv arXiv 2004
-
[19]
D-branes on Calabi-Yau manifolds,
P. S. Aspinwall, “D-branes on Calabi-Yau manifolds,” inTheoretical Advanced Study Institute in Elementary Particle Physics (TASI 2003): Recent Trends in String Theory, pp. 1–152. 3, 2004.arXiv:hep-th/0403166
Pith/arXiv arXiv 2003
-
[20]
A Comprehensive Survey of Brane Tilings,
S. Franco, Y.-H. He, C. Sun, and Y. Xiao, “A Comprehensive Survey of Brane Tilings,” Int. J. Mod. Phys. A32no. 23n24, (2017) 1750142,arXiv:1702.03958 [hep-th]
Pith/arXiv arXiv 2017
-
[21]
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention Is All You Need,”arXiv e-prints(June, 2017) arXiv:1706.03762,arXiv:1706.03762 [cs.CL]
Pith/arXiv arXiv 2017
-
[22]
A Formal Basis for the Heuristic Determination of Minimum Cost Paths,
P. E. Hart, N. J. Nilsson, and B. Raphael, “A Formal Basis for the Heuristic Determination of Minimum Cost Paths,”IEEE Transactions on Systems Science and Cybernetics4no. 2, (1968) 100–107
1968
-
[23]
B. T. Lowerre,The Harpy speech recognition system. PhD thesis, Carnegie Mellon University, Pennsylvania, Apr., 1976
1976
-
[24]
Proficient Graph Neural Network Design by Accumulating Knowledge on Large Language Models,
J. Wang, H. Liu, S. Di, Z. Wang, J. Wang, L. Chen, and X. Zhou, “Proficient Graph Neural Network Design by Accumulating Knowledge on Large Language Models,” arXiv e-prints(Aug., 2024) arXiv:2408.06717,arXiv:2408.06717 [stat.ML]
arXiv 2024
-
[25]
DualityCert: Verifier-Gated Language-Model Repair of Broken Duality Claims in Quantum Field Theory,
X. Yu, “DualityCert: Verifier-Gated Language-Model Repair of Broken Duality Claims in Quantum Field Theory,”arXiv:2607.23614 [cs.CR]
-
[26]
Dynamical Supersymmetry Breaking in Supersymmetric QCD,
I. Affleck, M. Dine, and N. Seiberg, “Dynamical Supersymmetry Breaking in Supersymmetric QCD,”Nucl. Phys. B241(1984) 493–534
1984
-
[27]
K. A. Intriligator and N. Seiberg, “The Runaway quiver,”JHEP02(2006) 031, arXiv:hep-th/0512347. 70
Pith/arXiv arXiv 2006
-
[28]
Seiberg Duality for Quiver Gauge Theories,
D. Berenstein and M. R. Douglas, “Seiberg Duality for Quiver Gauge Theories,” arXiv:hep-th/0207027
-
[29]
Exceptional Collections and del Pezzo Gauge Theories,
C. P. Herzog, “Exceptional Collections and del Pezzo Gauge Theories,”JHEP04 (2004) 069,arXiv:hep-th/0310262
Pith/arXiv arXiv 2004
-
[30]
Seiberg Duality is an Exceptional Mutation,
C. P. Herzog, “Seiberg Duality is an Exceptional Mutation,”JHEP08(2004) 064, arXiv:hep-th/0405118
Pith/arXiv arXiv 2004
-
[31]
D-Branes on Vanishing del Pezzo Surfaces,
P. S. Aspinwall and I. V. Melnikov, “D-Branes on Vanishing del Pezzo Surfaces,” JHEP12(2004) 042,arXiv:hep-th/0405134
Pith/arXiv arXiv 2004
-
[32]
Duality walls, duality trees and fractional branes,
S. Franco, A. Hanany, Y.-H. He, and P. Kazakopoulos, “Duality walls, duality trees and fractional branes,”arXiv:hep-th/0306092
-
[33]
Neural Message Passing for Quantum Chemistry,
J. Gilmer, S. S. Schoenholz, P. F. Riley, O. Vinyals, and G. E. Dahl, “Neural Message Passing for Quantum Chemistry,”arXiv e-prints(Apr., 2017) arXiv:1704.01212, arXiv:1704.01212 [cs.LG]
Pith/arXiv arXiv 2017
-
[34]
Recipe for a General, Powerful, Scalable Graph Transformer,
L. Ramp´ aˇ sek, M. Galkin, V. P. Dwivedi, A. T. Luu, G. Wolf, and D. Beaini, “Recipe for a General, Powerful, Scalable Graph Transformer,”arXiv e-prints(May, 2022) arXiv:2205.12454,arXiv:2205.12454 [cs.LG]
Pith/arXiv arXiv 2022
-
[35]
Attention Is All You Need,
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention Is All You Need,” inAdvances in Neural Information Processing Systems, vol. 30, pp. 5998–6008. Curran Associates, Inc., 2017
2017
-
[36]
Crystal Melting and Black Holes,
J. J. Heckman and C. Vafa, “Crystal Melting and Black Holes,”JHEP09(2007) 011, arXiv:hep-th/0610005
Pith/arXiv arXiv 2007
-
[37]
BPS Quivers and Spectra of Complete N=2 Quantum Field Theories,
M. Alim, S. Cecotti, C. Cordova, S. Espahbodi, A. Rastogi, and C. Vafa, “BPS Quivers and Spectra of Complete N=2 Quantum Field Theories,”Commun. Math. Phys.323(2013) 1185–1227,arXiv:1109.4941 [hep-th]
Pith/arXiv arXiv 2013
-
[38]
Transforming Calabi-Yau Constructions: Generating New Calabi-Yau Manifolds with Transformers,
J. H. T. Yip, C. Arnal, F. Charton, and G. Shiu, “Transforming Calabi-Yau Constructions: Generating New Calabi-Yau Manifolds with Transformers,” arXiv:2507.03732 [hep-th]
-
[39]
Sampling string vacua using generative models,
M. Walden and M. Larfors, “Sampling string vacua using generative models,”Mach. Learn. Sci. Tech.7no. 1, (2026) 015018,arXiv:2509.16029 [hep-th]
arXiv 2026
-
[40]
Generating Special Triangulations with Transformers,
C. Arnal, J. H. T. Yip, F. Charton, and G. Shiu, “Generating Special Triangulations with Transformers,” in . 6, 2026.arXiv:2606.26660 [hep-th]
Pith/arXiv arXiv 2026
-
[41]
Exploring Line Bundle Standard Models with Transformers,
J. H. T. Yip, A. Mininno, and G. Shiu, “Exploring Line Bundle Standard Models with Transformers,”arXiv:2607.00078 [hep-th]. 71
-
[42]
Reconstructing conformal field theoretical compositions with Transformers,
H. Cao, G. Merz, K. Cranmer, and G. Shiu, “Reconstructing conformal field theoretical compositions with Transformers,”arXiv:2605.01072 [hep-th]
-
[43]
Holographic Complexity Equals Bulk Action?,
A. R. Brown, D. A. Roberts, L. Susskind, B. Swingle, and Y. Zhao, “Holographic Complexity Equals Bulk Action?,”Phys. Rev. Lett.116no. 19, (2016) 191301, arXiv:1509.07876 [hep-th]
Pith/arXiv arXiv 2016
-
[44]
Comments on Holographic Complexity,
D. Carmi, R. C. Myers, and P. Rath, “Comments on Holographic Complexity,”JHEP 03(2017) 118,arXiv:1612.00433 [hep-th]
Pith/arXiv arXiv 2017
-
[45]
Quantum Complexity of Time Evolution with Chaotic Hamiltonians,
V. Balasubramanian, M. Decross, A. Kar, and O. Parrikar, “Quantum Complexity of Time Evolution with Chaotic Hamiltonians,”JHEP01(2020) 134,arXiv:1905.05765 [hep-th]
Pith/arXiv arXiv 2020
-
[46]
Generalized Complexity Distances and Non-Invertible Symmetries,
J. J. Heckman, R. J. Hicks, and C. Murdia, “Generalized Complexity Distances and Non-Invertible Symmetries,”arXiv:2604.14275 [hep-th]
-
[47]
Finite Heisenberg groups from nonAbelian orbifold quiver gauge theories,
B. A. Burrington, J. T. Liu, and L. A. Pando Zayas, “Finite Heisenberg groups from nonAbelian orbifold quiver gauge theories,”Nucl. Phys. B794(2008) 324–347, arXiv:hep-th/0701028
Pith/arXiv arXiv 2008
-
[48]
Vafa–Witten Invariants from Exceptional Collections,
G. Beaujard, J. Manschot, and B. Pioline, “Vafa–Witten Invariants from Exceptional Collections,”Commun. Math. Phys.385no. 1, (2021) 101–226,arXiv:2004.14466 [hep-th]
Pith/arXiv arXiv 2021
-
[49]
Layer Normalization,
J. Lei Ba, J. R. Kiros, and G. E. Hinton, “Layer Normalization,”arXiv e-prints(July,
-
[50]
Rectifier Nonlinearities Improve Neural Network Acoustic Models,
A. L. Maas, “Rectifier Nonlinearities Improve Neural Network Acoustic Models,” in . 2013
2013
-
[51]
Demystifying Oversmoothing in Attention-Based Graph Neural Networks,
X. Wu, A. Ajorlou, Z. Wu, and A. Jadbabaie, “Demystifying Oversmoothing in Attention-Based Graph Neural Networks,”arXiv e-prints(May, 2023) arXiv:2305.16102,arXiv:2305.16102 [cs.LG]. 72
Pith/arXiv arXiv 2023
-
[2016]
arXiv:1607.06450,arXiv:1607.06450 [stat.ML]
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.