Pith. sign in

REVIEW 4 major objections 4 minor 53 references

MDTransformer: A Hardware-Software Co-Design of Mode-Division Photonic Transformer Accelerator with Inverse-Designed Coherent Crossbar

T0 review · 4 major / 4 minor · reviewed 2026-08-01 · deepseek-v4-flash

Pith's one-line read MDTransformer claims that a photonic transformer accelerator can drop wavelength-division multiplexing, microcombs, and phase shifters entirely, using four spatial light modes in a single waveguide as independent compute lanes, and still cu

desk verdict Interesting architecture, but the arithmetic gap between the derived 1-bit operation and the claimed 4-bit precision is load-bearing; the savings numbers are not supported. read the letter →

arxiv 2607.26016 v1 pith:EAEPDLCV submitted 2026-07-28 cs.AR cs.AIcs.DC

classification cs.ARcs.AIcs.DC
keywords siliconphotonicsmode-divisionmultiplexingphotonictransformeracceleratorinversedesigncoherentcrossbarIQmodulationdot-productunithardware-softwareco-design
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper is trying to establish a cheaper way to build photonic accelerators for transformer inference. Instead of spreading computation across many wavelengths generated by a power-hungry microcomb, MDTransformer keeps everything on one continuous-wave 1550 nm laser and splits it into four orthogonal spatial modes (TE0–TE3) traveling in the same waveguide, each acting as an independent dot-product lane. The load-bearing piece is a compact inverse-designed coherent mixer, the mode-division dot-product unit, that produces fixed 0/90/180/−90 degree phase relations, letting balanced photodetection compute signed multiplications without any phase shifter in the core. The paper reports 40.4% area reduction, 63.6% power saving, and 40.6% energy saving against a same-size wavelength-based accelerator, with comparable latency across DeiT and BERT workloads, at about 4-bit effective precision. If these numbers hold, single-laser mode-division photonics becomes a practical route to low-power transformer inference.

What carries the argument

The mode-division dot-product unit (MDOT): an inverse-designed 8×8 µm² coherent mixer that takes two mode-encoded operands and outputs four interference states with fixed {0°, 90°, 180°, −90°} phase relationships. Its fixed phase basis means a balanced photodetector reads the signed dot product directly, eliminating the 90° phase shifter that prior coherent designs need. The companion piece is the inverse-designed four-mode MUX/DEMUX, solved by an adjoint-based objective that maximizes each mode's transmission to its own output while penalizing crosstalk to the other three; it turns a 2.5 µm multimode bus into four independent computational lanes at a single wavelength.

What would settle it

Inject all four TE modes at once into the fabricated 2.5 µm multimode bus and measure the field at each MDOT output port over the 1530–1560 nm band. If simultaneous inter-modal crosstalk exceeds −30 dB, or if balanced-detection MAC results at 1550 nm show effective precision below 4 bits, the central claim that mode-division lanes deliver four-fold single-wavelength parallelism at usable precision collapses.

Watch

Extended reading notes

Core claim

The paper's central claim is that spatial-mode interference can replace wavelength-based parallelism in a photonic transformer accelerator. MDTransformer encodes operands onto four guided TE modes in a 2.5 µm multimode bus; an inverse-designed mode MUX/DEMUX, optimized to maximize each mode's transmission to its designated port while penalizing leakage, routes them to an array of mode-division dot-product units (MDOTs). Each MDOT is an 8×8 µm² inverse-designed coherent mixer with output phase relations {0°, 90°, 180°, −90°}, so balanced detection directly yields the signed product of two full-range operands, with no phase shifter required. Four modes give four-fold MAC parallelism per wavegu

Load-bearing premise

The four guided modes must stay cleanly separated when they run at the same time, keeping inter-modal crosstalk below −30 dB and preserving enough modal purity for 4-bit effective multiply precision; the paper's evidence is single-mode-launch simulation plus a per-mode optimization objective, not a simultaneous four-mode measurement.

Editorial extensions

If this is right

  • Transformer inference can be accelerated by a photonic core that needs only one laser: no microcomb, no multi-wavelength generation, and no free-spectral-range limits.
  • Four TE modes in one waveguide provide four independent MAC lanes, so parallelism scales with modal count rather than wavelength count.
  • Because the MDOT's phase basis is fixed and the inverse design enforces 120 nm minimum features, the core is compatible with standard silicon-photonics fabrication and single-laser 1550 nm continuous-wave operation.
  • Across DeiT-Tiny/Small/Base and BERT-Base/Large, the design reports 40.4% lower area, 63.6% lower power, and 40.6% lower energy than the same-size state-of-the-art wavelength-based accelerator, at comparable latency.
  • The architecture is configurable in tiles, cores, and array size, so area, power, energy, and latency can be traded off against workload requirements.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The strongest test of the paper's thesis is simultaneous four-mode operation: the reported simulations and the optimization objective treat each mode independently, so the −30 dB crosstalk and 4-bit precision claims remain predictions until a multi-mode-launch measurement or simulation verifies them.
  • If simultaneous-mode crosstalk holds, the same inverse-designed MUX/DEMUX and MDOT building blocks could be extended to more TE modes or denser cores, multiplying lanes per waveguide without adding lasers—a scaling path the paper hints at but does not demonstrate.
  • Because the energy savings come mainly from removing the microcomb and phase shifters and shrinking the modulator, the inverse-designed coherent-mixer idea could transfer to other photonic linear-algebra accelerators beyond transformers, though each new context would need its own crosstalk and precision validation.
  • A testable extension would combine mode-division lanes with a small number of wavelengths for multiplicative parallelism, but only if MUX/DEMUX crosstalk stays below the precision target across the combined channel set.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper proposes MDTransformer, a photonic transformer accelerator built around spatial mode-division multiplexing instead of wavelength-division multiplexing. Four TE modes in a multimode waveguide are intended to act as independent computational lanes; inverse-designed MUX/DEMUX, a coherent MDOT mixer, and IQ modulators are combined into a Mode-Division Photonic Tensor Core (MPTC). The authors claim sub-4-bit/4-bit effective precision, below -30 dB inter-modal crosstalk, single-laser 1550 nm operation, and area/power/energy savings over the LT accelerator family, evaluated with Tidy3D and the LT hardware simulator.

Significance. If validated, replacing WDM lanes with spatial modes would be a meaningful architectural step: it could remove the microcomb, reduce modulator/phase-shifter overhead, and exploit compact inverse-designed components. The paper is transparent that the design is under fabrication, and it reports Tidy3D simulations of the optical primitives. However, the validation is currently insufficient: the arithmetic derivation does not support the stated precision, the simultaneous-mode crosstalk claim is not simulated, and the system-level numbers inherit unverified device parameters and the unsupported precision assumption.

major comments (4)
  1. [§III-B, Eqs. (5)–(7)] The MDOT derivation supports only 1-bit signed MAC. The text after Eq. (5) fixes s_x^(m), s_y^(m) ∈ {−1,+1} and φ_x=πp, φ_y=πq (p,q∈{0,1}); Eq. (6) then gives I_PC ∝ ∫ s_x s_y cos(Δφ) dt, i.e., a product of ±1 signs. No amplitude term from the IQ-modulator transfer functions Eqs. (9)–(11) enters the dot-product expression, and no bit-slicing or temporal summation over multiple bits is described. Therefore the derived operation is not complex-valued, full-range, 4-bit multiplication. The workload-level energy/latency results in §V assume a 4-bit-capable core and are thus unsupported.
  2. [§III-C, Eq. (8), Fig. 6] The below −30 dB inter-modal crosstalk and four-lane parallelism claims are not established. Eq. (8) is a sum of per-mode transmission terms with an unspecified penalty α, and Fig. 6 shows four single-mode-launch field panels. No simulation launches all four TE modes simultaneously and reports the output mode-purity/crosstalk matrix; Fig. 5 for the MDOT is likewise single-mode. Simultaneous-mode operation is exactly where inter-modal crosstalk and modal-purity degradation appear, so the crosstalk and parallelism claims are unverified.
  3. [§IV–V, Table I] The headline savings (40.4% area, 63.6% power, 40.6% energy vs LT-Custom) are computed with the LT hardware simulator using device parameters supplied by the authors, e.g., 6.2 dB IL/64 µm² for the 4-mode MUX/DEMUX and 6 dB IL/144 µm² for the coherent hybrid. These values are not extracted from the Tidy3D results or from measurements, and the simulator contains no accuracy model of the MDOT arithmetic. The evaluation therefore inherits both the unverified precision assumption and the unverified component parameters; the comparison to LT-Custom is not yet quantitative evidence for the claimed savings.
  4. [Abstract and §III-B] The claim of 'complex-valued arithmetic for full-range operations' is not supported by the derived photocurrent in Eqs. (6)–(7), which is real-valued; Eq. (4) merely lists hybrid output states. No complex MAC over arbitrary complex operands is derived. The precision statement is also inconsistent: the Abstract says 'sub-4-bit effective precision' while §I-C and the Conclusion say '4-bit effective precision.'
minor comments (4)
  1. [Abstract, §I-C, Conclusion] Reconcile the precision terminology: 'sub-4-bit effective precision' in the Abstract versus '4-bit effective precision' in §I-C and the Conclusion.
  2. [Figs. 5–6] Label axes and units fully; specify the wavelength ranges and the quantity plotted in the CMRR and phase-difference panels.
  3. [Eq. (4)] The output-state list contains repeated expressions; clarify the mapping to the stated {0°, 90°, 180°, −90°} phase basis and how the four outputs are used in balanced detection.
  4. [Table I and §IV] Explain how the DAC (8-bit) and ADC (6-bit) precisions relate to the claimed 4-bit effective MAC, and justify the MUX/DEMUX and coherent-hybrid insertion-loss/area values from simulation or literature rather than asserting them.

Circularity Check

0 steps flagged · score 1.0 of 10

No circular derivation; precision/crosstalk claims are unsupported evidence gaps, not circular steps.

full rationale

The central derivation chain is not circular. The MDOT arithmetic in Eqs. 4-7 is proposed as the device transfer function: the balanced photocurrent is proportional to the integral of s_x(t)s_y(t)cos(delta_phi(t)) dt, which by construction equals the intended signed dot product. This is the design equation of the device, not a prediction obtained by fitting or by renaming an input. The inverse-designed MUX/DEMUX is optimized with Eq. 8, which includes a crosstalk penalty; the reported below -30 dB crosstalk is a simulated outcome of that optimization, and the design is also benchmarked against reported inverse-designed devices. The hardware results are produced with the external LT simulator from [18], using MDTransformer's own configuration and component parameters; the reported area/power/energy savings are not the simulator's inputs renamed as outputs. Self-citations are present ([53] for a TIA parameter, [48]-[50] for SRAM sizing) but none is load-bearing; [53] is a measured receiver result and [48]-[50] are general memory-design references. The main weakness, the claim of 4-bit effective precision from equations that only show bipolar +/-1 multiplication, is an unsupported-evidence/correctness issue, not a circular-reasoning issue. Similarly, the unverified simultaneous-mode crosstalk is a validation gap, not a construction that reduces the result to its own inputs. Therefore no circular step is exhibited.

Assumptions & free parameters 4 free parameters · 4 assumptions · 2 invented entities

The central claims depend on several categories of unverified inputs: (i) simulation-based device parameters for the new inverse-designed components (MUX/DEMUX, coherent hybrid, crossings), (ii) the assumption that four TE modes are orthogonal enough to act as independent lanes in a dense crossbar, (iii) the validity of the LT hardware simulator for a mode-division architecture, and (iv) the existence of a multi-bit/complex-arithmetic mapping that the paper does not actually derive. These are not standard physical constants; they are domain assumptions that a fabricated chip must confirm.

free parameters (4)
  • Crosstalk penalty factor alpha (Eq. 8) = not reported
    Ad hoc weighting controlling crosstalk suppression in the inverse-design objective; no value or sensitivity study is given.
  • MUX/DEMUX insertion loss and area = 6.2 dB, 64 um^2
    Simulated values for the inverse-designed 4-mode MUX/DEMUX entered into the hardware simulator; no measured data.
  • Coherent hybrid insertion loss and area = 6 dB, 144 um^2
    Device parameters from Tidy3D simulation, used to compute area/power; no measurement.
  • MZI modulator calibration coefficients A, B, D, E (Eqs. 9-10) = unspecified
    Called empirical calibration coefficients but values are not given; they set the modulator transfer model used for power evaluation.
assumptions (4)
  • domain assumption Guided TE0-TE3 modes in a 2.5 um Si waveguide are orthogonal and can be independently excited/detected with negligible inter-modal coupling across the MPTC.
    Invoked in Section III-A/III-C to justify independent computational lanes; only single-mode simulations support it.
  • domain assumption The LT PTA hardware simulator [18] provides valid area/power/energy models for MDTransformer's architecture when supplied with the Table I parameters.
    All system-level results in Section V use this tool; the simulator was built for the WDM LT design, and its treatment of mode-division components is not independently validated.
  • domain assumption Tidy3D FDTD simulations accurately predict the fabricated performance of the inverse-designed structures.
    Device functionality and crosstalk are established solely through Tidy3D; no fabricated measurements are presented.
  • domain assumption Fabrication constraints (120 nm minimum feature size, e-beam lithography) are sufficient to realize the simulated behavior.
    The design is said to be under fabrication, but no post-fabrication measurements or process-variation analysis is included.
invented entities (2)
  • MDOT (Mode-Division Dot-Product Unit)
    purpose: Coherent mixer producing four interference outputs to perform signed multiplication of two mode-encoded operands.
    Introduced by this paper; only Tidy3D field/phase/CMRR simulations are shown, not measured characterization.
  • 4-mode inverse-designed MUX/DEMUX as a computational routing primitive
    purpose: Injects and recovers four TE modes into the MPTC crossbar.
    Simulated field plots only; no measured crosstalk matrix under simultaneous excitation is provided.

how reviews work

0 comments
Cite this review

Pith. "Pith review of MDTransformer: A Hardware-Software Co-Design of Mode-Division Photonic Transformer Accelerator with Inverse-Designed Coherent Crossbar." pith.science (2026). https://pith.science/paper/EAEPDLCV

@misc{pith2026260726016,
  author       = {Pith},
  title        = {Pith review of: MDTransformer: A Hardware-Software Co-Design of Mode-Division Photonic Transformer Accelerator with Inverse-Designed Coherent Crossbar},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/EAEPDLCV}},
  note         = {Machine review of arXiv:2607.26016}
}
read the original abstract

Recently, photonic transformer accelerators (PTAs) have successfully achieved significant speedup and energy efficiency improvements over electronic accelerators for expediting Transformer inference. However, state-of-the-art rely on expensive multi-wavelength light generation and large dot-product units due to active phase-shifter components, thus making their approach inefficient and impractical. To address this, we propose MDTransformer, a novel hardware-software co-design of PTA based on mode-division optical dataflow and operations. Specifically, MDTransformer performs complex matrix operations using spatial-mode interference, that leverages the inverse-designed multi-mode couplers, crossings, and Mach-Zehnder IQ modulators into a compact mode-division photonic tensor core (MPTC), capable of executing matrix multiplications in the optical domain. Its each guided mode (i.e., TE0-TE3) acts as an independent computational lane, enabling four-fold parallelism-per-waveguide without spectral filtering or free-spectral-range limitations. Moreover, its coherent detection and IQ modulation jointly encode amplitude and phase, realizing complex-valued arithmetic for full-range operations in transformers. MDTransformer offers analog multiplication with sub-4-bit effective precision and inter-modal crosstalk below -30 dB. Its inverse-designed approach also offers scalable and full compatibility with single-laser continuous-wave operation at 1550 nm. Experimental results show that MDTransformer achieves 40.4% area reduction, 63.6% power saving, 40.6% energy saving, and comparable latency over the state-of-the-art PTA across different workloads (i.e., DeiT-Tiny/Small/Base and BERT-Base/Large). These results show that MDTransformer offers a practical solution for high-performance and energy-efficient transformer-based systems.

Figures

Figures reproduced from arXiv: 2607.26016 by the authors.

Figure 1
Figure 1. (a) Transformer networks typically improve their performance at the cost of larger memory footprint; based on data from [6]. (b) The state-of-the￾art photonic tensor core for Transformer acceleration based on dynamically￾operated dot-product (DDot) unit from the LT accelerator [18]. diverse machine learning (ML) tasks, e.g., natural language processing (NLP) and computer vision [2]–[4], thereby paving the way toward… view at source ↗
Figure 2
Figure 2. Area breakdown of the state-of-the-art 4-bit LT accelerators: [PITH_FULL_IMAGE:figures/full_fig_p002_2.png] view at source ↗
Figure 3
Figure 3. Our proposed MDTransformer Accelerator: (a) MPTC with crossings, modulator, and MDOT; (b) architecture design. uation (e.g., area, power, energy, and latency) using the state￾of-the-art PTA hardware simulator from [18]. Furthermore, our design is also under fabrication. Experimental results show that, MDTransformer offers 4-bit effective precision for multiplication, low inter-modal crosstalk (i.e., below -30 dB), a… view at source ↗
Figures from the paper (7 more)
Figure 4
Figure 4. Figure 4: Inverse-designed MDOT design: (a) silicon projection of the design region. (b) Intensity distribution of light traveling through the structure from the left input. (a) (b) [PITH_FULL_IMAGE:figures/full_fig_p004_4.png]
Figure 5
Figure 5. Figure 5: MDOT properties: (a) simulated relative phase differences between the four output ports; and (b) CMRR. coherent mixer that produces four output interference states with fixed phase relationships. Each of the four mixer outputs can be described using a linear transforma…
Figure 6
Figure 6. Figure 6: Inverse-designed mode-based MUX/DEMUX: optical field distributions for the four input modes, showing selective routing to distinct single-mode [PITH_FULL_IMAGE:figures/full_fig_p005_6.png]
Figure 7
Figure 7. Figure 7: Cross-sectional optical field distributions for [PITH_FULL_IMAGE:figures/full_fig_p006_7.png]
Figure 8
Figure 8. Figure 8: Dataflow based on data tiling mechanism for [PITH_FULL_IMAGE:figures/full_fig_p007_8.png]
Figure 9
Figure 9. Figure 9: (a) Experimental setup and tools flow in this work. (b) Measurement setup for testing the fabricated chip. TABLE I SUMMARY OF DEVICE PARAMETERS Device Parameter Value DAC [51] Precision 8-bit Power 42 mW (@28 GSPS) Area 0.03 mm2 ADC [52] Precision 32 lines, 6-bit Power…
Figure 10
Figure 10. Figure 10: Experimental results for (a) area breakdown of MDTransformer; (b) power breakdown of MDTransformer; (c) comparison on area; (d) comparison on power; as well as energy consumption and latency for different workloads: (e.1) DeiT-T, (e.2) DeiT-S, (e.3) DeiT-B; (e.4) BERT…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

53 extracted references · 2 linked inside Pith

  1. [1]

    Attention is all you need,

    A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez et al., “Attention is all you need,”Advances in Neural Information Processing Systems (NIPS), vol. 30, no. 1, pp. 261–272, 2017

  2. [2]

    An image is worth 16x16 words: Trans- formers for image recognition at scale,

    A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An image is worth 16x16 words: Trans- formers for image recognition at scale,” inInternational Conference on Learning Representations (ICLR), 2021

  3. [3]

    Training data-efficient image transformers & distillation through attention,

    H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, and H. J ´egou, “Training data-efficient image transformers & distillation through attention,” inInternational Conference on Machine Learning (ICML). PMLR, 2021, pp. 10 347–10 357

  4. [4]

    Transformers in vision: A survey,

    S. Khan, M. Naseer, M. Hayat, S. W. Zamir, F. S. Khan, and M. Shah, “Transformers in vision: A survey,”ACM Computing Surveys (CSUR), vol. 54, no. 10s, pp. 1–41, 2022

  5. [5]

    Artificial general intelligence: Advancements, challenges, and future directions in agi research,

    G. Yenduri, R. Murugan, P. Kumar Reddy Maddikunta, S. Bhattacharya, D. Sudheer, and B. Bhushan Savarala, “Artificial general intelligence: Advancements, challenges, and future directions in agi research,”IEEE Access, vol. 13, pp. 134 325–134 356, 2025

  6. [6]

    A survey on vision transformer,

    K. Han, Y . Wang, H. Chen, X. Chen, J. Guo, Z. Liu, Y . Tang, A. Xiao, C. Xu, Y . Xu, Z. Yang, Y . Zhang, and D. Tao, “A survey on vision transformer,”IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), vol. 45, no. 1, pp. 87–110, 2023

  7. [7]

    Spatten: Efficient sparse attention architecture with cascade token and head pruning,

    H. Wang, Z. Zhang, and S. Han, “Spatten: Efficient sparse attention architecture with cascade token and head pruning,” in2021 IEEE International Symposium on High-Performance Computer Architecture (HPCA). IEEE, 2021, pp. 97–110

  8. [8]

    Transpim: A memory- based acceleration via software-hardware co-design for transformer,

    M. Zhou, W. Xu, J. Kang, and T. Rosing, “Transpim: A memory- based acceleration via software-hardware co-design for transformer,” in 2022 IEEE International Symposium on High-Performance Computer Architecture (HPCA). IEEE, 2022, pp. 1071–1085

Show all 53 references
  1. [9]

    Vitcod: Vision transformer acceleration via dedicated algorithm and accelerator co-design,

    H. You, Z. Sun, H. Shi, Z. Yu, Y . Zhao, Y . Zhang, C. Li, B. Li, and Y . Lin, “Vitcod: Vision transformer acceleration via dedicated algorithm and accelerator co-design,” in2023 IEEE International Symposium on High-Performance Computer Architecture (HPCA). IEEE, 2023, pp. 273–286

  2. [10]

    A survey on silicon photonics for deep learning,

    F. P. Sunny, E. Taheri, M. Nikdast, and S. Pasricha, “A survey on silicon photonics for deep learning,”ACM Journal of Emerging Technologies in Computing System (JETC), vol. 17, no. 4, pp. 1–57, 2021

  3. [11]

    Albireo: Energy- efficient acceleration of convolutional neural networks via silicon pho- tonics,

    K. Shiflett, A. Karanth, R. Bunescu, and A. Louri, “Albireo: Energy- efficient acceleration of convolutional neural networks via silicon pho- tonics,” in2021 ACM/IEEE 48th Annual International Symposium on Computer Architecture (ISCA). IEEE, 2021, pp. 860–873

  4. [12]

    Photonics for artificial intelligence and neuromor- phic computing,

    B. J. Shastriet al., “Photonics for artificial intelligence and neuromor- phic computing,”Nature Photonics, vol. 15, no. 2, 2021

  5. [13]

    Simphony: A device-circuit-architecture cross-layer modeling and simulation frame- work for heterogeneous electronic-photonic ai system,

    Z. Yin, M. Zhang, N. Gangi, R. Huang, J. Zhang, and J. Gu, “Simphony: A device-circuit-architecture cross-layer modeling and simulation frame- work for heterogeneous electronic-photonic ai system,” in2025 62nd ACM/IEEE Design Automation Conference (DAC). IEEE, 2025, pp. 1–7

  6. [14]

    Neuromorphic photonic networks using silicon photonic weight banks,

    A. N. Tait, T. F. De Lima, E. Zhou, A. X. Wu, M. A. Nahmias, B. J. Shastri, and P. R. Prucnal, “Neuromorphic photonic networks using silicon photonic weight banks,”Scientific Reports, vol. 7, no. 1, p. 7430, 2017

  7. [15]

    Crosslight: A cross- layer optimized silicon photonic neural network accelerator,

    F. Sunny, A. Mirza, M. Nikdast, and S. Pasricha, “Crosslight: A cross- layer optimized silicon photonic neural network accelerator,” in2021 58th ACM/IEEE design automation conference (DAC). IEEE, 2021, pp. 1069–1074

  8. [16]

    Deep learning with coherent nanophotonic circuits,

    Y . Shen, N. C. Harris, S. Skirlo, M. Prabhu, T. Baehr-Jones, M. Hochberg, X. Sun, S. Zhao, H. Larochelle, D. Englundet al., “Deep learning with coherent nanophotonic circuits,”Nature photonics, vol. 11, no. 7, pp. 441–446, 2017

  9. [17]

    Parallel convolutional processing using an integrated photonic tensor core,

    J. Feldmann, N. Youngblood, M. Karpov, H. Gehring, X. Li, M. Stap- pers, M. Le Gallo, X. Fu, A. Lukashchuk, A. S. Rajaet al., “Parallel convolutional processing using an integrated photonic tensor core,” Nature, vol. 589, no. 7840, pp. 52–58, 2021

  10. [18]

    Lightening-transformer: A dynamically- operated optically-interconnected photonic transformer accelerator,

    H. Zhu, J. Gu, H. Wang, Z. Jiang, Z. Zhang, R. Tang, C. Feng, S. Han, R. T. Chen, and D. Z. Pan, “Lightening-transformer: A dynamically- operated optically-interconnected photonic transformer accelerator,” in 2024 IEEE International Symposium on High-Performance Computer Archi...

  11. [19]

    Sprint: A high-performance, energy- efficient, and scalable chiplet-based accelerator with photonic intercon- nects for cnn inference,

    Y . Li, A. Louri, and A. Karanth, “Sprint: A high-performance, energy- efficient, and scalable chiplet-based accelerator with photonic intercon- nects for cnn inference,”IEEE Transactions on Parallel and Distributed Systems (TPDS), vol. 33, no. 10, pp. 2332–2345, 2022

  12. [20]

    Spacx: Silicon photonics-based scalable chiplet accelerator for dnn inference,

    ——, “Spacx: Silicon photonics-based scalable chiplet accelerator for dnn inference,” in2022 IEEE International Symposium on High- Performance Computer Architecture (HPCA), 2022, pp. 831–845

  13. [21]

    Tron: Transformer neural network acceleration with non-coherent silicon photonics,

    S. Afifi, F. Sunny, M. Nikdast, and S. Pasricha, “Tron: Transformer neural network acceleration with non-coherent silicon photonics,” in Great Lakes Symposium on VLSI (GSVLSI) 2023, 2023, pp. 15–21

  14. [22]

    A light-speed large language model accelerator with optical stochastic computing,

    S. Afifi, O. Alo, I. Thakkar, and S. Pasricha, “A light-speed large language model accelerator with optical stochastic computing,” inGreat Lakes Symposium on VLSI (GLSVLSI) 2025, 2025

  15. [23]

    Astra: A stochastic transformer neural network accelerator with silicon photonics,

    ——, “Astra: A stochastic transformer neural network accelerator with silicon photonics,”ACM Transactions on Embedded Computing Systems (TECS), 2025

  16. [24]

    Merit: A sustainable dnn accelerator design with photonic phase-change memory,

    Y . Li, A. Louri, and A. Karanth, “Merit: A sustainable dnn accelerator design with photonic phase-change memory,”IEEE Transactions on Sustainable Computing (TSUSC), vol. 10, no. 4, pp. 705–716, 2025

  17. [25]

    En- lighten: Lighten the transformer, enable efficient optical acceleration,

    H. Zhu, Z. Zhou, S. Ning, X. Wu, R. Chen, Y . Wan, and D. Pan, “En- lighten: Lighten the transformer, enable efficient optical acceleration,” arXiv preprint arXiv:2510.01673, 2025

  18. [26]

    Hyatten: Hybrid photonic-digital architec- ture for accelerating attention mechanism,

    H. Li, D. Chen, and T. Mitra, “Hyatten: Hybrid photonic-digital architec- ture for accelerating attention mechanism,” in2025 Design, Automation & Test in Europe Conference (DATE), 2025, pp. 1–7

  19. [27]

    P-dac: Power-efficient photonic accelerators for llm inference,

    W.-T. Chang, C.-F. Wu, and Y .-C. Lo, “P-dac: Power-efficient photonic accelerators for llm inference,” in2025 62nd ACM/IEEE Design Au- tomation Conference (DAC), 2025, pp. 1–7

  20. [28]

    Netcast: low-power edge computing with wdm- defined optical neural networks,

    R. Hamerly, A. Sludds, S. Bandyopadhyay, Z. Chen, Z. Zhong, L. Bern- stein, and D. Englund, “Netcast: low-power edge computing with wdm- defined optical neural networks,”Journal of Lightwave Technology, vol. 42, no. 22, pp. 7795–7806, 2024

  21. [29]

    Hybrid photonic-digital accelerator for attention mechanism,

    H. Li, D. Chen, and T. Mitra, “Hybrid photonic-digital accelerator for attention mechanism,”arXiv preprint arXiv:2501.11286, 2025

  22. [30]

    Photonic neural networks based on integrated silicon microresonators,

    S. Biasi, G. Donati, A. Lugnan, M. Mancinelli, E. Staffoli, and L. Pavesi, “Photonic neural networks based on integrated silicon microresonators,” Intelligent Computing, vol. 3, p. 0067, 2024

  23. [31]

    Microring weight banks,

    A. N. Tait, A. X. Wu, T. F. De Lima, E. Zhou, B. J. Shastri, M. A. Nahmias, and P. R. Prucnal, “Microring weight banks,”IEEE Journal of Selected Topics in Quantum Electronics, vol. 22, no. 6, pp. 312–325, 2016

  24. [32]

    Characterizing coherent integrated photonic neural networks under imperfections,

    S. Banerjee, M. Nikdast, and K. Chakrabarty, “Characterizing coherent integrated photonic neural networks under imperfections,”Journal of lightwave technology, vol. 41, no. 5, pp. 1464–1479, 2022. 10

  25. [33]

    Programmable photonic neural networks combining wdm with coherent linear optics,

    A. Totovic, G. Giamougiannis, A. Tsakyridis, D. Lazovsky, and N. Pleros, “Programmable photonic neural networks combining wdm with coherent linear optics,”Scientific reports, vol. 12, no. 1, p. 5605, 2022

  26. [34]

    Tidy3D: Next-generation electromagnetic simula- tion tool,

    I. Flexcompute, “Tidy3D: Next-generation electromagnetic simula- tion tool,” https://www.flexcompute.com/tidy3d/solver/, 2024, accessed: 2025-01-01

  27. [35]

    Inverse design of compact multimode multi-port photonic devices,

    N. V . e. a. Sapra, “Inverse design of compact multimode multi-port photonic devices,”Nature Communications, vol. 11, p. 6361, 2020

  28. [36]

    Adjoint-based inverse design of efficient, broadband mode conversion devices,

    J. S. e. a. Jensen, “Adjoint-based inverse design of efficient, broadband mode conversion devices,”ACS Photonics, vol. 7, pp. 1497–1506, 2020

  29. [37]

    Compact broadband directional couplers using inverse design,

    D. e. a. Vercruysse, “Compact broadband directional couplers using inverse design,”Optica, vol. 7, pp. 179–185, 2020

  30. [38]

    Mode division multiplexing on an inp membrane on silicon,

    Y . Wang, Y . Wei, V . Dolores-Calzadilla, K. Williams, M. Smit, D. Dai, and Y . Jiao, “Mode division multiplexing on an inp membrane on silicon,”Optics Letters, vol. 47, no. 16, pp. 4004–4007, 2022

  31. [39]

    Arbitrarily routed mode-division multiplexed photonic circuits for dense integration,

    Y . Liu, K. Xu, S. Wang, W. Shen, H. Xie, Y . Wang, S. Xiao, Y . Yao, J. Du, Z. Heet al., “Arbitrarily routed mode-division multiplexed photonic circuits for dense integration,”Nature communications, vol. 10, no. 1, p. 3263, 2019

  32. [40]

    Dynamic electro-optic analog memory for neuromorphic photonic computing,

    S. Lam, A. Khaled, S. Bilodeau, B. A. Marquez, P. R. Prucnal, L. Chrostowski, B. J. Shastri, and S. Shekhar, “Dynamic electro-optic analog memory for neuromorphic photonic computing,”arXiv preprint arXiv:2401.16515, 2024

  33. [41]

    Photonic-electronic integrated circuits for high-performance computing and ai accelerators,

    S. Ning, H. Zhu, C. Feng, J. Gu, Z. Jiang, Z. Ying, J. Midkiff, S. Jain, M. H. Hlaing, D. Z. Panet al., “Photonic-electronic integrated circuits for high-performance computing and ai accelerators,”Journal of Lightwave Technology, 2024

  34. [42]

    Temporal analog optical computing using an on-chip fully reconfigurable photonic signal processor,

    H. Babashah, Z. Kavehvash, A. Khavasi, and S. Koohi, “Temporal analog optical computing using an on-chip fully reconfigurable photonic signal processor,”Optics & Laser Technology, vol. 111, pp. 66–74, 2019

  35. [43]

    Wdm-compatible mode-division multi- plexing on a silicon chip,

    L.-W. Luo, N. Ophir, C. P. Chen, L. H. Gabrielli, C. B. Poitras, K. Bergmen, and M. Lipson, “Wdm-compatible mode-division multi- plexing on a silicon chip,”Nature communications, vol. 5, no. 1, p. 3069, 2014

  36. [44]

    Multi- dimensional data transmission using inverse-designed silicon photonics and microcombs,

    K. Y . Yang, C. Shirpurkar, A. D. White, J. Zang, L. Chang, F. Ashtiani, M. A. Guidry, D. M. Lukin, S. V . Pericherla, J. Yanget al., “Multi- dimensional data transmission using inverse-designed silicon photonics and microcombs,”Nature communications, vol. 13, no. 1, p. 7862, 2022

  37. [45]

    Integrated silicon nitride devices via inverse design,

    J. L. Pita Ruiz, N. Dalvand, and M. M ´enard, “Integrated silicon nitride devices via inverse design,”Nature Communications, vol. 16, no. 1, p. 9307, 2025

  38. [46]

    Ultra-compact scalable mode demultiplexers for high-speed optical interconnects via gpu-accelerated inverse design,

    J. Li, X. Li, L. Wu, M. Luo, Y . Li, Y . Wang, and Y . Qiu, “Ultra-compact scalable mode demultiplexers for high-speed optical interconnects via gpu-accelerated inverse design,”Optics Express, vol. 33, no. 21, pp. 44 908–44 924, 2025

  39. [47]

    Realization of an integrated coherent photonic platform for scalable matrix operations,

    S. Rahimi Kari, N. A. Nobile, D. Pantin, V . Shah, and N. Youngblood, “Realization of an integrated coherent photonic platform for scalable matrix operations,”Optica, vol. 11, no. 4, pp. 542–551, 2024

  40. [48]

    Drmap: A generic dram data mapping policy for energy-efficient processing of convolu- tional neural networks,

    R. V . W. Putra, M. A. Hanif, and M. Shafique, “Drmap: A generic dram data mapping policy for energy-efficient processing of convolu- tional neural networks,” in2020 57th ACM/IEEE Design Automation Conference (DAC), 2020, pp. 1–6

  41. [49]

    Romanet: Fine-grained reuse-driven off-chip memory access management and data organization for deep neural network acceler- ators,

    ——, “Romanet: Fine-grained reuse-driven off-chip memory access management and data organization for deep neural network acceler- ators,”IEEE Transactions on Very Large Scale Integration Systems (TVLSI), vol. 29, no. 4, pp. 702–715, 2021

  42. [50]

    Pendram: Enabling high-performance and energy-efficient pro- cessing of deep neural networks through a generalized dram data mapping policy,

    ——, “Pendram: Enabling high-performance and energy-efficient pro- cessing of deep neural networks through a generalized dram data mapping policy,”arXiv preprint arXiv:2408.02412, 2024

  43. [51]

    A 2x time- interleaved 28-gs/s 8-bit 0.03-mm 2 switched-capacitor dac in 16-nm finfet cmos,

    P. Caragiulo, O. E. Mattia, A. Arbabian, and B. Murmann, “A 2x time- interleaved 28-gs/s 8-bit 0.03-mm 2 switched-capacitor dac in 16-nm finfet cmos,”IEEE Journal of Solid-State Circuits, vol. 56, no. 8, pp. 2335–2346, 2021

  44. [52]

    A 12.8 gs/s time-interleaved adc with 25 ghz effective resolution bandwidth and 4.6 enob,

    Y . Duan and E. Alon, “A 12.8 gs/s time-interleaved adc with 25 ghz effective resolution bandwidth and 4.6 enob,”IEEE Journal of Solid- State Circuits, vol. 49, no. 8, pp. 1725–1738, 2014

  45. [53]

    64gb/s nrz/pam4 burst- mode optical receiver frontend with gain control, offset correction and gain decoupled from bandwidth,

    S. Serunjogi, M. Rasras, and M. Sanduleanu, “64gb/s nrz/pam4 burst- mode optical receiver frontend with gain control, offset correction and gain decoupled from bandwidth,” in2021 IEEE International Sympo- sium on Circuits and Systems (ISCAS). IEEE, 2021, pp. 1–4

Pith tools

Reviewed August 1, 2026 · model on record in the stance chip above.