REVIEW 3 major objections 4 minor 39 references
Who Needs DRAM? We Have Fiber
T0 review · 3 major / 4 minor · reviewed 2026-07-10 · grok-4.5
Pith's one-line read Datacenter fiber can serve as recirculating delay-line memory for LLM weights, cutting weight-delivery energy more than 70% versus HBM3e while eliminating redundant DRAM copies across thousands of accelerators.
desk verdict Solid first-order architecture paper: the MCF-ring + passive-tap + CPO-to-systolic idea for immutable LLM weights is new and the arithmetic is clean, but the 72% energy headline is an artifact of continuous-peak utilization and unproven OSNR regeneration. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The data-parallel optical broadcast delay-line: a 14-cable, 19-core multi-core fiber ring that stores 128 GB in flight at 25.6 TB/s, combined with 1:99 passive tap-and-amplify interfaces and regional PDFAs plus all-optical 2R regenerators that keep the circulating stream readable without electronic conversion at every hop.
What would settle it
Build a laboratory-scale multi-core O-band loop with the proposed tap density and measure whether OSNR stays above roughly 15 dB for 50 Gbaud PAM4 after twenty 50 km stages plus 1:99 chassis taps, or whether regenerator power and latency erase the projected energy advantage.
Extended reading notes
Core claim
Fiber Memory can eliminate redundant weight storage across 10,000 AI accelerators and reduce weight-delivery energy by over 70 percent (284.8 kW versus 1,024 kW) relative to HBM3e by streaming a full Llama-3-70B INT8 model plus slack continuously through a 1,000 km multi-core fiber ring that every chassis passively taps.
Load-bearing premise
The claim rests on cascaded O-band amplifiers and regional all-optical regenerators keeping the optical signal clean enough for high-speed PAM4 after many kilometers and many tiny taps; if noise grows faster or regenerators cost too much power, continuous recirculation fails.
Editorial extensions
If this is right
- Identical LLM weights need be stored and powered only once instead of once per accelerator, freeing hundreds of terabytes of HBM/DRAM capacity cluster-wide.
- Weight-delivery power for a 10,000-accelerator inference farm falls from over a megawatt to under 300 kW under the paper's numbers.
- Activations and KV cache remain the only data that must live in local high-bandwidth memory, shrinking the memory hierarchy for inference.
- Larger models are accommodated simply by lengthening the fiber loop and adding more amplifiers rather than by adding more HBM stacks.
Reading between the lines
- If the optical budget holds, the same ring could later carry immutable embeddings, codebooks, or frozen expert weights for mixture-of-experts models without further DRAM growth.
- The architecture implicitly pressures vendors to treat co-packaged optics and multi-core fiber as first-class memory interfaces rather than mere interconnects.
- A natural next measurement is whether residual bit flips that FEC leaves in the weight payload are absorbed by the known noise tolerance of quantized LLMs, further relaxing regenerator spacing.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes Fiber Memory: a recirculating optical delay-line architecture that stores immutable LLM weights in flight inside multi-core fiber loops and broadcasts them via passive 1:99 tap-and-amplify interfaces and co-packaged optics to thousands of accelerators. A Llama-3-70B INT8 case study (10 000 accelerators, 1 000 km of 19-core MCF, 25.6 TB/s aggregate) claims elimination of 700 TB of redundant HBM weight storage and a 72.1 % reduction in weight-delivery power (284.8 kW versus 1 024 kW) relative to continuous-peak HBM3e fetches. Supporting calculations cover bandwidth-delay product, O-band DWDM, PDFA amplification, regional all-optical 2R regeneration, and direct-detection receiver energy.
Significance. If the physical and utilization assumptions hold, the work would reframe hyperscale AI memory as a shared optical resource rather than per-node HBM, offering a concrete path to cut both capital DRAM demand and dynamic energy for weight delivery. The architecture is novel in combining space-division multiplexed MCF rings, asymmetric passive taps, CPO direct feed into systolic arrays, and Streamed Weight Packets. The quantitative case study is transparent about component power numbers and free parameters, making the claim falsifiable and a useful first-order feasibility study even if later refined.
major comments (3)
- Section 4.0.4 (Eq. 9) and the subsequent comparison compute P_HBM under continuous peak-bandwidth utilization (10 000 × 25.6 Tb/s × 4 pJ/bit = 1 024 kW) while Fiber Memory’s dominant terms (lasers 4.1 kW, loop PDFAs 7 kW, pod amplifiers/regenerators 94.5 kW) are largely static. The text acknowledges micro-stalls yet still reports a 72.1 % saving against this theoretical maximum and omits HBM static leakage. A utilization-aware model or sensitivity sweep over realistic duty cycles (e.g., u = 0.3–0.7 typical of memory-bound decoding) is required; otherwise the headline energy claim is an artifact of the continuous-peak assumption rather than a robust system property.
- Section 5 presents an approximate OSNR formula and asserts that regional all-optical 2R regenerators (cross-phase modulation in SOA or HNLF) fully reset OSNR after twenty 50 km stages plus 1:99 chassis taps, keeping the signal above the ~15 dB threshold for 50 Gbaud PAM4. No numerical evaluation of the cascaded OSNR, regenerator insertion loss, residual noise, or power/latency overhead is supplied. Because continuous recirculation and the energy budget both rest on this reset working with negligible cost, a quantitative noise-budget calculation (or citation of measured 2R performance under the stated O-band multi-wavelength conditions) is load-bearing.
- Sections 2.3 and 3.2 claim that co-packaged optics can feed raw weight parameters “directly to the systolic array registers” with “no electronic buffering” and that Streamed Weight Packets eliminate address translation. Real CPO receivers still require CDR, deserialization, FEC decoding, and clock-domain crossing before data can be presented to compute registers. The 0.7 pJ/bit receiver figure (Eq. 13) and the zero-buffering energy claim therefore need either a concrete micro-architecture sketch showing how these functions are absorbed without intermediate storage or an explicit residual energy/latency term.
minor comments (4)
- Figure 1 and the accompanying text describe a ring-and-pod topology; a short table listing total fiber length, number of spools, cores, and wavelengths would make the physical dimensioning easier to verify.
- The 45 % slack/replica fraction is introduced without a derivation of how much slack is required for typical KV-cache or batch-size jitter; a one-sentence bound would strengthen Section 3.3.
- Several references (e.g., [26], [12], [9]) appear as 2025–2026 preprints or product briefs; ensuring stable DOIs or arXiv identifiers would improve long-term citability.
- Typographical inconsistencies appear (e.g., “butfar worse”, mixed use of TB/s vs Tb/s in prose). A light copy-edit pass would help.
Circularity Check
No circularity: energy and capacity claims are direct arithmetic from stated design parameters and external literature constants, not reductions of inputs by construction.
full rationale
The paper's load-bearing quantitative claims (BDP capacity M = B imes L / v yielding 128 GB for the 1000 km / 25.6 TB/s loop; P_HBM = 10 000 imes 25.6 Tb/s imes 4.0 pJ/bit = 1024 kW; fiber-side sum of lasers + PDFAs + regenerators + receivers = 284.8 kW; 72.1 % reduction) are obtained by substituting explicitly chosen architectural numbers (14 imes 19-core MCF, 8 DWDM channels at 100 Gb/s, 1:99 taps, 125 pods, etc.) into standard formulas and literature energy figures. No parameter is fitted to a target quantity that is later re-presented as a prediction; no uniqueness theorem or ansatz is imported from the authors' own prior work; the derivation chain does not close on itself. The continuous-peak utilization assumption affects absolute numbers but is an open modeling choice, not a circular definition. The architecture is therefore self-contained against its own inputs.
Assumptions & free parameters
free parameters (6)
- HBM3e energy per bit =
4.0 pJ/bit
- IM-DD receiver energy per bit =
0.7 pJ/bit
- PDFA electrical power =
25 W
- DFB laser wall-plug efficiency and optical power =
5%, 100 mW
- Slack / replica fraction =
up to 45%
- Loop length and spool stages =
1000 km / 20 stages
assumptions (6)
- domain assumption Group velocity in silica fiber is c/1.5 = 200 km/ms
- domain assumption O-band attenuation is 0.32 dB/km and 1:99 tap insertion loss is 0.043 dB
- domain assumption PDFAs are immune to inter-channel cross-gain modulation because of long upper-state lifetime
- domain assumption LLM weight traffic is 90-99% of memory bandwidth for typical batch sizes
- ad hoc to paper All-optical 2R regenerators fully reset OSNR without electrical conversion
- domain assumption Direct-detection 50 Gbaud PAM4 over 1000 km is feasible inside the O-band zero-dispersion window with 100 GHz DWDM
invented entities (2)
-
Fiber Memory architecture (MCF ring + pod + asymmetric tap-and-amplify + CPO direct feed)
-
Streamed Weight Packets (SWPs) with preamble, FEC header, unrolled payload, CRC
Cite this review
Pith. "Pith review of Who Needs DRAM? We Have Fiber." pith.science (2026). https://pith.science/paper/RT75YJ4D
@misc{pith2026260708407,
author = {Pith},
title = {Pith review of: Who Needs DRAM? We Have Fiber},
year = {2026},
howpublished = {\url{https://pith.science/paper/RT75YJ4D}},
note = {Machine review of arXiv:2607.08407}
}
read the original abstract
The rising pressure on DRAM availability and contract pricing reflects generative AI's massive high-performance memory requirements. This pressure is heavily compounded by hyperscale data center expansion, which now consumes a significant portion of global DRAM output. In this work, we propose a new architecture: Fiber Memory, which reimagines the role of optical fiber in a hyperscale data center, deploying it as an active, recirculating delay-line memory for immutable data, such as large language model (LLM) weights. We present a data-parallel optical broadcast delay-line memory architecture that accounts for fiber's physical realities. By incorporating space-division multiplexed multi-core fibers (MCFs), passive optical tap-and-amplify interfaces, co-packaged optics (CPO), and regional all-optical regeneration, our case study evaluation demonstrates that Fiber Memory can eliminate redundant weight storage across 10,000 AI accelerators and reduce weight-delivery energy by over 70% compared to traditional HBM3e configurations.
Figures
Reference graph
Works this paper leans on
-
[1]
2012.Fiber-Optic Communication Systems: Fourth Edition
Govind Agrawal. 2012.Fiber-Optic Communication Systems: Fourth Edition. doi:10.1002/9780470918524
-
[2]
Ahmed, Mohamed Eladawy, John A
Mohamed G. Ahmed, Mohamed Eladawy, John A. Palmer, Ahmed El-Nozahi, and Pavan Kumar Hanumolu. 2021. A 16-Gb/s −11.6-dBm OMA Sensitivity 0.7-pJ/bit Optical Receiver in 65-nm CMOS Enabled by Duobinary Sampling.IEEE Journal of Solid-State Circuits56, 9 (2021), 2795–2805. doi:10.1109/JSSC.2021.3064248
-
[3]
Reisi Aminabadi, Samyam Rajbhandari, Minjia Zhang, Chuanho Li, et al. 2022. DeepSpeed-Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale. In2022 IEEE International Symposium on High-Performance Computer Architecture (HPCA). 255–
work page 2022
-
[4]
doi:10.1109/HPCA53960.2022.00027
-
[5]
K. Benyahya, Daniel Burridge, Daniel Cletheroe, Thomas Karagiannis, Brian Robertson, Ant Rowstron, Mengyang Yang, Arash Behziz, Jamie Gaudette, and Paolo Costa. 2025. MOSAIC: Breaking the Optics versus Copper Trade-off with a Wide-and-Slow Architecture and MicroLEDs. InProceedings of the ACM SIGCOMM 2025 Conference. ACM, Coimbra, Portugal
work page 2025
-
[6]
Stewart, Thang Pham, and Joseph M
Brandon Buscaino, Elizabeth Chen, James W. Stewart, Thang Pham, and Joseph M. Kahn. 2021. External vs. Integrated Light Sources for Intra-Data Center Co-Packaged Optical Interfaces.Journal of Light- wave Technology39 (2021), 1984–1996. doi:10.1109/jlt.2020.3043653
-
[7]
Kun-Zhi Chang, Linhai Zhang, Tiejun Wang, and Intel Communica- tions Group. 2022. Ultra-Low-Latency Hardware Architectures for Reed-Solomon Forward Error Correction in Next-Generation Optical and Electrical Interconnects.IEEE Transactions on Circuits and Sys- tems I: Regular Papers69, 8 (2022), 3245–3256. doi:10.1109/TCSI.2022. 3168941
-
[8]
Corning Incorporated. 2025. Extreme-Density Ribbons and Multi-Core Fiber Engineering. Corning Optical Product Catalogs
work page 2025
Show all 39 references
-
[9]
Fosco Connect. 2026. Optical Fiber Tutorial - Optic Fiber - Communi- cation Fiber. https://www.fiberoptics4sale.com/blogs/archive-posts/ 95146054-optical-fiber-tutorial-optic-fiber-communication-fiber. [Online; accessed 29-June-2026]
2026
-
[10]
Zhen Gao, Qiang Liu, Jie Deng, et al. 2026. Enabling the Use of Ap- proximate Memories in Quantized Large Language Models (LLMs): Dealing with Errors on the Scaling Factors. TechRxiv. Published February 27, 2026
2026
-
[11]
Randy Giles and Emmanuel Desurvire
C. Randy Giles and Emmanuel Desurvire. 1991. Modeling Erbium- Doped Fiber Amplifiers.Journal of Lightwave Technology9, 2 (1991), 271–283. doi:10.1109/50.65886
1991 doi
-
[12]
Ibrahim Guler, Chen Sun, and Rajeev J. Ram. 2023. Laser Source Constraints and Wall-Plug Efficiency Optimization for High-Density Silicon Photonics Interconnects.Journal of Lightwave Technology41, 8 (2023), 2411–2420. doi:10.1109/JLT.2023.3235612
2023 doi
-
[13]
Tingbo He. 2026. A Time Scaling Theory for Multi-Layer Electronic Systems. ChinaXiv. https://chinaxiv.org/abs/202605.00224
2026 arXiv
-
[14]
2026.Chapter 9: Photon- ics
IEEE Electronics Packaging Society. 2026.Chapter 9: Photon- ics. Heterogeneous Integration Roadmap (HIR), 2026 Edition. IEEE EPS. https://eps.ieee.org/wp-content/uploads/2026/05/HIR_ 9_Photonics_rev0.9.pdf
2026
- [15]
-
[16]
Kyomin Lee, Junyoung Kim, Taeyoung Oh, Kyoung-Hoi Sihn, et al
-
[17]
In2024 IEEE International Solid-State Circuits Conference (ISSCC)
A 1.2-TB/s HBM3E DRAM with Advanced Thermal Compression Non-Conductive Film Packaging and Optimized Core Power Delivery. In2024 IEEE International Solid-State Circuits Conference (ISSCC). 312–
-
[18]
doi:10.1109/ISSCC49657.2024.10454389
2024 doi
-
[19]
Takashi Matsui, Taiji Sakamoto, Fumihiro Hanzawa, Shigeru Tomita, Kyozo Tsujikawa, and Kazuhide Nakajima. 2017. Low-Loss and Low- DMD 6-Mode 19-Core Fiber With Cladding Diameter of Less Than 250 𝜇m.Journal of Lightwave Technology35, 3 (2017), 443–449. doi:10. 1109/JLT.2016.2608920
2017
-
[20]
Antonio Mecozzi, Mark Shtaif, and Cristian Antonelli. 2018. All- Optical 2R and 3R Regeneration in Regional Optical Fiber Communica- tion Systems.Journal of Lightwave Technology36, 7 (2018), 1411–1420. doi:10.1109/JLT.2017.2789104
2018 doi
-
[21]
2024.Micron HBM3E Product Brief: Introducing Memory Built for AI Innovation
Micron Technology, Inc. 2024.Micron HBM3E Product Brief: Introducing Memory Built for AI Innovation. Product Brief. Micron Technology, Inc. https://assets.micron.com/adobe/assets/urn:aaid:aem:b710d8f2-7f66- 44c1-a234-456e2b986347/renditions/original/as/hbm3e-product- brief.pdf
2024
-
[22]
Takashi Mori, Kazumasa Takada, and Takemi Hasegawa. 2022. Ultra- Low Latency O-Band Praseodymium-Doped Fiber Amplifiers for Data Center Interconnects.Journal of Lightwave Technology40, 8 (2022), 2445–2452. doi:10.1109/JLT.2022.3141108
2022 doi
-
[23]
Takashi Mori, Kazumasa Bro Takada, Hanawa Nishi, and Takemi Hasegawa. 2021. O-Band Praseodymium-Doped Fluoride Fiber Am- plifiers for Multi-Wavelength High-Speed Data Center Intercon- nects.Journal of Lightwave Technology39, 12 (2021), 3932–3940. doi:10.1109/JLT.2021.3061223
2021 doi
-
[24]
Nishi, K
M. Nishi, K. Takada, T. Mori, and T. Hasegawa. 2022. High-Gain O- Band Praseodymium-Doped Fluoride Fiber Amplifiers for High-Speed Data Center Interconnects.Journal of Lightwave Technology40, 14 (2022), 4689–4696. doi:10.1109/JLT.2022.3164401
2022 doi
-
[25]
NVIDIA. 2026. NVIDIA H100 Tensor Core GPU. https://www.nvidia. com/en-us/data-center/h100/ Accessed: 2026-07-08. 7 Hannah Atmer, Thiemo Voigt, Yuan Yao, and Stefanos Kaxiras
2026
-
[26]
NVIDIA Corporation. 2026. GPUDirect. NVIDIA Developer Portal. https://developer.nvidia.com/gpudirect Accessed June 12, 2026
2026
-
[27]
Yasutake Ohishi, Terutoshi Kanamori, Takeshi Kitagawa, Shiro Taka- hashi, Elias Snitzer, and George H. Sigel. 1998. Praseodymium-Doped Fiber Amplifiers for Optical Communications.IEEE Photonics Technol- ogy Letters10, 4 (1998), 516–518. doi:10.1109/68.662582
1998 doi
-
[28]
Puttnam, Ruben S
Georg Rademacher, Benjamin J. Puttnam, Ruben S. Luis, et al. 2020. High-Capacity Transmission Over Multi-Core Fibers in the O-Band. Journal of Lightwave Technology38, 2 (2020), 416–422. doi:10.1109/ JLT.2019.2947444
2020
-
[29]
J. Smith. 2025. The DRAM Bottleneck in Generative AI Systems.Forbes Tech Insights(October 2025)
2025
-
[30]
Richard Soref, Roberto De Rose, Ran Ding, et al. 2023. Co-Packaged Optics for Next-Generation Data Centers: Power Efficiency and Archi- tecture Analysis.IEEE Communications Magazine61, 4 (2023), 42–48. doi:10.1109/MCOM.001.2200385
2023 doi
-
[31]
Kristian Stubkjaer, Juerg Leuthold, and Charles H. Joyner. 2019. Power Consumption Dynamics and Thermal Biasing of All-Optical 2R Regen- erators Based on Semiconductor Optical Amplifiers.IEEE Photonics Technology Letters31, 14 (2019), 1143–1146. doi:10.1109/LPT.2019. 2918841
2019 doi
-
[32]
Kazumasa Takada, Takashi Mori, Takemi Hasegawa, and Hanawa Nishi. 2023. Power Consumption Analysis and Gain Optimization of O- Band Praseodymium-Doped Fiber Amplifiers for Optical Interconnects. Journal of Lightwave Technology41, 11 (2023), 3456–3464. doi:10.1109/ JLT.2023.3251104
2023
-
[33]
Min Tan, Jiang Xu, Siyang Liu, et al. 2023. Co-packaged optics (CPO): status, challenges, and solutions.Frontiers of Optoelectronics16, 1 (2023), 9. doi:10.1007/s12200-023-00058-7
2023 doi
-
[34]
Andrei-Alexandru Ulmămei and Cătălin Bîră. 2026. Reconfigurable SmartNICs: A Comprehensive Review of FPGA Shells and Heteroge- neous Offloading Architectures.Applied Sciences16, 3 (2026), 1476. doi:10.3390/app16031476
2026 doi
-
[35]
Webb, Tom Farrell, and Robert J
Benjamin C. Webb, Tom Farrell, and Robert J. Manning. 2011. All- Optical 2R Regeneration Using Cross-Phase Modulation in a Highly Non-Linear Fiber and an Optical Filter.IEEE Photonics Technology Letters23, 14 (2011), 983–985. doi:10.1109/LPT.2011.2148107
2011 doi
-
[36]
Jun Shan Wey, Xiang Liu, and Ed Harstead. 2020. Design and Opti- mization of O-Band DWDM Transmission Systems for High-Speed Mobile Fronthaul and Data Center Networks.Journal of Lightwave Technology38, 11 (2020), 2913–2921. doi:10.1109/JLT.2020.2978543
2020 doi
-
[37]
Wikipedia contributors. 2026. Delay-line memory — Wikipedia, The Free Encyclopedia. https://en.wikipedia.org/wiki/Delay-line_memory. [Online; accessed 26-June-2026]
2026
-
[38]
Junwen Zhang, Jianjun Yu, Jun Shan Wey, et al . 2020. SOA Pre- Amplified 100 Gb/s/𝜆 PAM-4 TDM-PON Downstream Transmission Using 10G-Class O-Band Transmitters.Journal of Lightwave Technol- ogy38, 2 (2020), 185–193. doi:10.1109/JLT.2019.2954992
2020 doi
-
[39]
K. Zhang. 2024. Quantifying Cable Lengths and Fiber Densities in Modern Mega-Datacenters.IEEE Communications Magazine62, 1 (2024), 45–52. 8
2024
Reviewed July 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.