REVIEW 4 major objections 3 minor 41 references
Hybrid Scandium Aluminum Nitride/Silicon Nitride Integrated Photonic Circuits
T0 review · 4 major / 3 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read A hybrid silicon nitride–ScAlN waveguide platform claims to cut propagation loss to 1.03 dB/cm while keeping ScAlN's nonlinear, ferroelectric, and piezoelectric functions.
desk verdict The photonics headline is abstract-only; the body is an unrelated diffusion paper, so the central claim cannot be checked and the manuscript deserves desk rejection. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is a hybrid evanescent platform: an etched Si3N4 core carries the optical field, and a ScAlN underlayer is kept close enough to preserve its electro-optic, ferroelectric, and piezoelectric functions without absorbing or scattering the guided light. The design's purpose is to decouple optical loss from functional performance in a monolithic, CMOS-compatible stack.
What would settle it
Fabricate the hybrid waveguide and measure propagation loss by a cutback method or by ring-resonator linewidth; if the intrinsic Q of the hybrid ring is below $3.35 \times 10^5$, or if the measured mode profile shows significant overlap with ScAlN while the loss is low but the ScAlN functional coefficients are degraded, the central claim fails. The provided text contains no such measurement details, so the claim can only be settled by the original experimental data.
Extended reading notes
Core claim
The central claim is that a hybrid Si3N4–ScAlN waveguide can have both low optical loss and functional ScAlN. Light is confined in the etched Si3N4 waveguide, so scattering and absorption in ScAlN are avoided, while the ScAlN layer underneath remains able to provide second-order nonlinearity, ferroelectricity, and piezoelectricity. The reported intrinsic quality factor $Q_{\mathrm{i}} = 3.35 \times 10^5$ corresponds to 1.03 dB/cm propagation loss, about 2.4 times better than the >2.4 dB/cm typically reported for ScAlN waveguides and comparable to commercial single-mode silicon-on-insulator waveguides.
Load-bearing premise
The load-bearing premise is that the measured $Q_{\mathrm{i}} = 3.35 \times 10^5$ really is propagation loss of the hybrid mode confined in the etched Si3N4 core, and that the ScAlN underlayer still performs its nonlinear, ferroelectric, and piezoelectric functions without reintroducing absorption or scattering.
Editorial extensions
If this is right
- ScAlN-based photonic integrated circuits could reach loss levels comparable to mature SOI platforms, removing a key obstacle for nonlinear and quantum photonics in this material.
- Ferroelectric and piezoelectric functions of ScAlN could be exploited in low-loss modulators, switches, and acousto-optic devices on the same chip as Si3N4 routing.
- The all-monolithic CMOS-compatible fabrication would allow scalable production of hybrid photonic circuits without bonding or transfer printing.
- The architecture suggests a path to integrating quantum light sources based on ScAlN's second-order nonlinearity with low-loss Si3N4 networks.
Reading between the lines
- If the 1.03 dB/cm figure is reproduced, the platform may make ScAlN more competitive with thin-film lithium niobate for on-chip nonlinear optics while remaining fully CMOS-compatible.
- A testable extension would be to vary the Si3N4 etch depth and ScAlN thickness and map how loss and the retained functional coefficients trade off as the optical mode overlap with ScAlN changes.
- The abstract compares the loss with commercial SOI waveguides, but whether the same loss holds at telecom wavelengths, at higher optical powers, and after repeated ferroelectric switching is not addressed there and would need direct measurement.
- As the provided manuscript body does not contain the mode analysis or loss-budget measurement behind the abstract's headline values, the quantitative claim currently stands only as an abstract-level assertion until the original device data is available.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript, as provided, consists of a photonics abstract and title ('Hybrid Scandium Aluminum Nitride/Silicon Nitride Integrated Photonic Circuits') followed by a full text and supplementary material that are an unrelated paper on personalized text-to-image diffusion models ('Steering Guidance for Personalized Text-to-Image Diffusion Models'). The abstract reports an intrinsic quality factor Qi = 3.35e5 and a propagation loss of 1.03 dB/cm for an etched Si3N4/ScAlN waveguide, with no accompanying device description, fabrication details, measurement procedure, mode analysis, or error analysis anywhere in the body of the text. The quantitative claims therefore cannot be checked against the manuscript content.
Significance. If the abstract's numbers were substantiated, the result would be significant: a monolithically integrated, CMOS-compatible ScAlN/Si3N4 platform with roughly 1 dB/cm propagation loss would be about 2.4 times lower loss than the cited ScAlN baseline, while retaining second-order nonlinearity, ferroelectricity, and piezoelectricity. That would plausibly open new applications in nonlinear, quantum, and electro-optic photonics. However, as submitted, the manuscript contains no support for these numbers: the body text is a different paper in a different field. There are no machine-checked proofs, reproducible code, parameter-free derivations, or falsifiable predictions in the photonics portion. The claimed advance is therefore unverifiable in the present document.
major comments (4)
- [Abstract and Full Text] The central quantitative claim of the paper, Qi = 3.35 × 10^5 and 1.03 dB/cm, appears only in the abstract. The full text is 'Steering Guidance for Personalized Text-to-Image Diffusion Models', which contains no mention of ScAlN, Si3N4, waveguides, or optical loss. No fabrication run, device cross-section, resonator geometry, measurement wavelength, or Q-extraction method is provided. A reviewer cannot check the claim because the evidence is absent.
- [Supplementary Material, Section D] The only appended limitation statement, Supplementary Discussion D ('Dependency on Fine-tuned Models'), concerns the quality of diffusion-model fine-tuning and is unrelated to the photonics claim. It therefore does not satisfy the requirement that the limitations of the optical-loss measurement be stated; in fact, it confirms that the supplementary text belongs to a different manuscript.
- [Abstract, loss conversion] The conversion from Qi to propagation loss is not actually justified in the text: the abstract states 'corresponding to' without specifying the wavelength, group index, or resonator type. Without these parameters, the conversion is not reproducible. A statement of the measurement platform and the formula used (e.g., alpha = 2*pi*n_g/(lambda*Q_i) converted to dB/cm) is needed before the headline numbers can be assessed.
- [Abstract, functional-preservation claim] The abstract asserts that the hybrid architecture preserves the functional properties of the ScAlN layer, such as ferroelectricity and piezoelectricity, but no characterization of these properties is presented. If the etched Si3N4 waveguide confines light while the ScAlN underlayer retains its functions, the paper must at least provide a mode-overlap analysis and, ideally, electro-optic or piezoelectric measurements showing that the desired functions survive the fabrication process. This is a load-bearing claim, not a background statement.
minor comments (3)
- [Abstract] The comparison 'comparable to that of commercial single-mode SOI waveguides' is not accompanied by a citation or a quantitative range; the benchmark needs a source or a defined comparison set to be meaningful.
- [Full Text] The manuscript has no Methods section describing sample fabrication, optical characterization, or data analysis; if the photonics content is to be reviewed, a complete methods description must be included.
- [Title and author list] The title and abstract concern a photonics platform, while the body text is authored by a different team and carries a different arXiv identifier; this mismatch must be resolved if the document is to be considered a coherent submission.
Circularity Check
No circular derivation chain is present: the photonics abstract reports a measured Q_i, and the body is an unrelated diffusion-model paper, so the claim is unsupported but not circular.
full rationale
The photonics abstract (arXiv:2508.00314) reports an intrinsic quality factor Q_i = 3.35e5 and converts it to 1.03 dB/cm, while the supplied full text is an unrelated text-to-image diffusion paper with no waveguide geometry, fabrication details, loss measurement, or mode analysis. That means the central claim cannot be verified from the provided document, but absence of support is not circularity. No equation in the abstract defines Q_i in terms of the loss it is said to predict, no parameter is fitted to a subset of data and then renamed a prediction, and no load-bearing step cites prior work by the same authors to forbid alternatives. The Q-to-loss conversion is the standard relationship between intrinsic quality factor and propagation loss, which is an independent definitional relation rather than a self-referential one. The body mismatch is a document-integrity and completeness problem, not a circular-derivation problem; under the hard rule requiring a quoted equation-level reduction, no circular step can be identified.
Assumptions & free parameters
assumptions (2)
- domain assumption The ScAlN layer retains its second-order nonlinearity, ferroelectricity, and piezoelectricity when integrated beneath an etched Si3N4 waveguide.
- domain assumption Loss in the hybrid waveguide is set by the low-loss Si3N4 core, and the evanescent field in the ScAlN underlayer does not reintroduce the loss the design is meant to avoid.
Cite this review
Pith. "Pith review of Hybrid Scandium Aluminum Nitride/Silicon Nitride Integrated Photonic Circuits." pith.science (2026). https://pith.science/paper/6M34QQLN
@misc{pith2026250800314,
author = {Pith},
title = {Pith review of: Hybrid Scandium Aluminum Nitride/Silicon Nitride Integrated Photonic Circuits},
year = {2026},
howpublished = {\url{https://pith.science/paper/6M34QQLN}},
note = {Machine review of arXiv:2508.00314}
}
abstract
Scandium-doped aluminum nitride has recently emerged as a promising material for quantum photonic integrated circuits (PICs) due to its unique combination of strong second-order nonlinearity, ferroelectricity, piezoelectricity, and complementary metal-oxide-semiconductor (CMOS) compatibility. However, the relatively high optical loss reported to date-typically above 2.4 dB/cm-remains a key challenge that limits its widespread application in low-loss PICs. Here, we present a monolithically integrated $\mathrm{Si}_3\mathrm{N}_4$-ScAlN waveguide platform that overcomes this limitation. By confining light within an etched $\mathrm{Si}_3\mathrm{N}_4$ waveguide while preserving the functional properties of the underlying ScAlN layer, we achieve an intrinsic quality factor of $Q_{\mathrm{i}} = 3.35 \times 10^5$, corresponding to a propagation loss of 1.03 dB/cm-comparable to that of commercial single-mode silicon-on-insulator (SOI) waveguides. This hybrid architecture enables low-loss and scalable fabrication while retaining the advanced functionalities offered by ScAlN, such as ferroelectricity and piezoelectricity. Our results establish a new pathway for ScAlN-based PICs with potential applications in high-speed optical communication, modulation, sensing, nonlinear optics, and quantum optics within CMOS-compatible platforms.
Reference graph
Works this paper leans on
-
[1]
Self-rectifying diffu- sion sampling with perturbed-attention guidance
Donghoon Ahn, Hyoungwon Cho, Jaewon Min, Wooseok Jang, Jungwoo Kim, SeonHwa Kim, Hyun Hee Park, Ky- ong Hwan Jin, and Seungryong Kim. Self-rectifying diffu- sion sampling with perturbed-attention guidance. In Euro- pean Conference on Computer Vision, pages 1–17. Springer,
-
[2]
In- structpix2pix: Learning to follow image editing instructions
Tim Brooks, Aleksander Holynski, and Alexei A Efros. In- structpix2pix: Learning to follow image editing instructions. In Proceedings of the IEEE/CVF conference on computer vi- sion and pattern recognition, pages 18392–18402, 2023. 3, 8
work page 2023
-
[3]
Emerg- ing properties in self-supervised vision transformers
Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerg- ing properties in self-supervised vision transformers. In Pro- ceedings of the International Conference on Computer Vi- sion (ICCV), 2021. 5
work page 2021
-
[4]
Improving subject-driven image syn- thesis with subject-agnostic guidance
Kelvin CK Chan, Yang Zhao, Xuhui Jia, Ming-Hsuan Yang, and Huisheng Wang. Improving subject-driven image syn- thesis with subject-agnostic guidance. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6733–6742, 2024. 3, 5, 6, 1
work page 2024
-
[5]
Para: Personalizing text-to-image diffusion via parameter rank reduction
Shangyu Chen, Zizheng Pan, Jianfei Cai, and Dinh Phung. Para: Personalizing text-to-image diffusion via parameter rank reduction. arXiv preprint arXiv:2406.05641, 2024. 2, 3
arXiv 2024
-
[6]
Subject-driven text-to-image generation via apprenticeship learning
Wenhu Chen, Hexiang Hu, Yandong Li, Nataniel Ruiz, Xuhui Jia, Ming-Wei Chang, and William W Cohen. Subject-driven text-to-image generation via apprenticeship learning. Advances in Neural Information Processing Sys- tems, 36:30286–30305, 2023. 2, 3
work page 2023
-
[7]
Cfg++: Manifold-constrained clas- sifier free guidance for diffusion models
Hyungjin Chung, Jeongsol Kim, Geon Yeong Park, Hyelin Nam, and Jong Chul Ye. Cfg++: Manifold-constrained clas- sifier free guidance for diffusion models. arXiv preprint arXiv:2406.08070, 2024. 1
arXiv 2024
-
[8]
Dream- sim: Learning new dimensions of human visual similar- ity using synthetic data
Stephanie Fu, Netanel Tamir, Shobhita Sundaram, Lucy Chai, Richard Zhang, Tali Dekel, and Phillip Isola. Dream- sim: Learning new dimensions of human visual similar- ity using synthetic data. arXiv preprint arXiv:2306.09344 ,
Show all 41 references
-
[9]
An image is worth one word: Personalizing text-to-image gener- ation using textual inversion
Rinon Gal, Yuval Alaluf, Yuval Atzmon, Or Patashnik, Amit Haim Bermano, Gal Chechik, and Daniel Cohen-or. An image is worth one word: Personalizing text-to-image gener- ation using textual inversion. In The Eleventh International Conference on Learning Representations. 2, 3, 1
-
[10]
Vico: Plug-and-play visual condition for personalized text-to-image generation
Shaozhe Hao, Kai Han, Shihao Zhao, and Kwan-Yee K Wong. Vico: Plug-and-play visual condition for personalized text-to-image generation. arXiv preprint arXiv:2306.00971,
-
[11]
Classifier-free diffusion guidance
Jonathan Ho and Tim Salimans. Classifier-free diffusion guidance. In NeurIPS 2021 Workshop on Deep Generative Models and Downstream Applications, 2021. 2, 3, 4, 5, 1
2021
-
[12]
Denoising diffu- sion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffu- sion probabilistic models. NeurIPS, 33:6840–6851, 2020. 2, 3
2020
-
[13]
Improving sample quality of diffusion models us- ing self-attention guidance
Susung Hong, Gyuseong Lee, Wooseok Jang, and Seungry- ong Kim. Improving sample quality of diffusion models us- ing self-attention guidance. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 7462– 7471, 2023. 3
2023
-
[14]
LoRA: Low-rank adaptation of large language models
Edward J Hu, yelong shen, Phillip Wallis, Zeyuan Allen- Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. LoRA: Low-rank adaptation of large language models. In ICLR, 2022. 2, 3, 5, 1
2022
-
[15]
Classdiffusion: More aligned personalization tuning with explicit class guidance
Jiannan Huang, Jun Hao Liew, Hanshu Yan, Yuyang Yin, Yao Zhao, and Yunchao Wei. Classdiffusion: More aligned personalization tuning with explicit class guidance. arXiv preprint arXiv:2405.17532, 2024. 2, 3, 5, 6, 1
2024 arXiv
-
[16]
In-context lora for diffusion transformers
Lianghua Huang, Wei Wang, Zhi-Fan Wu, Yupeng Shi, Huanzhang Dou, Chen Liang, Yutong Feng, Yu Liu, and Jin- gren Zhou. In-context lora for diffusion transformers. arXiv preprint arXiv:2410.23775, 2024. 8
2024 arXiv
-
[17]
Spatiotemporal skip guidance for enhanced video diffusion sampling
Junha Hyung, Kinam Kim, Susung Hong, Min-Jung Kim, and Jaegul Choo. Spatiotemporal skip guidance for enhanced video diffusion sampling. arXiv preprint arXiv:2411.18664,
-
[18]
Customizing text-to-image models with a single image pair
Maxwell Jones, Sheng-Yu Wang, Nupur Kumari, David Bau, and Jun-Yan Zhu. Customizing text-to-image models with a single image pair. In SIGGRAPH Asia 2024 Conference Papers, pages 1–13, 2024. 8
2024
-
[19]
Guiding a diffusion model with a bad version of itself
Tero Karras, Miika Aittala, Tuomas Kynkäänniemi, Jaakko Lehtinen, Timo Aila, and Samuli Laine. Guiding a diffusion model with a bad version of itself. Advances in Neural In- formation Processing Systems, 37:52996–53021, 2025. 2, 3, 4, 5, 1
2025
-
[20]
Autolora: Autoguid- ance meets low-rank adaptation for diffusion models
Artur Kasymov, Marcin Sendera, Michał Stypułkowski, Ma- ciej Zi˛ eba, and Przemysław Spurek. Autolora: Autoguid- ance meets low-rank adaptation for diffusion models. arXiv preprint arXiv:2410.03941, 2024. 6
2024 arXiv
-
[21]
Multi-concept customization of text-to-image diffusion
Nupur Kumari, Bingliang Zhang, Richard Zhang, Eli Shechtman, and Jun-Yan Zhu. Multi-concept customization of text-to-image diffusion. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages 1931–1941, 2023. 2, 3
1931
-
[22]
Aligning diffusion mod- els by optimizing human utility
Shufan Li, Konstantinos Kallidromitis, Akash Gokul, Yusuke Kato, and Kazuki Kozuka. Aligning diffusion mod- els by optimizing human utility. Advances in Neural Infor- mation Processing Systems, 37:24897–24925, 2025. 8
2025
-
[23]
Dreammatcher: appearance matching self-attention for semantically-consistent text-to- image personalization
Jisu Nam, Heesu Kim, DongJae Lee, Siyoon Jin, Seungry- ong Kim, and Seunggyu Chang. Dreammatcher: appearance matching self-attention for semantically-consistent text-to- image personalization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition...
2024
-
[24]
Scalable diffusion models with transformers
William Peebles and Saining Xie. Scalable diffusion models with transformers. In Proceedings of the IEEE/CVF inter- national conference on computer vision , pages 4195–4205,
-
[25]
Sdxl: Improving latent diffusion mod- els for high-resolution image synthesis
Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas Müller, Joe Penna, and Robin Rombach. Sdxl: Improving latent diffusion mod- els for high-resolution image synthesis. arXiv preprint arXiv:2307.01952, 2023. 2 9
2023 arXiv
-
[26]
Learning transferable visual models from natural language supervi- sion
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervi- sion. In International conference on machine learning, ...
2021
-
[27]
High-resolution image syn- thesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High-resolution image syn- thesis with latent diffusion models. In CVPR, pages 10684– 10695, 2022. 2, 4, 5, 1
2022
-
[28]
Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation
Nataniel Ruiz, Yuanzhen Li, Varun Jampani, Yael Pritch, Michael Rubinstein, and Kfir Aberman. Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation. In CVPR, pages 22500–22510, 2023. 2, 3, 5, 1
2023
-
[29]
Hyperdreambooth: Hypernetworks for fast personalization of text-to-image models
Nataniel Ruiz, Yuanzhen Li, Varun Jampani, Wei Wei, Tingbo Hou, Yael Pritch, Neal Wadhwa, Michael Rubinstein, and Kfir Aberman. Hyperdreambooth: Hypernetworks for fast personalization of text-to-image models. In Proceedings of the IEEE/CVF conference on computer vision and pat...
2024
-
[30]
Photorealistic text-to-image diffusion models with deep language understanding
Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L Denton, Kamyar Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, et al. Photorealistic text-to-image diffusion models with deep language understanding. NeurIPS, 35:36479–36494, 2022. 2
2022
-
[31]
In- stantbooth: Personalized text-to-image generation without test-time finetuning
Jing Shi, Wei Xiong, Zhe Lin, and Hyun Joon Jung. In- stantbooth: Personalized text-to-image generation without test-time finetuning. In Proceedings of the IEEE/CVF con- ference on computer vision and pattern recognition , pages 8543–8552, 2024. 2, 3
2024
-
[32]
Freeu: Free lunch in diffusion u-net
Chenyang Si, Ziqi Huang, Yuming Jiang, and Ziwei Liu. Freeu: Free lunch in diffusion u-net. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4733–4743, 2024. 3
2024
-
[33]
Denoising diffusion implicit models
Jiaming Song, Chenlin Meng, and Stefano Ermon. Denoising diffusion implicit models. arXiv preprint arXiv:2010.02502, 2020. 2, 3, 1
2010 arXiv
-
[34]
Leveraging previous steps: A training-free fast solver for flow diffusion
Kaiyu Song and Hanjiang Lai. Leveraging previous steps: A training-free fast solver for flow diffusion. arXiv preprint arXiv:2411.07627, 2024. 1
2024 arXiv
-
[35]
Diffusion model align- ment using direct preference optimization
Bram Wallace, Meihua Dang, Rafael Rafailov, Linqi Zhou, Aaron Lou, Senthil Purushwalkam, Stefano Ermon, Caiming Xiong, Shafiq Joty, and Nikhil Naik. Diffusion model align- ment using direct preference optimization. In Proceedings of the IEEE/CVF Conference on Computer Vision a...
2024
-
[36]
Elite: Encoding visual con- cepts into textual embeddings for customized text-to-image generation
Yuxiang Wei, Yabo Zhang, Zhilong Ji, Jinfeng Bai, Lei Zhang, and Wangmeng Zuo. Elite: Encoding visual con- cepts into textual embeddings for customized text-to-image generation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 15943–15953, 2023. 2, 3
2023
-
[37]
Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing in- ference time
Mitchell Wortsman, Gabriel Ilharco, Samir Ya Gadre, Re- becca Roelofs, Raphael Gontijo-Lopes, Ari S Morcos, Hongseok Namkoong, Ali Farhadi, Yair Carmon, Simon Ko- rnblith, et al. Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing ...
2022
-
[38]
Robust fine-tuning of zero-shot models
Mitchell Wortsman, Gabriel Ilharco, Jong Wook Kim, Mike Li, Simon Kornblith, Rebecca Roelofs, Raphael Gon- tijo Lopes, Hannaneh Hajishirzi, Ali Farhadi, Hongseok Namkoong, et al. Robust fine-tuning of zero-shot models. In Proceedings of the IEEE/CVF conference on computer vi- ...
2022
-
[39]
Stylealign: Analysis and applications of aligned stylegan models
Zongze Wu, Yotam Nitzan, Eli Shechtman, and Dani Lischinski. Stylealign: Analysis and applications of aligned stylegan models. arXiv preprint arXiv:2110.11323, 2021. 8
2021 arXiv
-
[40]
Sana: Efficient high-resolution image syn- thesis with linear diffusion transformers
Enze Xie, Junsong Chen, Junyu Chen, Han Cai, Haotian Tang, Yujun Lin, Zhekai Zhang, Muyang Li, Ligeng Zhu, Yao Lu, et al. Sana: Efficient high-resolution image syn- thesis with linear diffusion transformers. arXiv preprint arXiv:2410.10629, 2024. 2, 5, 6, 1
-
[41]
Scaling autoregres- sive models for content-rich text-to-image generation
Jiahui Yu, Yuanzhong Xu, Jing Yu Koh, Thang Luong, Gun- jan Baid, Zirui Wang, Vijay Vasudevan, Alexander Ku, Yin- fei Yang, Burcu Karagol Ayan, et al. Scaling autoregres- sive models for content-rich text-to-image generation. arXiv preprint arXiv:2206.10789, 2(3):5, 2022. 8 10...
2022 arXiv
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.