Pith. sign in

REVIEW 4 major objections 3 minor 41 references

Hybrid Scandium Aluminum Nitride/Silicon Nitride Integrated Photonic Circuits

T0 review · 4 major / 3 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read A hybrid silicon nitride–ScAlN waveguide platform claims to cut propagation loss to 1.03 dB/cm while keeping ScAlN's nonlinear, ferroelectric, and piezoelectric functions.

desk verdict The photonics headline is abstract-only; the body is an unrelated diffusion paper, so the central claim cannot be checked and the manuscript deserves desk rejection. read the letter →

arxiv 2508.00314 v1 pith:6M34QQLN submitted 2025-08-01 physics.optics cond-mat.mtrl-sci

classification physics.opticscond-mat.mtrl-sci
keywords scandiumaluminumnitrideScAlNsiliconwaveguidehybridintegratedphotonicspropagationlossintrinsicqualityfactorCMOS-compatiblequantumphotoniccircuits
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper proposes a monolithically integrated silicon nitride / scandium-doped aluminum nitride waveguide platform. It claims that by etching the Si3N4 core and letting it confine the optical mode while the ScAlN underlayer retains its ferroelectric, piezoelectric, and nonlinear properties, the platform reaches an intrinsic quality factor of $Q_{\mathrm{i}} = 3.35 \times 10^5$, corresponding to a propagation loss of 1.03 dB/cm. If this holds, ScAlN-based photonic circuits would no longer be limited by the roughly 2.4 dB/cm losses reported previously, and could combine low-loss routing with ScAlN's functional properties on a CMOS-compatible chip.

What carries the argument

The central mechanism is a hybrid evanescent platform: an etched Si3N4 core carries the optical field, and a ScAlN underlayer is kept close enough to preserve its electro-optic, ferroelectric, and piezoelectric functions without absorbing or scattering the guided light. The design's purpose is to decouple optical loss from functional performance in a monolithic, CMOS-compatible stack.

What would settle it

Fabricate the hybrid waveguide and measure propagation loss by a cutback method or by ring-resonator linewidth; if the intrinsic Q of the hybrid ring is below $3.35 \times 10^5$, or if the measured mode profile shows significant overlap with ScAlN while the loss is low but the ScAlN functional coefficients are degraded, the central claim fails. The provided text contains no such measurement details, so the claim can only be settled by the original experimental data.

Watch

Extended reading notes

Core claim

The central claim is that a hybrid Si3N4–ScAlN waveguide can have both low optical loss and functional ScAlN. Light is confined in the etched Si3N4 waveguide, so scattering and absorption in ScAlN are avoided, while the ScAlN layer underneath remains able to provide second-order nonlinearity, ferroelectricity, and piezoelectricity. The reported intrinsic quality factor $Q_{\mathrm{i}} = 3.35 \times 10^5$ corresponds to 1.03 dB/cm propagation loss, about 2.4 times better than the >2.4 dB/cm typically reported for ScAlN waveguides and comparable to commercial single-mode silicon-on-insulator waveguides.

Load-bearing premise

The load-bearing premise is that the measured $Q_{\mathrm{i}} = 3.35 \times 10^5$ really is propagation loss of the hybrid mode confined in the etched Si3N4 core, and that the ScAlN underlayer still performs its nonlinear, ferroelectric, and piezoelectric functions without reintroducing absorption or scattering.

Editorial extensions

If this is right

  • ScAlN-based photonic integrated circuits could reach loss levels comparable to mature SOI platforms, removing a key obstacle for nonlinear and quantum photonics in this material.
  • Ferroelectric and piezoelectric functions of ScAlN could be exploited in low-loss modulators, switches, and acousto-optic devices on the same chip as Si3N4 routing.
  • The all-monolithic CMOS-compatible fabrication would allow scalable production of hybrid photonic circuits without bonding or transfer printing.
  • The architecture suggests a path to integrating quantum light sources based on ScAlN's second-order nonlinearity with low-loss Si3N4 networks.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the 1.03 dB/cm figure is reproduced, the platform may make ScAlN more competitive with thin-film lithium niobate for on-chip nonlinear optics while remaining fully CMOS-compatible.
  • A testable extension would be to vary the Si3N4 etch depth and ScAlN thickness and map how loss and the retained functional coefficients trade off as the optical mode overlap with ScAlN changes.
  • The abstract compares the loss with commercial SOI waveguides, but whether the same loss holds at telecom wavelengths, at higher optical powers, and after repeated ferroelectric switching is not addressed there and would need direct measurement.
  • As the provided manuscript body does not contain the mode analysis or loss-budget measurement behind the abstract's headline values, the quantitative claim currently stands only as an abstract-level assertion until the original device data is available.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 3 minor

Summary. The manuscript, as provided, consists of a photonics abstract and title ('Hybrid Scandium Aluminum Nitride/Silicon Nitride Integrated Photonic Circuits') followed by a full text and supplementary material that are an unrelated paper on personalized text-to-image diffusion models ('Steering Guidance for Personalized Text-to-Image Diffusion Models'). The abstract reports an intrinsic quality factor Qi = 3.35e5 and a propagation loss of 1.03 dB/cm for an etched Si3N4/ScAlN waveguide, with no accompanying device description, fabrication details, measurement procedure, mode analysis, or error analysis anywhere in the body of the text. The quantitative claims therefore cannot be checked against the manuscript content.

Significance. If the abstract's numbers were substantiated, the result would be significant: a monolithically integrated, CMOS-compatible ScAlN/Si3N4 platform with roughly 1 dB/cm propagation loss would be about 2.4 times lower loss than the cited ScAlN baseline, while retaining second-order nonlinearity, ferroelectricity, and piezoelectricity. That would plausibly open new applications in nonlinear, quantum, and electro-optic photonics. However, as submitted, the manuscript contains no support for these numbers: the body text is a different paper in a different field. There are no machine-checked proofs, reproducible code, parameter-free derivations, or falsifiable predictions in the photonics portion. The claimed advance is therefore unverifiable in the present document.

major comments (4)
  1. [Abstract and Full Text] The central quantitative claim of the paper, Qi = 3.35 × 10^5 and 1.03 dB/cm, appears only in the abstract. The full text is 'Steering Guidance for Personalized Text-to-Image Diffusion Models', which contains no mention of ScAlN, Si3N4, waveguides, or optical loss. No fabrication run, device cross-section, resonator geometry, measurement wavelength, or Q-extraction method is provided. A reviewer cannot check the claim because the evidence is absent.
  2. [Supplementary Material, Section D] The only appended limitation statement, Supplementary Discussion D ('Dependency on Fine-tuned Models'), concerns the quality of diffusion-model fine-tuning and is unrelated to the photonics claim. It therefore does not satisfy the requirement that the limitations of the optical-loss measurement be stated; in fact, it confirms that the supplementary text belongs to a different manuscript.
  3. [Abstract, loss conversion] The conversion from Qi to propagation loss is not actually justified in the text: the abstract states 'corresponding to' without specifying the wavelength, group index, or resonator type. Without these parameters, the conversion is not reproducible. A statement of the measurement platform and the formula used (e.g., alpha = 2*pi*n_g/(lambda*Q_i) converted to dB/cm) is needed before the headline numbers can be assessed.
  4. [Abstract, functional-preservation claim] The abstract asserts that the hybrid architecture preserves the functional properties of the ScAlN layer, such as ferroelectricity and piezoelectricity, but no characterization of these properties is presented. If the etched Si3N4 waveguide confines light while the ScAlN underlayer retains its functions, the paper must at least provide a mode-overlap analysis and, ideally, electro-optic or piezoelectric measurements showing that the desired functions survive the fabrication process. This is a load-bearing claim, not a background statement.
minor comments (3)
  1. [Abstract] The comparison 'comparable to that of commercial single-mode SOI waveguides' is not accompanied by a citation or a quantitative range; the benchmark needs a source or a defined comparison set to be meaningful.
  2. [Full Text] The manuscript has no Methods section describing sample fabrication, optical characterization, or data analysis; if the photonics content is to be reviewed, a complete methods description must be included.
  3. [Title and author list] The title and abstract concern a photonics platform, while the body text is authored by a different team and carries a different arXiv identifier; this mismatch must be resolved if the document is to be considered a coherent submission.

Circularity Check

0 steps flagged · score 0.0 of 10

No circular derivation chain is present: the photonics abstract reports a measured Q_i, and the body is an unrelated diffusion-model paper, so the claim is unsupported but not circular.

full rationale

The photonics abstract (arXiv:2508.00314) reports an intrinsic quality factor Q_i = 3.35e5 and converts it to 1.03 dB/cm, while the supplied full text is an unrelated text-to-image diffusion paper with no waveguide geometry, fabrication details, loss measurement, or mode analysis. That means the central claim cannot be verified from the provided document, but absence of support is not circularity. No equation in the abstract defines Q_i in terms of the loss it is said to predict, no parameter is fitted to a subset of data and then renamed a prediction, and no load-bearing step cites prior work by the same authors to forbid alternatives. The Q-to-loss conversion is the standard relationship between intrinsic quality factor and propagation loss, which is an independent definitional relation rather than a self-referential one. The body mismatch is a document-integrity and completeness problem, not a circular-derivation problem; under the hard rule requiring a quoted equation-level reduction, no circular step can be identified.

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

The ledger is thin because the manuscript provides no body for the claimed photonics result: no free parameters or invented entities can be identified from the abstract alone, and a deeper audit is impossible. The two listed axioms are the abstract's implicit premises about the hybrid platform and its loss budget.

assumptions (2)
  • domain assumption The ScAlN layer retains its second-order nonlinearity, ferroelectricity, and piezoelectricity when integrated beneath an etched Si3N4 waveguide.
    The abstract promises 'preserving the functional properties of the underlying ScAlN layer,' which is what makes the hybrid platform valuable, but no measurement of these properties is present in the provided text.
  • domain assumption Loss in the hybrid waveguide is set by the low-loss Si3N4 core, and the evanescent field in the ScAlN underlayer does not reintroduce the loss the design is meant to avoid.
    The 1.03 dB/cm claim assumes this loss budget. No mode-overlap calculation, scattering estimate, or absorption analysis appears anywhere in the provided full text.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Hybrid Scandium Aluminum Nitride/Silicon Nitride Integrated Photonic Circuits." pith.science (2026). https://pith.science/paper/6M34QQLN

@misc{pith2026250800314,
  author       = {Pith},
  title        = {Pith review of: Hybrid Scandium Aluminum Nitride/Silicon Nitride Integrated Photonic Circuits},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/6M34QQLN}},
  note         = {Machine review of arXiv:2508.00314}
}
abstract

Scandium-doped aluminum nitride has recently emerged as a promising material for quantum photonic integrated circuits (PICs) due to its unique combination of strong second-order nonlinearity, ferroelectricity, piezoelectricity, and complementary metal-oxide-semiconductor (CMOS) compatibility. However, the relatively high optical loss reported to date-typically above 2.4 dB/cm-remains a key challenge that limits its widespread application in low-loss PICs. Here, we present a monolithically integrated $\mathrm{Si}_3\mathrm{N}_4$-ScAlN waveguide platform that overcomes this limitation. By confining light within an etched $\mathrm{Si}_3\mathrm{N}_4$ waveguide while preserving the functional properties of the underlying ScAlN layer, we achieve an intrinsic quality factor of $Q_{\mathrm{i}} = 3.35 \times 10^5$, corresponding to a propagation loss of 1.03 dB/cm-comparable to that of commercial single-mode silicon-on-insulator (SOI) waveguides. This hybrid architecture enables low-loss and scalable fabrication while retaining the advanced functionalities offered by ScAlN, such as ferroelectricity and piezoelectricity. Our results establish a new pathway for ScAlN-based PICs with potential applications in high-speed optical communication, modulation, sensing, nonlinear optics, and quantum optics within CMOS-compatible platforms.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

41 extracted references · 23 canonical work pages

  1. [1]

    Self-rectifying diffu- sion sampling with perturbed-attention guidance

    Donghoon Ahn, Hyoungwon Cho, Jaewon Min, Wooseok Jang, Jungwoo Kim, SeonHwa Kim, Hyun Hee Park, Ky- ong Hwan Jin, and Seungryong Kim. Self-rectifying diffu- sion sampling with perturbed-attention guidance. In Euro- pean Conference on Computer Vision, pages 1–17. Springer,

  2. [2]

    In- structpix2pix: Learning to follow image editing instructions

    Tim Brooks, Aleksander Holynski, and Alexei A Efros. In- structpix2pix: Learning to follow image editing instructions. In Proceedings of the IEEE/CVF conference on computer vi- sion and pattern recognition, pages 18392–18402, 2023. 3, 8

  3. [3]

    Emerg- ing properties in self-supervised vision transformers

    Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerg- ing properties in self-supervised vision transformers. In Pro- ceedings of the International Conference on Computer Vi- sion (ICCV), 2021. 5

  4. [4]

    Improving subject-driven image syn- thesis with subject-agnostic guidance

    Kelvin CK Chan, Yang Zhao, Xuhui Jia, Ming-Hsuan Yang, and Huisheng Wang. Improving subject-driven image syn- thesis with subject-agnostic guidance. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6733–6742, 2024. 3, 5, 6, 1

  5. [5]

    Para: Personalizing text-to-image diffusion via parameter rank reduction

    Shangyu Chen, Zizheng Pan, Jianfei Cai, and Dinh Phung. Para: Personalizing text-to-image diffusion via parameter rank reduction. arXiv preprint arXiv:2406.05641, 2024. 2, 3

  6. [6]

    Subject-driven text-to-image generation via apprenticeship learning

    Wenhu Chen, Hexiang Hu, Yandong Li, Nataniel Ruiz, Xuhui Jia, Ming-Wei Chang, and William W Cohen. Subject-driven text-to-image generation via apprenticeship learning. Advances in Neural Information Processing Sys- tems, 36:30286–30305, 2023. 2, 3

  7. [7]

    Cfg++: Manifold-constrained clas- sifier free guidance for diffusion models

    Hyungjin Chung, Jeongsol Kim, Geon Yeong Park, Hyelin Nam, and Jong Chul Ye. Cfg++: Manifold-constrained clas- sifier free guidance for diffusion models. arXiv preprint arXiv:2406.08070, 2024. 1

  8. [8]

    Dream- sim: Learning new dimensions of human visual similar- ity using synthetic data

    Stephanie Fu, Netanel Tamir, Shobhita Sundaram, Lucy Chai, Richard Zhang, Tali Dekel, and Phillip Isola. Dream- sim: Learning new dimensions of human visual similar- ity using synthetic data. arXiv preprint arXiv:2306.09344 ,

Show all 41 references
  1. [9]

    An image is worth one word: Personalizing text-to-image gener- ation using textual inversion

    Rinon Gal, Yuval Alaluf, Yuval Atzmon, Or Patashnik, Amit Haim Bermano, Gal Chechik, and Daniel Cohen-or. An image is worth one word: Personalizing text-to-image gener- ation using textual inversion. In The Eleventh International Conference on Learning Representations. 2, 3, 1

  2. [10]

    Vico: Plug-and-play visual condition for personalized text-to-image generation

    Shaozhe Hao, Kai Han, Shihao Zhao, and Kwan-Yee K Wong. Vico: Plug-and-play visual condition for personalized text-to-image generation. arXiv preprint arXiv:2306.00971,

  3. [11]

    Classifier-free diffusion guidance

    Jonathan Ho and Tim Salimans. Classifier-free diffusion guidance. In NeurIPS 2021 Workshop on Deep Generative Models and Downstream Applications, 2021. 2, 3, 4, 5, 1

  4. [12]

    Denoising diffu- sion probabilistic models

    Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffu- sion probabilistic models. NeurIPS, 33:6840–6851, 2020. 2, 3

  5. [13]

    Improving sample quality of diffusion models us- ing self-attention guidance

    Susung Hong, Gyuseong Lee, Wooseok Jang, and Seungry- ong Kim. Improving sample quality of diffusion models us- ing self-attention guidance. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 7462– 7471, 2023. 3

  6. [14]

    LoRA: Low-rank adaptation of large language models

    Edward J Hu, yelong shen, Phillip Wallis, Zeyuan Allen- Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. LoRA: Low-rank adaptation of large language models. In ICLR, 2022. 2, 3, 5, 1

  7. [15]

    Classdiffusion: More aligned personalization tuning with explicit class guidance

    Jiannan Huang, Jun Hao Liew, Hanshu Yan, Yuyang Yin, Yao Zhao, and Yunchao Wei. Classdiffusion: More aligned personalization tuning with explicit class guidance. arXiv preprint arXiv:2405.17532, 2024. 2, 3, 5, 6, 1

  8. [16]

    In-context lora for diffusion transformers

    Lianghua Huang, Wei Wang, Zhi-Fan Wu, Yupeng Shi, Huanzhang Dou, Chen Liang, Yutong Feng, Yu Liu, and Jin- gren Zhou. In-context lora for diffusion transformers. arXiv preprint arXiv:2410.23775, 2024. 8

  9. [17]

    Spatiotemporal skip guidance for enhanced video diffusion sampling

    Junha Hyung, Kinam Kim, Susung Hong, Min-Jung Kim, and Jaegul Choo. Spatiotemporal skip guidance for enhanced video diffusion sampling. arXiv preprint arXiv:2411.18664,

  10. [18]

    Customizing text-to-image models with a single image pair

    Maxwell Jones, Sheng-Yu Wang, Nupur Kumari, David Bau, and Jun-Yan Zhu. Customizing text-to-image models with a single image pair. In SIGGRAPH Asia 2024 Conference Papers, pages 1–13, 2024. 8

  11. [19]

    Guiding a diffusion model with a bad version of itself

    Tero Karras, Miika Aittala, Tuomas Kynkäänniemi, Jaakko Lehtinen, Timo Aila, and Samuli Laine. Guiding a diffusion model with a bad version of itself. Advances in Neural In- formation Processing Systems, 37:52996–53021, 2025. 2, 3, 4, 5, 1

  12. [20]

    Autolora: Autoguid- ance meets low-rank adaptation for diffusion models

    Artur Kasymov, Marcin Sendera, Michał Stypułkowski, Ma- ciej Zi˛ eba, and Przemysław Spurek. Autolora: Autoguid- ance meets low-rank adaptation for diffusion models. arXiv preprint arXiv:2410.03941, 2024. 6

  13. [21]

    Multi-concept customization of text-to-image diffusion

    Nupur Kumari, Bingliang Zhang, Richard Zhang, Eli Shechtman, and Jun-Yan Zhu. Multi-concept customization of text-to-image diffusion. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages 1931–1941, 2023. 2, 3

  14. [22]

    Aligning diffusion mod- els by optimizing human utility

    Shufan Li, Konstantinos Kallidromitis, Akash Gokul, Yusuke Kato, and Kazuki Kozuka. Aligning diffusion mod- els by optimizing human utility. Advances in Neural Infor- mation Processing Systems, 37:24897–24925, 2025. 8

  15. [23]

    Dreammatcher: appearance matching self-attention for semantically-consistent text-to- image personalization

    Jisu Nam, Heesu Kim, DongJae Lee, Siyoon Jin, Seungry- ong Kim, and Seunggyu Chang. Dreammatcher: appearance matching self-attention for semantically-consistent text-to- image personalization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition...

  16. [24]

    Scalable diffusion models with transformers

    William Peebles and Saining Xie. Scalable diffusion models with transformers. In Proceedings of the IEEE/CVF inter- national conference on computer vision , pages 4195–4205,

  17. [25]

    Sdxl: Improving latent diffusion mod- els for high-resolution image synthesis

    Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas Müller, Joe Penna, and Robin Rombach. Sdxl: Improving latent diffusion mod- els for high-resolution image synthesis. arXiv preprint arXiv:2307.01952, 2023. 2 9

  18. [26]

    Learning transferable visual models from natural language supervi- sion

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervi- sion. In International conference on machine learning, ...

  19. [27]

    High-resolution image syn- thesis with latent diffusion models

    Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High-resolution image syn- thesis with latent diffusion models. In CVPR, pages 10684– 10695, 2022. 2, 4, 5, 1

  20. [28]

    Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation

    Nataniel Ruiz, Yuanzhen Li, Varun Jampani, Yael Pritch, Michael Rubinstein, and Kfir Aberman. Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation. In CVPR, pages 22500–22510, 2023. 2, 3, 5, 1

  21. [29]

    Hyperdreambooth: Hypernetworks for fast personalization of text-to-image models

    Nataniel Ruiz, Yuanzhen Li, Varun Jampani, Wei Wei, Tingbo Hou, Yael Pritch, Neal Wadhwa, Michael Rubinstein, and Kfir Aberman. Hyperdreambooth: Hypernetworks for fast personalization of text-to-image models. In Proceedings of the IEEE/CVF conference on computer vision and pat...

  22. [30]

    Photorealistic text-to-image diffusion models with deep language understanding

    Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L Denton, Kamyar Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, et al. Photorealistic text-to-image diffusion models with deep language understanding. NeurIPS, 35:36479–36494, 2022. 2

  23. [31]

    In- stantbooth: Personalized text-to-image generation without test-time finetuning

    Jing Shi, Wei Xiong, Zhe Lin, and Hyun Joon Jung. In- stantbooth: Personalized text-to-image generation without test-time finetuning. In Proceedings of the IEEE/CVF con- ference on computer vision and pattern recognition , pages 8543–8552, 2024. 2, 3

  24. [32]

    Freeu: Free lunch in diffusion u-net

    Chenyang Si, Ziqi Huang, Yuming Jiang, and Ziwei Liu. Freeu: Free lunch in diffusion u-net. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4733–4743, 2024. 3

  25. [33]

    Denoising diffusion implicit models

    Jiaming Song, Chenlin Meng, and Stefano Ermon. Denoising diffusion implicit models. arXiv preprint arXiv:2010.02502, 2020. 2, 3, 1

  26. [34]

    Leveraging previous steps: A training-free fast solver for flow diffusion

    Kaiyu Song and Hanjiang Lai. Leveraging previous steps: A training-free fast solver for flow diffusion. arXiv preprint arXiv:2411.07627, 2024. 1

  27. [35]

    Diffusion model align- ment using direct preference optimization

    Bram Wallace, Meihua Dang, Rafael Rafailov, Linqi Zhou, Aaron Lou, Senthil Purushwalkam, Stefano Ermon, Caiming Xiong, Shafiq Joty, and Nikhil Naik. Diffusion model align- ment using direct preference optimization. In Proceedings of the IEEE/CVF Conference on Computer Vision a...

  28. [36]

    Elite: Encoding visual con- cepts into textual embeddings for customized text-to-image generation

    Yuxiang Wei, Yabo Zhang, Zhilong Ji, Jinfeng Bai, Lei Zhang, and Wangmeng Zuo. Elite: Encoding visual con- cepts into textual embeddings for customized text-to-image generation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 15943–15953, 2023. 2, 3

  29. [37]

    Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing in- ference time

    Mitchell Wortsman, Gabriel Ilharco, Samir Ya Gadre, Re- becca Roelofs, Raphael Gontijo-Lopes, Ari S Morcos, Hongseok Namkoong, Ali Farhadi, Yair Carmon, Simon Ko- rnblith, et al. Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing ...

  30. [38]

    Robust fine-tuning of zero-shot models

    Mitchell Wortsman, Gabriel Ilharco, Jong Wook Kim, Mike Li, Simon Kornblith, Rebecca Roelofs, Raphael Gon- tijo Lopes, Hannaneh Hajishirzi, Ali Farhadi, Hongseok Namkoong, et al. Robust fine-tuning of zero-shot models. In Proceedings of the IEEE/CVF conference on computer vi- ...

  31. [39]

    Stylealign: Analysis and applications of aligned stylegan models

    Zongze Wu, Yotam Nitzan, Eli Shechtman, and Dani Lischinski. Stylealign: Analysis and applications of aligned stylegan models. arXiv preprint arXiv:2110.11323, 2021. 8

  32. [40]

    Sana: Efficient high-resolution image syn- thesis with linear diffusion transformers

    Enze Xie, Junsong Chen, Junyu Chen, Han Cai, Haotian Tang, Yujun Lin, Zhekai Zhang, Muyang Li, Ligeng Zhu, Yao Lu, et al. Sana: Efficient high-resolution image syn- thesis with linear diffusion transformers. arXiv preprint arXiv:2410.10629, 2024. 2, 5, 6, 1

  33. [41]

    Scaling autoregres- sive models for content-rich text-to-image generation

    Jiahui Yu, Yuanzhong Xu, Jing Yu Koh, Thang Luong, Gun- jan Baid, Zirui Wang, Vijay Vasudevan, Alexander Ku, Yin- fei Yang, Burcu Karagol Ayan, et al. Scaling autoregres- sive models for content-rich text-to-image generation. arXiv preprint arXiv:2206.10789, 2(3):5, 2022. 8 10...

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.