Pith. sign in

REVIEW 3 major objections 5 minor 46 references

Frontier multimodal agents, guided by a deterministic geometric verifier, can convert images of photonic components into editable parametric programs, exceeding 0.9 mean IoU on eight targets and enabling cross-stack retargeting and verifier

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-04 01:19 UTC pith:AVPNVSHZ

load-bearing objection The headline IoU numbers are only as solid as Gemini-generated reference masks that were never validated against original GDS data; still, this is a transparent, well-engineered systems paper that deserves peer review. the 3 major comments →

arxiv 2608.00084 v1 pith:AVPNVSHZ submitted 2026-07-29 cs.CV physics.optics

From Pixels to PCells: A Neurosymbolic Approach to Photonic Component Creation

classification cs.CV physics.optics
keywords photonic integrated circuitsparametric cells (PCell)multimodal agentsgeometric verificationdomain-specific languagevisual program inductiondesign retargetingverifier-based reinforcement learning
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper aims to show that visual-to-parametric reconstruction of photonic components is not only possible but practical. PixCell pairs a small domain-specific language of geometric primitives with a deterministic render-and-compare verifier, so that frontier multimodal agents can write and revise programs that reproduce a target image at physical scale. On eight published photonic components, the best configurations exceed 0.9 mean IoU (up to 0.974) while satisfying a contract that forbids raw polygon output and prebuilt cells; in contrast, the same models without the agentic interface average only 0.416 best-turn IoU. The recovered programs are live parametric models: they can be retargeted across silicon, silicon-nitride, and thin-film-lithium-niobate stacks, carried through full-wave simulation, and used as a reward signal to train a smaller open-weight model. A sympathetic reader would care because photonic integrated-circuit design is fundamentally visual-geometric, and a reliable image-to-PCell pathway could let engineers turn figures into editable, process-portable components.

Core claim

The paper claims that a neurosymbolic pipeline—a small domain-specific language of geometric primitives, a frozen binary-mask target with physical calibration, and a deterministic IoU verifier—flips the economics of photonic component creation: checking a candidate is cheap, and frontier multimodal coding agents can use it to write and revise programs that exceed 0.9 mean IoU on all eight benchmark targets while passing a source contract that forbids raw polygons and prebuilt cells. It further claims that the recovered parametric programs are not just visually faithful but functionally useful: they can be retargeted across SOI, SiN, and TFLN stacks to meet an 8.0 nm free-spectral-range targe

What carries the argument

The carrying mechanism is the PixCell loop: a frozen target (binary mask plus physical footprint) defines ground truth; a fixed DSL of geometric primitives (regions, port-bearing waveguides, paths and cross-sections, boolean operations, references, and routing) constrains the program space; and a deterministic verifier executes each candidate program, renders it at target-derived pixel-per-micron scales without rescaling, and scores IoU, Dice, and squared error. The verifier makes evaluation asymmetric with generation—checking is cheap relative to proposing—so it can rank independent seeds, drive revision with spatial residuals, and supply a scalar reward for reinforcement learning. A source

Load-bearing premise

The load-bearing premise is that the binary masks and physical footprints extracted from published figures are a faithful ground truth for each component; the paper never validates these extractions against the original design data, and one target is admitted to be provenance-incomplete, so a systematic mask error would silently shift every reported IoU, retargeting, and training number.

What would settle it

Obtain the original GDS layouts or validated masks for the eight cited component papers, re-run the pipeline's extraction from the published figures, and compare the frozen masks against those originals; if the masks differ by more than a small tolerance in shape or footprint calibration—or if independent extraction runs disagree substantially—then the ground-truth anchor of the benchmark is unstable and the high IoU scores no longer establish faithful reconstruction.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Published photonic-component figures can become editable PDK cells without access to the original design source, opening a route to converting literature into reusable design libraries.
  • A reconstructed program's live parameters let a design be retargeted across material stacks (SOI, SiN, TFLN) while honoring spectral and footprint constraints—something fixed-interface library PCells can fail to do when the required geometry exceeds the footprint.
  • The same deterministic verifier can serve as a reward for reinforcement learning, improving a smaller open-weight model's program generation on held-out targets without supervised demonstrations, suggesting a path to reproducible design agents.
  • The frozen masks, footprints, and acceptance gates create a controlled benchmark for measuring visual-to-code capability, and the reported cost spread shows a two-orders-of-magnitude trade-off between compute and reconstruction quality.
  • Editable routing freedom, not just geometric similarity, is what enables functional retargeting; programs that preserve topology as named parameters can adapt where pixel-trace or fixed-PCell representations cannot.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The mask- and footprint-extraction step is an unvalidated link in the chain; if the frozen targets deviate systematically from the original device layouts, every IoU score, retargeting gate, and training reward inherits that error—an independent re-derivation of targets from original design files would settle the foundation.
  • The core asymmetry (cheap verification against expensive generation) is generic; domains beyond photonics where visual structure maps to parametric executables—microfluidics, MEMS, metamaterial unit cells—could reuse the same neurosymbolic loop, though each needs its own DSL and physical constraints.
  • The trained model's gain (roughly 0.42 to 0.49 champion IoU after revision) is real but far from frontier-agent performance; scaling the curriculum and reward shaping might close that gap, or might reveal a ceiling for small models on this task, which is itself a useful empirical question.
  • The paper's 'research contracts' make the framework self-measuring: by pre-registering shape-aware verification, process-faithful 2D-to-3D retargeting, dataset scale, and the smallest qualifying open model as versioned tests, the authors enable future claims to be settled by executable evidence rather than narrative.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper introduces PixCell, a neurosymbolic pipeline in which multimodal coding agents convert a binary silhouette and a physical footprint into a GDSFactory program over a small geometric DSL. The central empirical claim is that the best coding-agent configurations reach mean IoU above 0.9 across eight photonic targets, with best scores 0.974 and 0.955, while satisfying source-compliance gates. The paper further reports that recovered PCells can be retargeted across 220-nm SOI, 400-nm SiN, and 400-nm TFLN stacks; that a tiered simulation audit checks some of these results; and that the same geometric verifier can supply a GRPO training reward to improve a smaller open-weight model. The manuscript emphasizes that all targets, programs, transcripts, simulation inputs, and datasets are frozen and released.

Significance. If the reference masks are valid, the reconstruction results are notable: the paper demonstrates a working image-to-parametric-PCell loop with frontier agents, and the explicit separation of generation from deterministic verification is a useful contribution. The release of 208 archived programs, conductor-readback artifacts, simulation records, and a synthetic curriculum is unusually complete and should allow independent re-scoring. The physical studies are honestly tiered, and the fabrication-sensitivity experiments in Fig. 11 are a valuable caution against over-reading 2D masks. The main caveats are that every headline IoU is measured against automatically extracted, unvalidated reference masks, and that the benchmark does not estimate run-to-run variability for configuration-target pairs. Both issues affect the strength of the central claim rather than the quality of the released infrastructure.

major comments (3)
  1. [II.B, Figs. 3-4] The frozen reference targets are produced by Gemini 3 Pro Image/vision from published figures, and F8 is explicitly flagged as provenance-incomplete. No validation against original GDS layouts or independent manual re-digitization is reported for the other seven targets. Since Eq. (2) and Eq. (3) are computed against these masks, a systematic bias in mask extraction or in the pixel-per-micron calibration propagates to all IoU values, usable-IoU scores, retargeting footprint checks, and training rewards. The fixed-calibration rule (candidates are never rescaled) makes the results especially sensitive to footprint errors. I request (a) an independent validation of at least a subset of targets, e.g., by manual re-digitization or comparison with available GDS; (b) a sensitivity analysis over kappa_x and kappa_y; and (c) explicit extraction-confidence statements for F1-F7 analogous to the F8
  2. [III.C, Fig. 8] Sec. III.C states that each configuration is evaluated once on each target and that the benchmark does not estimate run-to-run variability. The bootstrap intervals resample targets, not repeated runs. The iterative API campaign in Sec. III.A reports a mean best-worst seed spread of 0.246 IoU and that 84% of experiments span at least 0.10, so single draws are high-variance. Consequently, the ordering Fable 5 max 0.974 vs. Opus 5 max 0.955, and the statement that twenty-two of 26 configurations are source-compliant on all eight targets, are not supported with any stochastic uncertainty. I request repeated runs for at least the top configurations, or a more careful claim such as 'observed mean' instead of 'consistently exceed.'
  3. [IV.B, IV.C, Table II] The fixed-representation pass/fail matrix in Table II is computed with the tier-0 analytic evaluator. The paper then shows in Sec. IV.C that this evaluator overestimates the directional-coupler coupling ratio by 5.5x (0.9439 vs. 0.1711 full-wave), and it explicitly notes that coupler entries in Table II are tier-0 outcomes. Moreover, the Opus-retargeted SOI MZI has an analytic FSR of 7.9995 nm but a full-wave fringe spacing of 7.77 nm, a 2.9% discrepancy that puts the design outside the stated 8.0 nm +/-2% target. This means the retargeting conclusions rest on a gate that is known to be inaccurate for at least one device class and may be outside tolerance for the headline MZI case. I recommend full-wave auditing of the disputed coupler/ring/spiral cells, or relabeling Table II as 'analytic reachability' rather than physical pass/fail retargeting.
minor comments (5)
  1. [Abstract, Fig. 8] The phrase 'consistently exceed 0.9 mean IoU' is ambiguous. Fig. 8(b) shows that per-target means for F3 and F5 are 0.695 and 0.660 averaged across configurations, and even top configurations have per-target lows. Recommend rewording to 'top configurations achieve mean-over-targets IoU above 0.9' or reporting per-target ranges for the top rows.
  2. [Eq. (4a), (4b)] The reward-shaping coefficients, the 0.05 floor, the Jrect threshold, and the chamfer scale 0.05d are not justified by ablations or sensitivity analysis. Since the training result is secondary to the reconstruction claim, this is not blocking, but the authors should either provide a short sensitivity study or state that these constants were chosen without tuning.
  3. [IV.A, Fig. 10] The off-diagonal FSR values are reported to four decimal places (e.g., 16.058, 8.817). Given the known analytic/full-wave discrepancy, this precision overstates the certainty of the underlying tier-0 model. Please round to a precision consistent with the evaluated model or add an uncertainty estimate.
  4. [Fig. 11(d)] The fabrication-sensitivity section notes that low-angle SOI fixtures have monitored output sums up to 1.084 and that normalization is used. This is an honest limitation, but it should also be mentioned in the conclusion or abstract so readers do not interpret the full-wave outputs as calibrated absolute transmissions.
  5. [Sec. V.B] The text says 'none of its 16 L4 responses executes on this probe draw' but also reports the 'All' row mean IoU 0.179. This is understandable, but the presentation could be clearer: the L4 column of Fig. 13(b) shows 0.00 for all checkpoints, which should be explicitly discussed as a failure mode of the training curriculum rather than left as an apparent artifact.

Circularity Check

0 steps flagged

No circular dependency found: benchmark targets are frozen external references and all claimed predictions are either held-out evaluations or explicitly tiered/audited outcomes.

full rationale

The paper's derivation chain is anchored to frozen target masks and physical footprints that are prepared once from published figures and reused across all configurations; every IoU score is computed by executing candidate programs against those frozen targets at fixed calibration, and the paper states explicitly that 'the candidate geometry is never rescaled to fit the target.' No parameter is fitted to the benchmark targets, and the headline 0.9+ IoU results are measured, not constructed, outcomes of agent search against these fixed references. The cross-stack retargeting claims are presented as tier-0 analytic evaluations, and the paper independently audits that same analytic tier with full-wave simulation, disclosing a 2.9% MZI discrepancy and a 5.5× coupler overestimate rather than presenting the analytic result as an unexamined prediction. The training experiment uses the geometric verifier as a reward, which is the intended training signal, but the reported evaluation is on held-out F1–F8 targets under the same frozen metric; this is standard RL evaluation, not circular. The use of the co-author's prior publication as the source of one target (F5) is a minor self-citation but is not load-bearing: F5 is one of eight published geometries, and the central claims do not reduce to that citation. The RC-01 limitation on shape-aware visual verification is a caveat about metric fidelity, not a demonstration that any result is equivalent to its input by construction. No specific equation, fitted parameter, or self-citation chain was found that reduces a claimed prediction to the benchmark inputs.

Axiom & Free-Parameter Ledger

3 free parameters · 5 axioms · 0 invented entities

The central claims rest on the GDSFactory toolchain semantics, the Gemini-derived target masks, the tier-0 analytic evaluator for retargeting gates, and the sufficiency of IoU as a parametric-correctness proxy. The paper is transparent about some of these, but they are load-bearing assumptions rather than derived results. No new physical entities are introduced; the system, DSL, and research contracts are methodological constructs, not scientific postulates.

free parameters (3)
  • Reward shaping coefficients (Eq. 4a/4b) = 0.05 + 0.40J + 0.15D + 0.40B(c) (weights 0.40, 0.15, 0.40; offset 0.05)
    Hand-chosen trade-offs between overlap and boundary terms in the GRPO reward; not derived from first principles.
  • Rectangle-baseline threshold = Jrect < 0.95
    Threshold deciding which reward branch applies; set by hand.
  • Chamfer distance scale = 0.05 * d (footprint diagonal)
    Defines T(c) in the boundary reward; a free scale chosen without independent justification.
axioms (5)
  • domain assumption GDSFactory 9.20.7 implements the DSL correctly
    All rendered IoU, retargeting geometry, and simulation inputs depend on this library's semantics; a tool bug would invalidate many results. Invoked in Section II.A.
  • domain assumption Gemini-3-Pro-prepared masks and footprints are accurate ground truth
    Targets are isolated from published figures and converted to binary masks by a vision model without validation against original GDS or manual reconstruction. F8 is provenance-incomplete. See Section II.B.
  • domain assumption The tier-0 analytic evaluator is sufficient for FSR and footprint gates
    Retargeting passes in Table II are judged by the same local-mode/supermode evaluator later shown to overestimate coupler ratio by 5.5× vs full-wave; MZI fringe spacing agrees within 2.9%. See Sections IV.A and IV.C.
  • domain assumption IoU on binary masks plus a single-variable perturbation is an adequate parametricity check
    The paper claims executable parametric programs, but independence of parameters is not tested. The perturbation test only confirms that changing one variable changes geometry. See Section II.D.
  • domain assumption The source-contract static checks correctly identify non-parametric programs
    Rejecting raw polygons and prebuilt cells is a proxy for authoring; it may not catch all degenerate programs. Invoked in Section II.D.

pith-pipeline@v1.3.0-alltime-deepseek · 16949 in / 17527 out tokens · 169099 ms · 2026-08-04T01:19:57.569570+00:00 · methodology

0 comments
read the original abstract

We present PixCell, a neurosymbolic system in which multimodal agents convert a visually presented photonic component into a parametric program over a small domain-specific language (DSL) of geometric primitives. A system enabling deterministic visual verification renders evaluation asymmetrically cheaper than the generation attempt. While models using multi-seed sampling and iterative revision reach a mean best-turn IoU of only 0.416, multimodal agents through PixCell's interface and verifier consistently exceed 0.9 mean IoU, with scores reaching 0.974 and 0.955 across eight component targets while also satisfying source contracts. These results demonstrate that frontier multimodal agents can reliably understand and render executable parametric representations from visual targets. Using these live parameters, cross-stack studies on an interferometer reconstruct primitive programs that satisfy an 8.0 nm free spectral range target and the original footprint constraint on modeled 220-nm SOI, 400-nm SiN, and 400-nm TFLN stacks. PixCell further carries a paper-derived splitter from visual reconstruction through SOI full-wave simulation, producing symmetric propagation and balanced outputs. Finally, the same executable verifier supplies a training reward and dataset used to train a Qwen3.6-35B-A3B model with LoRA and GRPO without supervised demonstrations. On eight training-excluded paper figures, its mean champion IoU rises from 0.422 after eight initial attempts to 0.491 after three verifier-guided revision rounds. These results therefore establish a controlled framework for measuring, retargeting, and improving visual-to-parametric photonic component design.

Figures

Figures reproduced from arXiv: 2608.00084 by Aadarsh Agarwal, Dirk Englund, Kenaish Al Qubaisi.

Figure 1
Figure 1. Figure 1: and Table I [10]. The language contains geomet￾ric regions, port-bearing waveguide elements, parameter￾ized paths and cross-sections, Boolean operations, and [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: FIG. 2 [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: FIG. 3 [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: FIG. 4 [PITH_FULL_IMAGE:figures/full_fig_p005_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: FIG. 5 [PITH_FULL_IMAGE:figures/full_fig_p006_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: FIG. 6 [PITH_FULL_IMAGE:figures/full_fig_p006_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: FIG. 7 [PITH_FULL_IMAGE:figures/full_fig_p007_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: FIG. 8 [PITH_FULL_IMAGE:figures/full_fig_p008_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: FIG. 9 [PITH_FULL_IMAGE:figures/full_fig_p009_9.png] view at source ↗
Figure 10
Figure 10. Figure 10: FIG. 10 [PITH_FULL_IMAGE:figures/full_fig_p010_10.png] view at source ↗
Figure 11
Figure 11. Figure 11: FIG. 11 [PITH_FULL_IMAGE:figures/full_fig_p011_11.png] view at source ↗
Figure 12
Figure 12. Figure 12: FIG. 12 [PITH_FULL_IMAGE:figures/full_fig_p012_12.png] view at source ↗
Figure 13
Figure 13. Figure 13: FIG. 13 [PITH_FULL_IMAGE:figures/full_fig_p013_13.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

46 extracted references · 11 linked inside Pith

  1. [1]

    Verification guides reconstruction, selection, and training throughout PixCell, and so better verification metrics are crucial

    RC-01, shape-aware visual verification, ex- tends geometric scoring so structurally correct re- constructions can be separated from visually sim- ilar shortcuts. Verification guides reconstruction, selection, and training throughout PixCell, and so better verification metrics are crucial

  2. [2]

    This connects the parametric freedom demonstrated in Sec

    RC-02, process-faithful 2D-to-3D retarget- ing, carries editable programs into fuller stack and fabrication models. This connects the parametric freedom demonstrated in Sec. IV to physical con- clusions beyond simplified two-dimensional geome- try

  3. [3]

    This supports controlled study of what a model learns from pixels, calibration, and program structure

    RC-03, representation and scale dataset, ex- pands the data linking topology, physical-scale evi- dence, and executable construction. This supports controlled study of what a model learns from pixels, calibration, and program structure

  4. [4]

    This extends the single open-weight training lineage studied here towards smaller, reproducible systems

    RC-04, smallest qualifying open model, maps how model scale affects reconstruction across the representation curriculum. This extends the single open-weight training lineage studied here towards smaller, reproducible systems. The current versioned terms, evidence requirements, and executable verdict logic for these directions are re- leased under research...

  5. [5]

    S. Yao, J. Zhao, D. Yu, N. Du, I. Shafran, K. Narasimhan, and Y. Cao, ReAct: Synergizing reasoning and acting in language models, in International Conference on Learn- ing Representations (2023) arXiv:2210.03629

  6. [6]

    J. Yang, C. E. Jimenez, A. Wettig, K. Lieret, S. Yao, K. Narasimhan, and O. Press, SWE-agent: Agent- computer interfaces enable automated software engineer- ing, in Advances in Neural Information Processing Sys- tems, Vol. 37 (2024)

  7. [7]

    T. Kwa, B. West, J. Becker, A. Deng, K. Garcia, M. Hasin, S. Jawhar, M. Kinniment, N. Rush, S. von Arx, et al., Measuring AI ability to complete long tasks, arXiv:2503.14499 (2025)

  8. [8]

    Novikov, N

    A. Novikov, N. V˜ u, M. Eisenberger, E. Dupont, P.- S. Huang, A. Z. Wagner, et al. , AlphaEvolve: A coding agent for scientific and algorithmic discovery, 15 arXiv:2506.13131 (2025)

  9. [9]

    OpenAI, GPT-4 technical report, arXiv:2303.08774 (2023)

  10. [10]

    C. Si, Y. Zhang, R. Li, Z. Yang, R. Liu, and D. Yang, Design2Code: Benchmarking multimodal code genera- tion for automated front-end engineering, in Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguis- tics: Human Language Technologies (2025) pp. 3956– 3974

  11. [11]

    J. Yang, C. E. Jimenez, A. L. Zhang, K. Lieret, J. Yang, X. Wu, O. Press, N. Muennighoff, G. Synnaeve, K. R. Narasimhan, D. Yang, S. I. Wang, and O. Press, SWE- bench multimodal: Do AI systems generalize to vi- sual software domains?, in International Conference on Learning Representations(2025)

  12. [12]

    Chrostowski and M

    L. Chrostowski and M. Hochberg, Silicon Photonics De- sign: From Devices to Systems (Cambridge University Press, Cambridge, 2015)

  13. [13]

    Bogaerts and L

    W. Bogaerts and L. Chrostowski, Silicon photonics cir- cuit design: methods, tools and challenges, Laser & Pho- tonics Reviews 12, 1700237 (2018)

  14. [14]

    Matres et al., GDSFactory: a Python library for chip design, https://github.com/gdsfactory/gdsfactory (2026), version 9.20.7; accessed July 28, 2026

    J. Matres et al., GDSFactory: a Python library for chip design, https://github.com/gdsfactory/gdsfactory (2026), version 9.20.7; accessed July 28, 2026

  15. [15]

    Y. Wu, X. Yu, H. Chen, Y. Luo, Y. Tong, and Y. Ma, PICBench: Benchmarking LLMs for photonic integrated circuits design, in Design, Automation & Test in Europe Conference (2025) pp. 1–6

  16. [16]

    Sharma, Y

    A. Sharma, Y. Fu, V. Ansari, R. Iyer, F. Kuang, K. Mis- try, R. I. Aishy, S. Ahmad, J. Matres, D. R. Englund, and J. K. S. Poon, AI agents for photonic integrated circuit design automation, APL Machine Learning 3, 046113 (2025)

  17. [17]

    Kharel, A

    P. Kharel, A. Khavasi, X. Chen, and T. W. Hughes, Au- tonomous agentic design for photonics, arXiv:2606.00915 (2026)

  18. [18]

    T. W. Hughes, M. Minkov, V. Liu, Z. Yu, and S. Fan, A perspective on the pathway toward full wave simu- lation of large area metalenses, Applied Physics Let- ters 119, 150502 (2021), the Tidy3D solver: https: //www.flexcompute.com/tidy3d/

  19. [19]

    Chaudhuri, K

    S. Chaudhuri, K. Ellis, O. Polozov, R. Singh, A. Solar- Lezama, and Y. Yue, Neurosymbolic programming, Foundations and Trends in Programming Languages 7, 158 (2021)

  20. [20]

    Ellis, D

    K. Ellis, D. Ritchie, A. Solar-Lezama, and J. B. Tenen- baum, Learning to infer graphics programs from hand- drawn images, in Advances in Neural Information Pro- cessing Systems (2018)

  21. [21]

    C. Chen, J. Wei, T. Chen, C. Zhang, X. Yang, S. Zhang, B. Yang, C.-S. Foo, G. Lin, Q. Huang, and F. Liu, CADCrafter: Generating computer-aided design mod- els from unconstrained images, inIEEE/CVF Conference on Computer Vision and Pattern Recognition(2025) pp. 11073–11082

  22. [22]

    Z. Shao, P. Wang, Q. Zhu, R. Xu, J. Song, X. Bi, H. Zhang, M. Zhang, Y. K. Li, Y. Wu, and D. Guo, DeepSeekMath: Pushing the limits of math- ematical reasoning in open language models (2024), arXiv:2402.03300

  23. [23]

    D. Guo, D. Yang, H. Zhang, et al., DeepSeek-R1 incen- tivizes reasoning in LLMs through reinforcement learn- ing, Nature 645, 633 (2025)

  24. [24]

    Cobbe, V

    K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, C. Hesse, and J. Schulman, Training verifiers to solve math word problems, arXiv:2110.14168 (2021)

  25. [25]

    Brown, J

    B. Brown, J. Juravsky, R. Ehrlich, R. Clark, Q. V. Le, C. R´ e, and A. Mirhoseini, Large language mon- keys: Scaling inference compute with repeated sampling, arXiv:2407.21787 (2024)

  26. [26]

    Snell, J

    C. Snell, J. Lee, K. Xu, and A. Kumar, Scaling LLM test- time compute optimally can be more effective than scal- ing parameters for reasoning, in International Conference on Learning Representations(2025) arXiv:2408.03314

  27. [27]

    Wei, Asymmetry of verification and veri- fier’s rule, https://www.jasonwei.net/blog/ asymmetry-of-verification-and-verifiers-law (2025)

    J. Wei, Asymmetry of verification and veri- fier’s rule, https://www.jasonwei.net/blog/ asymmetry-of-verification-and-verifiers-law (2025)

  28. [28]

    Y. Chen, Y. Shen, W. Huang, S. Zhou, Q. Lin, X. Cai, Z. Yu, J. Bu, B. Shi, and Y. Qiao, Learning only with im- ages: Visual reinforcement learning with reasoning, ren- dering, and visual feedback, arXiv:2507.20766 (2025)

  29. [29]

    J. Li, Y. Luo, Y. Lou, and X. Zhou, ReCAD: Re- inforcement learning enhanced parametric CAD model generation with vision-language models, Proceedings of the AAAI Conference on Artificial Intelligence 40, 6190 (2026)

  30. [30]

    V. H. Nguyen, I. K. Kim, and T. J. Seok, Low-loss and broadband silicon photonic 3-dB power splitter with en- hanced coupling of shallow-etched rib waveguides, Ap- plied Sciences 10, 4507 (2020)

  31. [31]

    Malka, Y

    D. Malka, Y. Danan, Y. Ramon, and Z. Zalevsky, A pho- tonic 1 × 4 power splitter based on multimode interfer- ence in silicon–gallium-nitride slot waveguide structures, Materials 9, 516 (2016)

  32. [32]

    D. Mao, Y. Wang, E. El-Fiky, L. Xu, A. Kumar, M. Jaques, A. Samani, O. Carpentier, S. Bernal, M. S. Alam, J. Zhang, M. Zhu, P.-C. Koh, and D. V. Plant, Adiabatic coupler with design-intended splitting ratio, Journal of Lightwave Technology 37, 6147 (2019)

  33. [33]

    Huang, K

    P. Huang, K. Chen, and L. Liu, Fabrication-tolerant di- rectional couplers on thin-film lithium niobate, Optics Letters 48, 1264 (2023)

  34. [34]

    Al Qubaisi and M

    K. Al Qubaisi and M. A. Popovi´ c, Photonic resonators with microring-like behavior based on standing wave cav- ity pairs with opposite-symmetry modes, in Frontiers in Optics / Laser Science(2020) p. FTu8E.2

  35. [35]

    Chandran, M

    S. Chandran, M. Dahlem, Y. Bian, et al., Beam shaping for ultra-compact waveguide crossings on monolithic sil- icon photonics platform, Optics Letters 45, 6230 (2020)

  36. [36]

    Q. Deng, A. H. El-Saeed, A. Elshazly, et al., Low-loss and low-power silicon ring based WDM 32 × 100 GHz filter enabled by a novel bend design, Laser & Photonics Reviews 19, 2401357 (2025)

  37. [37]

    Google DeepMind, Gemini 3 Pro image model card, https://storage.googleapis.com/deepmind-media/ Model-Cards/Gemini-3-Pro-Image-Model-Card.pdf (2025)

  38. [38]

    Google DeepMind, Gemini 3 Pro model card, https://deepmind.google/models/model-cards/ gemini-3-pro/ (2026), first published November 2025

  39. [39]

    Anthropic, Claude code, https://claude.com/ claude-code (2026), accessed July 28, 2026

  40. [40]

    OpenAI, Codex cli, https://github.com/openai/codex (2026), accessed July 28, 2026. 16

  41. [41]

    com/api/docs/pricing (2026), accessed July 27, 2026

    OpenAI, API pricing, https://developers.openai. com/api/docs/pricing (2026), accessed July 27, 2026

  42. [42]

    claude.com/docs/en/about-claude/pricing (2026), accessed July 27, 2026

    Anthropic, Claude API pricing, https://platform. claude.com/docs/en/about-claude/pricing (2026), accessed July 27, 2026

  43. [43]

    Agarwal, K

    A. Agarwal, K. Al Qubaisi, and D. Englund, PixCell Dataset: Representation-first image-to-code curriculum, Hugging Face dataset (2026), version 2.0.0

  44. [44]

    6-35b-a3b (2026)

    Qwen Team, Qwen3.6-35B-A3B: Agentic coding power, now open to all, https://qwen.ai/blog?id=qwen3. 6-35b-a3b (2026)

  45. [45]

    E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen, LoRA: Low- rank adaptation of large language models, in Interna- tional Conference on Learning Representations (2022) arXiv:2106.09685

  46. [46]

    Thinking Machines Lab, Tinker: a training API for researchers, https://thinkingmachines.ai/tinker (2025)