{"id":"efd62032-57f4-47f6-a935-3cefe4de876a","arxiv_id":"2607.08392","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":8,"one_line_summary":"A joint discrete–continuous flow-matching model designs open-vocabulary multilayer optical coatings from query-time n/k curves and spectra, validated on 224 tasks and four fabricated coolers.","lead":"IrisFlow designs multilayer optical coatings from a target spectrum, wavelength grid, and material optical curves supplied at query time, without fixed material IDs or thickness bins. A single trained model spans 2–100 layers, generalizes to held-out materials, and produced four fabricated color-displaying coolers.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"Oracle best-of-N TMM ranking and post-hoc selection rules carry more of the open-vocabulary claim than the generative model alone.","rationale":"The reader correctly isolates the load-bearing soft spot: oracle best-of-N under exact TMM, plus post-hoc cooling/practicality selection, is what makes the open-vocabulary and fabrication claims look strong. The interface itself (n/k tokens, continuous thickness, joint discrete–continuous flow, query wavelength grid) is well supported by the multi-tier benchmark and the process-corrected cooler loop, so the paper remains accept-shaped. Confidence stays moderate because code/data are request-only and the classical comparison is single-target, but those are secondary. No stronger internal inconsistency appears: OOD materials enter only as curves, wavelength extrapolation is cleanly isolated, and the mean-field CTMC gaps are acknowledged. The concrete test above would settle whether the generative model, rather than the TMM ranker and selection rules, is doing the work the abstract attributes to IrisFlow. Verdict therefore stays CONDITIONAL with the same caveats the reader already listed.","tokens_in":66807,"tokens_out":607,"duration_ms":7210,"concrete_test":"Re-score Tier 1 FULL cells, the 11 OptoGPT cases, and the four cooler pools under (i) N=1 greedy decode and (ii) model-only ranking without TMM (e.g., predicted clean-material posterior mass × thickness residual). If median RMSE rises by >2× or cooler ∆E00 exceeds ~6–8 under either protocol, the amortized open-vocabulary claim is overstated relative to oracle selection.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim is that a single query-conditioned model designs open-vocabulary coatings that transfer to hardware. Almost every headline number is the best of N=100–500 independent draws ranked by exact TMM re-simulation (Methods §4.10; Appendix §G.2), not a single-shot or model-internal ranking. Tier 1/2 medians, Tier 3 application wins, the OptoGPT 10–1 record, and the cooler pool all use this oracle. Appendix §D.2 shows a single deterministic draw is far worse (mean RMSE ~0.59 vs ~0.11 at N=100), so reported fidelity is selection quality as much as generative quality. The fabrication loop compounds this: designs are chosen by a cooling-aware / practicality rule outside the model, sit at the ∆E00≤3 boundary with no deposition headroom, magenta measures 5.2, and yellow is swapped for a simpler in-pool stack (Results §2.8; Appendix §N). Process-corrected n/k is a genuine open-vocabulary strength, but the hardware transfer still depends on TMM ranking plus human/process filters the model does not supply.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"IrisFlow is a single 136M-parameter query-conditioned generative model for multilayer optical coatings that treats materials as wavelength-aware n(λ),k(λ) tokens (not fixed IDs), samples material sequences by discrete flow matching over a query-local candidate bank, and samples thicknesses by continuous flow matching without discretization. A staged curriculum trains one checkpoint for 2–100 layers. On a 224-task suite the model reconstructs in-distribution targets (median combined R,T RMSE ~4.6×10⁻²), retains same-order accuracy on a 15-material held-out bank without retraining (median matched-cell OOD/in-distribution ratio ~1.67×), extrapolates wavelength support beyond the 380–1400 nm training envelope, designs against analytic application targets, and outperforms OptoGPT on an 11-case shared subset using OptoGPT’s library as OOD curves. With process-corrected optical constants, four color-displaying coolers were designed and fabricated; the three chromatic devices reach CIEDE2000 3.1–5.2 with 93–95% solar NIR reflectance.","tokens_in":67096,"tokens_out":1306,"duration_ms":22348,"significance":"If the results hold, this is a substantial interface contribution for amortized inverse design in photonics: open-vocabulary material entry via optical curves, query-conditioned wavelength grids, continuous thicknesses jointly generated with discrete material choices, and a single model spanning 2–100 layers. The evaluation is unusually thorough for the area (matched-cell OOD bank, nearest-training-sample audit against ~114M spectra, wavelength extrapolation, classical-optimizer budget comparison, partial-structure clamping, angle/polarization via bank tilt, and a fabrication loop). Appendix A’s consistency account of the joint discrete–continuous flow, the diversity analysis, and the process-corrected n/k fabrication path are genuine strengths. The work is relevant beyond coatings to other coupled discrete–continuous design problems with facility-local component banks.","major_comments":[{"comment":"Methods §4.10 and Appendix §G.2: essentially every headline fidelity number (Tiers 1–4, OptoGPT head-to-head, cooler pools) is best-of-N under oracle TMM re-simulation (N=100 or 500). Appendix §D.2 shows a single deterministic draw is far worse (mean RMSE ~0.59 vs ~0.11 at N=100). This is defensible for a one-to-many inverse map and is partially framed in Discussion as generator + cheap TMM filter, but the Abstract and Results still read as if the generative model alone delivers the reported RMSE. Please state the evaluation mode explicitly in the Abstract/Results, and report single-draw (or small-N) medians alongside best-of-N for at least Tier 1 and Tier 2 so readers can separate generative quality from selection quality.","section":null},{"comment":"Results §2.8 and Appendix §N: the hardware claim is that open-vocabulary design is “carried through to fabricated coatings,” yet (i) designs are selected by a cooling-aware / practicality rule outside the model, (ii) selected designs sit at the ∆E00≤3 boundary with no deposition headroom, (iii) measured magenta is ∆E00=5.2 (above the stated tolerance), and (iv) yellow was swapped for a simpler in-pool stack. Process-corrected n/k is a real open-vocabulary strength; please qualify the fabrication claim to match the measured outcomes (e.g., three chromatic coolers with stated color/cooling metrics, one limiting near-black case) and separate model generation from post-hoc selection more clearly in the main text.","section":null},{"comment":"Appendix §K / Table 2: the OptoGPT comparison is valuable as an OOD-material transfer test, but total draw budgets differ (500 for OptoGPT vs 19×500 for IrisFlow’s layer sweep) and models are not parameter-matched. The paper argues per-configuration N=500 matching and reports sizes; still, the 10–1 “wins” framing in Results §2.5 overstates a protocol that favors IrisFlow’s controllable depth axis. Soften the win language and lead with median/mean RMSE and the OOD-bank aspect rather than pairwise case counts.","section":null}],"minor_comments":[{"comment":"Figure 2 heatmaps: cell text is RMSE×10²; state this explicitly in the caption (it is easy to misread as raw RMSE).","section":null},{"comment":"Table 1 vs Appendix Table I1: AR-5 handling (median of two variants) should be footnoted in the main-text table for consistency with §G.2.","section":null},{"comment":"Methods §4.6: λ_st=0.4 and λ_soft=0.1 are fixed from a pilot; a one-sentence sensitivity note (or pointer to any ablation) would help reproducibility.","section":null},{"comment":"Appendix §P: the p-polarization admittance proxy residual is quantified in SI but only briefly mentioned in Discussion; a short main-text sentence on residual magnitude would help practitioners.","section":null},{"comment":"Code/data availability is “upon reasonable request.” For a methods paper of this type, releasing at least the benchmark case registry, effective n/k libraries, and evaluation scripts would strengthen the contribution.","section":null},{"comment":"Minor notation: “h4 (LaTiO3)” appears without first expansion in some early figures; define once in main text.","section":null}],"recommendation":"minor_revision","confidential_remarks":"Solid methods paper with unusually complete evaluation and a real fabrication loop. The best-of-N and fabrication-selection issues are presentation/framing problems, not core technical failures; minor revision is appropriate. Scope fits optics/photonics inverse-design venues well. I would not block on the OptoGPT protocol if the authors soften the win language."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The real news is the interface, not another black-box spectrum-to-stack net. Materials enter as wavelength-aware n(λ),k(λ) tokens on the user’s grid, thicknesses stay continuous under flow matching, materials use uniform-CTMC discrete flow over a query-local bank, and one 136M checkpoint covers 2–100 layers. That combination is new relative to OptoGPT-style AR tokenizers, tandem/MDN/cVAE/cINN, and classical TMM optimizers, and the paper positions the prior art cleanly.\n\nWhat they did well: a 224-task suite (in-distribution, held-out 15-material bank with matched-cell ~1.7× median ratio, analytic application targets with a 114M nearest-neighbor audit, wavelength extrapolation past 1400 nm), an OptoGPT head-to-head on OptoGPT’s own library as OOD curves, partial-stack clamping, angle/polarization via bank tilt plus exact angled TMM, and four process-calibrated coolers actually deposited. Appendix A is careful about mean-field reverse kernels and what the loss recovers. Fabrication with chamber-remeasured GSST-a is the right kind of evidence for open vocabulary.\n\nSoft spots, in proportion: almost every headline number is oracle best-of-N (100 or 500) under exact TMM re-simulation. Appendix D shows a single deterministic draw is much worse, so reported fidelity is generative quality plus selection. Cooler selection sits at the ∆E00≤3 boundary; magenta measures 5.2, yellow was swapped for a simpler in-pool stack, black is far by construction. Classical comparison is one target. Code/data are request-only. None of that erases the interface result; it bounds how turnkey you should treat the system.\n\nThis is for people who care about amortized inverse design interfaces and facility-facing photonics. Math and citation pattern look solid. I would send it to referees and would cite the open-vocabulary + joint flow framing. Engage.","headline":"Solid methods paper: open-vocabulary n/k tokens + joint discrete–continuous flow matching for coatings, with real fabrication; best-of-N TMM ranking is the main caveat, not a collapse of the claim.","tokens_in":67791,"tokens_out":524,"would_cite":true,"duration_ms":8824,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"A single open-vocabulary flow model designs 2–100-layer optical coatings from any material curves and wavelength grid, without retraining.","keywords":["query-based inverse design","open-vocabulary models","flow matching","joint discrete–continuous flow matching","multilayer optical coatings","photonics","transfer-matrix method"],"falsifier":"Supply a genuinely novel chamber-measured n,k curve never near the training neighborhood, freeze the same checkpoint and best-of-N budget, and check whether the best TMM-verified design still lands within roughly 2–3× the matched in-distribution RMSE and, after deposition, within the claimed color and solar-NIR tolerances without swapping stacks or relaxing the selection rule.","tokens_in":67611,"feed_emoji":"🔬","tokens_out":731,"duration_ms":8496,"temperature":0.7,"pith_summary":"Amortized neural inverse design usually freezes the material menu, the wavelength grid, and a thickness quantization into the model, so a new film or a new band forces retraining. This paper argues that multilayer optical coatings need not pay that closed-world price. IrisFlow is a single 136M-parameter model that accepts, at query time, a target reflectance/transmittance spectrum, a wavelength grid, a local bank of material n(λ),k(λ) curves, and a requested layer count. Materials enter as wavelength-aware optical tokens rather than learned IDs; material sequences are generated by discrete flow matching over the query bank and thicknesses by continuous flow matching, jointly under one shared denoising backbone. The same checkpoint reconstructs in-distribution targets, keeps same-order accuracy on a held-out 15-material bank, handles bands beyond the training envelope, and designs four color-displaying radiative coolers that were fabricated and measured. If the interface works, a coating facility can drop in chamber-calibrated optical constants and get usable stacks without rebuilding the model.","feed_headline":"One model designs any coating from material curves you plug in","feed_subtitle":"No fixed material menu or wavelength grid: open-vocabulary flow matching reaches fabricated color coolers.","key_machinery":"Joint discrete–continuous flow matching under a shared denoising backbone: materials are scored by discrete flow matching (uniform CTMC) against a query-local bank of wavelength-aware n,k tokens; thicknesses are integrated by continuous flow matching without discretization; both heads condition on the full joint noisy state and the query.","core_discovery":"IrisFlow shows that open-vocabulary, query-conditioned joint discrete–continuous flow matching can amortize inverse design of multilayer optical coatings: one 136M-parameter checkpoint designs 2–100-layer stacks from user-supplied spectra, wavelength grids, and candidate n(λ),k(λ) curves, reconstructs in-distribution targets at median combined R,T RMSE about 4.6×10⁻², retains same-order accuracy on a held-out material bank without retraining, and produces process-calibrated cooler designs that were fabricated.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["IrisFlow designs multilayer coatings from any material curves you supply","Open-vocab joint flows invent optical stacks for spectra you query","One model samples materials and thicknesses for 2-100 layer coatings","Query-time optical constants yield fabricated cooler designs","Joint discrete-continuous flow matching amortizes coating inverse design"],"cache_read_input_tokens":49280,"weakest_assumption_plain":"That ranking many random draws by exact transfer-matrix re-simulation, plus human or cooling-aware selection from that pool, is a fair stand-in for what the generative model itself delivers when a new material curve is simply plugged in.","fun_headline_variants_meta":{"raw":{"variants":["IrisFlow designs multilayer coatings from any material curves you supply","Open-vocab joint flows invent optical stacks for spectra you query","One model samples materials and thicknesses for 2-100 layer coatings","Query-time optical constants yield fabricated cooler designs","Joint discrete-continuous flow matching amortizes coating inverse design"]},"model":"grok-4.5","effort":"low","cost_usd":0.006894,"raw_usage":{"total_tokens":1789,"prompt_tokens":864,"num_sources_used":0,"completion_tokens":84,"cost_in_usd_ticks":68940000,"prompt_tokens_details":{"text_tokens":864,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":841,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":864,"tokens_out":84,"duration_ms":7910,"temperature":1.0,"reasoning_tokens":841,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-10T08:11:34.083908+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Supply a genuinely novel chamber-measured n,k curve never near the training neighborhood, freeze the same checkpoint and best-of-N budget, and check whether the best TMM-verified design still lands within roughly 2–3× the matched in-distribution RMSE and, after deposition, within the claimed color and solar-NIR tolerances without swapping stacks or relaxing the selection rule.","supporting_citations":[],"review_version":1}