{"id":"38404fe5-d66a-41c9-96af-cd44c9820967","arxiv_id":"2607.03256","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.5,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A three-layer perturbation probe shows latent selectivity is a near-binary rectified-flow fingerprint that survives ADD distillation, while score selectivity tracks distillation objective across 23 T2I models.","lead":"A training-free probe perturbs prompt, latent, and score layers of text-to-image diffusion models and measures which responses change. Across 23 teachers and few-step students it finds the latent layer nearly binary-detects rectified-flow prediction type, while the score layer tracks distillation objective.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"The RF latent detector rests on an ad-hoc band rule and a single ε-pred DiT control; residual architecture confounds remain unruled-out.","rationale":"The paper is careful, estimator-matched, and correctly scopes secondary score findings. The strongest claim is the latent RF split, which holds under the stated band rule across the released tables and is strengthened by the PixArt control and teacher/student RF pairs. The reader's weakest_assumption correctly identifies the soft spot: the band criterion is ad-hoc, non-UNet ε-pred coverage is a single control, and no non-ADD RF student is present, so residual confounds (architecture class, conditioning stack, resolution) are not fully isolated. That does not overturn the within-sweep empirical observation, but it keeps the causal reading of 'prediction type' provisional. No stronger internal inconsistency appears; the concern is design thinness rather than contradiction. Verdict therefore stays CONDITIONAL, matching the reader.","tokens_in":23991,"tokens_out":621,"duration_ms":5695,"concrete_test":"Apply the identical matched estimator and sustained-band rule to one additional public ε-pred DiT/MMDiT checkpoint that is not PixArt (e.g., a non-T5 or differently-conditioned ε-pred transformer if available) and, if possible, one non-ADD RF student or RF teacher under a different sampler; if any ε-pred model forms the band or an RF model loses it, the causal claim to prediction type weakens.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim (latent layer as near-binary RF detector) is defined by the sustained-band criterion in §3.1: R_lo > 1 on ≥3 of {0.05,0.1,0.2,0.3} including one s≥0.2. That rule cleanly separates the four RF cases (SD3.5/FLUX teachers + ADD students) from all ε-pred models in the 23-model tables, and PixArt-α (T5, ε-pred) rules out wide-T5 alone. However, the rule itself is post-hoc and not derived from flow-matching geometry; the only non-UNet ε-pred control is one DiT (PixArt), while RF cases are MMDiT/hybrid-DiT with different conditioning stacks, VAE resolutions, and training recipes. No non-ADD RF student exists, so survival under distillation is shown only for pure ADD. Thus the attribution to prediction type (vs residual architecture/conditioning) is still under-determined by the design in §4.3/§5.2, even though the empirical split is clean inside the sweep.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper proposes a training-free, decomposable probe for diffusion text-to-image models that injects controlled mean/variance/scale perturbations at three forward-pass layers (prompt encoder, denoiser input/latent, denoiser output/score) and summarises each cell by a bootstrap-median Bures W2^2 selectivity ratio on Inception-v3 features. Under one matched estimator on a 23-model sweep (five teachers, 18 few-step students; SDXL/SD1.5/SD3.5/PixArt-α/FLUX; UNet/DiT/MMDiT; five distillation paradigms), the three layers are reported to track three empirically separable factors: a universal prompt-mean response, prediction type (rectified-flow vs ε-prediction) on the latent layer, and distillation objective on the score layer. The main claim is that, within this sweep, a sustained low-to-mid latent band (R_lo>1 on a stated multi-strength rule) appears only for rectified-flow backbones (SD3.5, FLUX) as both teachers and ADD students; PixArt-α (T5, ε-prediction) is used to rule out wide-T5 conditioning alone. Two narrower score-layer findings are a 4-step ADD-vs-rest contrast on UNet families and a CI-separated early-strength spike on trajectory-rollout students (UNet and DiT). Per-cell CI tables and the estimator are released.","tokens_in":24300,"tokens_out":1706,"duration_ms":30655,"significance":"If the probe is reliable, it is a useful diagnostic instrument for a literature that still reports few-step quality almost exclusively via end-to-end FID/CLIP scalars. The matched estimator, bootstrap-median Bures ratios, explicit scoping of secondary findings, and public release of per-cell tables are real methodological strengths and make the empirical grid citable. The latent-layer RF fingerprint is an interesting, falsifiable empirical pattern even if its causal attribution remains incomplete. Downstream uses (recipe auditing, training-time signals) are correctly left open. The contribution is primarily an instrument plus carefully scoped readings rather than a mechanistic theory of flow matching or distillation; that is still valuable for cs.CV if the main separation is robust under the stated design.","major_comments":[{"comment":"§3.1 and §4.3: The main RF detector is the sustained-band rule (R_lo>1 on at least three of {0.05,0.1,0.2,0.3}, including one s≥0.2). This rule cleanly separates the four RF cases from all ε-prediction models in the released tables, but it is not derived from flow-matching geometry and appears tuned to the observed RF shape (vs isolated low-s excursions such as PixArt-LCM). Because the central claim is defined by this criterion, the paper needs either (i) a short sensitivity analysis over nearby band definitions / thresholds, or (ii) explicit language that the detector is an empirical fingerprint criterion chosen for this sweep, not a pre-specified or theory-derived test. Without that, the near-binary claim is harder to evaluate outside the current grid.","section":null},{"comment":"§4.3, Table 3, Table 8, and §5.2: Attribution of the latent band to prediction type (rather than residual architecture/conditioning/training confounds) rests on a single ε-prediction DiT control (PixArt-α) that holds T5 fixed. RF cases are MMDiT / hybrid-DiT with different conditioning stacks, resolutions, VAEs, and training recipes; the 2×2 in Table 8 has no RF+CLIP-only cell and only one non-UNet ε cell. The PixArt control rules out wide-T5 alone, which is useful, but does not fully isolate prediction type. The main claim should either soften from “prediction type” to “RF backbone family as instantiated in this sweep” or add a dedicated limitations paragraph that lists the remaining confounds as first-order, not secondary.","section":null},{"comment":"§4.3 and §5.1: Survival of the latent fingerprint is shown only for pure-ADD RF students (SD3.5-Turbo, FLUX-schnell); there is no non-ADD RF student. The body is careful (“survives ADD”), but the abstract and introduction still frame a more general “survives distillation” reading. Align abstract/intro with the precise §5.1 statement, and treat “no non-ADD RF student” as a load-bearing scope limit on the main result rather than a minor coverage gap.","section":null},{"comment":"§3.1 Eq. (1) and Appendix A: The headline statistic is amp_mean / max(amp_var, amp_scale) under a Gaussian Bures W2^2 on Inception pool3. The bootstrap-median fix for plug-in bias is well motivated, but the paper does not show that the mean-vs-higher-moment ratio (as opposed to raw amplitudes, other feature spaces, or non-Gaussian OT) is the quantity that isolates prediction type. A short ablation—e.g., reporting whether the RF band survives under DINOv2 features, under amp_mean alone, or under a diagonal-covariance control already mentioned as biased—would make the instrument choice less free-parameter-like for the central claim.","section":null}],"minor_comments":[{"comment":"Figure 2 is information-dense; annotating the sustained-band RF rows and the s=0.5 ADD column more explicitly (or splitting latent vs score panels) would help readers verify the three patterns without the appendix tables.","section":null},{"comment":"Several reported intervals collapse to a single two-decimal value (e.g., 0.66[0.66,0.66]). The footnote explains rounding, but stating n_resample and effective n per cell once in the main text (not only Appendix) would reduce the appearance of zero-width CIs.","section":null},{"comment":"Table 1 / paradigm labels: progressive-adv vs ADD vs mixed are operationally clear in §2.1, but a one-line mapping from vendor checkpoint names to loss-family labels in the table caption would reduce cross-referencing.","section":null},{"comment":"§4.6 prompt-collapse observation is correctly marked as non-law; consider moving it fully to appendix or discussion so it does not compete with the three structured findings.","section":null},{"comment":"Notation: R_sel, R_point, R_lo, R_hi are introduced cleanly; keep the equality line R=1, the band criterion, and the R_lo≥2 visual marker visually distinct in all figures (the text already warns against conflating them).","section":null}],"recommendation":"major_revision","confidential_remarks":"This is a careful empirical methods paper with honest scoping and a real release commitment. The main risk for the journal is over-reading a post-hoc band rule as a causal prediction-type detector; if the authors tighten attribution language and add a short sensitivity/ablation package, it becomes a solid accept-level diagnostic contribution. Fit is good for a CV/ML venue that values instruments and systematic sweeps; novelty is in the probe and the matched multi-family grid rather than theory of rectified flow."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The useful thing here is not another FID number. It is a training-free, three-layer perturbation probe (prompt / latent / score × mean-var-scale) with a single bootstrap-median Bures W2² selectivity ratio run identically on 23 public checkpoints. Inside that grid the layers actually separate: prompt is universal mean-selectivity (sanity channel), latent tracks rectified-flow vs ε-prediction as a sustained low-to-mid band, and score tracks distillation objective (ADD lowest at matched 4-step; trajectory-rollout early spike). That decomposition is new relative to the distillation, CFG-geometry, and FreeU/PAG literature they cite, and they are unusually explicit about scope.\n\nWhat they do well: matched estimator across SDXL/SD1.5/SD3.5/PixArt/FLUX, full per-cell CIs, PixArt as the T5-but-ε control that kills the obvious conditioning confound, and four RF cases (teachers + ADD students on MMDiT and DiT) that keep the latent band. Secondary score claims are correctly narrowed to a 4-step binary cut and a step/backbone-dependent spike. They release the tables and estimator; that is real work.\n\nSoft spots, in proportion. The RF detector is defined by a post-hoc sustained-band rule (R_lo>1 on ≥3 of four low-mid strengths including one ≥0.2). It cleanly partitions the sweep, but it is not derived from flow-matching geometry. Non-UNet coverage is three students; there is no non-ADD RF student, so “survives distillation” is only “survives ADD.” Residual architecture/VAE/conditioning differences between PixArt and the RF models are not fully closed. Those are design limits they mostly own in §5.2, not hidden contradictions. The Gaussian-Inception Bures choice is conventional and audited against plug-in bias; free parameters (strength grid, band rule) are stated.\n\nThis is for people who build or audit few-step students and want something more diagnostic than FID/CLIP. It is not a theory paper and not a universal detector claim. I would bring it to reading group, cite the probe and the RF-band observation when discussing distillation diagnostics, and send it to peer review. A referee should push on the band criterion and ask for one more control if possible, but the empirical instrument and the scoped findings already earn the time.","headline":"A careful, estimator-matched diagnostic that cleanly separates prediction type from distillation objective inside a 23-model sweep; the RF latent band is real in the data, but the causal attribution still rests on an ad-hoc rule and thin non-UNet controls.","tokens_in":24917,"tokens_out":611,"would_cite":true,"duration_ms":6179,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"A three-layer probe of few-step diffusion models separates prediction type from distillation objective across 23 text-to-image checkpoints.","keywords":["few-step diffusion","distillation","rectified flow","selectivity probe","Bures Wasserstein","text-to-image","latent diffusion","adversarial distillation"],"falsifier":"An epsilon-prediction model that forms a sustained latent band (R_lo > 1 on at least three of the four strengths in {0.05, 0.1, 0.2, 0.3}, including one s ≥ 0.2) under the same matched estimator, or a non-ADD rectified-flow student that loses that band, would falsify the prediction-type detector.","tokens_in":24835,"feed_emoji":"🔬","tokens_out":1101,"duration_ms":17072,"temperature":0.7,"pith_summary":"Few-step distilled diffusion models cut text-to-image sampling from dozens of network evaluations to a handful, but the quality gap is usually reported as a single FID or CLIP number that cannot say which part of the conditioning response changed. This paper replaces that scalar with a training-free probe that injects controlled mean, variance, and scale perturbations at three sites—the prompt encoder, the denoiser input, and the denoiser output—and summarizes each cell by a bootstrap-median Bures W₂² selectivity ratio on Inception features. Under one matched estimator applied to five teachers and eighteen students spanning UNet, DiT, and MMDiT backbones and five distillation paradigms, the three layers track three separable factors: the prompt layer is a universal mean response (a sanity channel), the latent layer reads prediction type, and the score layer reads the distillation objective. The main result is that only rectified-flow models form a sustained elevated latent selectivity band, and that band survives pure adversarial distillation on both teachers and students; a matched epsilon-prediction T5 control rules out wide text conditioning as the cause. Secondary score-layer patterns, under narrower scope, flag adversarial students at four steps and reveal an early-strength spike for trajectory-rollout students on both UNet and DiT.","feed_headline":"Probe finds rectified-flow fingerprint that survives distillation","feed_subtitle":"Across 23 text-to-image models, three layers separately track prediction type and training recipe.","key_machinery":"A decomposable layer-/mode-resolved probe: controlled mean, variance, and scale perturbations of six strengths injected at prompt encoder, denoiser input (latent), and denoiser output (score), summarized by the bootstrap-median Bures W₂² selectivity ratio R = amp_mean / max(amp_var, amp_scale) on Inception features under one matched estimator across all models.","core_discovery":"Within this 23-model sweep the latent-layer selectivity ratio exceeds 1 across a sustained low-to-mid strength band only for rectified-flow backbones (SD3.5, FLUX), as both teachers and adversarially distilled students; no epsilon-prediction model forms that band. A T5-conditioned epsilon-prediction control (PixArt-α) does not reproduce the band, attributing the fingerprint to prediction type rather than wide conditioning, and the fingerprint survives ADD distillation. The score layer separately tracks distillation objective via a 4-step ADD-versus-rest contrast and a CI-separated early-strength spike on trajectory-rollout students.","pith_inferences":["If the latent band truly isolates prediction type, it could audit released checkpoints whose training recipe is undisclosed without relying on architecture metadata.","The same per-layer readings could be used as an online training signal to steer distillation away from collapse on latent or score selectivity relative to the teacher.","A non-adversarial rectified-flow student would test whether the latent fingerprint is objective-invariant or only known to survive ADD.","An orthogonal diversity channel (for example Vendi on DINOv2 features) could separate sample-diversity collapse from the conditioning-response changes the current probe measures."],"forward_implications":["The latent-layer sustained band can serve as an empirical fingerprint of rectified-flow prediction type that survives pure adversarial distillation.","At fixed step count, score-layer ratios can distinguish adversarial-dominated distillation from other paradigms without training logs.","Trajectory-rollout objectives leave a detectable CI-separated early-strength score spike on both UNet and DiT.","Layer-resolved selectivity ratios can diagnose which axis of conditioning response changed after distillation, unlike single FID/CLIP scalars.","The released per-cell tables and matched estimator enable direct, CI-citable cross-model comparison under identical statistics."],"fun_headline_variants":["Latent layer flags rectified-flow backbones past distillation","Probe isolates rectified-flow fingerprint across 23 models","Latent selectivity band appears only in rectified-flow students","Three-layer probe maps prediction type separate from recipe","Rectified-flow ratio >1 band survives ADD in teachers and students"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"The claim that the sustained latent band isolates prediction type rests on one epsilon-prediction DiT control and no non-adversarial rectified-flow student, so residual architecture or conditioning factors could still drive the pattern.","fun_headline_variants_meta":{"raw":{"variants":["Latent layer flags rectified-flow backbones past distillation","Probe isolates rectified-flow fingerprint across 23 models","Latent selectivity band appears only in rectified-flow students","Three-layer probe maps prediction type separate from recipe","Rectified-flow ratio >1 band survives ADD in teachers and students"]},"model":"grok-4.5","effort":"low","cost_usd":0.005992,"raw_usage":{"total_tokens":1738,"prompt_tokens":1013,"num_sources_used":0,"completion_tokens":87,"cost_in_usd_ticks":59920000,"prompt_tokens_details":{"text_tokens":1013,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":638,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":1013,"tokens_out":87,"duration_ms":6318,"temperature":1.0,"reasoning_tokens":638,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-12T03:45:37.010347+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"An epsilon-prediction model that forms a sustained latent band (R_lo > 1 on at least three of the four strengths in {0.05, 0.1, 0.2, 0.3}, including one s ≥ 0.2) under the same matched estimator, or a non-ADD rectified-flow student that loses that band, would falsify the prediction-type detector.","supporting_citations":[],"review_version":1}