{"id":"bd9c6fc1-7017-42ba-9513-f74eefc5abc6","arxiv_id":"2607.24605","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.5,"correctness_risk":"low","formal_verification":"none","parameter_count":4,"one_line_summary":"MMT efficiently models CMB beam systematics via map-domain Mueller mixing, quantifying T-to-B leakage across eight TDM readout schemes and N_eff bias from time-constant errors.","lead":"Map Multi-Tool (MMT) is a map-domain simulator that turns beam distortions and detector crosstalk into Stokes mixing matrices, then into power spectra and parameter biases for CMB experiments. It lets instrument teams compare readout schemes and time-constant tolerances without full time-ordered-data runs.","discovery_kind":"new_method","skeptic_critique":{"model":"moonshotai/kimi-k3","headline":"The N_eff requirement curve (Fig. 11) rests on a single uniform τ rescaled by (1-p) with perfect cross-linking and a magnitude-only response; real τ distributions and anisotropic smear change the residual transfer function's shape at the damping tail, where the claimed bias lives.","rationale":"The reader's weakest_assumption — that the toy array, flat-sky, perfect-cross-linking setup may not carry over to real instruments — identifies the correct soft spot, and my concern is a sharpened instance of it applied specifically to the N_eff half of the strongest claim. I agree the crosstalk ranking results are internally robust because idealizations cancel in relative comparisons; the exposure is concentrated in the absolute calibration-requirement comparison of Fig. 11, where the residual transfer function's shape at ℓ~2000–4000 is doing all the work and is sensitive to the τ-distribution and cross-linking idealizations. This does not overturn the paper: it is a methods paper, the algebra is transparent, the code is public, and the qualitative conclusion (percent-level τ calibration matters for N_eff targets) survives order-of-magnitude scrutiny. The proposed test is directly executable with the released codebase and would either bound the idealization error or confirm the curve. Accordingly the CONDITIONAL verdict stands: the framework and qualitative findings are sound, but the quantitative requirement curve should be treated as conditional until the realism check is run. No formal verification is present, but none is needed for this kind of numerical methods work; the shipped reproducible pipeline counts in the paper's favor.","tokens_in":14211,"tokens_out":5095,"duration_ms":192748,"concrete_test":"Re-run the §3.2 pipeline twice using the public MMT code, propagating both cases through the same COBAYA N_eff fit: (1) replace the uniform τ with a per-detector log-normal distribution of 15% width, deconvolve with the single best-fit effective τ, and compute ΔN_eff vs. nominal fractional error; (2) replace perfect cross-linking with a single-direction scan (retain the complex phase of Eq. 11's Fourier transform rather than the magnitude-only Eq. 12) and compute ΔN_eff. If either variant shifts ΔN_eff at fixed p by more than ~50% relative to Fig. 11, the requirement curve should be labeled illustrative only; if both stay within that band, the Fig. 11 claim is substantially strengthened.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The strongest claim has two halves: (a) the crosstalk leakage rankings (Figs. 5–8) and (b) the statement that few-percent fractional time-constant errors bias N_eff at levels comparable to σ_Neff ≈ 0.045 (SO) / 0.03 (CMB-S4). Part (a) is a relative-ranking result computed with a single pipeline, so common-mode idealizations (flat sky, 4×4 toy array) largely cancel in the comparison; I find no load-bearing hole there beyond what the reader flagged. Part (b) is an absolute, requirement-setting comparison, and it is more exposed. Three compounding choices in §3.2 determine the result: (i) all detectors share one τ, and the error is modeled as a pure rescaling τ→τ(1-p) (Eq. 11); (ii) perfect cross-linking is assumed to convert the one-sided exponential smear into a uniform radial suppression; (iii) Eq. 12 keeps only the magnitude of the exponential's Fourier transform, discarding its phase (a scan-directional lag). Choice (i) matters because real arrays have a distribution of τ (typically 10–20% width); a superposition of exponentials deconvolved by a single effective τ leaves a residual transfer function whose ℓ-dependence is not in the family of Eq. 13, and the scatter itself suppresses the damping tail even at zero mean error. Choices (ii)–(iii) matter because realistic scans are imperfectly cross-linked, leaving an anisotropic residual; the phase term then does not cancel, and the effective suppression at ℓ~2000–4000 — precisely the lever arm for N_eff — can differ in amplitude from the radial model. Since the entire Fig. 11 bias comes from a ~0.5%-level residual modulation at ℓ~3000 (for p~2%, τ=10 ms, v=1°/s), factor-of-two-level distortions of that residual would move the bias across the σ_Neff threshold the paper compares against. The qualitative point (percent-level τ errors are worth worrying about) is robust and consistent with order-of-magnitude estimates, but the quantitative requirement curve is not yet established. Note also the polarization channel (E→B)","agreement_with_reader":"agree"},"referee_report":{"model":"moonshotai/kimi-k3","summary":"The paper presents Map Multi-Tool (MMT), a map-domain simulation framework in which beam systematics are encoded as a 3×3 Mueller matrix of maps M convolved with simulated Stokes sky maps (Eqs. 1–2), then propagated to flat-sky power spectra and cosmological parameter estimation. Two demonstration studies are given. First, crosstalk in a time-division multiplexed readout is modeled via per-detector coupling coefficients γ_ij inserted into a mapmaking-derived response matrix (Eqs. 3–10), and T-to-B leakage spectra are computed for eight TDM readout schemes on a simplified 4×4 dichroic dual-polarization array behind a 5 m f/2.8 telescope, with fixed 0.1% (row-switching) and 0.3% (inductive) couplings; the schemes rank differently, with two-pass schemes (3, 4) suppressing low-ℓ leakage. Second, a fractional error p in the measured detector time constant is modeled as a residual low-pass filter (Eqs. 11–13) under perfect cross-linking, and the resulting suppression of the high-ℓ TT damping tail is propagated through COBAYA fits to show N_eff biases comparable to SO/CMB-S4 statistical targets for few-percent p (Figs. 10–11). Code and outputs are public [30].","tokens_in":14681,"tokens_out":3599,"duration_ms":130052,"significance":"If the framework performs as described, MMT fills a real gap between analytic beam-leakage calculations (e.g., Hu et al. 2003; Shimon et al. 2008) and full TOD simulations: it is cheap enough for design iteration yet propagates systematics through mapmaking, power spectra, and parameter inference in one chain. Particular strengths: the mixing-matrix algebra (Eqs. 3–10) is transparent and verifiable; all results are forward simulations from stated instrumental inputs with no fitted-then-represented circularity; the crosstalk study yields an immediately usable comparative ranking of eight TDM readout schemes; and the codebase plus simulation products are publicly released [30], making the work reproducible and extensible. The N_eff example, while idealized, demonstrates the full systematics-to-parameters pipeline that design studies need.","major_comments":[{"comment":"The N_eff bias curve in Fig. 11 is the manuscript's only absolute, requirement-setting result (it is compared directly to σ_Neff = 0.045 for SO and 0.03 for CMB-S4 in §3.2), yet the COBAYA fits behind it are unspecified. Please state: the multipole range and weighting (noise levels or noiseless?), which ΛCDM parameters are varied versus fixed, whether lensing is included, and the exponential fit parameters shown as the dashed curve. The numerical value of the bias at the damping tail depends sensitively on ℓ_max and the effective noise weighting, so without this information Fig. 11 cannot be evaluated or reproduced.","section":"§3.2, Fig. 11"},{"comment":"Three compounding idealizations determine the magnitude of the Fig. 11 result, and all act at the damping tail (ℓ~2000–4000) where the N_eff lever arm lives: (i) all detectors share a single τ, with the error a pure rescaling τ→τ(1−p) (Eq. 11); real arrays have τ distributions of 10–20% width, and a superposition of exponentials deconvolved by one effective τ leaves a residual transfer function outside the family of Eq. 13, with nonzero tail suppression even at zero mean error. (ii) Perfect cross-linking converts the one-sided smear into a uniform radial filter. (iii) Eq. 12 keeps only the magnitude of the exponential's Fourier transform, discarding the scan-directional phase (a pointing lag), which does not cancel for imperfectly cross-linked or single-direction scans. These are defensible for an illustrative example, but since the text sets the bias against SO/CMB-S4 statistical target","section":"§3.2, Eqs. 11–13"},{"comment":"The leakage spectra in Figs. 5–8 are computed under the flat-sky approximation (§2), yet the ranking claim that schemes 3 and 4 are 'better suited for instruments targeting inflationary B-modes' rests on their steeply decreasing leakage at low ℓ, precisely where flat-sky E/B decomposition is least reliable (the primordial signal at r=0.001 peaks at ℓ≲100, and the paper itself notes leakage spectra peaking at ℓ=1 for monopole-like beams). Please either add a curved-sky check (e.g., against the analytic beam-leakage formalism of [13]/[14] at low ℓ) or explicitly restrict the scheme-ranking conclusions to the multipole range where the flat-sky E/B separation is valid.","section":"§3.1, Figs. 5–8 and §2"}],"minor_comments":[{"comment":"The fractional T-to-B leakage is defined as the ratio B̃_obs/Ĩ_sky of Fourier amplitudes, but Figs. 5–8 plot spectra against theoretical BB power curves, implying a power ratio C_ℓ^BB/C_ℓ^TT. Please state the definition precisely, including any normalization (e.g., per-ℓ binning, ensemble averaging over sky realizations).","section":"§2"},{"comment":"State the angular-units convention behind ℓ_cutoff = 180/vτ (e.g., ℓ = 180°/θ with vτ in degrees). For τ=10 ms and v=1°/s this gives ℓ_cutoff ≈ 1.8×10^4; making this explicit would help readers check Fig. 9.","section":"§3.2, Eq. 12"},{"comment":"The leakage power scales as γ², so the readout-scheme rankings are robust to the assumed 0.1%/0.3% amplitudes while the absolute levels are not; a one-line note to this effect (with the scaling) would clarify how to rescale the results for other instruments.","section":"§3.1"},{"comment":"Typos: 'with and a perfect cross-linking scan speed' (Fig. 10 caption); 'speed of of 1°/sec' (§3.2); 'opposite frequencies frequency' (§3.1); 'measured verses true τ' (§3.2); Fig. 8 caption repeats 'indicated by the gray dotted line' for the lensing curve; abstract begins a sentence with lowercase 'this framework enables'. Also 'LCDM' should be 'ΛCDM' in §3.2.","section":"Various"},{"comment":"The checkerboard pattern of 0°/90° and 45°/−45° pixels (Fig. 2) interacts with the readout ordering in producing the leakage beams; a brief note on how the results change if the polarization-angle pattern or pixel pitch is varied would help readers judge the generality of the 4×4 toy array.","section":"§3.1, Fig. 2"}],"recommendation":"minor_revision","confidential_remarks":"The manuscript is a good fit for the journal's instrumentation/methods scope. The claims are appropriately framed as framework demonstrations, and the public code/data repository is a genuine asset. My requested revisions concern caveat placement and fit specification rather than new simulations; I would not insist on full curved-sky or TOD validation for this version, though it would substantially strengthen a follow-up."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The useful bit here is a packaged, public map-domain Mueller mixer that lets you rank beam systematics without a full TOD run, plus two concrete worked examples. The eight TDM wiring comparisons (Figs. 5–8) are new numerical results that instrument teams will actually look at; the algebra that builds the mixing maps from detector offsets and γ_ij (Eqs. 3–10) is transparent and the relative leakage ordering is robust because common-mode idealizations largely cancel.\n\nWhat the paper does well: it ships code, keeps the formalism simple, and shows how to push the output spectra into COBAYA. That combination is more useful than another pure analytic leakage paper. The crosstalk section is the stronger half—different row-switching vs inductive patterns produce visibly different T-to-B shapes, and the single- vs dual-frequency distinction is cleanly illustrated.\n\nSoft spots are real but bounded. The array is a 4×4 toy with hand-set 0.1%/0.3% crosstalk; fine for ranking, not for absolute forecasts on a multi-thousand-detector focal plane. The time-constant half is thinner: one uniform τ, pure rescaling by (1-p), perfect cross-linking that turns the exponential into a radial low-pass, and magnitude-only Fourier response. Real τ scatter and residual scan anisotropy change the damping-tail residual at the ~0.5% level that drives Fig. 11, so the quantitative comparison to σ_Neff ≈ 0.03–0.045 is not yet established. The qualitative warning (few-percent τ error matters) still stands. Flat-sky is used throughout; no curved-sky check.\n\nCitations are appropriate (Hu/Shimon analytic lineage, existing TOD codes, COBAYA). No circularity—everything is forward simulation from stated inputs.\n\nThis is for CMB instrumentation and design groups who need fast relative rankings and requirement-setting intuition, not for people who need end-to-end map realism. It deserves a serious referee. I would engage with the tool and the crosstalk rankings; I would treat Fig. 11 as illustrative until the residual transfer function is stress-tested with τ distributions and realistic scans.","headline":"Useful open map-domain tool with solid relative rankings of TDM schemes; the absolute N_eff requirement curve is more idealized than the abstract implies.","tokens_in":15531,"tokens_out":539,"would_cite":true,"duration_ms":10570,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"A map-domain tool turns beam distortions into leakage spectra and parameter biases so CMB instruments can set design and calibration limits before full time-stream sims.","keywords":["CMB systematics","beam leakage","Mueller matrix","map-based simulation","electrical crosstalk","detector time constant","T-to-B leakage","N_eff bias"],"falsifier":"Rebuild the same eight readout schemes and the residual-tau suite inside a full time-ordered-data simulation of a realistic multi-thousand-detector focal plane; if the relative T-to-B leakage ordering or the N_eff-versus-p curve changes materially, the map-based claims do not transfer.","tokens_in":15239,"feed_emoji":"📡","tokens_out":909,"duration_ms":18104,"temperature":0.7,"pith_summary":"CMB experiments need quantitative control of beam-related systematics that can leak temperature into polarization or suppress the high-multipole damping tail. This paper presents Map Multi-Tool (MMT), a map-based pipeline that encodes those effects as a Mueller mixing matrix of beams, convolves them with simulated skies, and yields observed maps and power spectra that feed directly into cosmological estimators. Two worked examples show the payoff: eight different time-division multiplexed readout wiring schemes produce measurably different T-to-B leakage spectra for both row-switching and inductive crosstalk, and fractional errors of a few percent in detector time constants bias N_eff at levels comparable to the statistical targets of next-generation surveys. The framework is offered as a computationally light way to rank design choices and set calibration requirements before committing to expensive end-to-end simulations.","feed_headline":"Map tool ranks CMB beam leaks before full simulations","feed_subtitle":"Eight readout schemes and residual time-constant errors turn into leakage spectra and N_eff bias curves","key_machinery":"The Mueller mixing matrix M in the observed-map equation S_obs = M ⊛ S_sky + N: each element is a beam map that can encode crosstalk couplings or residual time-constant smear, so a single convolution yields the leaked Stokes maps and their spectra.","core_discovery":"Map Multi-Tool shows that beam-related systematics can be modeled efficiently in the map domain by convolving simulated skies with a 3x3 Mueller matrix of intensity and polarization beams (including leakage terms), producing power spectra and parameter biases that distinguish instrument designs. Applied to eight TDM readout schemes, the method ranks them by T-to-B leakage; applied to residual time-constant errors, it maps fractional tau uncertainty onto N_eff bias.","pith_inferences":["Once the Mueller-matrix construction is automated for arbitrary wiring graphs, the same ranking exercise could be run overnight for every proposed readout architecture of a Stage-4 experiment.","The residual-tau bias curve suggests that time-constant metrology may need to reach the percent level or better if N_eff is a primary science target; that requirement can be budgeted against other calibration terms.","Because the method is map-based, it can be inserted as a fast pre-filter before expensive TOD campaigns, concentrating full simulations only on the schemes that already look dangerous in MMT."],"forward_implications":["Instrument teams can rank candidate TDM wiring schemes by predicted T-to-B leakage before hardware is frozen.","Calibration requirements on detector time-constant precision can be set directly from the N_eff bias curve rather than from ad-hoc margins.","Multiple beam systematics (crosstalk, tau error, ellipticity, pointing) can be stacked inside one Mueller matrix to study their combined leakage.","The same pipeline supplies contaminated spectra as inputs to cosmological parameter estimators, closing the loop from design choice to science bias."],"fun_headline_variants":["Map Multi-Tool ranks TDM crosstalk by T-to-B leakage spectra","MMT maps residual tau error onto N_eff bias without full sims","3x3 Mueller beam convolution flags CMB systematics in map domain","Eight readout schemes sorted by power-spectrum leakage via MMT","Map-domain beams turn crosstalk and tau residuals into parameter bias"],"cache_read_input_tokens":128,"weakest_assumption_plain":"The rankings and bias curves drawn from a simplified 4x4 array, fixed crosstalk amplitudes, flat-sky Fourier transforms, and perfect cross-linking are taken to remain representative for real large-aperture, multi-thousand-detector instruments.","fun_headline_variants_meta":{"raw":{"variants":["Map Multi-Tool ranks TDM crosstalk by T-to-B leakage spectra","MMT maps residual tau error onto N_eff bias without full sims","3x3 Mueller beam convolution flags CMB systematics in map domain","Eight readout schemes sorted by power-spectrum leakage via MMT","Map-domain beams turn crosstalk and tau residuals into parameter bias"]},"model":"grok-4.5","effort":"low","cost_usd":0.004112,"raw_usage":{"total_tokens":1251,"prompt_tokens":794,"num_sources_used":0,"completion_tokens":78,"cost_in_usd_ticks":41124000,"prompt_tokens_details":{"text_tokens":794,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":379,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":794,"tokens_out":78,"duration_ms":6897,"temperature":1.0,"reasoning_tokens":379,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-31T10:45:25.002640+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Rebuild the same eight readout schemes and the residual-tau suite inside a full time-ordered-data simulation of a realistic multi-thousand-detector focal plane; if the relative T-to-B leakage ordering or the N_eff-versus-p curve changes materially, the map-based claims do not transfer.","supporting_citations":[],"review_version":1}