{"id":"02336591-4fb5-4935-8aa1-c902e113a61d","arxiv_id":"2607.19921","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"CATKit2-HCI is a new collaborative software layer that standardizes wavefront control, calibration, and diagnostics across six high-contrast coronagraph testbeds in the US and Europe.","lead":"This paper describes CATKit2-HCI, a shared software layer that six coronagraph labs use to pool wavefront-control, calibration, and monitoring tools instead of rebuilding them per testbed. It matters because reusable, comparable testbed software can speed development of coronagraphs for the Habitable Worlds Observatory.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The non-enforced 'rebasing' norm is the load-bearing premise; without evidence that labs actually migrate to shared code, duplication reduction is asserted, not demonstrated.","rationale":"The reader's weakest assumption correctly identifies the non-enforced 'rebasing' norm as the load-bearing premise. The central claim that CATKit2-HCI reduces duplicated effort directly depends on teams actually migrating local code to the shared layer, and the paper explicitly states this is a shared commitment rather than an enforcement mechanism. The early examples are consistent with feasibility but do not establish that migration occurs at scale; they only show that similar outputs can be produced. No repository-level evidence, such as removed local implementations or import provenance, is presented. The private nature of the repository further prevents external validation. This concern is not an objection to the framework's helpfulness; it is a precise gap between the claimed outcome (reduced duplication) and the evidence provided (two examples, no adoption metrics). Therefore, the CONDITIONAL verdict remains appropriate: accept conditionally pending evidence of actual rebasing practice. My read does not change the reader's verdict; it reinforces the stated condition.","tokens_in":6318,"tokens_out":3649,"duration_ms":43377,"concrete_test":"Obtain read access (or an anonymized diff summary) for the six participating testbed repositories over a defined 12-month window after a reference capability (e.g., the shared EFC analysis routine) is merged into CATKit2-HCI. For each testbed, count the number of local source files that implement equivalent functionality and remain in the repository versus the number that have been refactored to import from CATKit2-HCI. If the majority of testbeds still contain duplicate local implementations, the 'rebasing' norm is not producing convergence and the duplication-reduction claim fails. Additionally, verify that the outputs in Fig. 2 originate from a single shared CATKit2-HCI module by checking commit provenance, rather than from two similar but independent local scripts.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is functional: CATKit2-HCI reduces duplicated effort across six testbeds. This reduction occurs only if participating teams actually replace local implementations with shared CATKit2-HCI code after a capability is merged. The paper itself concedes in §2.3 that the 'rebasing' norm is 'not intended primarily as an enforcement mechanism, but as a shared commitment.' That concession is the load-bearing weak spot. If even a subset of testbeds continues to maintain local forks after features are merged, the shared repository accumulates code without reducing duplication — the exact failure mode the norm is designed to prevent. The only cross-testbed evidence (§3, Figs. 2–3) consists of plotting and dashboard outputs from the two founding teams (HiCAT, THD2); these display similar visual results but do not demonstrate that identical shared code is being invoked, nor that any local code has been removed. Moreover, no metrics from the other four testbeds (CAPSULE, SEAL, HCST, ExoSPEC) are provided, and the repository is private (§2.1), so an external reader cannot audit adoption. Thus the strongest claim rests on a behavioral hypothesis about a voluntary social norm, unsupported by measured migration data.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper describes CATKit2-HCI, a shared software layer for high-contrast coronagraph testbeds, built on the open-source CATKit2 hardware-control framework. It proposes a three-layer architecture (CATKit2, CATKit2-HCI, private testbed repositories), a collaboration model with a 'rebasing' norm, and lists shared capabilities in wavefront sensing/control, calibration, modeling, and visualization. Early examples from HiCAT and THD2 show common EFC plotting and dashboard concepts. The conclusions claim the framework reduces duplicated effort across six testbeds while preserving laboratory autonomy.","tokens_in":6525,"tokens_out":3292,"duration_ms":37169,"significance":"If validated, the framework would address a genuine bottleneck in laboratory astrophysics: many high-contrast testbeds independently reimplement similar wavefront sensing and control tools, calibration utilities, and monitoring dashboards. The open-source CATKit2 foundation is a concrete, citable asset, and the collaboration model for shared review and staff mobility is timely. The early HiCAT/THD2 examples are suggestive but not yet demonstrative of the central claim that duplication is actually reduced; stronger evidence of adoption and code migration is needed. The paper is valuable as a description of an ongoing community-infrastructure effort, but its functional claims currently outrun the presented evidence.","major_comments":[{"comment":"The load-bearing premise of the 'reduces duplicated effort' claim is the rebasing norm, but the paper states this norm 'is not intended primarily as an enforcement mechanism, but as a shared commitment.' No evidence is provided that any testbed has actually removed or refactored local code after a capability was merged into CATKit2-HCI, or that multiple testbeds invoke the same shared implementation for a given capability. Without adoption metrics or a concrete before/after case study, §4's conclusion that duplication is reduced is an assertion. Please provide at least one documented example of rebasing, or explicitly reframe the paper as a proposal rather than an achieved outcome.","section":"§2.3"},{"comment":"The two early examples show visually similar EFC plots and dashboards for HiCAT and THD2, but they do not demonstrate that identical shared code is being invoked. The reader cannot tell whether these outputs come from CATKit2-HCI functions or from parallel local scripts that merely look alike. Please specify the exact modules/functions used by each testbed to produce these figures, and clarify what code was contributed to CATKit2-HCI versus what remains testbed-specific. Without this, the examples illustrate a concept, not a shared implementation.","section":"§3, Figs. 2 and 3"},{"comment":"The repository is private and no information is given about the status of the other four participating testbeds (CAPSULE, SEAL, HCST, ExoSPEC). Since the central claim concerns six testbeds, the absence of any qualitative or quantitative evidence from those benches makes the 'six testbeds' framing unsupported. Please include a table or narrative describing, per testbed, which CATKit2-HCI capabilities are planned, under evaluation, or already in use. A private repository also prevents external audit; consider at least a public README or capability list.","section":"§2.1 and §2.4"},{"comment":"The premise that 'much of the software architecture is not fundamentally testbed-specific' is asserted but not supported by the examples, which focus on generic plotting and dashboards. To justify the claim of an HCI-specific shared layer, provide examples of more specialized algorithms (e.g., EFC, pairwise probing, dOTF) that have actually been implemented in CATKit2-HCI and used by more than one testbed. If such examples are not yet available, the paper should be careful to distinguish between planned capabilities and demonstrated ones.","section":"§1.1 and §4"}],"minor_comments":[{"comment":"Heading contains 'CA TKit2' instead of 'CATKit2' (typo). Similar spacing issues appear in the author list and abstract.","section":"§1.2"},{"comment":"The caption says 'See text for details' but does not describe the three layers in the caption itself. A self-contained caption would improve readability.","section":"Figure 1"},{"comment":"The list of shared capabilities (EFC, pairwise probing, dOTF, etc.) is useful, but a table mapping each capability to its maturity status (planned, under development, mature) would help the reader assess the framework's current state.","section":"§3"},{"comment":"The paper does not provide a link to the repository or a DOI beyond the CATKit2 reference [3]. Since software is the subject, a persistent identifier or public-facing description should be given.","section":"General"},{"comment":"Reference [3] is a Zenodo software release; consider including the version number and access date. Some SPIE references are formatted inconsistently with duplicated publisher names.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper is authored entirely by members of the collaboration it describes, which is not unusual for a software paper but compounds the need for independent evidence. The editor may want to consider whether a private repository and absence of adoption metrics are acceptable for a claim of 'reduced duplicated effort'; the revision should either supply such evidence or temper the claim to a proposal. The topic is appropriate for the journal and the early open-source foundation is a strength."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know this paper is a software-architecture and collaboration-governance description, not a physics or methods result. It describes CATKit2-HCI, a shared private code layer meant to sit between the public CATKit2 hardware-control framework and individual testbed repositories. What's actually new is the three-layer split, the specific estimator/controller/observer abstractions in §3, and the reciprocal \"rebasing\" norm in §2.3 where teams are expected to refactor local code onto shared implementations after a merge.\n\nThe paper does several things well. It is honest about its own stage: it says the framework is \"under active development,\" and it shows only two early cross-testbed examples (shared EFC plotting and dashboard concepts for HiCAT and THD2). The architecture description is internally consistent and plausible. The authors are also upfront that rebasing is \"not intended primarily as an enforcement mechanism.\" That is a real concession, not hidden.\n\nThe soft spots are exactly where the reader's report puts them. The central effectiveness claim — that CATKit2-HCI reduces duplicated effort — depends on teams actually replacing local code with shared implementations. The paper offers no metric of duplication before/after, no adoption data from the other four testbeds, and the repository is private, so an outsider can't audit whether the shared code is genuinely being used. The figures show similar outputs but don't demonstrate that identical shared code is invoked. These are not fatal flaws for a status paper, but they do mean the headline claim is a design goal, not a demonstrated outcome. The stress-test note is right to flag this; it's the load-bearing premise.\n\nI don't think this is a case where the central argument collapses. The paper is a proposal plus early evidence, and it reads as such. If the collaboration later publishes migration metrics or opens the repository, the claim will be testable. For now, the paper's value is as a clear description of a governance model that the HCI testbed community may want to adopt or critique.\n\nWho is this for? Groups running coronagraph testbeds, and people working on research-software sustainability in astronomy. It deserves a serious referee: the architecture and governance norms are concrete enough to evaluate, and the community needs this kind of paper to compare approaches. A reviewer should push for either public code or honest adoption metrics, but desk-rejecting it would be wrong.\n\nRecommendation: send to peer review. Conditional acceptance with a request for a limitations section is the right outcome.","headline":"A clear, honest architecture paper for shared coronagraph-testbed software; the duplication-reduction claim is a promise, not yet a demonstrated result.","tokens_in":7189,"tokens_out":1965,"would_cite":true,"duration_ms":19720,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A shared software layer now links six coronagraph testbeds.","keywords":["CATKit2","coronagraph testbeds","high-contrast imaging","wavefront sensing and control","reusable research software","Habitable Worlds Observatory","laboratory astrophysics","electric field conjugation"],"falsifier":"Inspect the private repositories of the six testbeds one year after a mature algorithm has been merged into CATKit2-HCI: if most still contain substantial maintained copies of that algorithm, or if the shared repository shows little import activity from testbed-specific code, the central claim that the collaboration reduces duplicated effort fails.","tokens_in":6169,"feed_emoji":"🔭","tokens_out":3960,"duration_ms":40025,"temperature":0.7,"pith_summary":"The paper argues that coronagraph testbeds building similar wavefront-sensing, calibration, and diagnostic tools separately are wasting effort, and that a shared layer — CATKit2-HCI — can capture mature, reusable tools without forcing benches into identical operation. Built on the open CATKit2 hardware-control framework, CATKit2-HCI sits between generic device handling and each lab's private experiment code. The collaboration's rebasing norm asks testbeds to fold merged capabilities back into local use, keeping the common codebase coherent. Early examples from two testbeds show shared electric-field-conjugation analysis and dashboard concepts working across different optical benches. A sympathetic reader would care because, if the model holds, it accelerates the laboratory development path for the Habitable Worlds Observatory and ground-based coronagraphs.","feed_headline":"Six coronagraph testbeds now share one software layer","feed_subtitle":"Reusable wavefront-control and calibration tools cut duplicated effort across labs, speeding the path to Habitable Worlds Observatory.","key_machinery":"The mechanism is a three-layer architecture: CATKit2 supplies testbed-agnostic hardware service control; CATKit2-HCI supplies the shared middle layer of HCI algorithms, calibration utilities, DM tools, modeling helpers, and dashboards; private repositories hold each bench's experiments and tuning. Within the shared layer, an estimator/controller/observer separation lets testbeds swap components without rewriting whole experiments. The load-bearing social mechanism is the rebasing norm — refactoring local code onto merged shared implementations to keep the shared layer a living common codebase rather than a set of forks.","core_discovery":"The paper's central claim is that CATKit2-HCI fills the missing software layer between public hardware control and private testbed-specific experiments, collecting mature high-contrast-imaging tools into a shared, reviewed repository. It states that the architecture separates estimators, controllers, observers, and testbed interfaces, so the same algorithm — electric field conjugation, pairwise probing, differential optical transfer function phase retrieval — can run in different optical environments. With six testbeds participating and a rebasing commitment, the authors argue the collaboration reduces duplicated development, improves code quality through shared review, enables direct cross-","pith_inferences":["The paper's strongest unproven premise is that teams will actually rebase: if testbeds merge tools but keep running their own forks, the shared repository accumulates code without reducing duplication. A quantitative duplication metric across the six repositories would test this.","The common-needs assumption is demonstrated for only two of six testbeds; the framework's generality would be better supported by examples from the other four, such as a vortex coronagraph bench or a broadband spectroscopy bench.","If the model works, it could generalize beyond coronagraphy to other distributed instrument-software collaborations, wherever hardware diversity coexists with shared algorithm needs.","A concrete near-term success criterion the paper leaves implicit: the fraction of each testbed's daily operational code that imports from CATKit2-HCI rather than from local implementations."],"forward_implications":["Bug fixes, calibration tools, and improved algorithms merged into the shared layer benefit all participating testbeds at once.","Researchers moving between labs inherit a common operational vocabulary and familiar plotting and monitoring tools, shortening onboarding.","Cross-testbed comparisons become cleaner because the same estimator or observer can be exercised on different optical benches.","Standardized implementations of classical algorithms such as EFC and pairwise probing reduce duplicated code across the collaboration.","Maturation of the framework would accelerate coronagraph technology development for future exoplanet imaging facilities including the Habitable Worlds Observatory."],"fun_headline_variants":["One software layer unites six exoplanet testbeds","Shared framework ends duplicated coronagraph code","Coronagraph labs adopt common software backbone","Six testbeds, one codebase: CATKit2-HCI scales","Collaborative code accelerates coronagraph readiness"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The framework only reduces duplication if each testbed actually refactors its local code to use the shared implementation once a capability is merged — the paper's rebasing norm is a shared commitment, not an enforcement mechanism.","fun_headline_variants_meta":{"raw":{"variants":["One software layer unites six exoplanet testbeds","Shared framework ends duplicated coronagraph code","Coronagraph labs adopt common software backbone","Six testbeds, one codebase: CATKit2-HCI scales","Collaborative code accelerates coronagraph readiness"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000123,"raw_usage":{"total_tokens":936,"prompt_tokens":745,"completion_tokens":191,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":489,"completion_tokens_details":{"reasoning_tokens":116}},"tokens_in":489,"tokens_out":191,"duration_ms":2754,"temperature":1.0,"reasoning_tokens":116,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T11:17:14.984318+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Inspect the private repositories of the six testbeds one year after a mature algorithm has been merged into CATKit2-HCI: if most still contain substantial maintained copies of that algorithm, or if the shared repository shows little import activity from testbed-specific code, the central claim that the collaboration reduces duplicated effort fails.","supporting_citations":[],"review_version":1}