{"id":"af7b33f8-a8e4-4a00-935c-3ac92712de8d","arxiv_id":"2508.20332","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"SPHEREx's survey planning software, an online greedy scheduler with a tuned figure of merit, achieves 99.54% planned all-sky voxel completeness and meets deep-field sensitivity requirements in mission simulations.","lead":"SPHEREx, a NASA infrared sky-mapping mission in low-Earth orbit, uses scheduling software to choose which part of the sky to observe next while balancing power, heat, stray light, downlinks, and radiation zones. In simulated full-mission runs the scheduler reaches about 99.5% planned sky coverage and meets deep-field sensitivity requirements with margin, and it is being used in flight.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 99.54% planned voxel completeness and 29.3e-6 deep-field sensitivity are simulation-only; the first public data release is invoked as in-flight proof but no achieved coverage numbers are reported.","rationale":"I read the paper as an engineering validation of the SPS: the central claim is that the greedy online target-selection algorithm produces a plan meeting SPHEREx's coverage and deep-field sensitivity requirements with margin. The evidence is a simulation using the actual post-launch trajectory, and the paper is careful to call the headline numbers 'planned' coverage. The weakest load-bearing step is not the algorithm logic or the constraint modeling but the translation of planned coverage into delivered coverage. Section 5 itself lists unmodeled loss terms and explicitly defers them. The abstract nevertheless asserts in-flight success based on the first public data release, yet no attained metric is reported. That makes the release anecdotal rather than quantitative. I do not see an internal inconsistency in the planner description or the sensitivity formalism; the issue is an unvalidated link between simulation and operations. A direct comparison at the epoch of the first release would settle it. This agrees with the reader's weakest_assumption, though I would sharpen it to require a same-epoch simulation baseline rather than comparing the early release to the full-mission 99.54% figure. Because the reader already conditioned on this gap, I recommend no change to the verdict.","tokens_in":10140,"tokens_out":10250,"duration_ms":116003,"concrete_test":"Run the SPS with the actual ephemeris and the same target list and parameters over the interval covered by the first public data release (July 2025); compute planned voxel completeness and deep-field Nhit maps for that interval. Compute the same metrics from the release data (doi:10.26131/IRSA629) using the Sec. 5 voxel definition. Compare achieved vs. planned completeness and achieved vs. planned Nhit for the deep fields. If achieved completeness remains above the 98% requirement and the deep-field sensitivity recomputed from actual Nhit (Eqs. 5-7) stays below 40e-6, the concern is resolved; if the gap approaches or exceeds the 1.54-point all-sky margin, the planned-coverage claim is not yet validated in flight.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's quantitative case is a mission-long simulation: planned voxel completeness 99.54% vs. 98% requirement, and deep-field sensitivity 29.3e-6 vs. 40e-6 (Sec. 5). Section 5 explicitly limits these numbers to 'planned coverage delivered by the SPS' and says other losses (downlink errors, spacecraft anomalies, defective pixels) are tracked elsewhere. The abstract's claim that the first public data release demonstrates in-flight performance is therefore the only bridge from planned to delivered, but the paper provides no achieved voxel completeness, no achieved deep-field Nhit maps, and no comparison to the simulation at the corresponding epoch. The margin over the all-sky requirement is only 1.54 percentage points; if actual losses or a higher glitch rate consume that margin, the capability claim fails even though the planner itself worked as designed. This is a validation gap, not an internal inconsistency.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper describes the SPHEREx Survey Planning Software (SPS), an algorithm for scheduling observations from a low-Earth orbit under time-varying pointing constraints (power, thermal, stray light, shuttle glow, downlink, and South Atlantic Anomaly outages). The SPS uses an online target-selection heuristic: it builds an allowable pointing zone from avoidance angles, prioritizes deep-field observations with a tunable probability threshold, and otherwise selects All-Sky targets via a figure of merit that balances completing partially observed target groups against observing targets about to leave the allowable zone. The paper reports mission-long simulations using the actual trajectory that yield 99.54% planned voxel completeness for the All-Sky Survey (requirement 98%) and a deep-field sensitivity of 29.3e-6 against a 40e-6 requirement. It also asserts in-flight success, citing the first public data release.","tokens_in":10363,"tokens_out":3089,"duration_ms":34626,"significance":"If the reported margins hold, the paper demonstrates a practical solution to a difficult sensor-scheduling problem: a LEO all-sky spectral survey with 102 channels and two deep fields, subject to multiple interacting time-varying constraints. The described algorithm and coverage metrics could inform future space survey missions. The paper is strong in defining quantitative coverage metrics (voxel completeness, deep-field power-spectrum sensitivity via Eqs. (5)-(7)) and in presenting a mission-long simulation with a concrete orbit and constraint geometry. The claimed 1.54 percentage-point margin over the all-sky requirement and ~27% margin in deep-field sensitivity are meaningful, if the planned-coverage metrics translate to delivered coverage. However, the paper's central quantitative performance claims are simulation-only; the in-flight evidence is qualitative. The tuning of the two main algorithm parameters on the same coverage metrics used for validation also limits the strength of the 'optimal' claim.","major_comments":[{"comment":"The paper's quantitative headline results (99.54% voxel completeness, 29.3e-6 deep-field sensitivity) are explicitly 'planned coverage delivered by the SPS' from mission-long simulations. The abstract and conclusions nevertheless state that the first SPHEREx public data release demonstrates that the approach 'is performing well in flight' and that 'our approach is performing well in flight,' but no in-flight achieved voxel completeness, deep-field hit-count maps, or epoch-matched comparison to the simulation are reported. The margin over the 98% requirement is only 1.54 percentage points, so a small degradation from real-world losses (downlink errors, anomalies, glitch rates above the assumed 2%) could consume it. This is a validation gap in the central capability claim. Please either add quantitative in-flight coverage numbers for the first data release or clearly reword the abstract/co","section":"Section 5 / Abstract / Conclusions"},{"comment":"The algorithm's two key parameters are tuned on the same mission-long simulations used to produce the headline coverage numbers: F is set to ~0.8 'based on a series of mission-long simulations' and the deep-field priority threshold is configured to select the deep field about 85% of the time to 'yield all-sky and deep coverage that meets requirements.' This is a self-consistency risk: the reported margins are optimized rather than independent. Please quantify sensitivity of the results to F and the 85% threshold, or validate on a withheld period or perturbed orbit. Without this, the 'optimal' characterization and the reported margins are weaker than they appear.","section":"Section 4.2, Eq. (4)"},{"comment":"The 2-degree inboard avoidance margin and the 2% glitch reobservation rate are load-bearing assumptions for translating planned coverage to delivered coverage. The paper states these values are based on pre-launch estimates (Section 3.5) and current glitch flagging (Section 4.2), but it does not report the actual orbit-predict error, attitude error, or glitch-rate statistics after launch, nor how they compare to the assumed values. If the real values are larger, the 99.54% planned completeness would not be achieved operationally. Please include an in-flight assessment of these margins, or clearly state that the capability claim depends on untested assumptions.","section":"Section 3.5 / Section 5"}],"minor_comments":[{"comment":"The text before Eq. (6) says 'We integrate this first against the area elements from the coverage maps in Figure 3' but the deep-field coverage maps are in Figure 6. This is a typo and should be corrected.","section":"Section 5"},{"comment":"The caption of Figure 5 says the mission-long simulation uses 'a nominal orbit predict,' while the body text says 'Using our actual trajectory, we performed mission-long simulations.' Please clarify whether the simulation uses the actual trajectory with a nominal orbit-predict model, or a purely nominal trajectory. This distinction matters for the credibility of the planned-coverage number.","section":"Section 5"},{"comment":"The sentence 'The resulting surface brightness sensitivity, if counting only fully spectrally sampled region as SPHEREx defines it, is therefore increased and is deeper than what can be achieved by the sum of small-step integrations' is unclear. Please rewrite for clarity, specifying what is being compared.","section":"Section 2.2"},{"comment":"The paper would benefit from stating explicitly that the 99.54% and 29.3e-6 numbers are simulation-only in the abstract, rather than only in Section 5. This would avoid overstating in-flight validation to readers who read only the abstract.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The paper is a solid engineering description of a complex scheduling system, and the simulation framework is credible. However, the abstract and conclusions overreach by citing the first public data release as evidence of in-flight performance without reporting any achieved coverage metrics. The tuning-circularity issue also needs to be addressed. I recommend major revision rather than rejection, because the issues are fixable within the manuscript's scope: add a quantitative in-flight comparison (or soften the claims), report sensitivity to tuning parameters, and clarify the trajectory/orbit-predict setup."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a legitimate, readable account of the SPHEREx survey planner. The core quantitative claims—99.54% planned voxel completeness against a 98% requirement, and 29.3e-6 combined deep-field sensitivity against 40e-6—are clearly defined simulation products, produced by a described algorithm on actual trajectory data. That part holds up.\n\nWhat's new: the figure of merit in Eq. (4), with its tunable weight F and the explicit trade between finishing partially observed target groups and finishing groups about to leave the pointing zone; the online target-group selection that interleaves deep-field and all-sky observing; and the practical handling of downlink, SAA, and the newly discovered shuttle glow constraint. The algorithm description is clear enough that a referee can see how the optimization is done, and the authors are honest that this is an online/greedy approximation to a high-dimensional problem. They also explicitly state in Sec. 5 that the 99.54% is planned coverage, and that other losses (defective pixels, downlink errors, anomalies) are tracked elsewhere. Good.\n\nSoft spots, in order of importance. First, the abstract and conclusions claim the first public data release demonstrates in-flight performance, but the paper reports no achieved voxel completeness, no Nhit maps, no comparison to the simulation at the corresponding epoch. The margin over the all-sky requirement is only 1.54 points; after accounting for the 2% reobservation rate that already goes into the simulation, any extra real-world loss could eat that margin. So there's a validation gap. It's not an internal contradiction, but it's exactly what a referee should push on.\n\nSecond, the \"optimal\" label is doing heavy lifting for a greedy heuristic with two user-tuned parameters (F and the 85% deep-field threshold). The tuning is done on the same mission-long simulation that is then used as the demonstration—a mild circularity. It's not fatal, but the sensitivity of the results to F should be shown, or at least the parameter ranges where the conclusion holds.\n\nThird, no code or data release. The target list and parameter settings would let someone reproduce the simulation; their absence limits the paper to a description rather than a reproducible result.\n\nBottom line: the central capability claim—that the SPS can plan a mission meeting both all-sky and deep-field requirements with margin—is well supported by the simulation evidence. The paper deserves a serious referee, and would be accepted after a revision that either provides real in-flight coverage numbers or softens the demonstration language, and ideally releases the target-list generation and parameter settings.","headline":"Solid engineering description of the SPHEREx survey planner with credible mission-long simulation results; the only real hole is the unquantified leap from planned to delivered coverage.","tokens_in":10976,"tokens_out":2155,"would_cite":true,"duration_ms":23099,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"SPHEREx's survey planning software claims that a greedy, one-target-at-a-time scheduler can meet the mission's all-sky and deep-field coverage requirements from low-Earth orbit, with planned voxel completeness of 99.54% against a 98% requir","keywords":["SPHEREx","low-Earth orbit","survey planning","target selection","all-sky spectral survey","voxel completeness","deep field","avoidance constraints"],"falsifier":"Compute the delivered voxel completeness for the completed surveys from the public data releases; if any of the four all-sky surveys falls below 98%, the claimed margin does not hold in flight. A second check is to count actual loss events: if real downlink errors, spacecraft anomalies, and glitch rates push reobservation demand above the simulated 2%, planned completeness overstates delivered coverage.","tokens_in":10046,"feed_emoji":"🛰️","tokens_out":6939,"duration_ms":74731,"temperature":0.7,"pith_summary":"This paper describes the survey planning software that decides, moment by moment, where the SPHEREx infrared telescope points during its 25-month low-Earth-orbit mission. The central claim is that the software's greedy target-selection rule, which only optimizes the next attitude rather than the full week-long schedule, still produces all-sky and deep-field coverage that meets mission requirements with real margin. In a mission-long simulation using the actual orbit, all four all-sky surveys reach at least 99.54% planned voxel completeness, and the combined deep fields reach a sensitivity of 29.3e-6 against a 40e-6 requirement. The paper argues this matters because it lets SPHEREx deliver its full spectral survey—102 near-infrared bands over the whole sky—despite the many sun, earth, moon, ram, power, and thermal constraints that would otherwise force a simpler scan pattern.","feed_headline":"SPHEREx scheduler hits 99.54% planned sky coverage","feed_subtitle":"All-sky and deep-field plans both beat mission requirements while juggling orbit constraints.","key_machinery":"The central mechanism is the optimal target selection algorithm, an online greedy scheduler. At each step it reduces the high-dimensional scheduling problem to a single three-dimensional choice: which target group to slew to next. The choice is driven by a figure of merit, FoM = (1-F)(17-Nobs)/17 + F(Δθ/θref), which balances completing partially observed target groups against catching groups that are about to leave the allowable pointing zone. The algorithm also interleaves deep-field priority, downlink passes, and safe-pointing fallbacks, and plans 2 degrees inside every avoidance angle to absorb orbit-predict, attitude-control, and fault-protection error.","core_discovery":"The paper's central claim is that a tractable online optimization—choosing one target at a time with a hand-tuned figure of merit—can satisfy a set of interacting, time-varying pointing constraints well enough to complete an all-sky spectral survey from low-Earth orbit. The target list is organized into groups of 17 pointings spaced one spectral channel apart, and the scheduler first checks which groups lie inside the current allowable pointing zone, then favors groups with few completed observations and groups about to rotate out of the zone. Deep-field observations are prioritized when available, with a tunable randomness to keep all-sky coverage balanced. The claimed payoff is quantified","pith_inferences":["The paper reports planned completeness from simulations, not delivered completeness from flight data; the true proof of the margin will be measured voxel completeness from the public data releases, which the paper does not quantify.","The greedy online formulation suggests a general template for other low-Earth-orbit all-sky missions: separate targets into spectrally stepped groups and use a coverage-deficit term paired with an urgency term, rather than solving the full week-long schedule. The paper does not make this generalization.","Because deep-field priority is controlled by a tunable random threshold set near 85%, the scheduler is effectively trading a small amount of all-sky margin to protect the deep fields; the same knob could be re-tuned if a future mission's science case shifts.","A direct batch optimization over a full observing period would provide a bound on how much performance the greedy approximation sacrifices; the paper does not compare against such a bound."],"forward_implications":["All four all-sky surveys, covering two position-angle orientations over two half-mission periods, are planned to exceed 98% voxel completeness, with the limiting survey at 99.54%.","The combined deep fields reach a planned power-spectrum sensitivity of 29.3e-6, below the 40e-6 requirement; the northern field alone reaches 37.9e-6.","The roughly 2% glitch reobservation rate is absorbed naturally by the coverage-to-date feedback, so flagged observations are re-targeted without a separate replan.","The scheduler's in-flight performance is evidenced by the first public data release, which the paper says was planned successfully with the same software."],"supporting_citations":[{"why":"Describes the SPHEREx instrument and its 102-band linear-variable-filter design, which fixes how target groups are stepped on the sky.","marker":"[3]"},{"why":"Defines the mission science requirements, including the deep-field sensitivity threshold the planner's coverage is measured against.","marker":"[4]"},{"why":"The first public data release, cited as evidence that the planned observations are being executed successfully in flight.","marker":"[5]"},{"why":"Introduces the earlier all-sky observing-scenario strategy that this software implements and extends with the current target-selection algorithm.","marker":"[19]"},{"why":"Supplies the orbit-predict engine whose accumulating error sets the 2-degree inboard avoidance margin used in planning.","marker":"[21]"},{"why":"Provides the power-spectrum sensitivity formalism that converts deep-field coverage maps into the 29.3e-6 requirement metric.","marker":"[22]"}],"fun_headline_variants":["SPHEREx's scheduler plans on the fly, hits 99.54% sky coverage","SPHEREx's online optimizer beats all-sky survey targets","SPHEREx's algorithm juggles orbit limits to max sky coverage","SPHEREx scheduler balances orbit constraints for full-sky success"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The coverage and sensitivity numbers are planned values from simulations using a nominal orbit predict; they only become real if the flight system points where the plan says, with orbit and attitude errors inside the 2-degree margin and observation losses no worse than the simulated 2% glitch rate.","fun_headline_variants_meta":{"raw":{"variants":["SPHEREx's scheduler plans on the fly, hits 99.54% sky coverage","SPHEREx's online optimizer beats all-sky survey targets","SPHEREx's algorithm juggles orbit limits to max sky coverage","SPHEREx scheduler balances orbit constraints for full-sky success"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.0005,"raw_usage":{"total_tokens":2270,"prompt_tokens":718,"completion_tokens":1552,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":462,"completion_tokens_details":{"reasoning_tokens":1485}},"tokens_in":462,"tokens_out":1552,"duration_ms":14839,"temperature":1.0,"reasoning_tokens":1485,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T15:06:04.446739+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute the delivered voxel completeness for the completed surveys from the public data releases; if any of the four all-sky surveys falls below 98%, the claimed margin does not hold in flight. A second check is to count actual loss events: if real downlink errors, spacecraft anomalies, and glitch rates push reobservation demand above the simulated 2%, planned completeness overstates delivered coverage.","supporting_citations":[],"review_version":1}