{"id":"5d69a523-70d0-4699-a825-8159420febe0","arxiv_id":"2507.21541","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A systematic mapping of 128 sun sensor calibration studies provides a taxonomy of model representations and feature extraction techniques, along with a gap analysis and future research directions.","lead":"This paper organizes 128 studies on sun sensor calibration into a taxonomy of model types, feature extraction methods, sensor architectures, and performance levels. It gives spacecraft engineers a structured way to choose calibration algorithms and lists the field's current gaps, such as missing public datasets and limited adversarial threat research.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'comprehensive' claim rests on a single-database, single-query search with a subjective 20-page cutoff; an independent search could reveal omitted studies and bias the taxonomy.","rationale":"The paper is a useful and well-organized systematic mapping, but its core claim of being the first comprehensive survey depends critically on the completeness of the 128-paper corpus. The reader's weakest-assumption analysis correctly identifies the literature search in Section 3.2 as the most vulnerable step. My independent reading confirms that the search is narrow (one database, one query phrase) and the stopping rule ('after twenty pages' and 'until no further relevant papers could be found') is subjective. A more subtle but equally important issue is inclusion criterion I4, which excludes any study that does not explicitly report sensor task, final performance metrics, and architectural configuration. For a systematic mapping, this is an unusually strict requirement and could systematically omit algorithm-focused papers that do not report full experimental details, thereby biasing the taxonomy and the gap analysis. The survey's internal structure, case studies, and pseudocode are valuable and not undermined by this methodological shortcoming, but the 'comprehensive' and 'first systematic survey' claims should be tempered until the search is made reproducible and demonstrably broader. The correct verdict remains CONDITIONAL, as the reader already recommended, so no change to the verdict is needed.","tokens_in":584,"tokens_out":3752,"duration_ms":106933,"concrete_test":"Independently reproduce the literature search in Scopus and Web of Science using an expanded Boolean query, e.g., TITLE-ABS-KEY(('sun sensor' OR 'solar sensor' OR 'sun position sensor' OR 'sun angle sensor') AND (calibrat* OR model* OR 'feature extraction')), limited to 2002–2024 and English. Apply the same inclusion criteria I1–I5 and compare the resulting set with the 128 papers in Tables 2 and 3. If the independent search yields more than a small number (e.g., >5) of additional relevant papers, or reveals a model or architecture category absent from the survey's taxonomy, then the literature set is materially incomplete. Also count how many retrieved papers are excluded solely by I4 to quantify whether that criterion biases the mapping.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that this is the first systematic mapping and comprehensive survey of sun sensor calibration algorithms, based on 128 studies. The load-bearing assumption is that the compiled study set represents the relevant literature. Section 3.2 does not establish this: the initial set comes from a single search in one library database (Ex Libris Primo) using only the phrase 'sun sensor calibration', with title screening of only the first twenty pages. The search is then extended by backward/forward snowballing in Google Scholar until 'no further relevant papers could be found.' This protocol is not reproducible and is likely to miss relevant work that uses other terminology (e.g., 'solar sensor calibration', 'sun position sensor', 'sun angle sensor', 'analog sun sensor modeling'), is not indexed by that database, or appears beyond the twentieth page. Snowballing cannot compensate for an incomplete seed set, because papers disconnected from the seed will not appear in reference lists or citation graphs. Additionally, inclusion criterion I4 requires explicit reporting of sensor task, final performance metrics, and architectural configuration; this excludes calibration studies that omit one of these details, potentially biasing the taxonomy and the gap analysis. If relevant papers are missing, the trend statistics, Sankey flows, and the claimed 'research gaps' may not be representative, and the 'comprehensive' claim is overstated.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents itself as the first systematic mapping and comprehensive survey of sun sensor modeling and calibration algorithms, based on an analysis of 128 studies. It develops a taxonomy across five dimensions (model representation, sensor goals, feature extraction, architecture, and performance), visualizes relationships among these attributes with Sankey diagrams, and discusses representative case studies with equations and pseudocode. The paper also identifies five research gaps and recommends future directions. The central claim is that no prior systematic survey exists in this specific niche and that the compiled dataset represents the relevant literature.","tokens_in":53999,"tokens_out":3897,"duration_ms":45691,"significance":"If the search methodology were sound, this would be a genuinely useful resource: the taxonomy is coherent, the case studies are detailed, the pseudocode for feature extraction methods (Algorithms 1–22) gives reproducible implementation detail, and the authors make the compiled dataset publicly available on Zenodo and Tableau. The gap analysis and practitioner recommendations are reasonable and largely follow from the surveyed material. However, the paper's central claim of comprehensiveness is not yet supported by the reported search protocol, which is the key load-bearing weakness.","major_comments":[{"comment":"The literature search protocol is not reproducible and does not justify the 'comprehensive' claim. The initial set was obtained from a single database (Ex Libris Primo) using the single query 'sun sensor calibration', with title screening of only the first twenty pages and a stopping rule described only as 'saturation'. Backward/forward snowballing cannot recover papers that are not in the reference lists or citation graphs of the seed set. Since the paper's central contribution is the claim of being 'the first systematic mapping' of 128 studies, and the gap analysis in Section 7 (RQ4) depends on the representativeness of this set, the protocol must be substantially strengthened: report all search strings, use multiple databases, give the search date, and provide a PRISMA-style flow diagram. Alternatively, the authors should soften the claim to a scoping review of a sample.","section":"Section 3.2"},{"comment":"Inclusion criterion I4 requires that papers explicitly report the sensor task, final performance metrics, and architectural configuration. This will systematically exclude many calibration studies that report only a subset of these details, potentially biasing the taxonomy and the performance and architecture Sankey analyses in Section 4. The manuscript does not report the number of papers screened, the number excluded by each criterion, or the number that failed I4. Without this information, the reader cannot judge whether the 128-study set is a biased subsample of the literature.","section":"Section 3.3, criterion I4"},{"comment":"The accuracy bins (coarse, fine, very-fine, ultra-fine) are central to the performance Sankey diagram and to the RQ1 findings, but the mapping of individual studies to these bins is not shown. The qualitative definitions are given (e.g., fine is better than 0.5°; very-fine is better than one arcminute), yet no table or appendix lists each study's reported accuracy and assigned bin. Without this traceability, the claimed correlations between architecture, model representation, and performance cannot be verified by the reader.","section":"Section 4, performance bins"},{"comment":"The full list of the 128 included studies is not present in the manuscript; the authors point to external repositories on Zenodo and Tableau. For a survey whose central claim is comprehensiveness and whose classifications are the main scientific output, the study list should be included as an appendix or supplementary table within the paper itself. Relying on external links makes it difficult for reviewers and readers to verify the inclusion set, and external links can become inaccessible over time.","section":"Appendix A and Section 3.4"}],"minor_comments":[{"comment":"There is a typo in the inclusion criteria: 'explicity' should be 'explicitly'.","section":"Section 3.3"},{"comment":"In the opening sentence, 'neutral network-based model' should be 'neural network-based model'.","section":"Section 5.6"},{"comment":"The figure captions list citations as '[2,15,16,18–144]'; this citation range is unusual because the dataset reference [16] is not a primary study, and the range skips 17. Please clarify which references are intended.","section":"Figures 3–6 captions"},{"comment":"The formula for the minimum number of thresholds, 'N_TH > sigma / sqrt(e) / Delta_x_m', is typeset ambiguously; please use parentheses to make the expression unambiguous.","section":"Algorithm 10, line 7"},{"comment":"The legend for the feature-extraction assessment table reads 'G #= partially provides feature', but the table rows use a combination of symbols that are not all explained; a clearer legend or explicit cell entries would improve readability.","section":"Table 4"}],"recommendation":"major_revision","confidential_remarks":"The paper has a coherent taxonomy and useful implementation-level detail, but the 'systematic' and 'comprehensive' claims require a substantially stronger search methodology. The stress-test concern about the single-database, single-query search lands: this is a load-bearing weakness, not a mere presentation issue. I see no evidence of misleading intent, but the manuscript should address the search protocol and the traceability of the 128-study set before it can be accepted as a systematic survey."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: this is the first systematic mapping of sun sensor calibration algorithms I know of, and it will probably become the standard entry point for that niche. The taxonomy (model representations, feature extraction, architecture, performance) is coherent and the paper does real work: 128 studies categorized, Sankey diagrams, case studies with pseudocode, and a public dataset on Zenodo/Tableau. The pseudocode is specific enough to reimplement. That reproducibility is a genuine plus.\n\nThe soft spots are mostly in the method section, and they are real but not fatal. Section 3.2 describes a single database (Ex Libris Primo), single query 'sun sensor calibration', screening only the first twenty pages by title, then snowballing until saturation. That is not a reproducible systematic review protocol, and it likely misses work using other terminology ('solar sensor', 'sun angle sensor', etc.) or outside that database. Snowballing cannot fix a biased seed. Inclusion criterion I4 also drops studies that do not report all three of task, metric, and architecture, which could skew the taxonomy and gap analysis. The paper would be stronger with a PRISMA-style flow diagram, a documented multi-database query, and explicit handling of terminology variants. Also Equation 29 has a clear typo: alpha_2 is assigned to itself plus a term, which cannot be right. Minor, but worth fixing.\n\nI do not think these issues sink the survey. The 'comprehensive' claim is slightly overstated, but the taxonomy and case studies stand on their own as useful synthesis. There is no new experimental result here, so the value is reference value, not discovery. The authors are honest about the search limitations, which counts for something. The paper is for practitioners and newcomers to sun sensor calibration, not for researchers expecting novel algorithms; for that audience it is genuinely helpful.\n\nMy recommendation: send it to peer review. A serious referee can require a tightened search description and the equation fix without a full rewrite. I would cite it as the entry point to this literature.","headline":"Useful first systematic map of a niche field, undermined by a thin search protocol; still worth refereeing and citing.","tokens_in":54575,"tokens_out":1613,"would_cite":true,"duration_ms":18971,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims to be the first systematic mapping of sun sensor modeling and calibration algorithms, organizing 128 studies into a taxonomy of model representations, feature extraction techniques, sensor architectures, and performance…","keywords":["sun sensor","calibration","systematic mapping","survey","spacecraft attitude determination","model representation","feature extraction","centroid detection"],"falsifier":"A concrete check would be to rerun the search across additional databases and query variants, such as 'digital sun sensor calibration', 'analog sun sensor error compensation', 'sun sensor modeling', and venues not indexed by the first database, screening with the same inclusion criteria; if any qualifying paper published from 2002 onward is absent from the 128-study dataset, the claim of systematic coverage fails.","tokens_in":53547,"feed_emoji":"☀️","tokens_out":6821,"duration_ms":72159,"temperature":0.7,"pith_summary":"This paper claims that the body of work on sun sensor calibration, 128 studies published since 2002, has grown large and fragmented enough to require a systematic map, and that no such map existed before. Its contribution is a taxonomy that organizes the field along five axes: model representation, sensor goals, feature extraction, sensor architecture, and reported performance, together with a decision flow that connects design requirements to recommended algorithm choices. The survey reports that geometric models dominate model representations, thresholded centroid detection dominates feature extraction, and accuracy is the most frequently prioritized sensor requirement. If the survey is correct, it gives spacecraft engineers a single entry point to a scattered literature and a research agenda for the gaps it identifies.","feed_headline":"128 studies mapped: sun sensor calibration gets its first taxonomy","feed_subtitle":"A systematic review sorts two decades of calibration algorithms by model, architecture, and error type.","key_machinery":"The machinery that carries the argument is the systematic mapping itself: a literature dataset of 128 papers, screened through stated inclusion criteria, classified into the taxonomy of Table 1, and visualized through Sankey diagrams that trace flows from error sources, performance bins, and design requirements to detectors, masks, model representations, and feature extraction techniques. The taxonomy and the associated decision flow of Figure 1 are the named working objects: they translate a qualitative survey into a structured selection guide and a gap analysis. The central mapping at work is the relationship between the five taxonomy dimensions and the requirements that drive sensor choice, namely accuracy, cost, field of view, latency, power, precision, and volume.","core_discovery":"The paper's central claim is that this is the first systematic mapping and survey of sun sensor modeling and calibration algorithms, based on 128 studies meeting explicit inclusion criteria. It asserts that the field can be structured by a five-part taxonomy, namely model representation, sensor goals, feature extraction, architecture, and performance, and that cross-analyzing these dimensions with Sankey diagrams reveals recurring decision patterns: accuracy-driven designs tend to pair CMOS detectors with multi-aperture or encoded masks and geometric or neural-network models; cost-driven designs favor photodiodes with single-aperture or maskless configurations; geometric models are the most widely implemented model representation; thresholded centroid detection is the most common feature extraction technique; and alignment, manufacturing, and optical errors dominate the literature while environmental and interference errors are underrepresented. The paper also claims to identify five concrete gaps: missing public datasets, architecture-tight models requiring manual error characterization, limited adaptive online calibration, diminishing returns in centroid feature extraction, and a lack of adversarial-attack research. On the strength of this mapping it recommends future directions spanning learned feature spaces, multiplexing masks, event-based sensors, hybrid Kalman-neural filters, and adversarial defense.","pith_inferences":["Because the survey couples algorithms to detector and mask configurations, its taxonomy could transfer to other optical attitude sensors, such as star trackers and Earth horizon sensors, where centroiding and model-representation choices follow the same shape.","The reported absence of public datasets suggests an immediately testable extension: a standardized benchmark suite of sun sensor images, voltages, and ground-truth angles, built from the 128 studies' reported accuracy bins, would allow subsequent calibration papers to be compared quantitatively for the first time.","If thresholded centroiding has indeed reached diminishing returns, a concrete next experiment is to compare time-domain energy filtering or learned feature extraction against basic thresholding variants on the same hardware and noise conditions; the survey points to this direction without claiming it.","The survey's own inclusion criteria imply that its trend statistics describe a specific slice of the literature, namely English-language, 2002-onward studies with detailed implementation reporting, so older or sparsely documented work may be underrepresented in those trends."],"forward_implications":["A newcomer can use the taxonomy and decision flow to choose a model representation and feature extraction method from reported sensor requirements, without reading all 128 studies.","The error-to-architecture Sankey gives practitioners a first-pass prediction of which error sources a given detector-and-mask combination will need to compensate, such as alignment errors for CMOS multi-aperture systems and electrical errors for photodiodes.","If the gap analysis is right, the next productive research targets are mask-agnostic models, deep feature extraction for digital sensors, adaptive online calibration, and adversarial-hardened digital sun sensors.","The survey's accuracy-bin definitions, from coarse at 0.5 degrees or worse down to arcsecond-level ultra-fine, provide a common vocabulary that future papers can use to report performance.","The finding that thresholded centroid detection dominates despite diminishing returns implies that classical single-frame centroiding is not where performance gains will come from."],"supporting_citations":[{"why":"The prior review of terrestrial sun position sensors that this survey extends and explicitly differentiates from.","marker":"[13]"},{"why":"The prior comparison of analog sun sensor architectures that motivates the claim of a gap for calibration-focused surveys.","marker":"[14]"},{"why":"The 2002 MEMS micro digital sun sensor prototype that sets the survey's 2002 publication cutoff.","marker":"[15]"},{"why":"Supports the motivating claim that sun sensors are the most common sensor for small satellite attitude determination.","marker":"[1]"},{"why":"Supports the claim that nearly all low-Earth orbiting small satellites employ sun sensors and serves as the slit-model case study.","marker":"[2]"},{"why":"Defines the coarse and fine accuracy bins used throughout the performance analysis.","marker":"[43]"},{"why":"Anchors the encoded-mask and multiplexing-model category with the varying and coded aperture configuration.","marker":"[106]"},{"why":"Case study for multiple centroid averaging and ANN-based mapping used in the model representation and feature extraction analyses.","marker":"[98]"},{"why":"Case study for the basic centroid thresholding method, the most common feature extraction technique identified.","marker":"[20]"}],"fun_headline_variants":["First sun sensor calibration taxonomy from 128 studies","128 studies, one taxonomy: sun sensor calibration surveyed","Sun sensor calibration: first systematic map finds five research gaps","Two decades of sun sensor calibration studies finally mapped","Sankey diagrams reveal sun sensor calibration trends and gaps"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The survey's completeness rests on one literature search of a single database using the query string 'sun sensor calibration', screening the first twenty pages of results by title, and then snowballing until no further relevant papers were found; if the query or the database missed relevant papers, the taxonomy, trend statistics, and gap analysis would be biased.","fun_headline_variants_meta":{"raw":{"variants":["First sun sensor calibration taxonomy from 128 studies","128 studies, one taxonomy: sun sensor calibration surveyed","Sun sensor calibration: first systematic map finds five research gaps","Two decades of sun sensor calibration studies finally mapped","Sankey diagrams reveal sun sensor calibration trends and gaps"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000587,"raw_usage":{"total_tokens":2779,"prompt_tokens":990,"completion_tokens":1789,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":606,"completion_tokens_details":{"reasoning_tokens":1713}},"tokens_in":606,"tokens_out":1789,"duration_ms":13456,"temperature":1.0,"reasoning_tokens":1713,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T12:38:28.079605+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete check would be to rerun the search across additional databases and query variants, such as 'digital sun sensor calibration', 'analog sun sensor error compensation', 'sun sensor modeling', and venues not indexed by the first database, screening with the same inclusion criteria; if any qualifying paper published from 2002 onward is absent from the 128-study dataset, the claim of systematic coverage fails.","supporting_citations":[{"cited_title":"Disorder correction to the minimal conductance of a nodal-point semimetal","cited_arxiv_id":"2002.10353","evidence_quote":"The 2002 MEMS micro digital sun sensor prototype that sets the survey's 2002 publication cutoff."},{"cited_title":"Bright high-purity quantum emitters in aluminium nitride integrated photonics","cited_arxiv_id":"2006.16276","evidence_quote":"Case study for the basic centroid thresholding method, the most common feature extraction technique identified."}],"review_version":1}