{"id":"5397b003-bb37-491a-ac5f-856fcc531e8f","arxiv_id":"2601.21946","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Re-analysis of the POSS1 technosignature data finds no Earth-shadow deficit in the vetted sample, shows the nuclear-test correlation is driven by the Palomar observing schedule, and documents dataset inconsistencies and circular reasoning in the original claims.","lead":"This paper re-examines the public datasets behind recent claims that old Palomar sky-survey photos contain glints of artificial satellites, and finds the alleged Earth-shadow deficit disappears in the most vetted sample while the nuclear-test correlation is nearly fully explained by the telescope's observing schedule. It also documents inconsistent dataset definitions and a circular argument in the original studies.","discovery_kind":"replication","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The most load-bearing concern is the proxy inference: the p=0.1 nuclear-test refutation is computed on R and the 312-day denominator is inferred from R/W, not verified on V; if V's sampling differs, the correlation may survive. Direct V access would settle it.","rationale":"The paper under review is best read as a critical evaluation of Villarroel et al.'s evidence, not as a positive claim that no technosignatures exist. Its strongest contributions are the documented dataset-definition inconsistencies, the circular-reasoning analysis, and the demonstration that the claimed nuclear-test overlap is 96% explained by the POSS1 observation schedule alone. The most vulnerable point is exactly the reader's weakest assumption: the quantitative p=0.1 result is not a direct test on V, and the 312-day denominator is inferred from R/W rather than verified from V. This is a genuine load-bearing concern because the central refutation of the nuclear-test correlation rests on that normalization. The concern does not invert the overall verdict: the target studies bear the burden of validating their unpublished dataset, and the schedule-overlap argument is largely independent of the exact V sampling. But a conditional verdict is appropriate until V is released and reanalyzed. I agree with the reader's identification of the proxy inference as the weakest assumption, and the proposed test — obtaining and rerunning on V — would settle whether the concern actually lands.","tokens_in":32087,"tokens_out":8425,"duration_ms":101042,"concrete_test":"Request the V dataset from the Villarroel authors, as their data-availability statement says it will be shared on reasonable request, and rerun the Section 6.2 contingency table on V using the actual set of POSS1 observation days sampled in V, with both 312 and 368 denominators. If the Fisher exact p-value for V remains ≥0.05 and SPFs are found on ~310 of 312 days, the proxy concern is resolved; if p<0.05, the critique's erasure of the nuclear-test correlation is weakened. If V is not released, compute a sensitivity bound: find the smallest denominator D for which the reported 54/56 overlap and 310 SPF days yield p<0.05, and assess whether D is plausible given V's reported size and the POSS1 plate list.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The quantitative refutation of the nuclear-test correlation (Section 6.2, Table 3) is computed on the public dataset R, not on the target dataset V. The 312-day normalization is an inference: R and W contain ~1.4% and ~0.07% southern-hemisphere features, and the V/V' size difference implies 1.4%, so the authors remove 56 southern-only observation days from 368. But V is unpublished and its construction is ambiguous (Appendix A documents at least five inconsistent definitions). If V sampled some southern-only plates, or excluded some northern plates/days, the effective denominator changes and the Fisher exact p≈0.1 could shift below 0.05. The 96% schedule-overlap argument is more robust because it uses the target study's own 54 overlap days against the 56 POSS1 observation days that overlap nuclear-test windows. However, the formal 'loses significance' claim is V-dependent. Section 5's shadow comparison has the same proxy structure: f_obs=1.36% vs f_exp=0.82% is computed on R and reported without uncertainty, so it cannot directly rule out a V-specific deficit if V's plate selection differs. The paper's assertion that access to V is unnecessary (Section 3.1) is too strong: it is unnecessary for exposing dataset inconsistencies, but it is necessary for the quantitative normalization claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper re-examines claims by Villarroel et al. (2025c) and Bruehl & Villarroel (2025) that unidentified features in POSS1-E plates show an Earth-shadow deficit, linear clusters, and correlations with nuclear tests. Using two public datasets, R (N=5,399, the most aggressively filtered set) and W (N=171,753), plus the MAPS celestial-object set M, the authors argue that the assumed uniform-random background is false: SPF density rises toward plate edges, and large-scale stripes and voids appear in W and R. They report no shadow deficit in R (f_obs=1.36% vs f_exp=0.82%), and find that the nuclear-test correlation becomes insignificant (p≈0.1) when normalized by 312 actual POSS1 observation days rather than 2,718 calendar days; they further show that 54 of the 56 days overlapping nuclear-test windows are pure schedule overlap. They also document inconsistent definitions of the unpublished dataset V and identify a circular argument in the target papers.","tokens_in":32426,"tokens_out":9653,"duration_ms":95851,"significance":"If the quantitative claims hold, the paper materially weakens the technosignature interpretation of the POSS1-E feature sets and provides a useful cautionary case study for archival-plate searches. Its strengths are the use of public datasets, the careful documentation of dataset inconsistencies in Appendix A, and the transparent schedule-overlap argument, which is simple and robust. The historical GRB-plate review is also well placed and supports the demand for independent validation. However, the strongest quantitative refutations—the p=0.1 nuclear-test result and the shadow comparison—are computed on the public dataset R, not on the unpublished target dataset V, and rely on inferred sampling properties of V. The paper would be more convincing with a direct test on V or a thorough sensitivity analysis; as written, the normalization claims are conditional on an unverified assumption.","major_comments":[{"comment":"The 'loses significance' claim (p=0.1) is computed with the public dataset R, while the target correlation is for the unpublished V. The 312-day denominator is inferred from the 1.4% southern-hemisphere fraction in R and W (§3.2) and then used to remove 56 southern-only observation days. Appendix A shows that V's construction is inconsistently defined (five different statements), so its effective observation window is not known. If V sampled southern-only plates, or excluded some northern plates, the denominator and p-value could change. The statement in §3.1 that access to V is 'not necessary' is therefore too strong: it is sufficient for exposing dataset inconsistencies, but not for the quantitative normalization claim. Please obtain V (one author had a copy) or provide a sensitivity analysis over plausible V constructions and denominators.","section":"§6.2, Table 3"},{"comment":"The in-shadow test is reported as f_obs=66/4866=1.36% versus f_exp=0.82%, with no uncertainty, confidence interval, or significance test. The expected fraction depends on simulation choices (30×30 cell grid, 8.5° in-shadow radius, plate-assignment radius) that are not varied. Although f_obs exceeds f_exp, the paper's claim that the reported deficit is 'not present' needs a formal test (e.g., Poisson with plate-level overdispersion) to be comparable to the 2.5–22σ claims in the target study. Moreover, the calculation is on R, not V; because V is unpublished and ambiguously defined, it cannot exclude a V-specific deficit if V's plate selection differs from R's. Please add a significance/sensitivity analysis and limit the scope of the conclusion to R (or V once obtained).","section":"§5"},{"comment":"The argument that the uniform-random null is false relies heavily on set W, which is a subset of S selected by proximity to NeoWISE objects. The paper calls W 'a dense, uniform-random sampling of S' without demonstrating that the NeoWISE positions used for matching are spatially uniform. Clustering in W could in principle reflect structure in the infrared catalog or in the matching procedure. The MAPS set M is a useful control for celestial sources but is not matched to NeoWISE. Please verify the uniformity of the sampling positions (e.g., compare W to a random subset of NeoWISE positions) or re-derive the nonuniformity conclusion directly from R's intra-plate density profiles (which are shown but not formally tested). This matters because the invalidation of the Poisson null is load-bearing for rejecting the target's shadow significance.","section":"§4.2–4.4"}],"minor_comments":[{"comment":"V is given as N=107,185, whereas Table 1 and §6 use N=107,875. Please correct the typo.","section":"§3.2, 'SetV' paragraph"},{"comment":"'solid angle (in degrees)' should read 'square degrees'.","section":"§5"},{"comment":"Plate 090R has 2,149 SPFs with an average of 257.5 per plate (§4.1); it is therefore ~8.3 times the average, not 'nearly 20 times the average'.","section":"§4.2 and §5"},{"comment":"The statement 'SPFs are found on 310 of the 312 observations days' conflicts with §6.2, where R has 289 unique days and W has 307; state that this refers to V and reconcile the numbers.","section":"§8, item 5"},{"comment":"Specify which statistical test produced p=0.1 (Fisher exact, chi-square with/without Yates) and report the test statistic; the text mentions Chi-Square=6.94 for the original but not for the recalculation.","section":"§6.1, Table 3"},{"comment":"Consider releasing the plate-assignment and shadow-simulation code to make the analysis fully reproducible; the current text lists parameters but not the algorithms.","section":"§4.1 and §6.1"}],"recommendation":"major_revision","confidential_remarks":"The paper is strongly worded but stays within scientific norms. The most defensible contribution is the schedule-overlap argument and the documentation of dataset inconsistencies. The referee concerns center on the proxy use of R/W instead of V; I would ask the editor to encourage the authors to pursue direct access to V (one author had a copy in July 2025) or to reframe the quantitative normalization claims as conditional on V's inferred sampling. If V cannot be obtained, a careful sensitivity analysis would still make the paper publishable, though with weaker claims."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my read. The paper is a systematic re-analysis of the POSS1 plate features that Villarroel and colleagues have been pushing as evidence for pre-Sputnik artificial satellites. Taking the public datasets (R and W) seriously, the authors show three things worth knowing. First, the alleged deficit of features in Earth's shadow does not appear in the most vetted dataset R: 66 of 4,866 unambiguously assigned features land in shadow, 1.36% versus the expected 0.82%. Second, the much-publicized correlation between feature dates and nuclear tests is essentially a product of the Palomar observing schedule: 56 of the 368 actual observation days overlap nuclear test windows, and 54 of those are claimed as coincidence, so 96% of the overlap is geometry and weather, not physics. Third, the target papers rely on a dataset V that is unpublished, defined inconsistently across the papers, and built by re-filtering a set that had already been cleaned to R; the original papers also use circular reasoning, treating the very correlations under test as validation of the data. These are serious points, and they land.\n\nWhat's genuinely new: the observation-schedule normalization, the R-based shadow test, the documentation of RA-stripe voids and edge/corner excesses, and a careful inventory of dataset inconsistencies. The manual inspection of 540 R features (4–5% residual contamination) is honest empirical work. The GRB plate-history review is context, not a result, but it is well chosen.\n\nThe soft spots are real but not fatal. The shadow comparison on R comes without uncertainty quantification; 66 versus an expectation of 40-ish sounds suggestive, but we are not given a p-value or confidence interval, and only ten plates contribute. More importantly, the quantitative refutation of the nuclear correlation—the p=0.1 Fisher exact result—is computed on R with a 312-day denominator inferred from R and W, not verified on V. V is unpublished, and if its plate selection differs in declination or in which northern plates were included, the denominator changes and the p-value could shift. The authors claim access to V is unnecessary; that is true for exposing inconsistencies and circularity, but it is false for the precise quantitative normalization claim. They should soften that. There is also a small independence issue: the acknowledgments state that Watters received a copy of V while collaborating on an earlier version of one of the target papers, then withdrew. He may well have prior knowledge of V's contents even if it was not used analytically. That does not invalidate the public-data analysis, but it undercuts the clean-room claim.\n\nCentral argument holds. The target papers' evidence does not survive contact with the public data, and the burden is on them to show V behaves differently. I'd send this to a serious referee, asking specifically for error bars on the shadow statistic and a more careful statement about what the R-based normalization can and cannot establish. For a reading group on archival-plate methods and null-model design, it is a good case study.","headline":"Solid critique of the Villarroel et al. shadow/technosignature claims, with a few inferential steps resting on the unpublished V dataset.","tokens_in":32917,"tokens_out":3774,"would_cite":true,"duration_ms":38226,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that the reported deficit of features inside Earth's shadow and the correlation with nuclear tests vanish when the most vetted public dataset is used and the survey's true observation days are counted.","keywords":["technosignatures","POSS1-E plates","optical transient candidates","Earth shadow deficit","plate artifacts","nuclear-test correlation","replication analysis","data validation"],"falsifier":"A direct re-release of the exact unpublished dataset V with plate IDs and timestamps would settle the question: if, after matching to plates, feature days appear on substantially fewer than ~290 of the 312 northern observation days, or if the in-shadow fraction in V falls below the 1.36% seen in R, the artifact explanation would need revision. Conversely, microscopic inspection of a random sample of V features showing star-like 3-D silver-grain structure and no catalog counterpart could revive the transient interpretation.","tokens_in":31987,"feed_emoji":"🔭","tokens_out":4828,"duration_ms":53305,"temperature":0.7,"pith_summary":"This paper re-examines the evidence behind claims that unidentified features on pre-Sputnik photographic plates are glinting artificial objects near Earth. Using the most aggressively filtered public subset of the feature catalog, the authors find no deficit of features inside Earth's shadow: 66 of 4,866 features fall in the shadow (1.36%), versus 0.82% expected. They also find that the reported correlation between feature observation days and nuclear tests disappears (p≈0.1) once the analysis is normalized by the 312 nights the survey actually observed the northern sky, and that 96% of the reported overlap with test windows is just the observing schedule. The paper documents that the unpublished dataset used in the original studies is inconsistently defined, largely unseparated from catalog stars and plate artifacts, and that the features' spatial patterns point to plate defects rather than sky sources. A sympathetic reader would care because the technosignature conclusion rests on these vanishing signals, and the paper shows how to test such archival claims properly.","feed_headline":"Earth-shadow and nuclear-test signals vanish under reanalysis","feed_subtitle":"Counting only the 312 nights the survey truly observed erases the reported correlation; the features sit near plate edges, not in the sky.","key_machinery":"The load-bearing tool is the 'remainder' set R, the most aggressively filtered public subset of the larger feature catalog, built by removing features near catalog stars, asteroids, variables, objects seen in only one scan, and other artifacts. Because R and the unpublished dataset share the same ~1.4% southern-hemisphere fraction, the authors use R as a stand-in for the unpublished data. Against R they run three quantitative tests: a cell-based shadow simulation comparing in-shadow fractions, a Clark-Evans nearest-neighbor ratio to measure clustering against complete spatial randomness, and radial density profiles to show edge and corner excesses. The temporal test uses the plate metadata t","core_discovery":"Working from the most heavily filtered public subset of the POSS1-E feature catalog (set R, N=5,399), the authors find 66 of the 4,866 unambiguously plate-assigned features inside the simulated geosynchronous shadow — 1.36%, versus 0.82% expected from the shadowed fraction of the survey area — so the reported 30-75% deficit is absent. After matching features to plates and normalizing by the 312 nights on which the survey actually exposed northern-sky plates (not the 2,718 calendar days of the study window), the feature/nuclear-test correlation drops to p≈0.1 with a relative risk of 1.07; 96% of the overlap between feature days and test windows is just the observing schedule. The paper also s","pith_inferences":["Beyond the paper's own claims, the results imply that any future archival-plate search for artificial satellites should require independent per-object validation (e.g., microscopic inspection of the emulsion) rather than relying on statistical correlations over unvalidated catalogs.","The conspicuous deficit band in right ascension (roughly 87°-107°) that survives aggressive filtering suggests a systematic pipeline effect; identifying its cause would provide a corrected background model for any reanalysis of these plates.","A testable extension would compare feature detection days against other weather-dependent activities, such as aerial surveys or artillery tests, to see whether the apparent seasonal coupling with nuclear tests is generic rather than specific.","Since even the vetted R set still contains 4-5% clear stars and artifacts, the cleanest available catalog needs per-object morphology screening before any single candidate can be used as evidence of a real optical transient."],"forward_implications":["The reported 30-75% deficit of features inside Earth's shadow, cited as evidence for geosynchronous glinting objects, is not reproduced with the vetted remainder set; the in-shadow fraction (1.36%) actually exceeds the expected 0.82%.","The claimed feature/nuclear-test correlation (χ²=6.94, p=0.008 in the original study) becomes p≈0.1 with relative risk 1.07 when normalized by the 312 true northern-sky observation days.","Because features appear on 93-99% of actual observation days, their occurrence is almost completely determined by when the survey observed the sky, not by nuclear testing.","A third of the candidate aligned-cluster features match catalog stars within 2 arcseconds, so they were not confidently distinguished from known objects.","The spatial patterns — edge/corner excess, right-ascension stripes, and plate-boundary-related clusters — indicate plate and digitization artifacts rather than sky sources."],"fun_headline_variants":["Reanalysis finds no occultation deficit in POSS1-E features","Nuclear-test link vanishes after proper normalization","POSS1-E technosignature evidence: flaws in shadow and schedule","Circular logic and plate defects undermine POSS1-E claims","No Earth-shadow deficit: POSS1-E features controlled by plate edges"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The analysis assumes the unpublished dataset V, on which the original claims were based, has the same sampling properties as the public R and W sets — particularly that southern-hemisphere plates were skipped and that the 312 northern observation days define the relevant exposure window; if the original authors sampled V differently, the normalization that erases the nuclear-test correlation would be weakened.","fun_headline_variants_meta":{"raw":{"variants":["Reanalysis finds no occultation deficit in POSS1-E features","Nuclear-test link vanishes after proper normalization","POSS1-E technosignature evidence: flaws in shadow and schedule","Circular logic and plate defects undermine POSS1-E claims","No Earth-shadow deficit: POSS1-E features controlled by plate edges"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000249,"raw_usage":{"total_tokens":1451,"prompt_tokens":871,"completion_tokens":580,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":615,"completion_tokens_details":{"reasoning_tokens":493}},"tokens_in":615,"tokens_out":580,"duration_ms":6080,"temperature":1.0,"reasoning_tokens":493,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T06:44:06.546889+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A direct re-release of the exact unpublished dataset V with plate IDs and timestamps would settle the question: if, after matching to plates, feature days appear on substantially fewer than ~290 of the 312 northern observation days, or if the in-shadow fraction in V falls below the 1.36% seen in R, the artifact explanation would need revision. Conversely, microscopic inspection of a random sample of V features showing star-like 3-D silver-grain structure and no catalog counterpart could revive the transient interpretation.","supporting_citations":[],"review_version":1}