{"id":"386631c3-3144-463d-8a83-a61ab9990d58","arxiv_id":"2608.07821","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Processing choices change pulsar noise model preferences; the 2023 reprocessing raises solar-wind support from four to ten pulsars, with densities matching the longer DR3 data set.","lead":"This paper compares two data-processing versions of the same Parkes pulsar timing observations and shows that the 2023 processing changes which noise models are preferred, making a solar-wind component appear in the top model for ten pulsars instead of four. The result matters because pulsar timing array searches for nanohertz gravitational waves depend on these noise models, and unaccounted model uncertainty can bias the search.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The solar-wind emergence count is conditional on a 13-model set with no chromatic competitor to DMv/SW; if unmodeled FD/profile-evolution systematics from the 2023 pipeline are absorbed by SW, the 10-vs-4 result and the DR3 'validation' (which uses the same reprocessed data) do not establish a…","rationale":"The reader's weakest_assumption identifies the fixed 13-model candidate set as the key vulnerability, and I agree that this is the most load-bearing concern. The paper's central observation that processing changes posterior model support is credible and useful, and the run-to-run stability of the MorphZ estimates is a genuine strength. However, the stronger conclusion that the solar wind is physically present and that the 2023 pipeline increases sensitivity to it depends on the SW component not absorbing an unmodeled frequency-dependent process. The 2023 pipeline explicitly adds FD parameters and new template construction to handle frequency-dependent profile evolution, which makes this omission concrete rather than hypothetical. The DR3 comparison is not an independent check because the 2023-reprocessed DR2 data are part of DR3; hence it cannot certify that the SW signal is real. The proposed expanded-model rerun would settle whether the SW-top count survives when a flexible chromatic competitor is present. If the count remains stable, the concern is resolved and the SW interpretation is much stronger. If it does not, the paper's headline claim about SW is conditional on model-set choice even though the broader processing-dependence message may still stand. Because the reader's conditional verdict already captures the need for such checks, my analysis does not move the verdict.","tokens_in":17794,"tokens_out":17150,"duration_ms":147292,"concrete_test":"For each of the ten pulsars whose 2023 top model contains SW, recompute posterior model probabilities with the candidate set expanded to include (i) EC+SW, (ii) EC+EQ+SW, (iii) a chromatic red-noise process with a free spectral index (in addition to DMv), and (iv) white-noise-only models. If a non-SW chromatic model becomes the top model for more than a couple of these pulsars, or if the SW-top-model count drops substantially below ten, the SW emergence is an artifact of the omitted chromatic process rather than a robust processing-induced detection.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative result, SW appearing in the highest-posterior model for ten pulsars under the 2023 pipeline versus four under the 2020 pipeline, is a statement about a fixed candidate set, not about the data-generating process. Section IV restricts the SW family to four models that always contain RN and/or DMv, and the set contains no chromatic noise process other than DMv and SW. The 2023 reprocessing (Section III) introduces wide-band-pulse-portrait templates and FD parameters to absorb frequency-dependent profile-evolution systematics; any residual frequency-dependent systematics from this step can be absorbed by the only remaining flexible chromatic terms, namely DMv and SW. The nearth agreement with PPTA DR3 [10] does not resolve this degeneracy because the '2023 DR2' data used here are the same reprocessed observations included in DR3 (Section III), so the agreement is not an independent validation. Consequently the 10-vs-4 count and the 'solar-wind signal already present' interpretation are conditional on the absence of an unmodeled chromatic process. The paper does acknowledge that model probabilities are conditional on the candidate space, but the physical SW interpretation goes beyond that conditional statement and needs the missing-process assumption to be tested.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper compares Bayesian single-pulsar noise-model inference for 22 pulsars from the Parkes Pulsar Timing Array second data release (PPTA DR2), processed with the original 2020 pipeline and the reprocessed 2023 pipeline. A fixed set of 13 candidate noise models is evaluated for each pulsar and each pipeline, with marginal likelihoods estimated using the authors' MorphZ method. The paper reports that posterior model probabilities change markedly between pipelines, with a solar-wind component appearing in the highest-posterior model for ten pulsars under the 2023 pipeline versus four under the 2020 pipeline, and that the recovered solar-wind electron densities at 1 au are consistent with the longer PPTA DR3 analysis. The authors conclude that single-pulsar noise-model selection is processing-dependent and argue that model uncertainty should be propagated into gravitational-wave background searches.","tokens_in":18184,"tokens_out":4755,"duration_ms":45460,"significance":"If the central claims hold, the paper makes a useful and timely contribution by quantifying how preprocessing choices can alter PTA noise-model selection and by demonstrating that many pulsars do not have a decisively preferred noise model. The explicit comparison of two processing chains on identical observations, the use of publicly available data, and the careful reporting of evidence-estimation precision are strengths. The paper also provides a concrete illustration of the differences between component-level support and complete-model selection, which is directly relevant to current PTA gravitational-wave analyses. However, the headline solar-wind result and the DR3 'validation' rest on assumptions about the candidate model set and the independence of the comparison that need to be tested before the physical interpretation can be accepted.","major_comments":[{"comment":"The 10-versus-4 solar-wind count is conditional on a candidate set that contains no chromatic noise process other than DMv and SW; Section IV restricts SW-containing models to four combinations that always include RN and/or DMv. The 2023 reprocessing (Section III) introduces FD parameters to absorb frequency-dependent profile-evolution systematics; residual frequency-dependent systematics from this step could be absorbed by the only flexible chromatic terms in the set, namely DMv and SW. The interpretation in Section V.D that the solar-wind signal is 'already present' therefore requires the assumption that no unmodeled chromatic process in the 2023 residuals mimics SW. Please test this assumption by adding a chromatic noise component with a free spectral index (or an explicit FD-like process) to the candidate set, and/or by showing that the recovered nearth values correlate with ecliptic latitude in the physically expected way and agree with in-situ solar-wind measurements.","section":"Section IV and Section V.B"},{"comment":"The DR3 comparison is not an independent validation because the '2023 DR2' data used here are the reprocessed subset of the observations included in the PPTA DR3 analysis (Section III, refs [13,37]). The agreement of n_earth with the DR3 values therefore does not establish that the recovered signal is a physical solar wind; both analyses can share the same processing-induced systematic. To support the physical interpretation, please validate against an external dataset or a differently processed sample (for example, in-situ OMNI solar-wind measurements, or NANOGrav/EPTA solar-wind results obtained with different pipelines), or otherwise demonstrate that the SW parameters cannot be explained by the FD/profile-evolution systematics introduced in the 2023 reprocessing.","section":"Section V.D and Table III"},{"comment":"The candidate model set excludes white-noise-only models, with the assertion that 'these excluded models have a zero posterior model probabilities for every pulsar in both datasets'; no supporting calculation or reference is provided. Because the reported posterior model probabilities, the 10-versus-4 SW count, and the 75% non-decisive fraction are all conditional on the chosen 13-model set, the exclusion needs to be justified quantitatively. Please report the evidence values for the white-noise-only models for at least a representative subset of pulsars, or state clearly that the reported probabilities are not directly comparable to analyses that include such models. This is load-bearing because the conclusion that model uncertainty should be marginalized over depends on the completeness of the candidate space.","section":"Section IV and Section VI"}],"minor_comments":[{"comment":"References [13] and [37] appear to be the same paper (Zic et al. 2023, PASA 40, e049); please merge them.","section":"References [13] and [37]"},{"comment":"The phrase 'assuming apriori equally probable' should read 'assuming a priori equally probable'.","section":"Section II"},{"comment":"The phrase 'a custom prior model probabilities' is ungrammatical; it should be 'custom priors on the model probabilities'.","section":"Section IV"},{"comment":"The sentence describing Table II is awkward and missing punctuation: 'The median and standard deviation of log(bz) the standard deviation...' Please rephrase to clarify that one quantity is the median of the run-to-run standard deviations and another is the standard deviation of those standard deviations.","section":"Section V.A and Table II"},{"comment":"The caption says the violin profiles show 'posterior summaries from Table III', but the markers already show medians and credible intervals. Please clarify whether the violins are smoothed full posteriors or are schematic, and how they relate to the tabulated intervals.","section":"Figure 4"},{"comment":"The statement that the DR3 values were 'independently obtained' is misleading in light of the shared reprocessed data; consider rewording to 'obtained with the longer DR3 dataset' or similar.","section":"Section V.D"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the pipeline comparison is genuinely new and worth having, but the solar-wind headline is conditional on a 13-model set that lacks any chromatic competitor beyond DMv and SW, so the physical reading should be softened.\n\nWhat is actually new here is the direct comparison of full posterior model probabilities for the same PPTA DR2 observations under the 2020 and 2023 processing pipelines. I don't know of another paper that does that. The MorphZ evidence estimates look numerically stable (median sigma of log z around 0.016 and 0.013), and the decisive/competitive/diffuse classification is a useful way to summarise model uncertainty. The 10-vs-4 count for solar-wind-containing top models is a fresh quantitative result, and the J1909-3744 example shows a nice separation between stochastic-process family and white-noise variant. The practical message that single-model summaries hide real uncertainty, and that noise models should be marginalized over in GWB searches, is sound.\n\nThe soft spots are real but not fatal. The candidate set is fixed at 13 models with only four SW-containing combinations, and white-noise-only models are excluded. The paper acknowledges in Section VI that the probabilities are conditional on this set, but the abstract and conclusion use language like 'the solar-wind signal is already present' that goes beyond what the conditional analysis can establish. The 2023 reprocessing introduced FD parameters to absorb profile-evolution systematics; any residual frequency-dependent effect has nowhere to go except DMv and SW, the only flexible chromatic terms in the set. So the ten-vs-four result could partly be an artifact of unmodeled chromatic processes. The agreement with DR3 nearth values is consistent, but it is not independent confirmation because DR3 includes the same reprocessed DR2 data. The text calls those values 'independently obtained', which overstates it. I would also like to see injection-recovery tests of the evidence estimates on these models; without them, we don't know how the model probabilities behave under a known truth.\n\nWho is this for: PTA analysts doing single-pulsar noise characterization and anyone propagating noise-model uncertainty into GWB searches. The core observation that processing choices materially change model support is credible and useful.\n\nRecommendation: send it to peer review. It deserves referee time. A good referee should push for a more careful separation of conditional statements from physical interpretations, and ideally some injection tests. But this is a solid, honest paper that raises a real issue.","headline":"A genuinely useful pipeline comparison, but the solar-wind claim is conditional on a fixed model set with no chromatic competitor—send to review with revision.","tokens_in":18690,"tokens_out":3612,"would_cite":true,"duration_ms":32267,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Reanalyzing the same pulsar timing data with a newer pipeline changes which noise models are preferred, with solar wind rising to the top model for ten pulsars instead of four.","keywords":["pulsar timing arrays","noise model selection","solar wind","Bayesian model comparison","marginal likelihood","gravitational-wave background","data reprocessing","posterior model probabilities"],"falsifier":"Re-process the same raw observations with a third, independent pipeline and run the same 13-model comparison: if the solar-wind component does not appear in most of the ten pulsars, or the inferred electron densities disagree with the independent longer-release values, the claim that reprocessing reveals a real solar-wind signal would be falsified.","tokens_in":17618,"feed_emoji":"🌞","tokens_out":10164,"duration_ms":84580,"temperature":0.7,"pith_summary":"The paper claims that the way pulsar timing data are pre-processed—calibration, radio-frequency-interference removal, template construction, and arrival-time estimation—is not statistically neutral for noise-model selection. Taking the same set of 22 pulsars from a second data release, processed by an older and a newer pipeline, it finds that the posterior support for competing noise hypotheses shifts materially: a solar-wind component becomes the highest-probability model for ten pulsars under the newer pipeline, versus four under the older one, and the recovered solar-wind electron densities agree with the array's longer third data release. The paper also shows that most pulsars, about 75%, keep substantial posterior support for alternative noise models, so no single best model adequately summarizes the data. If this is right, gravitational-wave background searches that condition on one fixed single-pulsar noise model are understating their uncertainty and should instead average over the model space.","feed_headline":"Solar wind emerges in 10 of 22 pulsars after reprocessing","feed_subtitle":"Two pipelines on the same observations shift posterior noise support, so gravitational-wave searches should not fix a single model.","key_machinery":"The analysis is carried by posterior model probabilities computed from independently estimated marginal likelihoods. For each of 22 pulsars and each of 13 fixed noise models, a marginal-likelihood estimator runs separate sampling and evidence calculation, so the Bayes factor between any two models and the posterior probability of every model follow from one application of Bayes' theorem. The candidate set combines white-noise components, with a per-observation error scale factor present in all models, with power-law red noise, dispersion-measure variations, and a solar-wind term parametrized by electron density at 1 AU plus amplitude and spectral index; this design lets the data choose both the stochastic-process family and the white-noise variant simultaneously.","core_discovery":"The central discovery is a comparative map of the noise hypothesis space for each pulsar under two processing pipelines. The authors evaluate the same 13 candidate noise models—combinations of white-noise parameters, achromatic red noise, dispersion-measure variations, and solar-wind variations—using marginal-likelihood estimates for every model, and compute posterior model probabilities with uniform model priors. They find that reprocessing changes both the concentration and the composition of these posteriors, changing the classification of seven pulsars between decisive, competitive, and diffuse, and systematically increasing support for solar-wind-containing models. The solar-wind densities inferred for the ten pulsars are consistent with the longer, independently analyzed data release, which the authors take as evidence that the solar-wind signal was already present in the older observations and that improved processing, not longer baseline, revealed it.","pith_inferences":["If the candidate set were expanded to include white-noise-only models or additional chromatic noise components, the ten-versus-four solar-wind count could shift; the reported numbers are conditional on a specific 13-model space.","Because each model evidence is estimated independently, the same results can be reweighted under informed model priors without resampling, offering a direct way to test how sensitive gravitational-wave-background evidence is to prior assumptions about noise components.","A natural extension is to apply the same hypothesis-space comparison to data from other pulsar timing arrays; if the solar-wind emergence is a generic reprocessing effect, the same pattern should appear in independent data."],"forward_implications":["Single-pulsar noise uncertainty should be propagated into gravitational-wave searches, for example by Bayesian model averaging, rather than conditioning on one selected model; otherwise common-process inference will overstate confidence.","For the roughly 75% of pulsars whose model posteriors are not decisive, reported noise parameters should be quoted as model-averaged quantities rather than as properties of a single top model.","Because reprocessing alone increased the number of pulsars with a top solar-wind model from four to ten, future data releases should treat processing choices as part of the systematic-error budget.","The agreement between the reprocessed shorter-baseline solar-wind densities and the longer-data-release values implies that improved processing can recover weak chromatic signals that older pipelines miss."],"supporting_citations":[{"why":"Supplies the original second-release timing models and arrival times used as the 2020 pipeline dataset.","marker":"[36]"},{"why":"Documents the reprocessed second-release dataset that forms the 2023 pipeline.","marker":"[37]"},{"why":"Describes the third data release pipeline and its processing changes, including RFI excision, templates, and ToA estimation, that the comparison relies on.","marker":"[13]"},{"why":"Provides the independent longer-data-release solar-wind electron density estimates used as the consistency benchmark.","marker":"[10]"},{"why":"Defines the noise components and the earlier sequential model-selection procedure that this paper contrasts with.","marker":"[21]"},{"why":"Introduces the marginal-likelihood estimator used to compute all model evidences.","marker":"[34]"},{"why":"Provides the software implementation of the noise models and likelihoods in which the analysis is carried out.","marker":"[42]"}],"fun_headline_variants":["Reprocessing exposes solar wind in 10 pulsars","Pipeline shift reshapes pulsar noise posterior","Data processing reveals solar wind in 10 pulsars","New pipeline changes pulsar noise model support"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The analysis assumes that the fixed set of 13 candidate noise models includes the true noise processes for every pulsar; if a real process is missing, the reported posterior probabilities, the 75% non-decisive fraction, and the ten-versus-four solar-wind count are conditional artifacts.","fun_headline_variants_meta":{"raw":{"variants":["Reprocessing exposes solar wind in 10 pulsars","Pipeline shift reshapes pulsar noise posterior","Data processing reveals solar wind in 10 pulsars","New pipeline changes pulsar noise model support"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00019,"raw_usage":{"total_tokens":1300,"prompt_tokens":870,"completion_tokens":430,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":486,"completion_tokens_details":{"reasoning_tokens":370}},"tokens_in":486,"tokens_out":430,"duration_ms":4341,"temperature":1.0,"reasoning_tokens":370,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T04:13:02.848562+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-process the same raw observations with a third, independent pipeline and run the same 13-model comparison: if the solar-wind component does not appear in most of the ten pulsars, or the inferred electron densities disagree with the independent longer-release values, the claim that reprocessing reveals a real solar-wind signal would be falsified.","supporting_citations":[{"cited_title":"The data are referred to the TT(BIPM2018) timescale and the JPL DE436 solar sys- tem ephemeris","cited_arxiv_id":null,"evidence_quote":"Supplies the original second-release timing models and arrival times used as the 2020 pipeline dataset."},{"cited_title":"Ellis and R","cited_arxiv_id":null,"evidence_quote":"Documents the reprocessed second-release dataset that forms the 2023 pipeline."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Describes the third data release pipeline and its processing changes, including RFI excision, templates, and ToA estimation, that the comparison relies on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the independent longer-data-release solar-wind electron density estimates used as the consistency benchmark."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the noise components and the earlier sequential model-selection procedure that this paper contrasts with."},{"cited_title":"Antoniadis, P","cited_arxiv_id":null,"evidence_quote":"Introduces the marginal-likelihood estimator used to compute all model evidences."}],"review_version":1}