{"id":"a9c421db-b92d-4003-a8c3-7a992d61ffcf","arxiv_id":"2505.04990","paper_version":2,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A METM-based calibration using archetype frequency functions and Gaussian processes improves polarization calibration and timing quality for the full Nançay NUPPI pulsar archive.","lead":"This paper improves how pulsar observations from the Nançay Radio Telescope are calibrated for polarization. A new technique makes historical data from 2011 to 2019 nearly as clean as modern data, boosting the precision of pulsar timing used in gravitational wave searches.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"METM calibration assumes J0953+0755's polarized profile is stable over 2011–2024, but this is not independently tested; if false, the apparent homogenization and timing gains could reflect a consistent but incorrect calibration.","rationale":"The paper is careful and the evidence is directionally consistent: the METM-based profiles are visibly more homogeneous, S/N ratios generally improve, and the timing noise parameters mostly move in the expected direction. The reader's weakest-assumption analysis correctly identified the archetype shape and reference-pulsar stability as the key assumption. I agree that this is the load-bearing point, but I would place more weight on the absence of any independent check of J0953+0755's polarization stability. The internal self-consistency of the archetype construction (Fig. 8) cannot distinguish a correct instrument model from a stable but wrong reference profile, and the paper explicitly notes that METM can fail when the reference pulsar's profile is time-variable. Because the method is intended for adoption in EPTA/IPTA data releases, the missing validation is consequential. A split-half cross-validation using independent subsets of J0953 observations would settle whether the derived calibration is robust or whether the reported improvements partly reflect a consistent calibration error. If that test passes, the paper's conclusion should stand; if it fails, the central claim would need substantial revision. I therefore recommend CONDITIONAL rather than unconditional ACCEPT: accept the analysis and methodology, but require the proposed cross-check before relying on the pre-2019 calibration for precision-timing applications.","tokens_in":34740,"tokens_out":8792,"duration_ms":97490,"concrete_test":"Perform a split-half cross-validation on J0953+0755: after fixing the segment boundaries as in Sect. 3, split the J0953 observations within each time segment into two independent subsets (e.g., even/odd epochs). Derive archetypes and Gaussian-process predictions separately from each subset, apply both calibrations to the same full set of MSP observations, and compare the resulting polarimetric profiles and the Table 4 Wrms values (per pulsar and as a median). If the two subsets yield systematically different MSP profiles or significantly different Wrms, reference-profile evolution or overfitting is contaminating the calibration; if they agree within the expected statistical scatter, the assumption of J0953 stability is supported.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim that the METM-based calibration significantly improves pre-November 2019 NUPPI data (Sect. 4, Table 4) rests on two linked assumptions: (i) within each manually defined time segment, every calibration parameter's frequency dependence is a scaled/offset version of a single archetype (Eqs. 7–9), and (ii) the polarized reference profile of PSR J0953+0755 is intrinsically stable over the full 2011–2024 span. The paper demonstrates internal consistency: scaled/offset parameters match the archetypes (Fig. 8), and GP fits track the time variations (Fig. 9). However, profile homogeneity is not by itself an accuracy metric. A stable but incorrect reference profile would be propagated by METM to all MSPs, making their calibrated profiles mutually consistent while imprinting the reference pulsar's polarization errors. The timing improvement is more independent, but it is modest (median Wrms 1.146→1.078 µs for FDM; 0.842→0.818 µs for MTM) and per-pulsar changes are often of order the reported uncertainties. The paper cites Dey et al. (2024) showing METM can fail when the reference pulsar's profile is unstable, and it discards J1136+1551 because of strong parameter variability, yet it provides no quantitative stability test for J0953+0755. Since pre-2019 NUPPI data have no independent MEM calibration, this unverified reference-stability assumption is the most load-bearing point for the paper's headline claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents two main results. First, using three sessions of special rotating-feed observations of bright pulsars at different declinations, the authors test whether the polarimetric response of the Nançay Radio Telescope (NRT) depends on hour angle or declination. From a model-comparison analysis with a modified pcm tool and validation on normal-mode millisecond pulsar (MSP) observations, they conclude that the response does not appear to vary with these parameters. Second, to calibrate pre-November 2019 NUPPI data, they develop a new calibration scheme based on the Measurement Equation Template Matching (METM) technique, using PSR J0953+0755 as the reference pulsar. Within manually determined time segments, the frequency dependence of each calibration parameter is represented by an archetype function, and the time variation of per-observation scale and offset parameters is modeled with Gaussian processes. Applying this calibration to 12 MSPs, they find more homogeneous polarimetric profiles, higher signal-to-noise ratios in most cases, and improved timing quality as quantified by reduced weighted RMS residuals and lower white- and red-noise parameters (median Wrms drops from 1.146 to 1.078 microseconds for FDM and from 0.842 to 0.818 microseconds for MTM).","tokens_in":35086,"tokens_out":7351,"duration_ms":75638,"significance":"If the calibration improvement is real, it will enable consistent use of the NRT's 2011-2019 NUPPI data in pulsar timing array analyses and polarimetric studies, extending the improvements previously demonstrated for post-November 2019 data. The paper's main strengths are the validation of the calibration on MSPs that were not used to derive the calibration solutions, the multi-metric assessment (profile homogeneity, S/N, TOA uncertainties, and noise parameters), and the careful null test of direction dependence. The work also provides a reproducible analysis path through public PSRCHIVE extensions. The conclusions are consistent with those of Rogers et al. (2024) and appropriately nuanced in their wording, e.g., that the response 'does not appear to vary' with direction. The manuscript is well written and the figures are informative.","major_comments":[],"minor_comments":[{"comment":"The choice of PSR J0953+0755 as the METM reference is justified by the smoothness of the derived parameters compared with J1136+1551, but a quantitative stability test would strengthen the paper. For example, deriving independent calibration solutions from disjoint subsets of J0953+0755 observations (e.g., early versus late epochs) and comparing them would directly address the reference-profile-stability failure mode discussed for Dey et al. (2024).","section":"Sect. 3, choice of METM reference pulsar"},{"comment":"For three MSPs (J0613-0200, J1022+1001, J2124-3358) the median S/N ratio R1 is below 1, so the statement that S/N increases in 'almost all' tested MSPs is correct but could be more explicit. It would be helpful to state which pulsars show degradation and whether the effect is statistically significant relative to the scatter in the ratios.","section":"Sect. 4.1 and Table 3"},{"comment":"The reported improvements in Wrms are modest (a few percent) and per-pulsar changes in Table C.1 are often comparable to the quoted uncertainties. A paired statistical test across the 12 pulsars (e.g., a Wilcoxon signed-rank test on the Wrms ratios or on the EQUAD values) would better quantify the significance of the global improvement and would address the concern that some changes may be within noise.","section":"Sect. 4.2 and Table 4"},{"comment":"The quasi-AIC metric is introduced with appropriate caveats about the asymmetric sampling of declination versus hour angle, and the conclusion of no direction dependence ultimately rests on the S/N comparison in Fig. 5. It would be informative to report the effective number of observations that probe extreme hour angles for each declination, or to show a direct comparison of calibration solutions (e.g., differential gain) evaluated at the extreme hour angles, to make the null result more transparent.","section":"Sect. 2, qAIC metric"}],"recommendation":"minor_revision","confidential_remarks":"The paper is suitable for A&A and the central claims are defensible. The reference-stability concern raised in the stress-test is real but does not invalidate the paper: the authors have selected the smoother of two candidate reference pulsars and validated the calibration on independent MSP data. I recommend requesting a quantitative stability test and a statistical significance statement, but these are local additions that do not require reworking the analysis. I support publication after minor revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a careful instrumentation paper that does what it claims: it recovers the pre-2019 Nançay NUPPI archive for precision timing. The headline result, the METM-based calibration using J0953+0755 as reference with per-segment archetypes and Gaussian-process interpolation, genuinely improves the data. Across 12 MSPs not used to build the calibration, profiles become more homogeneous, S/N generally rises, and TOA noise drops (median Wrms 1.146 to 1.078 microseconds for FDM, 0.842 to 0.818 microseconds for MTM). The effect is modest but consistent, and validation on held-out pulsars keeps the circularity burden low.\n\nWhat is actually new: the joint multi-pulsar MEM fitting in pcm with 2D polynomials in declination and hour angle, the archetype-plus-GP procedure for the historical data, and the parallactic-angle sign fix in PSRCHIVE for NRT, which itself improves fit quality. The paper also handles the literature honestly, citing Dey et al. (2024) as a case where METM failed and Rogers et al. (2024) where it worked, and explains why Nançay lands in the Rogers camp.\n\nSoft spots, in proportion. The conclusion that the NRT response is direction-independent is an indirect null. The qAIC numbers alone are ambiguous; several polynomial models beat the 0/0 model, so the argument rests on the S/N comparisons and the overfitting demonstration (higher-order models degrade S/N in about 80 percent of cases). That is reasonable, but the claim is a bit more confident than the evidence strictly allows.\n\nThe stress-test concern about reference-profile stability is real but lands on the absolute polarimetry claims more than on the timing claims. A stable-but-wrong reference would imprint a static distortion on all MSPs, harming absolute polarization accuracy while still removing epoch-dependent distortion, which is precisely what lowers TOA noise. So the timing improvement is fairly robust to this worry. The more homogeneous profiles metric, however, cannot distinguish correct from consistently applied, and the paper never quantitatively tests J0953+0755's long-term stability. It did discard J1136+1551 for instability, which is some evidence of care, but for future RM or position-angle work an external check is needed.\n\nSmaller items: time-segment boundaries were set by visual inspection, no data or scripts are released (though the pcm branch is available), a couple of pulsars show no S/N gain or slight regression, and Table C.1 labels its R column METM+MEM where it means METM+MTM.\n\nWho this is for: anyone using NUPPI, EPTA, or IPTA data, and anyone doing pulsar polarimetric calibration. It deserves serious refereeing; the main referee questions should be the reference-stability test and the qAIC-versus-S/N tension. My recommendation: engage, and accept with the stability caveat made explicit.","headline":"A careful, transparent calibration paper that convincingly recovers the pre-2019 Nançay archive for precision timing, with the main caveats being an unquantified reference-pulsar stability test and an indirect null result on direction dependence.","tokens_in":35614,"tokens_out":5563,"would_cite":true,"duration_ms":57037,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The Nançay Radio Telescope's response is direction-independent, and a template-based calibration recovers accurate polarimetry for pre-2019 pulsar data, cutting timing noise.","keywords":["polarization calibration","pulsar timing","measurement equation template matching","matrix template matching","Nançay Radio Telescope","millisecond pulsars","radio polarimetry","Gaussian process interpolation"],"falsifier":"Re-derive the pre-2019 calibration without the archetype constraint, fitting each epoch's calibration parameters freely, and compare the resulting millisecond-pulsar timing residuals with the archetype-based ones; materially lower white or red noise in the free fits would show the archetype assumption is biasing the solutions. Independently, compare METM-predicted Stokes $Q$ and $U$ for J0953+0755 at an early epoch against a well-calibrated observation of a different bright pulsar from the same epoch; systematic growth of the residuals would falsify the assumed decade-long profile stability.","tokens_in":34553,"feed_emoji":"📡","tokens_out":15514,"duration_ms":140930,"temperature":0.7,"pith_summary":"This paper establishes that the Nançay Radio Telescope's polarimetric response is effectively independent of where in the sky it points: the apparent hour-angle dependence seen in earlier work came from a sign error in the parallactic-angle calculation, and a constant calibration model performs best once that is fixed. It then shows how to extend the improved polarization calibration to NUPPI observations taken between 2011 and 2019, before the rotating-feed-horn calibration scheme existed, using measurement equation template matching on a bright, polarization-stable reference pulsar. The method describes each calibration parameter's frequency dependence as a scaled and offset version of a segment-specific archetype function, with the scale and offset factors interpolated in time by Gaussian processes. Applied to twelve millisecond pulsars, this calibration yields more homogeneous polarized profiles, higher signal-to-noise ratios in most cases, and timing data with lower white and red noise; the best results come from combining it with matrix template matching for time-of-arrival extraction. If right, it makes the first eight years of NUPPI data usable at modern precision, which matters for pulsar timing array searches for gravitational waves.","feed_headline":"Recalibrated Nançay pulsar data cut timing noise across a decade","feed_subtitle":"Archetype-based METM calibration sharpens pre-2019 profiles; adding matrix template matching gives the best TOAs.","key_machinery":"The central object is the calibration archetype, used inside the METM (measurement equation template matching) procedure that derives instrumental calibration by comparing observations of a reference pulsar with a well-calibrated polarized template. Within each manually identified time segment in which the NRT's response was stable, the frequency variation of every calibration parameter is represented by a single archetype function, and each individual observation's parameter values are assumed to be a scaled and possibly offset copy of that function (Eqs. 7 to 9). The time evolution of the scale and offset factors is then modeled with Gaussian processes, yielding predicted calibration solutions at arbitrary epochs. The direction-dependence test uses a modified measurement-equation model in which differential gain and phase are two-dimensional polynomials of declination and hour angle, fit jointly to several pulsars' rotating-horn observations; the winning model is the one with constant parameters.","core_discovery":"On the paper's own terms, the central discovery is twofold. First, the NRT's polarimetric response does not appear to vary measurably with hour angle or declination: a joint analysis of rotating-feed-horn observations of seven pulsars spanning declinations from roughly $-28^\\circ$ to $+56^\\circ$, with differential gain and phase modeled as polynomials in hour angle and declination, selects the constant $0/0$ model, and applying higher-order solutions to normal-mode MSP observations degrades signal-to-noise ratios. Second, a calibration procedure built on measurement equation template matching recovers accurate polarization calibration for pre-November 2019 data, where no rotating-horn observations exist. Using the bright pulsar J0953+0755 as a reference, the authors define time segments of stable instrumental response, construct archetype functions for the frequency dependence of each calibration parameter, and model the time evolution of scale and offset factors with Gaussian processes. On twelve millisecond pulsars this raises median signal-to-noise ratios (for example $1.035$ for J1730$-$2304 and $1.130$ for J1744$-$1134), lowers time-of-arrival uncertainties, and reduces both white noise and red noise in timing residuals. Combined with matrix template matching for TOA extraction, the calibration gives the lowest median weighted-rms residuals among the four dataset types tested, with the median dropping from $1.146$ to $1.078~\\mu\\mathrm{s}$ for standard FDM extraction and from $0.842$ to $0.818~\\mu\\mathrm{s}$ for MTM extraction.","pith_inferences":["The same archetype-plus-Gaussian-process recipe should transfer to the older BON backend data, which the paper flags as noisy and possibly poorly calibrated; success there would extend high-quality NRT timing back toward 2004.","The contrast between an earlier result where the simplest feed model won and the results here and at another telescope where model-based calibration won suggests the best calibration method is set by each telescope's reference-source and feed stability rather than by a universal rule; a portable comparison protocol could test this across observatories.","Because the procedure only requires a bright, frequently observed, polarization-stable pulsar and regular noise-diode measurements, other transit telescopes with narrow parallactic-angle coverage could adopt it directly.","Correcting the parallactic-angle sign changes the interpretation of the earlier apparent hour-angle dependence and may require revisiting published NRT position angles from analyses that used the uncorrected convention."],"forward_implications":["The 2011-2019 NUPPI archive can be calibrated to the same standard as post-2019 data, so pulsar timing array analyses no longer need to treat the earlier epoch as a separate, noisier regime.","Combining the new calibration with matrix template matching for TOA extraction gives the lowest median weighted-rms residuals and the lowest additional white noise among the four dataset combinations tested, so future NRT-based timing analyses should adopt both together.","Because the polarimetric response is independent of hour angle and declination, a single calibration solution per stable epoch is sufficient for normal-mode observations; no pointing-dependent correction is needed beyond the known variation of absolute gain with declination.","More homogeneous and higher-signal-to-noise polarimetric profiles should improve rotation-measure determinations and wide-band template matching on NRT data."],"supporting_citations":[{"why":"Prior MEM calibration from rotating-horn observations of the bright pulsar J0742$-$2822; the baseline this work extends and the source of the apparent hour-angle dependence.","marker":"Guillemot et al. 2023"},{"why":"Introduces the METM method that derives calibration solutions by matching observations to a well-calibrated reference profile.","marker":"van Straten 2013"},{"why":"Defines the measurement equation modeling formalism and the parameterization used for the rotating-horn analyses.","marker":"van Straten 2004"},{"why":"Introduces matrix template matching, which uses all four Stokes parameters when forming time-of-arrival estimates.","marker":"van Straten 2006"},{"why":"A comparison at another radio telescope showing that METM plus matrix template matching improves pulsar timing; motivates the same combination here.","marker":"Rogers et al. 2024"},{"why":"A comparison at another radio telescope where the simplest feed model outperformed MEM/METM; the contrasting result this work is set against.","marker":"Dey et al. 2024"},{"why":"Supplies the Gaussian-process machinery used to interpolate the scale and offset parameters in time.","marker":"Matthews et al. 2017"},{"why":"Flux-density measurements used to choose the reference pulsar for the METM analysis.","marker":"Jankowski et al. 2018"}],"fun_headline_variants":["Stable polarimetric response found at Nançay, plus improved pulsar calibration","New METM calibration cuts timing noise in pre-2019 Nançay pulsar data","Nançay pulsar timing improved: no sky-dependent polarization, new calibration","Recalibration of Nançay data yields sharper pulsar timing across a decade","Polarimetric response constant at Nançay; calibration boosts pulsar timing"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The procedure assumes that inside each manually chosen time segment the frequency shape of every calibration parameter stayed constant up to a per-epoch scaling and offset, and that the polarization profile of the reference pulsar J0953+0755 was intrinsically stable over the whole decade.","fun_headline_variants_meta":{"raw":{"variants":["Stable polarimetric response found at Nançay, plus improved pulsar calibration","New METM calibration cuts timing noise in pre-2019 Nançay pulsar data","Nançay pulsar timing improved: no sky-dependent polarization, new calibration","Recalibration of Nançay data yields sharper pulsar timing across a decade","Polarimetric response constant at Nançay; calibration boosts pulsar timing"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001502,"raw_usage":{"total_tokens":6170,"prompt_tokens":1233,"completion_tokens":4937,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":849,"completion_tokens_details":{"reasoning_tokens":4831}},"tokens_in":849,"tokens_out":4937,"duration_ms":33767,"temperature":1.0,"reasoning_tokens":4831,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T23:16:46.420386+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-derive the pre-2019 calibration without the archetype constraint, fitting each epoch's calibration parameters freely, and compare the resulting millisecond-pulsar timing residuals with the archetype-based ones; materially lower white or red noise in the free fits would show the archetype assumption is biasing the solutions. Independently, compare METM-predicted Stokes $Q$ and $U$ for J0953+0755 at an early epoch against a well-calibrated observation of a different bright pulsar from the same epoch; systematic growth of the residuals would falsify the assumed decade-long profile stability.","supporting_citations":[{"cited_title":"2023, , 678, A79","cited_arxiv_id":null,"evidence_quote":"Prior MEM calibration from rotating-horn observations of the bright pulsar J0742$-$2822; the baseline this work extends and the source of the apparent hour-angle dependence."},{"cited_title":"2013, , 204, 13","cited_arxiv_id":null,"evidence_quote":"Introduces the METM method that derives calibration solutions by matching observations to a well-calibrated reference profile."},{"cited_title":"2004, , 152, 129","cited_arxiv_id":null,"evidence_quote":"Defines the measurement equation modeling formalism and the parameterization used for the rotating-horn analyses."},{"cited_title":"F., van Straten , W., Gulyaev , S., et al","cited_arxiv_id":null,"evidence_quote":"A comparison at another radio telescope showing that METM plus matrix template matching improves pulsar timing; motivates the same combination here."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the Gaussian-process machinery used to interpolate the scale and offset parameters in time."},{"cited_title":"F., et al","cited_arxiv_id":null,"evidence_quote":"Flux-density measurements used to choose the reference pulsar for the METM analysis."}],"review_version":1}