{"id":"be6ddc55-bddb-482b-8b0c-be886ca593ab","arxiv_id":"2502.06946","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":2.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A decade of sub-arcsecond imaging with the International LOFAR Telescope is reviewed, covering calibration, scientific highlights, and future surveys.","lead":"This review explains how the International LOFAR Telescope makes very sharp radio images at low frequencies, covers the calibration tricks and science results from the past decade, and argues the telescope will stay unique even after the Square Kilometre Array is built. It is a useful reference for astronomers planning low-frequency surveys.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'unparalleled even in the SKAO era' claim depends on a not-yet-public comparison (Shimwell et al., submitted) and on unstated SKAO-Low assumptions; a reproducible sensitivity/confusion calculation is needed.","rationale":"I read the manuscript as a community review whose central scientific value is the demonstration that the ILT now enables sub-arcsecond, wide-field, low-frequency imaging, and whose headline claim is the uniqueness of that parameter combination. That claim is not a measurement reported here; it is a comparative judgment about future facilities. The review is honest about many of its own limitations: Section 2.1 details reduced calibrator S/N, ionospheric isoplanatic patch sizes, and LBA difficulties; Section 4 acknowledges the low-frequency calibration bottlenecks. These limitations do not undermine the main scientific content. The weakest link is the forward-looking comparison to SKAO-Low. The text itself flags the missing support by citing 'Shimwell et al. submitted' (a paper not yet available) for the decisive confusion-limit statement, and Figure 8's Tb comparison is qualitative. This is not an internal inconsistency, and it is not a disagreement with the community consensus that ILT is unique; it is a verification gap in a strong universal claim. The reader's conditional verdict correctly identifies this. My stress-test agrees: the condition should be a public, quantitative, reproducible comparison, or a tempered statement such as 'unparalleled in resolution among low-frequency instruments.' The rest of the review is unaffected. No formal verification exists, but none is expected for an instrument review; the detailed references to pipeline software and published images are appropriate independent support.","tokens_in":31128,"tokens_out":6815,"duration_ms":60418,"concrete_test":"Once Shimwell et al. (submitted) is public, reproduce its SKAO-Low vs ILT comparison: using the SKAO-Low system specifications (station effective area, Tsys, bandwidth, 8-h integration, max baseline ~65 km) compute the expected rms point-source sensitivity at 144 MHz, the synthesized beam solid angle, the corresponding Tb sensitivity, and the confusion noise from 144-MHz source counts. Compare these numbers with the ILT values underlying Figure 8 (0.3 arcsec beam, 8-h rms, Tb sensitivity). If SKAO-Low's Tb sensitivity is within a factor of ~2 of the ILT's, or if its confusion limit permits comparable faint-source surveys, the 'unparalleled' claim needs qualification. If the submitted paper remains unavailable, an equivalent calculation in the review, with equations and input parameters, would satisfy the condition.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's headline assertion is that the ILT's combination of sub-arcsecond resolution, ~6 deg^2 field of view, and <200 MHz observing band makes it unique, 'and it will remain unparalleled even in the era of the Square Kilometre Array Observatory' (Abstract). The load-bearing step is Section 3's claim that 'Sub-arcsecond resolution also easily beats the confusion limit compared to the SKAO-Low at the same frequencies,' and the Figure 8 implication that SKAO-Low cannot reach comparable Tb sensitivity. The only quantitative anchor offered is 'Shimwell et al. submitted' plus a general citation to Sabater et al. (2021), which is a LoTSS deep-field paper and does not itself present an SKAO-Low comparison. No numbers appear in this manuscript for SKAO-Low's point-source sensitivity, synthesized beam, Tb sensitivity, or confusion noise. The assertion may well be correct: SKAO-Low's longest baselines are far shorter than the ILT's ~2000 km, so it cannot match 0.3 arcsec resolution at 144 MHz. But the review does not demonstrate this with the specifications, and the universal 'No other current or planned radio instrument is capable of this' is exactly the kind of forward-looking claim that needs a reproducible calculation. If the submitted comparison is delayed, revised, or wrong, the central uniqueness claim is undercut. The rest of the review—calibration strategy, technical challenges, and science highlights—is well-supported and independent of this issue.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This is a review article describing the technical methods, calibration strategies, and scientific results enabled by sub-arcsecond imaging with the International LOFAR Telescope (ILT) over the past decade. It covers the challenges of mismatched station fields of view, time and bandwidth smearing, u-v coverage, calibration signal-to-noise, and ionospheric effects, then summarizes the calibration pipeline (LINC and the LOFAR-VLBI pipeline) and highlights science applications including multi-wavelength studies, brightness-temperature-based AGN identification, spectral modelling, rare-object discovery, strong lensing, and the dynamic range of spatial scales. The manuscript also discusses ongoing work on wide-area surveys (LoTSS-HR), WEAVE-LOFAR, wide-field imaging, polarisation, and low-frequency (LBA) observations. The central claim is that the ILT's combination of sub-arcsecond resolution, large field of view, and low observing frequency is unique and will remain unparalleled even in the SKAO era.","tokens_in":31395,"tokens_out":5404,"duration_ms":44976,"significance":"If the central uniqueness claim is accepted, this review provides a timely and useful community reference for the technical and scientific state of sub-arcsecond low-frequency radio astronomy. The manuscript is a well-organized synthesis of published work, with clear descriptions of the calibration pipeline, data products, and representative science cases. Its strengths include the documentation of publicly available pipelines (LINC, VLBI-cwl), the reproducibility of the described calibration strategy, and the use of peer-reviewed examples to illustrate each capability. The paper would be a valuable entry point for astronomers planning to use ILT high-resolution data. However, the forward-looking assertion that the ILT will remain 'unparalleled' in the SKAO era is the load-bearing claim and is currently supported by qualitative statements and a citation to an unpublished manuscript rather than by a quantitative comparison.","major_comments":[{"comment":"The central uniqueness claim is not quantitatively supported. The sentence 'Sub-arcsecond resolution also easily beats the confusion limit compared to the SKAO-Low at the same frequencies' (Section 3) is not accompanied by any numbers for SKAO-Low's synthesized beam, point-source sensitivity, brightness-temperature sensitivity, or confusion noise. The references given are 'Shimwell et al. submitted' (not yet public) and Sabater et al. (2021), which is a LoTSS deep-field paper that does not present an SKAO-Low comparison. Because the abstract states that the ILT 'will remain unparalleled even in the era of the Square Kilometre Array Observatory,' this comparison is load-bearing. The authors should either include a reproducible calculation based on SKAO-Low baseline specifications and expected sensitivity/confusion limits, or qualify the claim to state that a quantitative comparison is in preparation.","section":"Section 3 and Abstract"}],"minor_comments":[{"comment":"There is an inconsistency in the stated observing time for GOODS-N: the text says the GOODS-N survey used 17.5 hours, while the Figure 8 caption says both surveys use ~32 hours of data. Please correct one of these statements.","section":"Section 3.2 and Figure 8 caption"},{"comment":"The sentence 'the number of sources increases inversely with the limiting flux density, and hence is proportional to the inverse square root of the observing time' contains a sign error: for a Euclidean source count N(>S) ∝ S^{-1} and a noise level S ∝ t^{-1/2}, the number of sources grows as t^{1/2}, not t^{-1/2}.","section":"Section 3.4"},{"comment":"The claim that 'the area covered by ELAIS-N1 is 105 times larger than that of GOODS-N' should be verified. Using the stated GOODS-N central radius of 7.5 arcmin and ELAIS-N1 area of 6.6 deg^2, the area ratio is approximately 130, not 105.","section":"Section 3.2"},{"comment":"The statement that LoTSS-HR 'is set to form the highest-resolution sky survey with a comparably wide effective sky coverage by over an order of magnitude at any radio frequency' is another forward-looking claim that would benefit from quantitative support or a reference to a more detailed forecasting paper.","section":"Section 4"},{"comment":"The manuscript contains several typographical and formatting issues: 'LOF AR' should be 'LOFAR', 'WEA VE' should be 'WEAVE', 'ELIAS-N1' should be 'ELAIS-N1', and 'William Hershel Telescope' should be 'William Herschel Telescope'.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The review is written by a group heavily involved in developing the LOFAR-VLBI pipeline and in many of the cited scientific results, which is natural for a review but results in extensive self-citation. More importantly, the central forward-looking uniqueness claim rests on an unpublished comparison ('Shimwell et al. submitted'), which is a weakness for a published review. The authors should be encouraged to either add a quantitative comparison or soften the claim. The paper otherwise fits the scope of astro-ph.IM and would be a useful reference."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a review, not a new result. It does a good job of collecting the technical and scientific state of the art for ILT sub-arcsecond imaging, and the illustrative simulations are fine. The soft spot is the forward-looking claim about SKAO-Low. I largely agree with the reader's take.\n\nWhat is actually good: the paper clearly explains the technical constraints—matched primary beams, time/bandwidth smearing, uv coverage, calibration SNR, ionosphere, LBA difficulties. These are the real barriers to exploiting this mode, and having them in one place is genuinely useful. The science examples are well chosen: matched-resolution multi-wavelength studies, Tb-based AGN identification, spectral modelling anchors, rare objects, and spatial-scale range. The comparison of ELAIS-N1 vs GOODS-N (Figure 8) is a nice quantitative demonstration of field-of-view gain. The Figure 2 smearing curves are simple but directly relevant.\n\nThe weak spot is also the headline. The abstract says ILT 'will remain unparalleled even in the era of the Square Kilometre Array Observatory.' Section 3 asserts that sub-arcsecond resolution 'easily beats the confusion limit compared to the SKAO-Low' and that 'No other current or planned radio instrument is capable of this.' These are the kind of claims that need a reproducible calculation: what is SKAO-Low's expected synthesized beam, Tb sensitivity, and confusion noise at 144 MHz? The paper gives no numbers. It cites Shimwell et al. (submitted) and Sabater et al. (2021), but Sabater is a LoTSS deep-field paper, not an SKAO comparison. The claim may well be true—ILT's ~2000 km baselines do give resolution that SKAO-Low will not match—but the review doesn't demonstrate it. This is a forward-looking statement that will age; if the comparison is delayed or revised, the central claim is undercut.\n\nThere is also a minor grammar issue in the abstract ('calibration methods sub-arcsecond imaging'), and the circularity burden is real but not disqualifying: the authors developed much of the pipeline, so self-citation is expected and the descriptions match the cited literature. I would not call this a flaw in itself.\n\nWho is this for? Anyone planning or interpreting low-frequency high-resolution observations, and students wanting to understand the calibration landscape. It deserves peer review as a reference article, but I would ask the authors to either temper the 'unparalleled' phrasing or supply a quantitative SKAO-Low comparison. The rest of the paper is solid and independent of that claim.","headline":"A competent, useful review of LOFAR sub-arcsecond imaging whose headline 'unparalleled in the SKAO era' claim is not yet quantitatively supported.","tokens_in":31950,"tokens_out":2067,"would_cite":true,"duration_ms":18961,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The International LOFAR Telescope can image the low-frequency sky at 0.3-arcsecond resolution over wide fields, and this review argues that this combination will remain unmatched even after SKAO begins operations.","keywords":["radio astronomy","extragalactic","high-resolution imaging","radio surveys","International LOFAR Telescope","sub-arcsecond imaging","active galactic nuclei","brightness temperature"],"falsifier":"A quantitative end-to-end simulation of SKAO-Low's expected resolution, confusion limit, and brightness-temperature sensitivity over a wide field at 144 MHz would settle the claim: if SKAO-Low can identify AGN via $T_b \\gtrsim 10^5$ K across a comparable area, the paper's central uniqueness claim is undercut.","tokens_in":2054,"feed_emoji":"📡","tokens_out":2219,"duration_ms":95674,"temperature":0.7,"pith_summary":"The paper is a status report and argument: after a decade of development, sub-arcsecond imaging with the International LOFAR Telescope (ILT) has moved from expert-only custom processing to a public, pipeline-driven capability. Its central claim is that the ILT is unique — 0.3-arcsecond resolution at 144 MHz over a roughly 6.6-square-degree field — and that no planned instrument, including SKAO-Low, will match this combination of resolution, field of view, and low frequency. That matters because the combination turns sub-arcsecond imaging from a tool for studying individual objects into a blind survey instrument: brightness-temperature measurements from these images can identify active galactic nuclei in large statistical samples, and the low frequencies anchor spectral-age modelling. A sympathetic reader should come away understanding that wide-area, low-frequency, sub-arcsecond radio science is now a practical enterprise, with LoTSS-HR and related surveys about to make it fully public.","feed_headline":"Sub-arcsecond LOFAR imaging stays unique","feed_subtitle":"A decade of calibration makes 0.3-arcsec, 144-MHz imaging a wide-area tool for AGN science.","key_machinery":"The load-bearing object is the International LOFAR Telescope itself: baselines up to roughly 2,000 km across Europe give 0.3-arcsecond resolution at 144 MHz, while the large primary beams of the stations keep a field of view of about 6.6 square degrees; an 8-hour observation fills the uv plane densely enough that the array behaves more like high-resolution interferometry than classical VLBI. On top of this sits a calibration chain — the LOFAR long-baseline calibrator survey, the LINC and LOFAR-VLBI pipelines, and facet-based direction-dependent calibration — that makes the data tractable. The scientific quantity that carries the uniqueness argument is brightness temperature, $T_b \\propto \\nu^{-1}\\Omega^{-1}$, the flux density per unit solid angle expressed as an equivalent temperature: at 144 MHz with a 0.3-arcsecond beam, ILT images reach the roughly $10^5$ K threshold where compact emission cannot be explained by star formation and must be AGN, turning sub-arcsecond imaging into a blind statistical galaxy survey tool.","core_discovery":"On the paper's own terms, the central claim is that the ILT's combination of sub-arcsecond resolution, large field of view, and low observing frequency is unique among current and planned facilities, and that it will remain unparalleled even in the era of the Square Kilometre Array Observatory. The evidence assembled is technical and observational: an 8-hour observation fills the uv plane densely out to about 1 Mλ; a sky survey of calibrators provides roughly one usable calibrator per square degree; and wide-field imaging has been demonstrated on the Lockman Hole and ELAIS-N1 fields at 0.3, 0.6, and 1.2 arcseconds. The scientific payoff is that high-resolution images reach brightness-temperature sensitivities around $10^5$ K, the threshold above which radio emission cannot be explained by star formation alone and must be AGN activity. The paper shows this capability in action — de-blending sources, matching HST and Chandra resolution, anchoring spectral-age fits, discovering a giant jet in a $z=4.9$ quasar, resolving strong lenses, and recovering emission from 0.3 arcsec to 80 arcsec in a single observation.","pith_inferences":["The paper leaves implicit that the strongest test of its uniqueness claim is a quantitative SKAO-Low simulation rather than another survey demonstration.","If LOFAR2.0's wider bandwidth and sidereal visibility averaging deliver the intended depth, ultra-deep sub-arcsecond surveys could push brightness-temperature AGN identification to fainter and more numerous objects than the current ELAIS-N1 result.","The polarisation work suggests that once full-Jones calibration is public and residual instrumental polarisation is quantified, sub-arcsecond RM-synthesis maps could become a standard product for LOFAR deep fields.","The success of LoTSS-HR will depend on whether roughly 20 to 30 automatic calibrators per wide field can be found reliably in all directions, not only in the fields demonstrated so far."],"forward_implications":["LoTSS-HR will reprocess the LoTSS archive with the international stations, publishing 0.3-arcsecond cutouts of all sources brighter than 10 mJy plus full-field 1.2-arcsecond images, making it the widest high-resolution radio survey at any frequency by an order of magnitude.","Brightness-temperature AGN identification becomes a statistical, wide-area method: the ELAIS-N1 demonstration covers 105 times more sky than the EVN GOODS-N survey and yields 51 times more AGN identifications, enabling the first radio luminosity functions split by physical process.","The combination with WEAVE-LOFAR will give every LoTSS-HR source a spectroscopic redshift, opening redshift-resolved studies of AGN morphology and host galaxy properties.","Matched-resolution multi-wavelength and spectral-age studies become possible for distant objects: the 54 and 144 MHz ILT data anchor the injection index in spectral-age fits for sources such as 4C 43.15 at $z=2.4$.","A single ILT observation can be imaged from roughly 0.3 arcsec to 80 arcsec, simultaneously recovering compact AGN cores and Mpc-scale diffuse emission, as demonstrated for the Perseus cluster."],"supporting_citations":[{"why":"Defines the foundational calibration strategy and public pipeline that made ILT sub-arcsecond imaging accessible.","marker":"Morabito et al. (2022)"},{"why":"Completes the Long Baseline Calibrator Survey, establishing the roughly one-per-square-degree calibrator sky density the strategy relies on.","marker":"Jackson et al. (2022)"},{"why":"Provides the first full-field-of-view sub-arcsecond ILT image and the resolved spectral-age modelling of 4C 43.15 used to demonstrate the low-frequency anchor.","marker":"Sweijen et al. (2022)"},{"why":"Delivers the deepest wide-field sub-arcsecond images of ELAIS-N1 and the brightness-temperature AGN comparison against GOODS-N.","marker":"de Jong et al. (2024)"},{"why":"Supplies the 6-arcsecond ELAIS-N1 LOFAR deep-field images that the high-resolution work builds on and is compared with.","marker":"Sabater et al. (2021)"},{"why":"Cited for the confusion-limit comparison showing that sub-arcsecond resolution beats SKAO-Low at the same frequencies.","marker":"Shimwell et al. (2025)"},{"why":"Establishes the physical surface-brightness limit of star formation that makes brightness temperature an unambiguous AGN diagnostic.","marker":"Condon (1992)"},{"why":"The EVN GOODS-N survey used as the comparison benchmark for area and the number of brightness-temperature-identified AGN.","marker":"Radcliffe et al. (2018)"},{"why":"The LOFAR telescope description supplying the stations, baselines, and bands that define the ILT.","marker":"van Haarlem et al. (2013)"}],"fun_headline_variants":["LOFAR sub-arcsecond imaging remains unparalleled","A decade of sub-arcsecond LOFAR views still unique","ILT's sub-arcsecond radio vision: unmatched, even with SKA","Sub-arcsecond LOFAR: a unique window for AGN science","LOFAR's sharp radio imaging stays ahead after a decade"],"cache_read_input_tokens":34048,"weakest_assumption_plain":"The claim that the ILT will remain unmatched depends on SKAO-Low not achieving comparable resolution and brightness-temperature sensitivity at low frequencies, something the paper supports mainly by qualitative comparison rather than a full quantitative analysis.","fun_headline_variants_meta":{"raw":{"variants":["LOFAR sub-arcsecond imaging remains unparalleled","A decade of sub-arcsecond LOFAR views still unique","ILT's sub-arcsecond radio vision: unmatched, even with SKA","Sub-arcsecond LOFAR: a unique window for AGN science","LOFAR's sharp radio imaging stays ahead after a decade"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000214,"raw_usage":{"total_tokens":1477,"prompt_tokens":1049,"completion_tokens":428,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":665,"completion_tokens_details":{"reasoning_tokens":336}},"tokens_in":665,"tokens_out":428,"duration_ms":4090,"temperature":1.0,"reasoning_tokens":336,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T14:14:39.535633+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A quantitative end-to-end simulation of SKAO-Low's expected resolution, confusion limit, and brightness-temperature sensitivity over a wide field at 144 MHz would settle the claim: if SKAO-Low can identify AGN via $T_b \\gtrsim 10^5$ K across a comparable area, the paper's central uniqueness claim is undercut.","supporting_citations":[],"review_version":1}