{"id":"b4faed0e-4764-4138-9c9c-ac0d80011622","arxiv_id":"2508.07998","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":4.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"A thesis exploring machine learning and Virtual Observatory methods for discovering and characterizing M dwarfs and ultracool dwarfs across large astronomical surveys.","lead":"Pedro Mas-Buitrago's thesis applies machine learning and Virtual Observatory database tools to discover and characterize M dwarfs and ultracool dwarfs, the smallest and coolest low-mass objects. The goal is to build scalable pipelines for finding and classifying these objects in current and upcoming surveys such as Euclid and LSST.","discovery_kind":"unclear","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Cross-survey generalization of the ML classifiers is asserted but not demonstrated in the accessible text.","rationale":"The reader identified exactly the same weakest assumption: ML models trained on existing spectroscopic labels must generalize to new surveys and to the full M-L-T-Y sequence. My concern is a concrete specification of that assumption, and the test would settle it. I do not recommend changing the reader's UNVERDICTED verdict because the body text is not reliably readable; the concern is real but cannot be confirmed as a fatal flaw without the full manuscript. If the clean text contains the cross-survey validation I propose, the concern would be resolved. Therefore the verdict stays UNCHANGED.","tokens_in":59752,"tokens_out":2145,"duration_ms":26251,"concrete_test":"Locate the clean PDF/source and run a transfer test: apply the trained classifiers to held-out simulated Euclid/LSST photometry generated from model spectra (e.g., BT-Settl or Sonora) for spectral types M0 through Y0 at representative distances and reddening values. Report per-class accuracy and completeness against input labels, and compare to the in-domain test performance from the thesis. Also run a cross-survey test: train on SDSS/2MASS data and test on Pan-STARRS1 or Gaia XP spectra. If late-L/T/Y accuracy is below 50% or drops by more than 10 percentage points relative to the training-domain test set, the generalization claim in the abstract is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that VO-enabled machine learning pipelines can support discovery and characterization of low-mass objects at survey scale, including in Euclid and LSST. This requires classifiers and regressors trained on existing photometric/spectroscopic samples to transfer to surveys with different filter systems, depths, resolutions, and coverage, and to perform across the full M-L-T-Y sequence. The abstract promises this, but the provided body text is almost entirely garbled, so no cross-survey validation, confusion matrices, spectral-type accuracies, or completeness estimates are visible. If the training labels are dominated by relatively bright M dwarfs—as is typical for current spectroscopic samples—then the rare late-L, T, and Y dwarfs that Euclid/LSST are expected to discover may be outside the training distribution. A pipeline that runs end-to-end could still return unreliable labels for these objects. This is not an internal inconsistency; it is an unverified domain-shift assumption. The claim would only be load-bearing if the thesis actually demonstrates transfer performance, which the accessible text does not show.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript (an apparent thesis abstract plus an unreadable body) claims to explore the discovery and characterisation of M dwarfs and ultracool dwarfs by combining Virtual Observatory technologies and protocols with machine and deep learning techniques, with an eye toward new-generation surveys such as Euclid and LSST. The abstract identifies real astrophysical challenges (faintness, strong molecular absorption, survey data volume) and promises data-driven, scalable methodologies. However, the full text is almost entirely corrupted by character-encoding errors: all body sections appear as garbled text, with no accessible derivations, datasets, experimental setups, figures, tables, or results. The only evaluable content is the abstract and the table of contents headings.","tokens_in":59959,"tokens_out":3166,"duration_ms":36459,"significance":"If the promised pipelines were actually demonstrated, this work could be relevant to survey-scale exploitation of low-mass-object data, particularly for pre-selection and characterisation of M and ultracool dwarfs in Euclid/LSST. The abstract correctly identifies the main difficulties and points to a plausible methodological direction. I credit the manuscript for being explicit about its data-driven approach, which is in principle falsifiable. However, as submitted, the paper contains no checkable evidence: no validation metrics, no comparison with existing surveys, no description of training data, and no reproducible code or catalogues. The significance of the claimed contribution therefore cannot be assessed.","major_comments":[{"comment":"The body text is unreadable due to pervasive character corruption from the front matter through the appendices. Equations, tables, and figure captions are similarly garbled. This is not a presentation nit: it prevents any verification of the methods, data, or results. The authors must resubmit a clean, machine-readable PDF (e.g., proper LaTeX or OCR with correct encoding) before the paper can be reviewed. As it stands, only the abstract and the list of section headings are accessible.","section":"Full text (all body sections)"},{"comment":"No validation metrics or quantitative results appear anywhere in the accessible text. The central claim—that the developed ML/DL methodologies achieve reliable discovery and characterisation of M and ultracool dwarfs—requires concrete evidence: classification accuracy, spectral-type scatter, completeness and contamination, or a comparison against established catalogues (e.g., 2MASS, WISE, Gaia, SDSS). Without such numbers, the claim of survey-scale capability is unsupported.","section":"Abstract; TOC sections on methods/results"},{"comment":"The abstract asserts that the methods will enable a 'quantum leap' with Euclid and LSST and will advance understanding across the M, L, T, Y sequence. This implicitly assumes that classifiers/regressors trained on existing photometric/spectroscopic samples will transfer to surveys with different filter systems, depths, and resolutions. No transfer-learning experiment, domain-shift analysis, or evaluation on the rare late-L/T/Y population is provided. Given that current spectroscopic training samples are dominated by relatively bright M dwarfs, this unverified generalization is load-bearing for the promised survey-scale impact.","section":"Abstract (Euclid/LSST outlook)"}],"minor_comments":[{"comment":"The phrase 'quantum leap' is rhetorical; if the authors intend a genuine step-change, it should be quantified (e.g., expected yield, depth, or sample size compared with previous surveys).","section":"Abstract"},{"comment":"Several section headings are visible but garbled, and the author list and abstract formatting are incomplete. A clean PDF with intact metadata is needed for bibliographic and review purposes.","section":"Table of contents / front matter"},{"comment":"The manuscript does not state whether the underlying data, trained models, or code are publicly available. Given the reproducibility standards in astroinformatics, a data/code availability statement should be added.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The uploaded file appears severely corrupted at the text-extraction level; I was unable to read any body section. I recommend returning the file to the authors to produce a clean PDF before assigning further review. If the clean version still consists only of a thesis abstract with no original results, the suitability for a research journal should be reconsidered."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"About arXiv:2508.07998: the useful thing to know is that this is a PhD thesis, not a typical results paper. The abstract frames an application program — ML and VO methods for M/ultracool dwarf discovery and classification — without claiming a new physical result or benchmark-beating algorithm. That framing is honest. What it does well: the topic is real; faint cool objects with molecular bands are a genuine classification bottleneck, and the thesis correctly points at Euclid/LSST as the next data frontier. The TOC suggests substantial practical chapters on photometric/spectroscopic classification and on VO infrastructure rather than a pure survey of the literature.\n\nThe problem is that I can't check the body. The provided full text is heavily garbled and truncated, so there are no validation metrics, confusion matrices, completeness values, or comparisons to existing surveys. The stress-test concern is fair as far as it goes: the abstract promises cross-survey transfer to Euclid/LSST but doesn't show it, and with training labels mostly from brighter M dwarfs, late-L/T/Y behavior is a real domain-shift question. But this is absence of evidence in what I was given, not evidence of a flaw. The chapter headings indicate the author does attempt applications; if those chapters contain proper held-out validation and completeness analysis, the thesis delivers what it promises. I can't tell from here.\n\nOne point where I'd adjust the reader's take: I wouldn't mark soundness low just because the abstract lacks metrics. Abstracts normally do. The fair score is 'unverifiable from accessible text', which is what the reader actually concluded. No circularity is visible; no invented entities; the citation pattern can't be assessed from the garble.\n\nRecommendation: if the full text is available, send it to someone working on ultracool dwarf selection and ask whether the validation matches the claims. That referee would get real value from the application chapters, and the central claim — that ML can transfer to new surveys — deserves a genuine check. As submitted to arXiv, the abstract alone is too thin to cite for a result, but the project is worth engaging with.","headline":"A well-framed thesis abstract on ML/VO for M and ultracool dwarfs, but with no checkable result in the accessible text, so it reads as a research program rather than a demonstrated pipeline.","tokens_in":60378,"tokens_out":4202,"would_cite":false,"duration_ms":44338,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This thesis sets out to show that data-driven classifiers operating on Virtual Observatory data can discover and characterise M dwarfs and ultracool dwarfs at the scale of next-generation surveys.","keywords":["M dwarfs","ultracool dwarfs","Virtual Observatory","machine learning","deep learning","photometric surveys","spectroscopic surveys","spectral classification"],"falsifier":"Run a trained classifier on Euclid or LSST photometry in a field with independent follow-up spectroscopy, then compare predicted spectral types against the follow-up types for a complete, magnitude-limited subsample. If accuracy collapses for faint or late-type objects, or if the predictions are systematically offset from the independent types, the claim of survey-scale reliability is refuted.","tokens_in":59662,"feed_emoji":"🔭","tokens_out":7254,"duration_ms":70322,"temperature":0.7,"pith_summary":"The thesis aims to establish a practical claim: M dwarfs and ultracool dwarfs—the Galaxy's most numerous but faintest constituents—can be discovered and characterised by machine-learning pipelines running on Virtual Observatory data, rather than by case-by-case spectroscopic analysis. This matters because M dwarfs make up roughly 75% of the stars within 10 parsecs of the Sun and are prime targets for Earth-like planet searches, yet their strong molecular absorption bands and faintness make them hard to classify. The author's route is data-driven: assemble labelled samples from spectroscopic surveys, use photometric and low-resolution spectral features as inputs, and train machine and deep learning models flexible enough to absorb the coming flood of data from Euclid and LSST. If the claim holds, the payoff is a much more complete census of the lowest-mass stars and substellar objects, and a clearer view of the boundary between them.","feed_headline":"Machine learning scales up discovery of the Galaxy's faintest objects","feed_subtitle":"Low-mass stars are faint and hard to classify; data-driven tools promise to find them across the new mega-surveys.","key_machinery":"The load-bearing mechanism is the combination of Virtual Observatory interoperability with supervised machine and deep learning. VO protocols allow heterogeneous archives to be queried and cross-matched as one system, effectively turning many surveys into a single training resource; the classifiers map photometric colours and low-resolution spectral features onto spectral-type labels. The mechanism does the work of replacing one-by-one manual classification with automated, scalable estimates that remain tied to the existing spectroscopic taxonomy of M, L, T, and Y types.","core_discovery":"On the paper's own terms, the central claim is methodological rather than a single numerical result: the discovery and characterisation of low-mass objects can be automated at survey scale using Virtual Observatory interoperability plus supervised machine and deep learning. The concrete content is a programme of building flexible classifiers from existing spectroscopic labels and applying them to photometric and spectroscopic survey data, with the aim of covering the whole M-to-Y sequence. A sympathetic reading is that the author is establishing the viability of this mode of work—that standard VO protocols plus modern machine learning are not just conveniences but essential, given the volume","pith_inferences":["The same recipe should be testable in reverse before Euclid and LSST launch: train on synthetic photometry for those surveys and validate on existing ground-based spectroscopic samples, giving an early and cheap check of the generalisation claim.","The argument implicitly makes label quality a first-order scientific variable; quantifying spectral-type label noise in the training sets, and modelling it in the loss function, would be a natural extension the abstract does not develop.","If the approach succeeds for low-mass objects, it becomes a template for finding rare, spectrally distinctive populations in any large survey, so the payoff is likely to extend beyond M, L, T, and Y dwarfs.","To turn classifications into a substellar mass function, the classifier's output probabilities need to be treated as calibrated measurements; the thesis's scientific yield therefore depends on uncertainty estimation as much as on raw accuracy."],"forward_implications":["Survey-scale catalogues of M dwarfs and ultracool dwarfs become feasible, substantially completing the local census of stars and substellar objects.","Target lists for exoplanet searches around M dwarfs can be generated automatically from the same photometric data, widening the pool of promising host stars.","Cross-matching through VO protocols turns dispersed archives into one large training set, so a new survey can be absorbed by retraining rather than by manual reclassification.","Spectral-type estimates from the pipelines can be updated as Euclid and LSST data arrive, keeping the census current with the surveys rather than lagging years behind them."],"supporting_citations":[],"fun_headline_variants":["Virtual Observatory and ML find the Galaxy's faintest objects","Machine learning automates the hunt for ultracool dwarfs","Data-driven discovery scales for low-mass stars and brown dwarfs","AI and VO open a new era for finding the dimmest stars","Machine learning unlocks survey-scale discovery of cool dwarfs"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The load-bearing premise is that models trained on existing spectroscopically labelled samples will generalise to new surveys such as Euclid and LSST and hold up across the full M-to-Y sequence; if those labels are biased, incomplete, or unrepresentative, the pipeline will produce confident but unreliable classifications.","fun_headline_variants_meta":{"raw":{"variants":["Virtual Observatory and ML find the Galaxy's faintest objects","Machine learning automates the hunt for ultracool dwarfs","Data-driven discovery scales for low-mass stars and brown dwarfs","AI and VO open a new era for finding the dimmest stars","Machine learning unlocks survey-scale discovery of cool dwarfs"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000208,"raw_usage":{"total_tokens":1282,"prompt_tokens":824,"completion_tokens":458,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":568,"completion_tokens_details":{"reasoning_tokens":373}},"tokens_in":568,"tokens_out":458,"duration_ms":5223,"temperature":1.0,"reasoning_tokens":373,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T22:12:44.982045+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a trained classifier on Euclid or LSST photometry in a field with independent follow-up spectroscopy, then compare predicted spectral types against the follow-up types for a complete, magnitude-limited subsample. If accuracy collapses for faint or late-type objects, or if the predictions are systematically offset from the independent types, the claim of survey-scale reliability is refuted.","supporting_citations":[],"review_version":1}