{"id":"464d1afb-5cda-4707-aff3-5920356b1e42","arxiv_id":"2501.15491","paper_version":1,"verdict":"UNVERDICTED","confidence":"HIGH","novelty_score":2.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"A field guide from the Knowing Machines project that turns existing critical dataset studies scholarship into lifecycle questions for practitioners, with no new empirical or formal results.","lead":"This paper is an educational field guide, not a research result. It compiles advice from critical data studies into practical questions for anyone who builds, uses, or transforms machine learning datasets, and it explains the concepts behind dataset work in plain language.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Central claim of improved dataset practice rests solely on reflection; no evidence that the Section 6 questions change decisions or outcomes.","rationale":"The reader's weakest_assumption identifies the same load-bearing concern: the guide promises that reflection alone will make practitioners avoid dataset problems and build more reliable systems, but it provides no evaluation of that promise. I agree that this is the central weakness. The abstract and Section 2 make causal claims, not merely descriptive ones, so the absence of evidence is a genuine gap rather than a category error. However, the appropriate verdict remains UNCHANGED. This document is explicitly a field guide and literature synthesis, not a research preprint reporting new measurements; under the UNVERDICTED rule for non-research documents, the lack of efficacy evidence supports UNVERDICTED rather than ACCEPT or REJECT. The guide's own hedges—'We are not lawyers and this is not legal advice' (Section 6.1), 'this guide cannot fully address here' (Section 7.3), and 'a starting point' (Section 8)—show epistemic care but do not substitute for evidence. The section-number swap in Section 1.1 and the uncontested COMPAS framing are real but secondary; they do not threaten the guide's overall argument as much as the unsupported efficacy premise. No internal inconsistency or fatal flaw emerged from the full-text review, and the guide's practical suggestions remain coherent and well-sourced even if their effectiveness is untested.","tokens_in":23542,"tokens_out":2966,"duration_ms":28647,"concrete_test":"Run a preregistered between-subjects study with dataset practitioners (e.g., 60–100 graduate students or industry ML engineers). Both groups receive the same realistic dataset-selection task with known pitfalls (missing consent, deprecated source, license mismatch). The treatment group gets the Section 6 questions; the control group gets a placebo checklist (e.g., formatting or style reminders) or no intervention. Blind expert raters score the resulting project plans on provenance review, consent and licensing checks, deprecation handling, and documentation. If the treatment group does not outperform the control on these pre-specified outcome measures, the central efficacy claim fails. A smaller think-aloud follow-up can reveal whether participants actually apply the questions or merely acknowledge them.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The guide's central value proposition is causal: a practitioner who works through the lifecycle questions in Section 6 will be 'more capable of avoiding the problems unique to datasets' and 'construct more reliable, robust solutions' (Abstract; Sections 1.1 and 2). The only mechanism supplied is reflection—asking questions about origins, usage, and stewardship. The document presents no user study, no behavioral outcome measure, and no comparison against alternative interventions; Section 8 explicitly hedges by calling itself 'a starting point.' That hedge mitigates but does not remove the overclaim, because the promises in the Abstract and Section 2 are unqualified. The concern is not that the advice is wrong; it is that the claimed effect is unsupported. For the central claim to hold, reading and working through the checklist would have to materially improve real dataset decisions—especially under the time, cost, and organizational pressures Section 1.1 acknowledges. The guide offers no evidence for that mechanism, and the checklist format alone does not guarantee adoption. This is the load-bearing weakness: if reflection alone does not change behavior, the guide is still a useful reference but its headline promise fails.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript is a field guide to working critically with machine learning datasets. It introduces key concepts (data, datasets, parts of datasets, types, transformations) and offers a lifecycle framework with questions about origins, usage, and stewardship. It draws on critical data studies, documents common pitfalls, and provides references to tools and practices such as datasheets, deprecation frameworks, and FAIR/CARE principles. The guide is written for a broad audience including researchers, journalists, artists, and developers.","tokens_in":23497,"tokens_out":4264,"duration_ms":38812,"significance":"If adopted, the guide could serve as a useful pedagogical and reference resource that bridges critical AI scholarship and practical data work. Its strengths are the breadth of synthesized literature, the concrete questions and checklists, the inclusion of case studies, and the clear distinction between technical and sociotechnical dimensions of dataset care. However, the paper makes strong efficacy claims about improved outcomes without empirical support, and some benefit claims (e.g., reduced legal liability) are overstated. With appropriate qualifications, the guide would be a valuable contribution to the critical data studies literature.","major_comments":[{"comment":"The guide repeatedly claims that working through its recommendations will make practitioners 'more capable of avoiding the problems unique to datasets' and 'construct more reliable, robust solutions.' The only mechanism proposed is reflection on a set of questions; no user study, behavioral outcome measure, or comparison against alternative interventions is presented. The conclusion (Section 8) hedges by calling the guide 'a starting point,' but the earlier statements are unqualified. This is a load-bearing issue because the guide's value proposition rests on the assumption that reflection changes dataset practice. Please either soften these claims to indicate potential benefit, or add an explicit discussion of the evidence status and limitations, noting that the efficacy of such checklists remains an open empirical question.","section":"Abstract; Sections 1.1, 2, 6; Section 8"},{"comment":"The claim that proactive attention to legal and ethical concerns yields 'INCREASED PROTECTION FROM LIABILITY' is stated as a benefit. Although the text disclaims that it is not legal advice, this is a factual/legal assertion that is not substantiated with legal analysis or citations to specific remedies. At a minimum, phrase this as 'may reduce legal and ethical risk' and advise readers to consult counsel, rather than implying guaranteed protection.","section":"Section 2"}],"minor_comments":[{"comment":"The overview states that readers will find TYPES of datasets in Section 5 and TRANSFORM in Section 4, but the Table of Contents and section headers place Types in Section 4 and Transforming in Section 5. Please correct the cross-references.","section":"Section 1.1"},{"comment":"The invented term 'Data subjectees' is defined, but it could be confused with 'data subjects' throughout the guide. Since the authors already cite the direct/indirect stakeholder distinction from Friedman and Hendry, consider using that established vocabulary to improve clarity.","section":"Section 3"},{"comment":"References [57] and [53] are duplicates of the same CARE Principles citation; please consolidate.","section":"References"},{"comment":"Long sequences of block characters appear to be intentional design elements from the spreadsheet origin of the guide. In a text-only version these render as garbled characters and may be inaccessible; consider replacing them with descriptive text or an accessible version.","section":"Section 7.3"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a field guide rather than a conventional research article; the editors should consider whether the journal's scope includes such artifacts. The authors and editors are affiliated with the Knowing Machines project, which is also the publisher and the source of several cited references (e.g., [22], [50], [82]); this is disclosed but worth noting in terms of potential self-citation bias."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First, what you should know: this is a field guide, not a research paper. It consolidates existing critical-dataset scholarship into a practical, well-organized resource. The core value is pedagogical; the novelty is low, and the authors mostly don't claim otherwise. The guide does what it does well: it clearly explains dataset concepts, offers thoughtfully designed lifecycle questions, credits its sources, and repeatedly acknowledges its own limits (\"we are not lawyers\", \"not a definitive source\"). The Section 6 questions are a genuinely useful checklist for teaching and for practitioners.\n\nThe main soft spot is the efficacy claim. The abstract and Section 2 promise that working through the guide will make readers \"more capable of avoiding the problems unique to datasets.\" That is an empirical claim, and there is no evidence—no user study, no behavioral measure—that reflection alone changes decisions. The stress-test note is right about that. But I don't think it sinks the document, because the guide is better read as an aspirational framing, and Section 8 explicitly hedges. Still, the authors should temper the promise or say plainly that the effectiveness is untested.\n\nA second issue: the COMPAS example (Section 1.1) is presented without acknowledging the substantial rebuttal literature—the ProPublica analysis has been criticized on technical and statistical grounds. For a guide emphasizing critical thinking, that omission undercuts its own message.\n\nMinor: there's a section-numbering error in the Section 1.1 overview (Parts, Types, and Transforming are listed out of order), and the new term \"data subjectees\" is defined but may confuse readers. These are easily fixed.\n\nOverall, this is a competent and honest educational synthesis. It's not a research result, and it shouldn't be evaluated as one. Its value is for teachers, students, journalists, and practitioners who want a starting point. I'd be happy to see it in a teaching syllabus or as a companion resource. If it were submitted as a peer-reviewed article, it would deserve review, but the authors should reframe it as a pedagogical contribution and address the efficacy claim and the COMPAS discussion.","headline":"A well-made teaching synthesis, not a research contribution; the efficacy promise is untested and should be read as aspiration.","tokens_in":24277,"tokens_out":2388,"would_cite":true,"duration_ms":22969,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that conscientious dataset stewardship—working through lifecycle questions on origins, usage, and stewardship—can help practitioners avoid technical, legal, and ethical harms and build more reliable machine learning…","keywords":["machine learning datasets","dataset lifecycle","data stewardship","dataset bias","consent","dataset deprecation","critical data studies","datasheets for datasets"],"falsifier":"A controlled study in which two groups of practitioners—one using the lifecycle questions, one not—build or select datasets for comparable tasks, with independent audit of resulting harms, errors, and legal or ethical issues, would settle the claim; if the question-using group shows no measurable improvement, the guide's central value proposition fails.","tokens_in":23106,"feed_emoji":"🧭","tokens_out":4728,"duration_ms":42553,"temperature":0.7,"pith_summary":"The paper argues that machine learning datasets are powerful but unwieldy resources whose problems—technical, legal, and ethical—can be managed through conscientious stewardship. It offers a practical field guide giving questions, suggestions, and strategies for every phase of a dataset's life. The central promise is that practitioners who ask the guide's lifecycle questions will be more capable of avoiding dataset-specific harms and constructing more reliable systems. A sympathetic reader would care because the guide translates critical AI scholarship into accessible, usable practices for students, journalists, artists, researchers, and developers.","feed_headline":"Ask three lifecycle question sets to avoid ML dataset harms","feed_subtitle":"A practical field guide walks practitioners through origins, usage, and stewardship of datasets from bias to consent to legal liability.","key_machinery":"The central mechanism is the Dataset Lifecycle framework, a set of critical questions organized into three stages: Origins (what is the dataset's story, who created and consented to it), Usage (what story will you tell with it), and Stewardship (what story will it keep telling after you). Each stage prompts reflection on provenance, consent, annotation, missing data, transformation, licensing, harm mitigation, documentation, and deprecation. The framework carries the argument by turning the abstract claim that 'datasets are not neutral' into a repeatable questioning practice.","core_discovery":"The guide's central claim is that no dataset is neutral or ready to use off the shelf: datasets are contingent on how they are made, who made them, and the settings in which they circulate, and they remain tied to the people they represent and affect. Working through the lifecycle framework—Origins, Usage, and Stewardship—makes these entanglements visible and gives practitioners concrete questions to ask before, during, and after a project. The paper asserts that such critical care yields more robust datasets, more reliable results, greater protection from legal and ethical liability, and more conscientious outcomes for those impacted.","pith_inferences":["The lifecycle questions could be turned into a routinized checklist embedded in dataset repositories or model registries, making the critical reflection the guide advocates a structural part of dataset access.","A testable extension would be to measure whether datasets accompanied by completed datasheets and lifecycle documentation are reused less frequently in inappropriate contexts, or attract more caution from downstream users.","The framework's logic generalizes beyond existing datasets to generated or synthetic data, where provenance and consent questions become murkier and the guide's emphasis on documenting transformations is even more salient."],"forward_implications":["Practitioners who work through the Origins questions are more likely to detect biased collection methods, missing consent, and licensing restrictions before a project starts.","Practitioners who apply the Usage questions are more likely to catch transformation choices that erase context or skew results, such as dropping missing data or binning continuous values.","Practitioners who follow the Stewardship questions are more likely to document derivatives, monitor for dataset deprecation, and plan for ethical archiving.","Widespread adoption of the lifecycle questions would push datasheets and deprecation frameworks toward field-standard practice, reducing the continued use of 'zombie datasets.'"],"supporting_citations":[{"why":"Introduces the datasheet documentation practice the guide adopts as a core tool for Origins and Stewardship questions.","marker":"[24]"},{"why":"Supplies the dataset deprecation framework the guide recommends for Stewardship, including reasons, timelines, and appeals.","marker":"[50]"},{"why":"Surveys common dataset pitfalls—spurious tasks, artifacts, sloppy annotation, and representation—that motivate the guide's critical questions.","marker":"[60]"},{"why":"Establishes data as local 'data settings' tied to their origins, a premise for the Origins and Usage questions.","marker":"[1]"},{"why":"Documents how classification choices in data work silence ambiguity, supporting the guide's caution about annotation and labeling.","marker":"[62]"},{"why":"Defines data as a relational category contingent on users and purposes, grounding the guide's central claim about datasets.","marker":"[8]"}],"fun_headline_variants":["No dataset is neutral: ask lifecycle questions before, during, after","Three question sets for critical ML dataset stewardship","Datasets are contingent: use lifecycle questions to avoid harms","Ask origins, usage, stewardship questions for safer ML datasets"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The guide assumes that reflecting on these questions will actually change practitioners' decisions and reduce harm, but it provides no evidence or user study that the questions alter behavior or outcomes.","fun_headline_variants_meta":{"raw":{"variants":["No dataset is neutral: ask lifecycle questions before, during, after","Three question sets for critical ML dataset stewardship","Datasets are contingent: use lifecycle questions to avoid harms","Ask origins, usage, stewardship questions for safer ML datasets"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000203,"raw_usage":{"total_tokens":1307,"prompt_tokens":788,"completion_tokens":519,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":404,"completion_tokens_details":{"reasoning_tokens":453}},"tokens_in":404,"tokens_out":519,"duration_ms":5122,"temperature":1.0,"reasoning_tokens":453,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T14:15:19.769510+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A controlled study in which two groups of practitioners—one using the lifecycle questions, one not—build or select datasets for comparable tasks, with independent audit of resulting harms, errors, and legal or ethical issues, would settle the claim; if the question-using group shows no measurable improvement, the guide's central value proposition fails.","supporting_citations":[{"cited_title":"The Data-Production Dispositif","cited_arxiv_id":"2205.11963","evidence_quote":"Documents how classification choices in data work silence ambiguity, supporting the guide's caution about annotation and labeling."}],"review_version":1}