{"id":"e93e9919-2755-456a-bf13-a653791f13a0","arxiv_id":"2506.22512","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Radical Questioning, a five-step pre-design ethics framework, redirects an AI human trafficking project from ad surveillance to survivor-centered evidence support.","lead":"This paper proposes Radical Questioning, a five-step framework for deciding whether an AI-for-good project should be built at all. The authors apply it to a human trafficking project and describe how it moved the design from online ad surveillance to a survivor-centered evidence tool.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Case study does not demonstrate pre-design use: T-NET detection system was already built and published, undermining the 'pre-project' claim central to RQ.","rationale":"After reading the paper, the most load-bearing concern is not only the dependence on stakeholder engagement (as the reader noted) but whether the case study actually exercises RQ in the pre-design role that is central to the paper's contribution. The paper repeatedly emphasizes 'pre-project,' 'pre-design,' and 'upstream,' and the abstract says RQ 'demonstrate[s]' a shift away from surveillance. Yet the authors are the same research group that built T-NET, a published detection system for escort ads. Their own initial framing matches T-NET, indicating that RQ was applied after a concrete AI system had been developed and peer-reviewed. This does not compromise the ethical value of reflecting post-hoc, but it does undermine the specific claim that RQ is a pre-project gate. If the framework is only shown to work as a mid-course or post-hoc reflection, the institutional recommendation to 'ask before you build' lacks direct evidential support from the paper's own case. The reader's CONDITIONAL verdict remains appropriate; the authors should either provide a true pre-design case study or soften the demonstration claims. There is no need to move the verdict to REJECT because the framework could still be useful and the paper is transparent about many limitations. However, the internal inconsistency is more concrete than the stakeholder-engagement concern and can be settled by a simple timeline audit.","tokens_in":8936,"tokens_out":7253,"duration_ms":82088,"concrete_test":"Audit the public record for T-NET (AAAI 2024 paper, any repository or demo) and the authors' statements to establish a timeline: when was the detection system designed, trained, and published relative to the RQ engagement described in the paper? Additionally, examine whether the paper provides any design logs, meeting notes, or dated artifacts showing that RQ steps were completed before any technical work on T-NET. If the timeline shows T-NET predates RQ, the 'pre-project' claim is not supported by the case study.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim of the paper is that Radical Questioning (RQ) is an upstream, pre-project ethical assessment tool that can prevent harmful AI-for-good deployments before technical design begins. The only supporting evidence is the authors' human-trafficking case study, in which RQ is said to have shifted the project from detection to documentation. However, the paper's own references show that the authors had already built and published T-NET, a weakly supervised graph learning system for detecting trafficking in escort advertisements (ref. [37], Nair et al., AAAI 2024). The initial framing described in Section 3.1—'How can we detect trafficking online using escort advertisements?'—is precisely the T-NET project. Thus, RQ was applied after a functional detection system existed, not before any design. The subsequent pivot is a retrospective course correction, not a pre-build gate. This does not necessarily invalidate RQ as a framework, but it means the abstract's claim to 'demonstrate how RQ... guides us away from surveillance-based interventions' is not a demonstration of pre-design use. The paper should either present a genuine pre-design case or explicitly reframe the contribution as a retrospective analysis and proposal, with the pre-design efficacy left as an open empirical question.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper introduces Radical Questioning (RQ), a five-step pre-design ethical assessment framework for AI-for-good projects, and illustrates it through the authors' human-trafficking case study. The authors argue that RQ shifted their project from automated detection of trafficking in online escort advertisements to a survivor-centered evidence management tool, and they claim the five-step structure can generalize to other high-stakes domains even though the specific questions must be contextual. The paper positions RQ relative to existing AI ethics, refusal, and participatory design frameworks, and it concludes with recommendations for institutionalizing pre-project ethical deliberation.","tokens_in":9145,"tokens_out":3433,"duration_ms":40464,"significance":"If the central claim were well supported, the paper would make a useful contribution by operationalizing the question 'should we build this at all?' as a structured, upstream deliberative step, and by candidly listing the conditions under which such questioning can work (genuine stakeholder engagement, time and institutional support, team openness). The stepwise questions are concrete, the framing is normatively important, and the paper explicitly acknowledges its own limitations, including the risk of performative ethics. However, the evidence base is a single first-person retrospective case study with no transcripts, engagement logs, or decision records, and the same case study that shaped RQ is then used to validate it. This limits the strength of the demonstration, though not necessarily the value of the proposal; the paper can be revised to reframe its contribution honestly.","major_comments":[{"comment":"The central claim that RQ is a 'pre-project' and 'before design' framework is not supported by the case study as presented. The initial framing quoted in §3.1 ('How can we detect trafficking online using escort advertisements?') is precisely the T-NET detection system that the same group published at AAAI 2024 (ref [37]). The paper supplies no timeline or contemporaneous evidence showing that RQ was applied before any technical design work began; the narrative in §5 reads as a retrospective reinterpretation of a system that had already been built and published. Because the abstract and §1 use the HT case to 'demonstrate' pre-project use, this is a load-bearing gap. Please either provide dated documentary evidence of the sequencing (e.g., workshop notes, decision logs, or an explicit project chronology) or reframe the contribution as a retrospective analysis and proposal, explicitly leaving pre-design efficacy as an open empirical question.","section":"§3.1, §5, ref [37]"},{"comment":"The validation of RQ is self-referential. Section 3 states that RQ was 'developed through its application in the HT domain' and that its questions were 'co-shaped by individuals differently situated in relation to harm, power, and intervention,' while §5 attributes the project's pivot to RQ's effects. The same stakeholder engagement that shaped the framework is the only evidence offered for its value, so the case study cannot distinguish the effect of RQ from the effect of the authors' prior ethical commitments, the advisory board, or the stakeholder input itself. Please separate framework formation (what was learned from the case) from framework evaluation (against independent criteria or a second, pre-registered case), or explicitly label the current evaluation as anecdotal and self-referential.","section":"§3 and §5"},{"comment":"The paper acknowledges that RQ's effectiveness hinges on 'genuine stakeholder engagement' and that when teams are unwilling to act, 'RQ risks becoming performative.' These are precisely the non-ideal conditions where a safeguarding framework is most needed, yet the paper does not show what RQ alone contributes in such settings. Since the abstract claims RQ is a generally applicable pre-project tool, the manuscript should either specify the boundary conditions more sharply (e.g., 'RQ is useful only when certain institutional and relational preconditions hold') or offer concrete strategies for building the needed engagement when gatekeepers or distrust block access. This is not a call to solve all real-world constraints, but the contingency weakens the demonstration of generalizability as currently worded.","section":"§4, limitations 2 and 3"}],"minor_comments":[{"comment":"The text says 'Gray boxes in the following section highlight actual questions raised during the design process,' but the manuscript shows plain bulleted lists rather than gray boxes; either restore the formatting or delete the reference to gray boxes.","section":"§3, figure/box formatting"},{"comment":"The ACM reference format and the 'Received 20 February 2007; revised 12 March 2009; accepted 5 June 2009' dates are clearly template placeholders and should be corrected before submission.","section":"Front matter"},{"comment":"In the sentence 'These shortcomings calls for upstream, human-centered frameworks,' the verb should agree with the subject: 'These shortcomings call for.'","section":"§1"},{"comment":"The comparison with existing tools such as the Situate AI Guidebook [24] is made in a single sentence; a short comparative table or an explicit differentiation criterion (e.g., 'pre-decision' versus 'during-development') would help readers verify the claimed novelty.","section":"§2 vs §4"},{"comment":"Section 4 lists 'RQ is not prescriptive' as a limitation, but Section 5 offers six prescriptive recommendations for practitioners; reconcile this tension by clarifying that the framework itself is non-prescriptive while its adoption recommendations are deliberately directive.","section":"§4 and §5"}],"recommendation":"major_revision","confidential_remarks":"This is a position/experience paper rather than an empirical study, and that is fine for the venue, but the gap between the abstract's 'demonstrate' language and the retrospective, self-referential evidence is significant. The most useful revision path would be to reframe the contribution as a proposal grounded in a retrospective case study, add explicit chronology and, if available, any artifacts that would corroborate the pre-design claim. The philosophical arguments are reasonable and the limitations are honestly stated, so the paper should not be rejected outright; it needs to match claims to evidence."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: RQ is a real contribution as a structured, named pre-design deliberation process, and the authors know the literature. But the case study is not evidence of pre-design use: their own T-NET detector was built and published (Nair et al., AAAI 2024) before RQ was applied. The paper frames RQ as having shifted the project from detection to documentation, which is a retrospective course-correction. That doesn't sink the framework, but it means the abstract's claim to 'demonstrate' pre-design guidance is unsupported.\n\nWhat is actually new: the five-step RQ sequence — define scope, identify stakeholders, understand contextual nuance, map ethical concerns, iterate with feedback — with a concrete set of probing questions grounded in the HT case. The positioning against RED, Situate AI Guidebook, Design Justice, and refusal literature is careful and fair. The limitations section is honest: non-prescriptive, dependence on genuine engagement, risk of performativity. The prose is clear.\n\nSoft spots, in proportion:\n\n1. No artifacts. There are no interview transcripts, engagement logs, decision records, team meeting notes, or before/after comparisons. The central claim that RQ 'reshaped our project' rests entirely on the authors' first-person account.\n\n2. The timeline problem. Section 3.1 gives the initial framing as 'How can we detect trafficking online using escort advertisements?' — that is exactly the T-NET project, already published. So RQ was applied after a functional detection system existed. This is a retrospective analysis, not a pre-build gate. The paper should either present a real pre-design case or reframe the contribution as a proposal with a structured validation agenda.\n\n3. Self-referential validation. Section 3 says RQ was 'developed through its application' in the HT domain, and Section 5 then uses that same application as the demonstration. That is not definitionally circular, but it is weak evidence.\n\n4. The engagement dependency is acknowledged but not solved. Limitation 2 admits that trust-building requires time and resources and may be blocked by gatekeepers. The paper doesn't show how to secure that engagement in settings where the authors' institutional position is absent.\n\nThe core argument — ask before you build, use reflective questioning to interrogate assumptions — is sound and important. The weakness is not the reasoning; it's the evidence for the central claim. That is fixable with reframing or external evaluation.\n\nRecommendation: this deserves a serious referee. I would want substantial revision to align the claims with what the case can actually show. I'd bring it to a reading group as a discussion piece on AI-for-good ethics, and I'd cite it if writing about pre-project ethical assessment frameworks.","headline":"A genuinely useful five-step pre-design ethics scaffold, but the only demonstration is a retrospective self-report of a system that was already built, so the paper's central pre-design claim is not yet supported.","tokens_in":9688,"tokens_out":1944,"would_cite":true,"duration_ms":20710,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that AI-for-good projects should undergo a five-step pre-design ethical assessment, Radical Questioning, that can halt or reframe interventions before harm occurs.","keywords":["AI ethics","responsible AI","radical questioning","human trafficking","pre-design assessment","techno-solutionism","survivor-centered design","AI for good"],"falsifier":"Run two comparable human-trafficking AI projects, one with and one without the full RQ process; if the RQ project, with genuine survivor engagement, still produces the original surveillance-based detection tool and comparable false positives, the central claim that RQ prevents or redirects harmful deployments would be falsified.","tokens_in":8726,"feed_emoji":"❓","tokens_out":6748,"duration_ms":65264,"temperature":0.7,"pith_summary":"This paper argues that AI-for-good projects, especially those addressing human trafficking, should begin by asking whether an AI system should be built at all. It introduces Radical Questioning (RQ), a five-step pre-design ethics process that defines the social problem, identifies stakeholders, understands contextual nuance, maps ethical concerns, and iterates with feedback. In the paper's case study, RQ shifted a planned automated trafficking-detection tool into a survivor-centered evidence-management tool, moving from surveillance to support. If this works, upstream questioning could prevent AI-for-good deployments that harm the very communities they claim to help.","feed_headline":"Five-step 'ask before you build' framework redirects AI-for-good","feed_subtitle":"Applied to human trafficking, the Radical Questioning process shifted a detection tool toward survivor support.","key_machinery":"The carrying mechanism is the five-step Radical Questioning (RQ) framework: (1) Define the Scope of the Problem, asking who gets to define the social issue; (2) Identify Stakeholders, asking who is impacted, who owns the tool, and whether marginalized voices are meaningfully involved; (3) Understand Contextual Nuance, probing contested notions of justice, consent, and success; (4) Map Ethical Concerns, covering accountability, privacy, fairness, and legitimacy; and (5) Iterate with Feedback, requiring continuous, deliberative stakeholder input and willingness to halt the project. RQ is a deliberative practice, not a checklist, and its questions were co-shaped with survivor-led organizations in the case study.","core_discovery":"The central discovery is a structured, pre-design method for questioning the legitimacy of an AI intervention before technical development begins. Radical Questioning does not replace principles-based ethics; it precedes it, creating a deliberative space to confront assumptions, map power, and consider harms. Applied to human trafficking, the method surfaced risks of over-surveillance, false positives, and retraumatization, and reoriented the project from detection to documentation and from surveillance to support. The authors claim the framework is transferable to other contested domains, provided its questions are re-grounded in local context.","pith_inferences":["The paper leaves open who has authority to halt a project when RQ surfaces serious harms; institutionalizing RQ would mean assigning that decision to a party without sunk costs in the build.","RQ's effectiveness could be tested empirically by comparing equivalent AI-for-good projects with and without the framework on downstream outcomes such as false-positive surveillance or retraumatization reports.","Because RQ depends on genuine engagement, its transferable core may be the reflective posture itself, with the specific questions acting as a scaffold that must be rebuilt for each domain."],"forward_implications":["Institutionalizing RQ as a standard pre-project step would require funding cycles and timelines that make room for reflection and the option to walk away.","Applying RQ in other high-stakes domains, such as child welfare or predictive policing, would require re-grounding its questions in local histories, laws, and power dynamics.","The case study's pivot from detection to documentation implies that success metrics in sensitive domains may be empowerment and harm reduction rather than optimization and scale.","Teams using RQ would establish survivor-led advisory structures and trauma-informed engagement practices as ongoing governance, not one-time consultations."],"supporting_citations":[{"why":"Supplies the critique that good intentions are not enough for AI-for-social-good, motivating the call to question whether to build.","marker":"[22]"},{"why":"Contributes the refusal-based orientation of interrogating power and problem framings that RQ converts into an upstream pre-design method.","marker":"[5]"},{"why":"Contributes the commitment to centering marginalized voices that anchors RQ's stakeholder and accountability steps.","marker":"[11]"},{"why":"Identifies ethical tensions in AI for human trafficking, the domain-specific risks RQ is designed to surface early.","marker":"[13]"},{"why":"Documents sex workers' views on AI policing of online ads, evidence for the harms of surveillance-based detection tools.","marker":"[10]"},{"why":"Offers a multi-stakeholder early-stage deliberation approach that RQ complements by operating even further upstream.","marker":"[24]"},{"why":"Establishes the critique of technology-driven anti-trafficking efforts that RQ builds on.","marker":"[36]"},{"why":"Provides the chilling-effect mechanism cited as a specific harm of ad-monitoring AIs that RQ helped the authors avoid.","marker":"[43]"}],"fun_headline_variants":["Ask before you build: five-step framework for AI-for-good","Radical Questioning redirects human-trafficking AI to survivor support","Five steps to question if AI should be built before building it","RQ framework moves AI from surveillance to support in trafficking"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The framework only works if a project can secure genuine, trusting engagement with affected communities and if the development team is willing to act on what it hears; without those, RQ becomes performative and provides no safeguard against harm.","fun_headline_variants_meta":{"raw":{"variants":["Ask before you build: five-step framework for AI-for-good","Radical Questioning redirects human-trafficking AI to survivor support","Five steps to question if AI should be built before building it","RQ framework moves AI from surveillance to support in trafficking"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000289,"raw_usage":{"total_tokens":1642,"prompt_tokens":845,"completion_tokens":797,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":461,"completion_tokens_details":{"reasoning_tokens":726}},"tokens_in":461,"tokens_out":797,"duration_ms":8559,"temperature":1.0,"reasoning_tokens":726,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T22:36:50.169420+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run two comparable human-trafficking AI projects, one with and one without the full RQ process; if the RQ project, with genuine survivor engagement, still produces the original surveillance-based detection tool and comparable false positives, the central claim that RQ prevents or redirects harmful deployments would be falsified.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the critique that good intentions are not enough for AI-for-social-good, motivating the call to question whether to build."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Contributes the refusal-based orientation of interrogating power and problem framings that RQ converts into an upstream pre-design method."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Identifies ethical tensions in AI for human trafficking, the domain-specific risks RQ is designed to surface early."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Documents sex workers' views on AI policing of online ads, evidence for the harms of surveillance-based detection tools."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Offers a multi-stakeholder early-stage deliberation approach that RQ complements by operating even further upstream."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes the critique of technology-driven anti-trafficking efforts that RQ builds on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the chilling-effect mechanism cited as a specific harm of ad-monitoring AIs that RQ helped the authors avoid."}],"review_version":1}