{"id":"4443b798-acb3-4fe5-9395-d92c2c77876c","arxiv_id":"2411.18908","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"DuetML adds multimodal LLM agents, one reactive and one proactive, to an interactive machine learning interface, and a small user study found outside evaluators rated its users' category definitions as more aligned with the task.","lead":"This paper presents DuetML, a system that embeds multimodal AI agents into a drag-and-drop machine learning interface to help people without technical backgrounds define image classification tasks. In a 12-person study, outside evaluators judged that DuetML users created category names that better matched the target task than users of a similar system without agents.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central claim hinges on one subjective, unreliability-unchecked rubric item (EQ1-a); without inter-rater agreement and a clear unit of analysis, p=0.002 may reflect rater noise or instruction-following rather than better task alignment.","rationale":"Read in good faith, the paper is a well-structured prototype and a legitimate user study; the system is actually implemented, the baseline is sensible, and the qualitative interview excerpts give plausible evidence that agents helped some participants. The central quantitative claim, however, is exactly the one the reader's verdict flags. The perfect-separation p-value is impressive in one sense, but with n=6 per arm and an unvalidated single-item outcome, it is fragile. The multiple-comparison worry is less severe than the reader suggests because p=0.002 would survive a Bonferroni correction for five metrics; the more damaging gaps are the absence of any inter-rater reliability estimate and the ambiguity in the unit of analysis. A mixed-effects re-analysis or raw-data release would settle this. I also note a smaller internal-consistency issue: the paper labels a between-subjects comparison as a 'Wilcoxon signed-rank test' in Sec 4.3.1, which should be a rank-sum test; this does not change the qualitative conclusions but should be corrected. The multimodality section (4.3.4) compares actual MLLM outputs with 'hypothetical' text-only responses, which is not a controlled comparison; this weakens the multimodality-specific claims but is not the load-bearing pillar for the EQ1-a result. On balance, the appropriate verdict remains conditional: the central claim is plausible and partially supported, but the quantitative proof depends on a reliability check that has not been reported.","tokens_in":20119,"tokens_out":7144,"duration_ms":65037,"concrete_test":"Obtain the raw per-rater EQ1-a ratings for all 12 participants (and, if possible, the other EQ items). Compute inter-rater reliability using ICC(2,k) or Krippendorff's alpha, and re-run the condition comparison with a mixed-effects model: rating ~ condition + (1 | participant) + (1 | rater), reporting the condition effect and its 95% CI. Also report the U statistic, effect size (e.g., rank-biserial correlation), and a cluster-robust or permutation test preserving the participant-rater nesting. If ICC(2,k) is below 0.6 or the 95% CI for the condition effect includes 0, the central claim is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"DuetML's quantitative superiority reduces to a single 5-point Likert item (EQ1-a: 'category names express object usage') rated by five hired practitioners on a researcher-authored rubric. Three problems make this load-bearing. (1) No inter-rater reliability is reported; if raters disagree about what 'express object usage' means, the average scores are not a stable measure of alignment. (2) The reported p=0.002 is exactly the minimum achievable for a two-tailed Mann-Whitney test with six observations per arm, indicating perfect separation of the six DuetML and six baseline participant means, but no effect size, confidence interval, or raw data are given; if the analysis instead pooled raters' scores as independent observations, the test is invalid because the five ratings of the same participant are not independent. The paper's phrasing 'comparing the evaluations from each of the five evaluators separately' is ambiguous about which unit was tested. (3) The directed-task instructions and the EQ1-a criterion were authored by the same researchers who designed DuetML, and the active/passive agent system prompt explicitly says 'Suggest category names'; the metric may therefore measure the LLM's instruction-following and the user's compliance, not general task-formulation ability. The abstract's broad 'better aligns with target tasks' claim also goes beyond the evidence, since only one of five primary rubric items reached significance.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"DuetML presents a framework that integrates multimodal LLM agents into an interactive machine-learning (IML) interface, with one reactive and one proactive agent that converse with non-expert users while they build image classifiers. The paper reports a between-subjects user study (N=12) comparing DuetML with a baseline IML system, plus a third-party evaluation by five ML practitioners who rated the participants' training data on five Likert-scale rubric items (EQ1-a through EQ5-a) and corresponding free-response items. The primary quantitative claim is that DuetML users produced category names better aligned with the directed task, based on EQ1-a (0.7 vs. -0.6, p=0.002). The paper also reports no significant usability differences, positive subjective ratings of the agents, qualitative interaction-log analyses, and a comparative analysis of multimodal versus text-only advice. The central claim is that human-LLM collaboration helps non-expert users formulate better ML tasks without increasing cognitive load.","tokens_in":20410,"tokens_out":5688,"duration_ms":52219,"significance":"If the central result holds, DuetML is a useful contribution to the IML and human-AI collaboration literature: it demonstrates a concrete way to combine user agency with proactive LLM guidance, and it provides a reproducible system description with full prompt text. The study design is appropriate for an initial comparison, and the qualitative analyses of interaction logs are informative. However, the load-bearing quantitative evidence is currently thin: the single significant result is based on one subjective rubric item, with no inter-rater reliability, an under-specified statistical test, and a small sample. The significance of the broader 'better task alignment' claim therefore rests on statistical and construct-validity analyses that the manuscript does not yet provide. The framework itself is sound as a proof-of-concept, and the reported qualitative observations support that the agents were useful, but the quantitative demonstration needs strengthening before the advertised claim can be accepted.","major_comments":[{"comment":"The statistical test for EQ1-a is under-specified. With six participants per arm, the smallest achievable two-tailed Mann-Whitney p is 2/924 ≈ 0.002, which is the exact reported value; this means the result is a perfect separation of the two groups regardless of effect magnitude. The phrase 'comparing the evaluations from each of the five evaluators separately' does not identify the unit of analysis: if participant-level means were used, the paper should report per-participant and per-rater data and an effect size with confidence interval (e.g., rank-biserial correlation); if the five raters' scores were pooled, the observations are not independent and the reported p-value would not be valid. The authors must state the unit of analysis, report raw data, and provide an effect size and confidence interval for EQ1-a. They should also report whether EQ1-a was a pre-registered primary endpoint or one of five endpoints tested; if the latter, a simple multiple-comparison bound (e.g., Bonferroni) should be reported alongside the individual p-values.","section":"§4.3.2"},{"comment":"No inter-rater reliability is reported for the five practitioners' ratings. Because the central significant result is a single Likert item ('category names express object usage') rated by five individuals, the manuscript must report agreement metrics (e.g., intraclass correlation coefficient for EQ1-a, or at least per-item agreement tables). Without such information, the average scores could reflect rater-specific interpretations of 'express object usage' rather than systematic differences between conditions. This is not a minor omission: it directly bears on whether the reported 0.7 vs. -0.6 difference is a stable property of the training data or an artifact of rater noise.","section":"§4.3.2"},{"comment":"The usability comparison uses the wrong statistical test. The text states that 'Statistical analysis (Wilcoxon signed-rank test)' was applied to compare DuetML and baseline on Q1-Q5, but these are between-subjects comparisons (different participants used each system), so the paired Wilcoxon signed-rank test is inappropriate; the Mann-Whitney U test should be used. Because the abstract's claim that DuetML works 'without increasing cognitive load' is supported only by the absence of significant differences on these items, the authors must rerun the analysis with the correct test and report the resulting p-values and effect sizes. If the test choice is changed, the conclusion may change.","section":"§4.3.1"},{"comment":"The construct validity of EQ1-a is questionable because the task instructions, the agent prompts, and the evaluation rubric encode the same normative view of good task formulation. The directed task explicitly instructs participants to classify objects 'considering how these objects are used,' and the active/passive system prompts tell the agent to 'Suggest category names' and to 'suggest testing with adversarial or ambiguous images.' The significant EQ1-a result may therefore measure the LLM's success in getting participants to follow the system's own recommended strategy rather than a general improvement in non-experts' task-formulation ability. The abstract's sweeping claim that DuetML enables users to define training data that 'better aligns with target tasks' is also broader than the evidence, since only one of the five primary rubric dimensions reached significance. Please temper the claim or provide an independent evaluation, such as a rubric dimension not present in the agent prompts, a comparison against a text-only LLM condition, or an evaluation of the final trained model's performance.","section":"§4.1, Appendix A.1 and B.1"}],"minor_comments":[{"comment":"The submitted text contains an apparent compilation artifact: around the '4.3 Results' heading, there is a stray passage with page headers '12 W. Kawabe et al.' and a figure caption 'Fig. 5: Overall trends in questionnaire responses' that appears to belong to a different manuscript. Please remove or correct it.","section":"§4.3"},{"comment":"Detailed metrics EQ1-b through EQ5-b are reported as descriptive means without statistical tests; if these are intended as supporting evidence, the authors should either add statistical tests or explicitly label them as descriptive trends.","section":"§4.3.2"},{"comment":"The word 'objectively' in 'we conducted a third-party evaluation' is misleading, because the evaluation is based on subjective Likert ratings by practitioners; it is better described as 'independent' or 'third-party' rather than 'objective.'","section":"§4.2"},{"comment":"The paper does not provide a data availability statement or supplementary raw data, which would allow readers to verify the reported p-values and to examine the distribution of ratings; please add such a statement or appendix.","section":"§4.3.2"},{"comment":"The familiarization task uses a vegetable image dataset, while the two main tasks use animal and Caltech-101 datasets, but the paper does not report whether participants' prior familiarity with these categories varied across conditions; this would be useful context for interpreting the results.","section":"§4.1"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope for an HCI/interactive-systems venue and the framework contribution is real, but the statistical reporting is not yet at journal standard. I would ask the authors to provide the raw participant-level and rater-level data, clarify the unit of analysis for EQ1-a, add inter-rater reliability and effect sizes, and correct the within-subjects test used for between-subjects usability comparisons. Also, the stray figure/passage from another paper in Section 4.3 suggests a compilation error that should be fixed before any resubmission."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The short version: DuetML is a real system contribution—a GUI-based IML tool with two complementary MLLM agents, one reactive and one proactive, both with access to the user's actual training images. That design is new and worth knowing about. But the empirical case for it rests on one fragile statistical result, and the paper's claims run ahead of what the data can support.\n\nWhat it does well: the system is thoughtfully designed. The active/passive split is a reasonable instantiation of mixed-initiative interaction, and giving the agents access to the user's images is a concrete step beyond text-only copilots. The prompts are in the appendix, implementation details are concrete, and the user study is a genuine between-subjects comparison with non-experts. The qualitative examples of image-aware advice versus text-only responses are illustrative and do show the MLLM using image content in a way a text-only system couldn't. The citation pattern looks fine: Horvitz, recent copilot work, and IML design guidelines are all represented.\n\nSoft spots: the entire quantitative case for 'better alignment' reduces to EQ1-a, one five-point Likert item rated by five hired practitioners. The reported p=0.002 for 6 vs 6 is perfect separation, which is fragile when five metrics were tested and no correction is reported. The paper says 'comparing the evaluations from each of the five evaluators separately'—that's ambiguous about the unit of analysis. If the five ratings per participant were pooled, non-independence invalidates the test. There is no inter-rater reliability, no effect size, no confidence intervals. This matters because the EQ1-a construct ('category names express object usage') is exactly what the system prompt tells the agent to suggest; the metric may partly measure instruction-following rather than general task-formulation ability. Also, the abstract claims 'without increasing cognitive load,' but the study measured subjective usability, not cognitive load, and a null result with six participants does not support an equivalence claim.\n\nNone of this means the system doesn't work. The contradiction between participants' self-reports and the third-party ratings is interesting, and the qualitative data suggest the agents did help people think harder. But the paper needs corrected statistics, inter-rater agreement, and a more careful framing to make the central claim stick.\n\nWho should read it: people designing IML tools with LLM assistance, and HCI researchers interested in mixed-initiative evaluation. It deserves a serious referee—an editor should send it to review, not desk reject. A good reviewer can sort out whether the effect is real.","headline":"A genuinely useful dual-agent MLLM/IML system whose headline result rests on one fragile Likert item; worth refereeing, not worth citing as established fact.","tokens_in":20948,"tokens_out":2624,"would_cite":true,"duration_ms":23962,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that multimodal LLM agents collaborating inside an interactive machine-learning interface help non-expert users define training data that better matches the task they actually want to solve.","keywords":["Interactive Machine Learning","Large Language Models","Multimodal LLM","Human-AI Collaboration","Task Formulation","Non-Expert Users","Image Classification","User Study"],"falsifier":"Re-run the directed task with more than six participants per condition and a pre-registered analysis: the reported $p = 0.002$ is the smallest p-value attainable for a 6-versus-6 Mann-Whitney comparison, so a larger replication either reproduces the EQ1-a effect or reveals it as a multiple-comparison artifact, and reporting inter-rater reliability would show whether the rubric judgments are stable enough to measure task alignment at all. A second decisive test is an ablation that keeps the agents and prompts identical while withholding image access, checking whether the benefit comes from multimodality or from conversational guidance alone.","tokens_in":19871,"feed_emoji":"🤖","tokens_out":14470,"duration_ms":112443,"temperature":0.7,"pith_summary":"This paper claims that the hardest step for a non-expert building a custom image classifier — turning a vague goal into a well-structured training task — can be substantially supported by multimodal large language model (MLLM) agents embedded in an interactive machine learning (IML) interface. The proposed system, DuetML, pairs the user with two complementary GPT-4o-based agents: a passive agent that answers on-demand questions and an active agent that proactively suggests improvements. In a between-subjects study with twelve non-experts, external ML practitioners rated training data produced with DuetML as better aligned with the directed task — category names expressing object usage scored 0.7 versus −0.6 for the baseline, $p = 0.002$ — while users reported no extra cognitive load. The paper argues this demonstrates a middle path between fully human-driven IML and fully machine-driven LLM prompting, in which users keep final authority over the task while machines contribute formulation expertise. If the claim holds, it offers a concrete recipe for democratizing task-specific model building.","feed_headline":"LLM agents improve non-experts' ML training data","feed_subtitle":"Users with the agent system named categories closer to the target task, at no added cognitive cost.","key_machinery":"The load-bearing mechanism is DuetML's two-agent design, implemented as two asynchronous GPT-4o models with complementary intervention styles. The passive agent responds reactively to explicit user requests — chat input and 'Ask the assistant' buttons placed in the training-data and evaluation sections — while the active agent monitors the user's overall interaction and volunteers suggestions on a 60-second interval, with a user toggle to disable it. Both agents receive prompts embedding the full dialogue history and the current training data rendered as composite images (each category shown as its name plus up to 50 randomly selected associated photographs), which is what lets the advice reference the user's actual state. The system prompt directs the agents to probe vague initial goals, propose category names, watch for misalignments such as inappropriate names or insufficient categories, and suggest adversarial test images to uncover overlooked categories. A MobileNet-plus-SVM classifier with fast training keeps the workflow interactive, and the study's third-party rubric measures precisely the behaviors the prompt targets — whether category names express object usage, whether images and names match, and whether categories cover and cleanly separate the task.","core_discovery":"The paper's central discovery claim is that human-LLM collaborative ML works: non-expert participants who built an image classifier with DuetML created training data that third-party ML practitioners rated as better aligned with the target task than training data created with an otherwise identical IML system lacking the agents. The significant difference appeared on the primary rubric item, EQ1-a ('category names express object usage', means 0.7 vs −0.6, $p = 0.002$ by Mann-Whitney U test), with the other four rubric metrics also trending in DuetML's favor without reaching significance. The authors further claim this gain came without increasing cognitive load, that users embraced the agents as collaborators rather than a burden, and that interaction-log and interview evidence shows participants thinking more deeply about their ML tasks — refining categories hierarchically, extracting domain knowledge from the agents, and adopting abstraction principles such as 'A' versus 'not A' categories. Underlying these results is the claim that the agents' access to the user's actual training images is what allowed their advice to be concrete, and that this collaborative paradigm preserves human agency while machine intelligence supplies formulation expertise.","pith_inferences":["A natural next design step, suggested by the logs but not tested in the paper, is adaptive proactivity: the active agent could scale its intervention frequency to the user's observed engagement, which might outperform the fixed 60-second cycle.","If the category-naming effect is real, the same collaborative pattern should transfer to other places where novices struggle to convert goals into machine-readable structure — prompt writing, data cleaning, schema design — making the paradigm testable well beyond image classification.","The paper's image-aware versus text-only comparison is illustrative rather than measured, so a controlled ablation that holds prompts and users fixed while toggling image access would be the clean way to prove that multimodality, not conversational scaffolding alone, drives the benefit.","The paradigm implies a rebalancing of ML prototyping roles: the machine carries formulation expertise while the human keeps final authority, a division of labor that future IML toolkits could standardize rather than treating automation and user control as opposites."],"forward_implications":["Task-formulation support can improve measurably without users noticing it: self-reported usability and success were statistically indistinguishable between DuetML and the baseline while external ratings favored DuetML.","The agents' image access appears to be what makes advice actionable — text-only responses to the same scenarios were generic, whereas image-aware responses pointed at specific misplaced images, mislabeled species, and missed categories.","Users with the agent system spontaneously adopted strategies that IML research has tried to teach novices — hierarchical refinement, abstraction into complementary pairs, and knowledge lookup — indicating the collaboration doubles as an on-demand tutor.","Reactive and proactive channels serve different needs: participants messaged the passive agent more in the directed task and used the ask buttons more in the open-ended task, suggesting both interaction styles are worth retaining."],"supporting_citations":[{"why":"Frames the interactive machine-learning paradigm that DuetML extends and that the baseline system instantiates.","marker":"[7]"},{"why":"Defines the democratization mission and the roles of humans in IML that motivate the paper.","marker":"[8]"},{"why":"Supplies prior evidence that non-expert users struggle to translate needs into ML tasks, the problem DuetML targets.","marker":"[9]"},{"why":"Adds further evidence of non-expert difficulty with task formulation in interactive ML.","marker":"[10]"},{"why":"Provides the mixed-initiative principles that justify balancing reactive and proactive agents.","marker":"[20]"},{"why":"Documents the GPT-4 model family whose multimodal image-and-language capabilities power the agents.","marker":"[41]"},{"why":"Supplies the Caltech-101 dataset used in the directed task that the third-party evaluation scores.","marker":"[55]"}],"fun_headline_variants":["LLM agents help non-experts refine ML training data","Non-experts create better ML data with LLM agents","Agent-aided ML data beats solo efforts for novices","LLM agents sharpen non-experts' ML task clarity","Human-LLM teamwork refines non-experts' ML data"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central claim rests on the assumption that the third-party rubric item 'category names express object usage' actually measures how well training data aligns with the target task — the rubric and the task instructions were written by the same researchers, no inter-rater reliability is reported, and only one of the five rubric metrics showed a statistically significant difference.","fun_headline_variants_meta":{"raw":{"variants":["LLM agents help non-experts refine ML training data","Non-experts create better ML data with LLM agents","Agent-aided ML data beats solo efforts for novices","LLM agents sharpen non-experts' ML task clarity","Human-LLM teamwork refines non-experts' ML data"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000775,"raw_usage":{"total_tokens":3446,"prompt_tokens":977,"completion_tokens":2469,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":593,"completion_tokens_details":{"reasoning_tokens":2386}},"tokens_in":593,"tokens_out":2469,"duration_ms":18146,"temperature":1.0,"reasoning_tokens":2386,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T10:45:40.468152+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the directed task with more than six participants per condition and a pre-registered analysis: the reported $p = 0.002$ is the smallest p-value attainable for a 6-versus-6 Mann-Whitney comparison, so a larger replication either reproduces the EQ1-a effect or reveals it as a multiple-comparison artifact, and reporting inter-rater reliability would show whether the rubric judgments are stable enough to measure task alignment at all. A second decisive test is an ablation that keeps the agents and prompts identical while withholding image access, checking whether the benefit comes from multimodality or from conversational guidance alone.","supporting_citations":[{"cited_title":"A review of user interface design for interactive machine learning","cited_arxiv_id":null,"evidence_quote":"Frames the interactive machine-learning paradigm that DuetML extends and that the baseline system instantiates."},{"cited_title":"Power to the people: The role of humans in interactive machine learning","cited_arxiv_id":null,"evidence_quote":"Defines the democratization mission and the roles of humans in IML that motivate the paper."},{"cited_title":"Image-to-text translation for interactive image recognition: A comparative user study with non-expert users","cited_arxiv_id":null,"evidence_quote":"Supplies prior evidence that non-expert users struggle to translate needs into ML tasks, the problem DuetML targets."},{"cited_title":"Use of machine learning by non-expert dhh people: Technological understanding and sound perception","cited_arxiv_id":null,"evidence_quote":"Adds further evidence of non-expert difficulty with task formulation in interactive ML."},{"cited_title":"One-shot learning of object categories","cited_arxiv_id":null,"evidence_quote":"Supplies the Caltech-101 dataset used in the directed task that the third-party evaluation scores."}],"review_version":1}