{"id":"f95045cd-95fd-4c41-a686-7008600b4375","arxiv_id":"1908.08931","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A position paper proposing expert-centric machine teaching and Massive Open Learning AI (MOLA) as a path to more humane, inclusive, and intelligent ML systems.","lead":"This paper argues that machine learning systems could be built directly by domain experts who teach concepts rather than label data, with communities of experts collaborating like open source developers. It is a vision paper that outlines challenges and potential benefits, not a tested system.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Even if machine-teaching interfaces mature, the MOLA claim that communities will build more intelligent systems needs the very inspectability and modularity the paper concedes ML models lack; the open-source analogy is ungrounded without a mechanism for contradiction handling.","rationale":"The reader correctly identified the immaturity of machine teaching as a central risk: if domain experts cannot teach without ML intermediaries, the whole expert-centric vision collapses. My stress-test goes one step downstream: even if the user-facing teaching interfaces of Sec. 3.3 become feasible, the MOLA extension in Secs. 4-7 requires a knowledge representation that is inspectable, modular, and contradiction-tolerant in ways that the paper itself says current ML models are not. The paper explicitly acknowledges this in Sec. 4, yet it does not offer even a sketch of how the model's knowledge state could be made visible and editable enough for communities to review, fork, and repair it. Without that, the analogy to open source software, which is the main evidence for the 'more intelligent' and 'high quality' claims, is an analogy rather than an argument. I do not think this concern changes the reader's verdict: the paper is honestly framed as a proposal, it flags many of its own open problems, and it should remain a conditional acceptance as a position paper rather than a demonstrated result. The concern is real but not a rejection; it sharpens why the conditions matter.","tokens_in":9277,"tokens_out":6218,"duration_ms":73782,"concrete_test":"Run a controlled repair study: build a minimal MOLA for one well-scoped domain (e.g., dietary advice for diabetic patients) with a transparent knowledge store that logs every expert contribution as an inspectable, revertible edit and automatically flags contradictions. Seed 10 known errors into the model, then ask 20 domain experts with no ML training to locate and fix those errors using only the interface. If fewer than 80% of seeded errors are correctly located and repaired without ML-expert help, the community-maintenance mechanism at the heart of MOLA fails, and the open-source analogy is not load-bearing.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that community-built MOLAs will produce more intelligent systems depends not only on the acknowledged immaturity of machine teaching (Sec. 3.3) but on an unexamined analogy to open source. Open source software works because contributions are inspectable, modular, revertible, and testable; code review, forking, and Linus's law all assume these properties. Section 4 explicitly concedes the opposite for ML knowledge: 'unlike the wikipedia page which is concrete, clearly visible, and easily reconfigured by the members of the community, the knowledge a ML system possesses, which determines how it actually behaves, is much more complicated to represent, change, and assess.' Section 6 then lists contradictions, simultaneity, and evaluation consistency as open challenges with no proposed solution, and Section 5 resolves contradictions by a human forum outside the machine. If community members cannot see what the model 'believes', identify which contribution caused an error, or fork and revert a change, the mechanisms that make open source communities high-quality do not transfer. The paper offers no technical proposal for making the model's knowledge state inspectable enough to support these community mechanisms; it merely asserts that a community will help. Thus the 'more intelligent' portion of the central claim is unsupported at exactly the point where the paper's novel contribution, MOLA, is strongest.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper is a vision/position piece arguing that current deep learning is approaching limits and that expert-centric machine teaching—where domain experts teach ML systems directly using concepts, definitions, and other high-level constructs, without ML-expert mediation—can make ML development more humane, more inclusive, and possibly more intelligent. The paper distinguishes teaching from training, lists four technical requirements for direct expert access, introduces the notion of Massive Open Learning AI (MOLA) as a community-built, openly taught ML system, illustrates the idea with a hypothetical nutritionist community building a diabetes assistant, and discusses challenges such as handling contradictions, simultaneity, varying expertise, and evaluation consistency. The paper concludes that MOLAs may unlock intelligent systems in complex domains such as health, law, and economics. It contains no experiments or formal analyses; its argument is based on analogy, illustrative examples, and citations of prior work.","tokens_in":9524,"tokens_out":3721,"duration_ms":43567,"significance":"If the paper's central claims could be substantiated, the work would be significant for human-centered ML and interactive machine learning: it reframes the development process from ML-experts to domain experts and offers a concrete vision of community-built ML systems. The paper is commendably honest about the technical gap: Section 3.3 explicitly calls the required machine-teaching experience a 'formidable challenge' for current ML and MT technologies. It also engages with the relevant literature on interactive machine learning, machine teaching, and open-source community dynamics, and it acknowledges risks such as sabotage (Tay), bias, and hijacking of community processes. The main value of the paper lies in posing a plausible research agenda rather than in demonstrating a result; its broad claims about 'more intelligent' systems are not yet supported.","major_comments":[{"comment":"The central premise of the paper depends on machine-teaching technology satisfying the listed requirements (transparency, explainability, multiple teaching patterns, and re-training in under 5 seconds), yet the paper itself states that this is a 'formidable challenge for both current ML and MT technologies.' The manuscript offers no concrete technical roadmap, no citation of a system that jointly meets these requirements, and no evidence that the gap is closing. Consequently, the conclusion in Section 8 that MOLAs 'may be key to unlock the deployment of intelligent systems in complex, diverse domains' rests on an acknowledged but unresolved feasibility assumption. The paper should either temper this claim explicitly to a hypothetical, or provide a substantive argument, with references or preliminary results, that the four requirements are jointly achievable.","section":"Section 3.3"},{"comment":"The analogy between MOLA communities and open-source software communities is not sufficiently grounded. Section 4 concedes that, unlike Wikipedia pages, the knowledge of an ML system 'is much more complicated to represent, change, and assess.' Yet the quality mechanisms of open-source communities—inspectability, modularity, forkability, revertibility, and testability—all presuppose exactly those properties. The paper never explains how model knowledge can be made inspectable or revertible enough for community review, and Section 6 lists contradictions, simultaneity, and evaluation consistency as unsolved problems, with no proposed mechanism. Section 7 then asserts community-based development yields 'high quality and reliability,' citing Linux and Bitcoin. That is an analogy, not an argument; without a concrete mechanism for attribution, inspection, and rollback of community contributions, the claim that MOLAs will be more intelligent is unsupported at the point where the paper's novelty is strongest.","section":"Sections 4 and 7"},{"comment":"The claim that expert-centric and community-built systems will be 'more intelligent' is never defined or operationalized. The paper convincingly argues for humaneness and inclusiveness, but 'improved intelligence' is asserted rather than analyzed. The nutritionist scenario in Section 5 is explicitly hypothetical and says 'let us forget for a moment the technical difficulties,' so it cannot serve as evidence. Section 7 attributes capabilities to communities (consensus on real needs, role of motivation, group understanding of the machine, high quality and reliability) without citing empirical studies connecting these properties to ML system performance. To make the central claim load-bearing, the authors need to either define 'intelligence' here in terms of measurable capabilities (e.g., handling evolving knowledge, diverse viewpoints, or rare edge cases) or explicitly reframe the paper as proposing testable hypotheses rather than asserting an outcome.","section":"Sections 1, 5, and 8"},{"comment":"The supporting evidence for unmediated direct access is thin. The example of intent-action chatbots is cited as a success because 'hundreds of thousands of chatbots' were created by non-ML-experts, but this is a measure of adoption, not of system quality or intelligence. The paper itself notes the intent-action model tends to 'make difficult the scale up of such efforts, thus limiting depth of knowledge and quality.' This evidence therefore does not support the stronger claim that direct expert access leads to better ML systems; at most it suggests ease of use. The manuscript should distinguish accessibility from capability when using this example.","section":"Section 3.2"}],"minor_comments":[{"comment":"There are several typos: 'LTSMs' should be 'LSTMs', 'an revolutionary technology' should read 'a revolutionary technology', and 'We are now almost close to a decade' is awkwardly phrased.","section":"Section 1"},{"comment":"The sentence 'or get if from human beings in “labeling farms”' contains a typo: 'get if' should be 'get it' or 'obtain it'.","section":"Section 2"},{"comment":"The phrase 'Teaching machines withing the context of a community' contains a typo: 'withing' should be 'within'.","section":"Section 3.1"},{"comment":"The hypothetical community is described as nutritionists, but later paragraphs refer to 'the physicians’ language' and 'a sub-group of physicians keeps monitoring the quality of the advice.' If the example is about nutritionists, the terms should be 'nutritionists' or 'clinicians' to avoid confusion.","section":"Section 5"},{"comment":"Some reference entries are incomplete: for example, the Amershi et al. 2015 entry and several conference proceedings entries lack full page ranges or publisher information. A consistency pass over the reference list is needed.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"This is a clearly written, thought-provoking position paper, and the author is appropriately candid about the immaturity of the technology. My main concern is the gap between the speculative nature of the argument and the strength of the claims in the abstract and Section 8 about 'better ML systems' and MOLAs unlocking intelligent systems in health, law, and economics. If the venue accepts vision papers, the manuscript is within scope; otherwise it may be better positioned as a Perspectives article. The central claims can be made acceptable by reframing them as testable research hypotheses and by either providing a mechanism for community inspectability/revertibility or explicitly flagging that mechanism as the open research problem it is."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's the short take: this is a clearly argued vision essay, not an empirical or formal result. The genuinely new thing is the MOLA label and the argument that machine teaching should be re-centered on domain experts, with a useful list of technical challenges. It reads well, and the author knows the interactive ML and machine teaching literature.\n\nWhere it earns its keep: the critique of label farms and the 'grinding-ungrinding' metaphor is fair; the distinction between teaching (high-level concepts) and training (examples) is a useful frame even if not new; and Section 6's list of challenges—contradictions, simultaneity, expertise levels, evaluation consistency—is a practical checklist for anyone trying to build such a system. The paper is honest that the technology is not there; 'formidable challenge' is stated in Section 3.3, and the nutritionist example is explicitly hypothetical.\n\nThe soft spot is the claim that community-built MOLAs will be more intelligent. The open-source analogy needs inspectability, modularity, revertibility, and testability. Section 4 concedes that ML knowledge is far less visible and reconfigurable than a Wikipedia page, and Section 6 leaves contradiction handling to a human forum outside the machine. So the very conditions that make open source communities high-quality are exactly what the paper says current ML lacks. The paper offers no mechanism to close that gap. That means 'more intelligent' is a hope, not a conclusion. It is fine to present it as a research hypothesis, but the abstract and final discussion state it too confidently.\n\nAlso, the teaching-vs-training dichotomy is somewhat stacked—teaching is defined as richer, so the claim that humans prefer it is almost tautological. And the 5-second retraining figure comes from one 2003 study, worth flagging as domain-specific.\n\nOverall judgment: as a position paper it is decent. A serious editor should send it out for review at a venue that publishes vision statements, with a requested revision to frame it as a proposal and add a roadmap. I would not desk-reject it. In my own work I wouldn't cite it as a technical result, but I might mention the MOLA framing in a related-work paragraph. I could see bringing this to a reading group that discusses human-centered AI; it will generate a good argument.","headline":"A well-written vision essay that names MOLA and lays out useful challenges, but the open-source analogy undercuts itself: the 'more intelligent' claim is a hope, not a result.","tokens_in":10049,"tokens_out":3606,"would_cite":false,"duration_ms":33534,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that machine learning systems should be built by domain experts teaching concepts directly, rather than by ML engineers mediating through labeled data.","keywords":["machine teaching","expert-centric machine learning","interactive machine learning","massive open learning AI","domain experts","community-based AI","humane AI","inclusive AI"],"falsifier":"A concrete test: build an expert-centric teaching system for a realistic domain, such as a conversational assistant for diabetic nutrition, and measure whether domain experts can successfully teach it without ML support. If expert-taught performance falls clearly below a conventionally trained baseline, or if experts abandon the tool because re-training exceeds the five-second limit, the central claim weakens. A second test: run a MOLA experiment with a community of experts on a knowledge domain with genuine disagreements and see whether contradictory teachings can be resolved into a coherent, reliable system; persistent incoherence would falsify the community-learning thesis.","tokens_in":9048,"feed_emoji":"🎓","tokens_out":6316,"duration_ms":54914,"temperature":0.7,"pith_summary":"The paper argues that current machine learning development, in which domain experts produce labeled examples and ML engineers turn them into models, bottlenecks knowledge and treats people as machine feeders. It proposes that ML systems should instead be directly teachable by domain experts using concepts, definitions, demonstrations, procedures, and tests—an approach the paper calls machine teaching. If such expert-centric machine teaching becomes feasible, experts could build and maintain AI systems themselves, and communities of experts could jointly construct intelligent systems in the same way open source communities build software. The paper claims this would make ML systems more humane, more inclusive of diverse viewpoints, and potentially more intelligent, because knowledge would no longer be reduced to input-output pairs.","feed_headline":"Let domain experts teach machines directly","feed_subtitle":"A vision for expert-driven machine teaching promises more humane, inclusive, and intelligent AI","key_machinery":"The load-bearing distinction is between training and teaching. Training shows a machine input-output examples so it can induce patterns; teaching transfers knowledge through high-level constructs such as concepts, definitions, demonstrations, procedures, and tests. The paper names four technical requirements for making teaching work for domain experts—transparency, explainability, multiple teaching patterns, and fast re-training (under five seconds)—and, for community-scale learning, introduces the MOLA (Massive Open Learning AI system), which must handle contradictions, simultaneity, varying levels of expertise, and evaluation consistency. This combination is what would let domain experts replace ML intermediaries and build intelligent systems themselves.","core_discovery":"The central claim is that shifting the center of gravity of machine learning development from ML experts to domain experts can lead to better ML systems. The paper proposes that machine teaching—transferring knowledge to a machine through declarative concepts, exemplars, definitions, demonstrations, procedures, and tests, rather than through labeled examples—can allow domain experts to teach ML systems without mediation. It then extends this to communities: a Massive Open Learning AI system (MOLA) learns directly from a large, diverse, mostly unmediated community of experts, and must explicitly handle contradictory facts and opinions instead of hiding them. The paper argues that expert-centric systems built this way are more humane (experts are no longer reduced to labeling data), more inclusive (community knowledge reflects multiple points of view), and potentially more intelligent (knowledge transfer is richer and can cover complex, changing domains such as health, law, and economics).","pith_inferences":["A testable next step is a bounded MOLA prototype in one specialty, such as diabetic nutrition advice, to compare whether community-taught systems capture more expert knowledge than a conventionally trained model on the same task.","If expert-centric teaching matures, the central design problem may shift from algorithms to community governance—how to weight conflicting teachers, prevent vocal subgroups from dominating, and manage authorship of aggregated knowledge.","The proposal implies new evaluation metrics: instead of a single held-out accuracy, systems would need measures of how faithfully they reflect the range and distribution of expert views, which do not yet exist.","One could use the teaching interface itself as a measurement instrument, logging the concepts and corrections experts provide across domains to test whether a general teaching language can be standardized."],"forward_implications":["Labeled data becomes a secondary resource: the primary knowledge source shifts to the experts who hold it, and ML systems can be updated by the people with domain knowledge rather than by ML engineers.","Communities of experts could build and maintain AI systems in the manner of open source software, producing higher quality and resilience through diverse contributors.","ML systems that must handle contradictory expert opinions would be forced to represent disagreement explicitly, making their uncertainty more visible than today's models.","Complex, diverse, and fast-changing domains—health, law, economics, technical support—would become viable targets for intelligent systems where data-hungry deep learning is impractical."],"supporting_citations":[{"why":"Defines the machine teaching paradigm and the focus on the teacher, which the paper builds on for expert-centric ML construction.","marker":"Simard et al., 2017"},{"why":"Coined interactive machine learning and supplied the finding that experts wait under five seconds for re-training, a key requirement.","marker":"Fails and Olsen, 2003"},{"why":"Shows people prefer teaching to training and employ multiple teaching acts, supporting the transparency and multiple-patterns requirements.","marker":"Thomaz and Breazeal, 2008"},{"why":"Demonstrates that users use feedback for multiple overlapping purposes, motivating the need for multiple teaching patterns.","marker":"Stumpf et al., 2009"},{"why":"Provides a formalization of machine teaching as an inverse problem to machine learning, a technical foundation for teaching algorithms.","marker":"Zhu, 2015"},{"why":"NELL is a success case of a machine learning from a large community, supporting the MOLA concept's feasibility.","marker":"Mitchell et al., 2015"},{"why":"Documents the Tay case as evidence of community sabotage risks, grounding the paper's discussion of MOLA dangers.","marker":"Neff and Nagy, 2016"}],"fun_headline_variants":["Teach machines directly: domain experts, not ML experts","Expert-led machine teaching for humane, inclusive AI","Community of experts: the future of machine learning","From labels to concepts: domain experts teach AI"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claim stands on machine teaching technology actually meeting the paper's own requirements—transparency, explainability, multiple teaching patterns, and re-training in under five seconds—so that domain experts can teach without ML intermediaries; the paper itself calls this a formidable challenge for current ML and MT technologies.","fun_headline_variants_meta":{"raw":{"variants":["Teach machines directly: domain experts, not ML experts","Expert-led machine teaching for humane, inclusive AI","Community of experts: the future of machine learning","From labels to concepts: domain experts teach AI"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000222,"raw_usage":{"total_tokens":1411,"prompt_tokens":863,"completion_tokens":548,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":479,"completion_tokens_details":{"reasoning_tokens":488}},"tokens_in":479,"tokens_out":548,"duration_ms":5887,"temperature":1.0,"reasoning_tokens":488,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T12:26:40.893915+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete test: build an expert-centric teaching system for a realistic domain, such as a conversational assistant for diabetic nutrition, and measure whether domain experts can successfully teach it without ML support. If expert-taught performance falls clearly below a conventionally trained baseline, or if experts abandon the tool because re-training exceeds the five-second limit, the central claim weakens. A second test: run a MOLA experiment with a community of experts on a knowledge domain with genuine disagreements and see whether contradictory teachings can be resolved into a coherent, reliable system; persistent incoherence would falsify the community-learning thesis.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Coined interactive machine learning and supplied the finding that experts wait under five seconds for re-training, a key requirement."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Shows people prefer teaching to training and employ multiple teaching acts, supporting the transparency and multiple-patterns requirements."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Demonstrates that users use feedback for multiple overlapping purposes, motivating the need for multiple teaching patterns."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides a formalization of machine teaching as an inverse problem to machine learning, a technical foundation for teaching algorithms."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"NELL is a success case of a machine learning from a large community, supporting the MOLA concept's feasibility."},{"cited_title":"and Nagy, P","cited_arxiv_id":null,"evidence_quote":"Documents the Tay case as evidence of community sabotage risks, grounding the paper's discussion of MOLA dangers."}],"review_version":1}