{"id":"acfad339-afa2-4b00-9b80-bf694643efdf","arxiv_id":"2412.07066","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"AI model terms of use are likely largely unenforceable because model weights and outputs lack copyright protection, and other legal doctrines do not fill the gap.","lead":"This legal analysis argues that AI model terms of use, especially restrictions on model weights and outputs, are probably unenforceable because the underlying artifacts are not protected by copyright. A generalist should read it because it challenges the common assumption that licenses can control harmful or competitive uses of AI, with consequences for AI safety policy and competition.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central 'mirage' claim depends on resolving the ProCD/Genius circuit split; in ProCD jurisdictions clickwrap terms over uncopyrightable data are commonly enforced, so the paper's conclusion is overbroad unless limited to preemption-friendly circuits.","rationale":"The reader's weakest-assumption analysis identifies the same concern I would stress-test: the article's central conclusion depends on courts adopting Genius- and X Corp.-style copyright preemption of contract claims over uncopyrightable AI outputs and weights, despite ProCD's no-preemption rule in several circuits. The paper is unusually candid about this dependency—Part III.C.3 explicitly notes that 'if Judge Alsup's conflict preemption analysis takes hold, even responsible AI terms may face strong preemption challenges'—but the title and abstract generalize the 'mirage' claim beyond that conditional. Because the article is a law-review-style doctrinal argument, the unresolved circuit split is the most load-bearing point: it determines whether the enforceable core of AI ToS is nearly empty or substantially intact. I considered whether the copyrightability of model weights is an even more fundamental concern. It is a real and unsettled question, but the paper's analysis there is more robust and heavily hedged, and a court would have to reject its human-authorship and functionality arguments before that concern becomes decisive. The preemption split is nearer and is explicitly acknowledged. The reader's CONDITIONAL verdict already captures this uncertainty, so my stress-test does not change the verdict. The concrete test—a systematic coding of post-ProCD preemption outcomes, or a focused brief on OpenAI's no-compete clause under Seventh Circuit law—would directly test whether the paper's unqualified conclusion survives in ProCD jurisdictions.","tokens_in":44352,"tokens_out":4580,"duration_ms":58465,"concrete_test":"Conduct a systematic doctrinal check: code every post-ProCD decision in the Seventh Circuit and other ProCD-following circuits that addresses copyright preemption of a clickwrap or browsewrap contract claim over uncopyrightable data (e.g., database or scraping cases), recording whether the contract was enforced. If a majority enforce without a copyright predicate, the article's 'mirage' conclusion must be narrowed to Second/Sixth/Ninth-style preemption jurisdictions. For a sharper test, brief the specific OpenAI 'no use output to train competing models' clause under Seventh Circuit law; if it survives a Rule 12(b)(6) preemption defense, the paper's central claim fails in that circuit.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's headline conclusion—that AI terms of use are largely unenforceable because model weights and outputs are uncopyrightable—does not follow unless contract claims over uncopyrightable subject matter are preempted. That is exactly the unresolved ProCD v. Zeidenberg / Genius / X Corp. v. Bright Data circuit split. In the Seventh Circuit and others following ProCD, a two-party clickwrap agreement supplies an 'extra element' that saves a state contract claim from preemption even when the underlying data is uncopyrightable; ProCD itself enforced a shrinkwrap restriction over a noncopyrightable database. The paper acknowledges this at Part III.C.1 and III.C.3 ('if Judge Alsup's conflict preemption analysis takes hold'), but its title, abstract, and policy recommendations state the mirage conclusion without a geographic qualifier. If a court in a ProCD-following circuit enforces an anti-competition or responsible-use clause as an ordinary contract, the enforceable core of AI ToS is substantially larger than the article claims. The empirical observation that no model creator has yet sued is also equivocal: it is consistent with the paper's legal conclusion, but equally consistent with companies avoiding adverse precedent. The article is carefully hedged and internally consistent, but the central claim is contingent on a doctrinal development that has not occurred. This is not a fatal flaw, but it is the load-bearing uncertainty.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper argues that the restrictive terms of use attached by AI model creators to model weights and model outputs are largely unenforceable. It develops three claims: first, model weights and model outputs are generally not copyrightable, so there is no copyright-based licensing hook; second, state-law contract and tort claims built on those restrictions face serious copyright preemption problems under recent cases like Genius and X Corp. v. Bright Data, while DMCA, CFAA, and trespass-to-chattels theories offer little recourse; and third, anti-competitive restrictions are especially vulnerable to preemption, misuse, and antitrust doctrine, whereas narrow responsible-use restrictions have a better, though still uncertain, chance of survival. The paper recommends that policymakers rely on statutory regulation rather than private licensing terms as the primary mechanism for controlling harmful uses of AI.","tokens_in":44655,"tokens_out":3983,"duration_ms":49277,"significance":"If the paper's central conclusion is correct, it would substantially narrow the enforceable core of AI terms of use, with direct implications for the California AI Transparency Act, NTIA policy discussions, open-weight model licensing, and the broader debate about private ordering as a substitute for AI regulation. The paper is valuable for its systematic survey of current model-provider terms, its careful doctrinal mapping of the preemption landscape, and its willingness to engage with contrary authority such as ProCD and the Seventh Circuit's approach. Its strengths are institutional rather than formal: it is a well-hedged, internally consistent legal analysis that identifies concrete test cases and policy levers. It does not claim machine-checked proofs or quantitative results; its contribution is doctrinal synthesis and policy argument.","major_comments":[{"comment":"The article's headline conclusion that AI terms of use are a \"mirage\" is stated without a geographic qualifier in the abstract, the introduction, and the conclusion, but the paper's own doctrinal analysis shows that the outcome depends on an unresolved circuit split. In ProCD v. Zeidenberg jurisdictions, a two-party clickwrap agreement supplies an \"extra element\" that saves a state contract claim from section 301 preemption even when the underlying data is uncopyrightable, and the paper concedes at Part III.C.3 that \"if Judge Alsup's conflict preemption analysis takes hold, even responsible AI terms may face strong preemption challenges.\" Because the article does not argue that ProCD is wrongly decided or inapplicable to AI terms of use, the enforceable core of AI terms is substantially larger if ProCD controls. This is load-bearing: the title, abstract, and policy recommendations should be conditioned on the preemption-friendly circuits, or the article should make an affirmative doctrinal argument for why AI terms should follow Genius and X Corp. rather than ProCD.","section":"Part III.C.1 and III.C.3"},{"comment":"The paper repeatedly relies on the observation that \"no model creator has actually tried to enforce these terms with monetary penalties or injunctive relief\" as circumstantial support for its legal conclusion (Introduction; Part I.D). This inference is equivocal. The absence of litigation is equally consistent with companies avoiding adverse precedent, settling quietly, relying on account suspensions and technical access control, or deciding that the reputational costs of suing researchers and users outweigh the benefits. The paper should either present the non-enforcement observation as a neutral motivating fact or identify what additional evidence would distinguish these explanations. As written, the passage risks treating the very doctrinal uncertainty the article identifies as if it were already a settled judicial rejection of enforcement.","section":"Introduction and Part I.D"},{"comment":"The discussion of \"system code\" dismisses the copyright hook too quickly. The paper acknowledges that inference-system code is likely copyrightable, but argues that users can simply switch to interchangeable software and that reverse engineering is protected by fair use. That response does not address the common distribution channel in which the model creator distributes the weights and the inference code together in a single package, and the license is presented as a condition on downloading that package. A user who copies the creator's code, rather than independently written compatible code, may be bound by copyright conditions associated with that code even if the weights themselves are uncopyrightable. The article should explicitly state whether, and why, this route is unavailable for the specific terms surveyed in Part I, or should narrow the \"no licensing hook\" conclusion accordingly.","section":"Part II.D"}],"minor_comments":[{"comment":"The phrase \"systematically assesses of the enforceability\" contains a typo; it should be \"systematically assesses the enforceability.\"","section":"Abstract"},{"comment":"The sentence ending \"illegality amount machine learning practitioners\" appears to be a typographical error; it should likely read \"among machine learning practitioners.\"","section":"Introduction"},{"comment":"There is a duplicated word: \"If enforced, these terms would would:\" should be \"would:\".","section":"Part III.C.3"},{"comment":"Footnote 308 contains an incomplete citation: \"cite statute\" appears where the Defend Trade Secrets Act reverse-engineering provision should be cited, and the following sentence \"Contra Bowers 320 F.3d at 1317\" is too cryptic to be useful.","section":"Part IV.A"},{"comment":"The appendix references four figures with captions, but the images themselves are not visible in the manuscript text; please confirm that the figures render in the submission format.","section":"Appendix"}],"recommendation":"major_revision","confidential_remarks":"This is a law review article submitted to a cs.CY venue; the fit with the journal's scope is worth confirming with the editor. The central legal claim is contingent on an unresolved circuit split, and I would recommend that the editor seek a copyright-law specialist as a second reviewer if one has not already been consulted. The paper is otherwise well researched and carefully argued, and the requested revisions are within the manuscript's own frame rather than requiring new empirical work."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague—\n\nThis is a genuinely useful piece, and the best systematic treatment I know of the enforceability of AI model terms of use. It does three things well: it lays out the copyright problem for model weights and outputs; it connects recent preemption cases (Genius, X Corp. v. Bright Data) to AI licensing; and it distinguishes anti-competitive terms from responsible-use terms, arguing the former are least likely to survive. The policy recommendation—statutes, not private fiat—is clear and properly argued. The paper is honest about uncertainty, citing contrary authority and noting circuit splits.\n\nThe soft spot is real but not fatal: the headline “mirage” conclusion depends on courts adopting the Second/Ninth Circuit preemption approach rather than ProCD. The paper acknowledges this in Part III.C, but the title and abstract sweep too broadly. In ProCD jurisdictions, clickwrap terms over uncopyrightable data are routinely enforced, so the enforceable core of AI ToS may be larger than the paper implies. The empirical point—that no model creator has yet sued—is equivocal; it is consistent with weak enforceability but also with companies avoiding adverse precedent. That should have been stated more carefully.\n\nThe article is not a formal or empirical result; it is a doctrinal synthesis. That is fine for what it is. The citations to the authors' own prior work are not problematic; those works stand independently. No equations, no fitted parameters, no circularity concerns.\n\nWho is this for? Anyone working on AI governance, open-weight licensing, or platform law. It deserves a serious referee—the analysis is careful even where the conclusion is contingent. My own preference would be to require the authors to qualify the abstract and conclusion to say “in preemption-friendly circuits” or “likely,” rather than asserting a global mirage. But as a piece of law review scholarship, it clears the bar.\n\nVerdict: conditional accept for peer review after revision.","headline":"A careful, genuinely useful synthesis of AI ToS enforceability; the mirage metaphor overstates a circuit-dependent conclusion, but the analysis is honest and should be published after qualification.","tokens_in":45118,"tokens_out":1333,"would_cite":true,"duration_ms":16419,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"AI model terms of use are largely unenforceable because the weights and outputs they protect are not copyrightable, and copyright preemption blocks contract claims that try to create exclusive rights where none exist.","keywords":["AI terms of use","copyright preemption","model weights","model outputs","open-weight models","responsible AI licensing","generative AI","contract enforceability"],"falsifier":"A federal court in a circuit that follows ProCD (for example, the Seventh Circuit) would falsify the paper's central claim if it enforced, with damages or an injunction, a terms-of-service clause forbidding a user from training a competing model on uncopyrightable AI outputs—without finding any copyright in the outputs or weights.","tokens_in":44198,"feed_emoji":"⚖️","tokens_out":6778,"duration_ms":67947,"temperature":0.7,"pith_summary":"The paper argues that the restrictive terms of use attached to AI models and their outputs—bans on training competing models, anti-scraping clauses, and \"responsible use\" restrictions—are mostly a legal mirage. AI companies likely own no copyright in model weights or model outputs, so there may be nothing to license and no infringement claim to condition on the license. Recent copyright preemption decisions indicate that contract claims seeking to control copying of this material may be preempted, while other statutes like the DMCA and the CFAA offer little recourse. If the paper is right, companies and policymakers who treat these terms as enforceable tools for preventing misuse are relying on a house of cards; the better route is statutory regulation of harmful uses, not private fiat.","feed_headline":"AI terms of use likely can't be enforced in court","feed_subtitle":"Without copyright in weights or outputs, license terms that try to control copying are preempted.","key_machinery":"The load-bearing mechanism is the combination of two doctrines. First, the human-authorship and functionality doctrines of copyright law remove model outputs and model weights from the set of protectable works, so there is no underlying exclusive right for a license to condition. Second, copyright preemption—both express preemption under § 301(a) and the recently revived conflict preemption from X Corp. v. Bright Data—converts that absence of copyright into a bar on state contract and tort claims that would give the same control over copying that copyright would have given had it existed. The paper's analysis pivots on how courts resolve the split between ProCD, which treats a two-party contract as an extra element that avoids preemption, and Genius/X Corp., which preempts contracts that protect copying of uncopyrightable material.","core_discovery":"The central discovery is that the enforceability of AI terms of use collapses because the two artifacts those terms purport to protect—model weights and model outputs—sit largely outside copyright law. Model outputs are generated by automated processes and accordingly fail the human-authorship requirement under Copyright Office guidance and recent case law; model weights are similarly machine-produced and are functional artifacts that § 102(b) excludes from protection. With no copyright in the underlying asset, the traditional mechanism for enforcing software licenses—conditioning a copyright grant on compliance—does not work, and contract claims that try to recreate copyright-like control over uncopyrightable material increasingly run into express and conflict preemption. The paper therefore concludes that anti-competitive restrictions are least likely to survive, that even narrow responsible-use terms may be preempted if Judge Alsup's conflict-preemption analysis takes hold, and that terms do not bind third parties who obtain outputs from the original user.","pith_inferences":["The authors leave implicit that the practical gap is largest for open-weight releases: closed API providers can still punish misuse through account revocation and access control even if their terms could not be enforced in court.","A testable extension of the paper's logic is to track the first litigated cases enforcing an output anti-distillation clause: if those cases are withdrawn, settled quietly, or dismissed on preemption grounds, the \"mirage\" claim is strengthened.","If the preemption reasoning generalizes, the same logic would undercut restrictive terms attached to other uncopyrightable public data, making it harder for platforms to use contract law to control scraping of facts and user-generated content.","The paper implies that the debate over whether restrictive open-weight licenses are \"truly open source\" is somewhat beside the point: if the weights are uncopyrightable, the licenses are legally hollow regardless of their classification."],"forward_implications":["If the paper is right, \"don't train a competing model on our outputs\" clauses are the most vulnerable terms, because they are closest to reproducing copyright's exclusive rights and can also be attacked as copyright misuse or anticompetitive.","Responsible-use restrictions will survive only when they are specific, focus on the purpose of use rather than copying, and target conduct with independent public-policy salience; vague or copying-centered restraints will face serious preemption challenges.","Even enforceable terms would not bind third parties who receive model outputs from the original user, because contracts do not run with uncopyrightable information.","Open-weight model licenses cannot rely on the copyleft enforcement mechanism of open-source software, since that mechanism depends on a copyright interest the model creator does not have.","Policymakers relying on mandated license terms, such as watermark-preservation duties in the California AI Transparency Act, should expect weak or nonexistent private enforcement and should legislate prohibited uses directly."],"supporting_citations":[{"why":"Seventh Circuit precedent that shrinkwrap contract terms are not preempted by copyright; the paper identifies it as the principal doctrinal obstacle to its preemption argument.","marker":"ProCD v. Zeidenberg"},{"why":"Second Circuit decision preempting a breach-of-contract claim that tried to control scraping and reproduction of content the plaintiff did not own; the paper uses it as the closest analog to anti-distillation restrictions.","marker":"Genius"},{"why":"District court decision holding that adhesive terms of use cannot create a private copyright regime over uncopyrightable public data; the paper's main conflict-preemption hook.","marker":"X Corp. v. Bright Data Ltd."},{"why":"Agency guidance stating that purely AI-generated material is not registrable because it lacks human authorship; supplies the premise that model outputs are largely uncopyrightable.","marker":"Copyright Office AI Authorship Guidance"},{"why":"Federal court decision refusing copyright registration for an autonomous AI-generated work; reinforces the no-copyright-in-outputs premise.","marker":"Thaler v. Perlmutter"},{"why":"Case establishing that open-source license terms can be enforced as copyright conditions; the paper argues this mechanism fails when the licensed weights and outputs are uncopyrightable.","marker":"Jacobsen v. Katzer"},{"why":"Supreme Court narrowing of the CFAA's \"exceeds authorized access,\" which the paper relies on to reject CFAA claims based on terms-of-service violations.","marker":"Van Buren v. United States"}],"fun_headline_variants":["AI terms of use: legally toothless without copyright","Why AI license limits likely won't hold in court","No copyright in weights, so AI terms are unenforceable","AI's machine-made outputs escape copyright, gutting contracts","AI terms of use: a mirage that private fiat can't fix"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire conclusion rests on courts adopting broad copyright preemption of contract claims over uncopyrightable outputs, so if ProCD's no-preemption approach dominates instead, the same clickwrap terms could be enforced even without copyright.","fun_headline_variants_meta":{"raw":{"variants":["AI terms of use: legally toothless without copyright","Why AI license limits likely won't hold in court","No copyright in weights, so AI terms are unenforceable","AI's machine-made outputs escape copyright, gutting contracts","AI terms of use: a mirage that private fiat can't fix"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000213,"raw_usage":{"total_tokens":1462,"prompt_tokens":1028,"completion_tokens":434,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":644,"completion_tokens_details":{"reasoning_tokens":349}},"tokens_in":644,"tokens_out":434,"duration_ms":5163,"temperature":1.0,"reasoning_tokens":349,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T19:10:31.541856+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A federal court in a circuit that follows ProCD (for example, the Seventh Circuit) would falsify the paper's central claim if it enforced, with damages or an injunction, a terms-of-service clause forbidding a user from training a competing model on uncopyrightable AI outputs—without finding any copyright in the outputs or weights.","supporting_citations":[],"review_version":1}