{"id":"73f66647-8c64-488d-afbe-29ed9a457883","arxiv_id":"2412.18337","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"AI-generated titles boost short-video viewership mainly by filling in missing titles, but human-revised AI titles outperform both pure AI and pure human titles.","lead":"A large field experiment on a short-video platform gave about one million video creators access to AI-generated titles and found that their videos were watched 1.6% more and for 0.9% longer. The gain was much larger when creators used the AI title, and largest when they revised it, pointing to human-AI co-creation as the best approach.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 7.1% LATE is not identified: the paper's own 'inspiration effect' means treatment can affect non-adopters directly, violating the exclusion restriction for the IV in Eqs. (2)-(3).","rationale":"The reader's weakest assumption concerns the Section 5 PSM comparison that drops 51.56% of treated titled videos. That is a real selection concern for the secondary 'AI titles are lower quality' and 'co-creation' claims, but the first-order headline claim of the paper—that AI-generated titles increase viewership—rests on the ITT in Table 4, which is well identified by the randomized treatment assignment. The more load-bearing problem is the LATE in Table 8, which is prominently reported in the abstract and conclusion. The IV exclusion restriction is violated by the paper's own documented 'inspiration effect': non-adopting treated producers who revise AI titles are directly affected by treatment. Because non-adopters are ~76.6% of the treatment group, the Wald estimator can be substantially biased even by small direct effects. This is an internal inconsistency, not a disagreement with consensus, and it directly undermines the 'adoption increases viewership by 7.1%' claim. The ITT remains credible, and the paper's main policy conclusion (AI metadata reduces sparsity and boosts consumption) survives, so the appropriate verdict stays CONDITIONAL. I recommend no change to the reader's verdict, but the LATE should be re-labeled as a reduced-form-plus-direct-effect estimate or re-identified using a design that isolates adoption from inspiration. The proposed test using the algorithmic-failure variation would settle whether the exclusion restriction actually fails.","tokens_in":29141,"tokens_out":5416,"duration_ms":55423,"concrete_test":"Use within-treatment variation in whether the GAI tool actually returned a title (the algorithmic-failure subsample in Appendix B). Among treatment-group videos whose producer did not exactly adopt, compare outcomes for videos where the AI title was displayed versus videos where the AI title was never generated, conditioning on the same controls as Eq. (1). If the displayed-but-not-adopted videos have significantly higher ValidWatch/WatchDuration than the never-generated videos, the exclusion restriction for the IV LATE is violated and the 7.1%/4.1% estimates in Table 8 are biased. The platform's logging of AI-title display would make this test feasible.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline LATE (7.1% valid watches, 4.1% watch duration, Table 8) is estimated by instrumenting Adopt with Treat. This requires that Treat affects outcomes only through exact adoption of the AI-generated title. But the paper's Section 5.2 survey documents an 'inspiration effect': producers who do not exactly adopt the AI title still use it as a creative catalyst to write better titles. These non-adopting treated producers are affected by treatment, so the exclusion restriction fails. With a first stage of only ~23.4% (mean Adopt in treatment = 0.234; control = 0), non-adopters are ~76.6% of the treatment group, so even a modest direct effect on non-adopters is multiplied by ~3.3 in the Wald ratio and can materially inflate the reported LATE. The authors acknowledge the inspiration effect for the co-creation analysis but do not address its implications for the IV validity of the main LATE. The ITT results in Table 4 are not affected by this critique, but the 'when producers adopted' claim in the abstract is not causally identified as stated. Appendix B's exclusion of 51.56% of treated titled videos due to algorithmic issues does not resolve this; it actually provides the within-treatment variation that could test exposure effects.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a large-scale randomized field experiment on a large short-video platform in Asia, in which roughly one million content producers were randomly assigned to receive access to AI-generated video titles on the posting page. The authors estimate that access to AI-generated titles increases valid watches by 1.6% and watch duration by 0.9% (intention-to-treat), and that adoption of AI-generated titles increases these outcomes by 7.1% and 4.1% (local average treatment effect estimated by instrumental variables). They attribute the effect to reduced metadata sparsity (a 41.4% increase in the likelihood of having a title) and to improved user-video matching accuracy, as measured by recommender-system AUC. Section 5 further reports that among videos that already had human titles, access to AI-generated titles is associated with a 37.9% decline in valid watches, but that videos whose producers substantially revised the AI title outperform human-titled controls, supporting a human-AI co-creation narrative.","tokens_in":29447,"tokens_out":5695,"duration_ms":54291,"significance":"If the headline results hold, this is a valuable contribution: it provides rare large-scale experimental evidence on the economic value of AI-generated metadata, a type of AIGC that does not directly engage viewers and whose impact operates through the recommender system rather than through user-facing content quality. The strength of the paper is the clean randomization: the ITT estimates in Table 4 are based on treatment assignment and external viewership outcomes, so they are internally valid. The paper also includes multiple robustness checks, alternative adoption thresholds (Table 28), alternative outcome measures (Tables 23-26), and a mechanism analysis using the platform's own recommendation predictions. However, as detailed below, the LATE and the quality/co-creation comparisons rely on stronger assumptions that are not fully supported by the evidence presented.","major_comments":[{"comment":"The instrumental-variable strategy for the LATE requires that treatment assignment affects viewership only through exact adoption of the AI-generated title. The paper's own survey evidence in Section 5.2 documents an 'inspiration effect' in which treated producers who do not exactly adopt the AI title still use it as a creative catalyst for writing better titles. Such producers are directly affected by treatment, violating the exclusion restriction. With a first-stage adoption rate of only 23.4% (Table 2 and Section 4.2), non-adopters constitute about 76.6% of the treatment group, so even a modest direct effect on non-adopters is amplified by roughly a factor of 3.3 in the Wald ratio. The ITT results are not affected, but the abstract's claim that 'when producers adopted these titles' viewership increased by 7.1% and 4.1% is not causally identified as stated. The authors should either re-estimate the LATE under weaker assumptions (e.g., bounds that allow for direct effects), use the within-treatment variation from algorithmic title-generation failures (Online Appendix B) to test for direct effects on non-adopters, or explicitly relegate the LATE to a secondary, descriptive role.","section":"Section 4.2, Eqs. (2)-(3) and Section 5.2"},{"comment":"The matched-sample analysis drops 2,226,922 treated titled videos (51.56% of the titled treatment group) because AI-generated titles were not produced due to 'algorithmic issues.' The PSM procedure can only balance observed covariates; it cannot address selection on unobservables if title-generation failure is correlated with video complexity, producer effort, or underlying video quality. The post-matching balance tests (Table 17) show very high p-values (e.g., 0.96, 0.99), which is unusual and may reflect over-matching or reduced effective sample size; more importantly, the balance statistics do not speak to unobservable drivers of the outcome. The estimated 37.9% decline in Table 12 and the co-creation gains in Table 14 would be biased if the excluded videos differ systematically from the included ones. The authors should provide evidence on the excluded videos' observable characteristics and, if possible, their viewership outcomes (e.g., by linking algorithmic failure to video features), or present bounding exercises that relax the missing-at-random assumption.","section":"Online Appendix B and Section 5.1 (Tables 12-14)"},{"comment":"The coefficient on Treat in the matched titled-video sample is an intention-to-treat effect among videos that had titles, not the effect of adopting AI-generated titles. In the treatment group, the titled videos include exact adopters, partial revisers, and producers who wrote their own titles after being inspired by the AI suggestion. The paper's interpretation that 'adopting the AI-generated title decreased its viewership' (abstract and Section 5.1) overstates what Table 12 identifies. To support the quality-comparison claim, the authors would need to isolate adoption within the titled subsample (e.g., by instrumenting exact adoption with treatment and estimating a LATE for titled videos) or at least explicitly frame the result as the average effect of access on already-titled videos.","section":"Section 5.1, Table 12"},{"comment":"The AUC comparison uses the platform's recommender-system predictions, and video titles are an input to that recommender (Section 3.1). Because the treatment increases the likelihood that a title exists, the observed improvement in prediction accuracy may partly reflect the recommender having access to more input features rather than a genuine improvement in matching quality. The authors should clarify whether the AUC is computed on out-of-sample predictions from a fixed model or from a retrained model, and should discuss the extent to which the AUC gain is mechanical. The ITT viewership results stand regardless, but the mechanistic claim of 'improved user-video matching accuracy' is not fully separated from the recommender's own dependence on the title feature.","section":"Section 4.3, Table 11"}],"minor_comments":[{"comment":"There is a typo in 'larege-scale experimental evidence' in the final paragraph of the introduction; it should read 'large-scale.'","section":"Section 1"},{"comment":"The experiment is described as running from July 20 to August 21, 2023, but AI-generated titles were only stored between August 8 and August 21 because of technical issues. This discrepancy should be acknowledged and discussed, as it means the pre-treatment period and the start of the treatment period are not contiguous.","section":"Section 3.2 and Section 3.3"},{"comment":"The variable Treatij is defined with both producer and video subscripts, but treatment is assigned at the producer level. Please clarify the unit of treatment and use a consistent notation (e.g., Treati).","section":"Table 1"},{"comment":"The statement that the main effect of Similarity 'would be absorbed by the interaction term (Treati × Similarityij) due to collinearity' is not strictly correct. In ordinary least squares, an interaction term does not absorb a main effect unless the main effect has zero variance in the control group. Please explain the coding of Similarity for control videos (e.g., whether it is set to zero or undefined) and justify the omission of its main effect.","section":"Section 5.1, Eq. (4)"},{"comment":"The conclusion states that the viewership-boosting effect 'was amplified for utilitarian videos and those produced by low-skilled creators,' but the results in Table 6 show a negative interaction for utilitarian-content videos (Treat × Utilitarian = -0.032 for valid watches), meaning the effect is smaller, not larger, for utilitarian videos. This is internally inconsistent and should be corrected.","section":"Section 7 (Conclusion)"},{"comment":"Some post-matching standardized biases remain relatively high (e.g., Experience with %Bias of 7.7 in Table 17 and 25.6 in Table 18, and LowSkill with 9.7 in Table 18). The authors state that the mean differences are no longer statistically significant, but with over a million observations per group, statistical significance is a weak criterion; reporting whether these biases are within conventional thresholds (e.g., <10%) and discussing the implications for the matched-sample estimates would be more informative.","section":"Online Appendix B, Table 17 and Table 18"}],"recommendation":"major_revision","confidential_remarks":"The paper's ITT result is clean, well-powered, and likely publishable in a leading empirical journal. The main concerns are the IV exclusion restriction for the LATE and the selection on unobservables in the matched-sample quality/co-creation analyses. These are fixable in revision by (i) reframing the LATE as secondary and emphasizing the ITT, (ii) providing additional bounds or within-treatment tests using the algorithmic-failure variation, and (iii) clarifying the interpretation of Table 12. The AUC mechanism is interesting but would benefit from a discussion of how much of the improvement is mechanical. I would not reject the paper; the contribution is solid and the issues are addressable within the scope of a revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nQuick read: this is a serious field experiment with a clean intention-to-treat result that will be useful for anyone working on metadata sparsity and recommender systems. The 1.6% valid-watch increase from giving producers access to AI-generated titles is measured on 10.7 million videos with tight standard errors, and the heterogeneity (larger gains for hedonic content and low-skilled producers) is plausible. That part is solid.\n\nThe problem is the \"when producers adopted\" claim. The 7.1% LATE in Table 8 is estimated by instrumenting adoption with treatment, which requires treatment to affect outcomes only through exact adoption. But the paper's own survey documents an inspiration effect: treated producers who do not exactly adopt the AI title still use it as a creative catalyst. That is a direct effect on outcomes through a channel other than the instrumented variable, so the exclusion restriction fails. With a first stage of only 23.4%, roughly three-quarters of treated producers are non-adopters; even a small direct effect on their titles can be scaled up by the Wald ratio and inflate the LATE. The authors acknowledge the inspiration effect in the co-creation section but never reconcile it with the IV assumptions. This is not a nitpick; it cuts the headline.\n\nThe co-creation finding is also fragile, though in a different way. The matched sample drops 51.56% of treated titled videos that never got an AI-generated title due to \"algorithmic issues.\" If those failures correlate with video tractability or unobserved quality, the 37.9% decline from adopting AI titles and the co-creation gains could be selection artifacts. The authors don't test this. The low-similarity PSM in Appendix C is a step, but it still conditions on an endogenous revision choice.\n\nWhat the paper does well beyond the ITT: the recommender-AUC mechanism analysis is a good attempt to open the black box, and the focus on metadata that users never directly see is a genuine contribution. The scale alone is unusual for this literature.\n\nVerdict: the ITT deserves attention, and the paper should go through peer review. But the authors need to either stop headlining the LATE or find a design (e.g., randomized encouragement to revise) that actually identifies adoption effects. As it stands, the abstract overstates what the data identify.\n\nRecommendation: accept for review with major revisions clearly flagged.","headline":"Solid ITT from a huge field experiment, but the headline LATE is not cleanly identified because the paper's own inspiration effect violates the exclusion restriction.","tokens_in":29942,"tokens_out":2493,"would_cite":true,"duration_ms":25130,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that AI-drafted video titles add value by fixing missing metadata: access raised valid watches 1.6% and watch time 0.9%, adoption raised them 7.1% and 4.1%, while already-titled videos lost viewership unless creators…","keywords":["generative AI","video metadata","user-generated content","short-video platform","field experiment","recommender systems","human-AI co-creation","metadata sparsity"],"falsifier":"A follow-up experiment could A/B-test identical videos whose titles are randomly assigned to be fully AI-generated, fully human-written, or a lightly revised AI draft, with no producer choice involved; the valid-watch difference between the first two arms would directly test the claim that AI titles underperform human titles on already-titled videos.","tokens_in":28950,"feed_emoji":"🎬","tokens_out":11022,"duration_ms":93651,"temperature":0.7,"pith_summary":"Generative AI is usually studied as content viewers see, such as ad copy or product descriptions. This paper asks whether AI can create value in metadata that users almost never notice: video titles that feed a recommendation engine. In a randomized field experiment with about two million short-video producers, giving uploaders access to an AI-drafted title raised valid watches (full or sufficiently long plays) by 1.6% and watch time by 0.9%; among producers who adopted the AI title, the increases were 7.1% and 4.1%. The paper attributes most of the gain to filling in missing metadata rather than to any change viewers see. It then argues that AI titles are not generally better than human titles, and that the real upside lies in human-AI co-creation, where producers revise the draft into a more specific, lexically rich title.","feed_headline":"AI titles lift short-video views 7.1% when adopted","feed_subtitle":"Field experiment with about 2 million creators shows gains come from filling missing metadata, and AI-only titles underperform human ones.","key_machinery":"The load-bearing object is the video title as structured metadata in a two-stage recommender system: candidate generation and ranking use title text, together with user profiles and engagement, to decide which videos to surface. The paper's empirical machinery is the randomized offer of an AI-generated title in the posting interface, used both as an intention-to-treat contrast and as an instrument for actual adoption. For the already-titled comparison, the paper relies on propensity-score matching (pairing treated and control videos on observed pre-treatment covariates). Separately, a proprietary log of 93,618,096 recommendation sessions supplies predicted engagement probabilities, letting the paper compare areas under the ROC curve for treated versus control videos as a measure of matching accuracy.","core_discovery":"On a user-generated-content platform, the value of AI-generated metadata is real but conditional. Using random assignment of about two million producers to access to AI-generated titles, the paper estimates that offering the tool increased the share of videos with a title by 41.4% (and tags by 72.4%), that this raised valid watches by 1.6% and watch duration by 0.9% on an intention-to-treat basis, and that adoption of an AI title raised the same outcomes by 7.1% and 4.1%. The effect is stronger among low-skilled producers and hedonic content, exactly the segments with sparse metadata. An analysis of 93.6 million recommendation sessions shows higher recommender AUCs for engagement for treatment videos, supporting the proposed mechanism: the added title text improves user-video matching rather than directly attracting viewers. The paper's second main claim is that, among videos that would have a human title anyway, adopting the AI title reduces valid watches by 37.9% and watch time by 32.6%, while substantial human revision of the AI draft flips the sign, with less similar titles performing better and showing higher lexical density, variation, and entropy.","pith_inferences":["If the paper's mechanism is sparsity relief, the observed 1.6% and 7.1% gains are a one-time catch-up effect: as AI titles become widespread, the marginal benefit of the tool should shrink toward the negative already-titled case.","The co-creation result suggests a testable design principle for other UGC platforms: the economically relevant output of an AI title tool may be the human revision it inspires, not the title itself, so interventions should measure revision rates rather than adoption rates.","The same logic should apply to other sparse metadata fields, such as product descriptions, hashtags, or image captions, but the magnitude will depend on how heavily the platform's recommender weights each field."],"forward_implications":["A platform can increase content consumption without changing what viewers see, purely by reducing metadata sparsity on the supply side.","The gains are largest for low-skilled producers and hedonic videos, so targeting the tool at those segments would capture most of the benefit.","Fully automated AI titles should not be treated as a replacement for human titles; for videos that already have a title, auto-adoption can reduce viewership.","Designing the creator flow to encourage revision of AI drafts, rather than one-click adoption, can convert an average negative effect into a positive one.","Improved matching accuracy should generalize across recommendation channels and to downstream engagement metrics such as likes, shares, and follows."],"supporting_citations":[{"why":"Supplies the data-augmentation and LLM-for-recommendation literature that motivates the claim that enriched metadata improves user-content matching.","marker":"Wei et al. 2024"},{"why":"Provides the large language model creative-work setting and human-AI collaboration comparison that the metadata experiment extends.","marker":"Chen and Chan 2023"},{"why":"Offers the ITT/LATE field-experiment estimation approach the paper mirrors for access versus adoption effects.","marker":"Sun et al. 2019"},{"why":"Describes the two-stage candidate-generation and ranking recommender architecture that the mechanism analysis relies on.","marker":"Davidson et al. 2010"},{"why":"Supplies evidence on generative AI and creativity used to interpret the survey's inspiration effect and the co-creation results.","marker":"Zhou and Lee 2024"},{"why":"Defines the contrast between user-facing AI-generated content and the non-user-facing metadata setting studied here.","marker":"Su et al. 2024"},{"why":"Provides the cosine-similarity text measure used to operationalize how much producers revise AI-generated titles.","marker":"Burtch et al. 2022"}],"fun_headline_variants":["AI titles lift views 7% when adopted, but hurt if humans would write","AI titles boost short-video views 7% when creators adopt","AI titles help videos with sparse metadata, but hurt when humans write","Co-creation with AI beats pure AI or human titles"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper's causal reading of the AI-versus-human-title results assumes that the 51.56% of treatment-group titled videos that received no AI title because of algorithmic issues are missing at random after matching; if failure correlates with hard-to-describe or lower-quality videos, the estimated decline and co-creation gains would partly reflect selection.","fun_headline_variants_meta":{"raw":{"variants":["AI titles lift views 7% when adopted, but hurt if humans would write","AI titles boost short-video views 7% when creators adopt","AI titles help videos with sparse metadata, but hurt when humans write","Co-creation with AI beats pure AI or human titles"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000789,"raw_usage":{"total_tokens":3560,"prompt_tokens":1106,"completion_tokens":2454,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":722,"completion_tokens_details":{"reasoning_tokens":2390}},"tokens_in":722,"tokens_out":2454,"duration_ms":16704,"temperature":1.0,"reasoning_tokens":2390,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T04:47:32.213133+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A follow-up experiment could A/B-test identical videos whose titles are randomly assigned to be fully AI-generated, fully human-written, or a lightly revised AI draft, with no producer choice involved; the valid-watch difference between the first two arms would directly test the claim that AI titles underperform human titles on already-titled videos.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Offers the ITT/LATE field-experiment estimation approach the paper mirrors for access versus adoption effects."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Describes the two-stage candidate-generation and ranking recommender architecture that the mechanism analysis relies on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies evidence on generative AI and creativity used to interpret the survey's inspiration effect and the co-creation results."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the cosine-similarity text measure used to operationalize how much producers revise AI-generated titles."}],"review_version":1}