{"id":"577ba522-1c42-414a-95c3-29ab2c13e746","arxiv_id":"2505.03769","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"Rewritten Reddit titles for shared YouTube videos are associated with higher engagement, especially when they are longer, emotionally charged, and matched to community norms.","lead":"This paper studies whether rewriting the title of a YouTube video before sharing it on Reddit changes how much engagement the post receives. It reports that rewritten titles, especially longer and more emotional ones, are associated with more upvotes, though the causal claim rests on matched observational data.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 0.5-hour pairing window does not actually neutralize the earlier post's extra exposure; because pairs are later re-ordered by score, apparent title effects may be timing effects.","rationale":"The reader's weakest assumption is unconfoundedness after matching, and the concrete threat I identify is the same concern, sharpened to a specific mechanism: residual exposure time within the 0.5-hour window combined with the decision to re-order pairs by engagement. This is load-bearing because the causal claim — \"title rewrites measurably improve engagement\" — depends entirely on the pairing design making title text the only systematic difference. The paper's own Figure 6 shows that timing effects exist, and Section 4.2 explicitly discards the variable that would let the analysis adjust for them. The proposed test is feasible with the existing dataset and would directly separate the timing explanation from the title explanation. I do not think this requires changing the reader's CONDITIONAL verdict: the concern is serious but addressable, and the paper has substantial strengths — a large new dataset, a thoughtful multi-phase pair construction, and a prediction experiment that at least shows title text carries information. However, unless the stratified analysis is run, the controlled experiments do not currently support the causal framing as strongly as the abstract suggests.","tokens_in":22606,"tokens_out":7234,"duration_ms":87017,"concrete_test":"Re-run the Section 4.4 feature comparisons (Table 4 and Table 5) and the Section 5.4 BERT ranking experiment on two strata of the Mix dataset: pairs where the more popular post is the earlier submission, and pairs where the more popular post is the later submission. Also report the fraction of pairs in each stratum and the Post-2-wins ratio within the 0.5-hour window. If the title-feature differences (word count, CTTR, sentiment, etc.) and the BERT accuracy appear only in the earlier-is-popular stratum and vanish or reverse in the later-is-popular stratum, then the headline result is a timing artifact rather than a title effect. If the effects persist in both strata and the Post-2-wins ratio is close to 0.5, the timing concern is empirically refuted.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central causal claim requires that within each matched pair the only systematic difference is the title text. The paper's own Exact Match phase (Section 4.1, Figure 6) shows that earlier posts receive extra exposure, with the later post outperforming the earlier one only about 44% of the time after four hours; a 30-minute window \"diminishes\" this bias but the paper never demonstrates it is zero. Then, in Section 4.2, the authors deliberately discard submission order: \"post pairs are ordered by engagement, with Post 1 being the more popular (Score 1 > Score 2), replacing earlier timing- or random-based assignments.\" Consequently, in every analyzed pair, Group 1 is defined as the higher-scoring post, which will disproportionately be the earlier submission if residual timing effects remain. If earlier posts also tend to have a different title style (e.g., user-written commentary rather than a verbatim copy of the YouTube title), then the reported differences in word count, lexical richness, and sentiment between Group 1 and Group 2 can emerge even if title text has no causal effect on engagement. The additional score-ratio filter (\"at least double the score\") in Section 4.2 selects pairs where any residual timing advantage has materialized, further amplifying this mechanism. The paper never conditions on which post was submitted first in the feature comparisons or in the BERT ranking experiment, so the timing confound is not tested. Unobserved poster identity and thumbnail choice are additional threats, but the timing mechanism alone is sufficient to undermine the causal interpretation.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies whether rewriting the titles of Reddit posts that share YouTube videos affects engagement. The authors build a large Reddit-YouTube dataset, report that 21% of titles are near-copies of video titles, and use a multi-phase matched-pair design (Exact, Similar, and Inverse Match) to compare titles of more- and less-popular posts. They find that more popular posts have longer, more lexically diverse, less neutral, and more community-specific titles. A fine-tuned BERT model predicts the more popular post in pairs with about 74% accuracy, outperforming random, time-based, view-based, and GPT-4o baselines. The paper interprets these results as evidence that title rewrites causally improve engagement and that the controlled dataset isolates text effects.","tokens_in":22885,"tokens_out":5251,"duration_ms":54968,"significance":"If the causal conclusion held, the paper would make a valuable contribution to cross-platform engagement research: it introduces a large dataset, a thoughtful matched-pair framework, and a plausible predictive demonstration that title text carries information about engagement within near-identical sharing contexts. The statistical work is careful in several respects — multiple tests, Bonferroni corrections, effect sizes, and robustness splits are reported. However, the central causal claim is not established by the design. The paper's own figure and baselines show residual timing effects, and the analysis reorders pairs by the outcome variable after pairing, so the reported feature differences and BERT accuracy can arise from exposure time or selection on the outcome rather than from title text. The contribution is still substantial as a descriptive and predictive study, but the manuscript overstates its causal status.","major_comments":[{"comment":"The matched-pair design does not neutralize the earlier post's extra exposure. Figure 6 shows that in Exact Match pairs the later post wins only about 44% of the time after four hours, and the text states that a 0.5-hour window 'diminishes' but does not eliminate this bias. Section 4.2 then explicitly replaces submission-order assignment: 'post pairs are ordered by engagement, with Post 1 being the more popular (Score 1 > Score 2), replacing earlier timing- or random-based assignments.' Because the order is now determined by the outcome, any residual timing advantage will disproportionately assign the earlier post to Group 1. The additional score-ratio filter ('at least double the score') selects pairs in which this residual advantage materialized. As a result, all Group 1-versus-Group 2 feature comparisons and the BERT ranking can reflect posting time even if title text has no effect. The authors should condition on which post was submitted first or provide a sensitivity analysis across time windows and submission-order strata.","section":"§4.1–§4.2"},{"comment":"The feature comparisons are descriptive contrasts between posts selected on the outcome: Group 1 is defined by Score 1 > Score 2 after enforcing a doubling of score and a minimum score difference of 20. Under this design, any title feature that is even weakly correlated with engagement can show a significant pairwise difference, so the t-tests, Wilcoxon signed-rank tests, and McNemar tests do not establish that title rewrites cause engagement. The paper should either present this part as a correlational analysis or use an identification strategy that accounts for the selection rule, such as pairing equally on posting order and applying within-pair causal methods rather than post-hoc ordering by score.","section":"§4.4–§4.5"},{"comment":"The time-based ranking baseline achieves 56.8% on the Date split and 53.3% on average, versus roughly 50% for random guessing. The text interprets this as 'validating our timing controls,' but if the time window had removed timing effects, ranking by earlier submission should perform near chance. The above-chance time baseline is direct evidence that a timing signal remains in the curated dataset, which is especially problematic because Section 4.2 reorders pairs by engagement rather than by submission time. At minimum, the authors should report the time-based baseline on the exact feature-comparison dataset and discuss why a 6.8-point deviation from chance is compatible with the claim that timing is controlled.","section":"§5.4, Table 7"},{"comment":"The manuscript repeatedly describes the design as a 'controlled experiment,' but title assignment is observational; the authors do not randomize which title appears on the earlier versus later post. The Conclusion itself acknowledges that thumbnails, descriptions, and other multimodal elements are not considered, and these are exactly the kind of unobserved confounders that can co-vary with title rewriting. The causal claim that 'title rewrites measurably improve engagement' therefore rests on an unconfoundedness assumption that is neither tested nor convincingly defended. The language should be softened to 'associated with' or the authors should add an explicit sensitivity analysis for unobserved confounders.","section":"Abstract and §4.1"},{"comment":"The Inverse Match phase selects pairs by outcome: it begins with cases where Score 1 > Score 2 and VVR1,2 ≤ 1, i.e., the more popular post links to a less popular video. This is a post-hoc selection on the outcome, not a controlled manipulation. The 'amplified text effects' in Figure 8 therefore may be an artifact of the selection rule: by requiring a large score contrast in a direction opposite to video popularity, the authors mechanically select pairs in which text features (or unobserved variables) happen to align with the score difference. The comparison of T-statistics between Similar and Inverse phases is not a valid demonstration that text effects become stronger when video influence is removed.","section":"§4.1 and §4.5, Figure 8"}],"minor_comments":[{"comment":"The heading 'Inverse Match Phrase' should read 'Inverse Match Phase'; the same misspelling appears in Table 8's column header, where 'Phrase' should be 'Phase.'","section":"§4.1"},{"comment":"The text refers to 'Pair H' and the caption lists pairs (A–H), but the figure displays only pairs A–G; either add Pair H or correct the caption and reference.","section":"§5.6.2 and Figure 9"},{"comment":"The prediction experiments use a 'strict 0.1-hour time window,' whereas Section 4 uses a 0.5-hour window; the manuscript does not explain this discrepancy or report sensitivity of the prediction results to the window choice.","section":"§5.3"},{"comment":"The max-margin loss is written as Loss(x1, x2) = max(0, x2 − x1) after defining xi as the model's score; as written this is minimized when both scores are arbitrarily negative. Please clarify whether a margin term or an additional 'score the opposite pair' term is used, or define the loss with reversed scores.","section":"§5.2, Eq. (1)"},{"comment":"There is a typo in 'gaming subredditsl ike r/Games' — it should read 'gaming subreddits like r/Games.'","section":"§5.6.1"},{"comment":"The sentence 'restricting the dataset to Reddit-YouTube interactions may limits the generalizability' contains a subject-verb agreement error ('may limits' should be 'may limit').","section":"Conclusion"}],"recommendation":"major_revision","confidential_remarks":"The manuscript has a strong empirical component, but the causal framing in the abstract and Section 4 is not supported by the design. A revision that reframes the contribution as a descriptive and predictive study, and that conditions on submission order or adds sensitivity analyses, could make this a solid paper. I would not recommend rejection, because the dataset and prediction results are valuable and the timing concerns are addressable within the manuscript's scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: this is a serious empirical paper with a genuinely new dataset and a thoughtful matched-pair design, but the causal language in the abstract and Section 4 overreaches. The timing confound the authors claim to neutralize is only diminished, not eliminated, and the decision to re-order pairs by engagement in Section 4.2 means residual timing effects can masquerade as title effects.\n\nWhat is actually new: the large integrated dataset (24.1M Reddit posts, 12.6M YouTube videos), the quantification of 21% minimally modified titles, and the three-phase matching including the inverse-match condition. That inverse-match idea—where the more popular post shares a less popular video—is clever and worth building on. The statistical work is mostly careful: multiple tests, Bonferroni correction, effect sizes. And the BERT pairwise ranking at ~74% accuracy is a solid benchmark result that stands on its own as a prediction experiment.\n\nThe soft spots are real but concentrated. The stress-test note is correct: the Exact Match phase itself shows later posts win only ~44% of the time after four hours; a 30-minute window reduces that bias but never demonstrates it is zero. Then in Section 4.2 the authors deliberately discard submission order and re-order each pair by score. So Group 1 is the higher-scoring post, which will be the earlier post whenever residual timing effects exist. If earlier posts also tend to have user-written commentary rather than verbatim YouTube titles—which seems likely, since the first sharer may be more invested—the reported differences in length, lexical richness, and sentiment could emerge even if title text has no causal effect. The score-ratio filter (\"at least double the score\") selects pairs where any timing advantage has materialized, amplifying the problem. The paper never conditions on which post was submitted first in the feature comparisons or the BERT experiment, so this mechanism is not tested.\n\nA few smaller issues: several thresholds (LD=70, VVR range, 0.5-hour window, score ratio) are chosen using the outcome or test accuracy, which is selection on the dependent variable; and no code or data is released, so the numbers cannot be independently checked. Using BERT accuracy as evidence that the curation succeeded is an interpretive leap, but that is minor.\n\nThe central descriptive claim—rewrites correlate with engagement—is plausible and the Section 3 observational analysis supports it. What does not hold up is the causal framing. The paper says \"controlled experiment\" but it is an observational matched-pair study with unobserved confounders.\n\nWho this is for: computational social scientists and content strategists working on cross-platform sharing. It deserves a serious referee and major revision, not a desk reject. If the authors re-run the feature comparisons conditioning on submission order, or reframe the claims as conditional associations, it could be a useful contribution. As written, I would not cite the causal claims, but I would cite the dataset and inverse-match framework.","headline":"A big, careful observational study whose title-effect conclusion is plausible but overstated: the 0.5-hour pairing window does not eliminate the timing confound, and re-ordering pairs by score lets timing masquerade as wording.","tokens_in":23427,"tokens_out":1941,"would_cite":true,"duration_ms":22500,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Rewriting a Reddit title when sharing a YouTube video measurably boosts engagement, and longer, lexically richer, less neutral, community-aware titles win in matched comparisons.","keywords":["cross-platform content sharing","user engagement","title rewriting","Reddit","YouTube","pairwise ranking","BERT","confounding control"],"falsifier":"A matched-pair replication on a fresh sample of Reddit-YouTube posts that applies the same filters but finds no significant difference, or a reversed direction, in the headline title features (length, CTTR/MTLD lexical diversity, neutral-sentiment share) between more and less popular posts would falsify the central claim. More directly, if a randomized field experiment that assigns rewritten versus copied titles to identical videos, subreddits, accounts, and posting times showed no engagement advantage for rewritten titles, the claim would collapse.","tokens_in":22402,"feed_emoji":"📈","tokens_out":5474,"duration_ms":50520,"temperature":0.7,"pith_summary":"This paper claims that when Reddit users share a YouTube video, rewriting the video's title into a new post title measurably improves engagement, even after holding the video, the subreddit, and the posting time fixed. On a large dataset of Reddit posts sharing YouTube videos, it finds that 21% of titles are near-copies of the original video title, and that heavily rewritten titles are associated with higher scores. In a multi-phase matched-pair design, more popular posts consistently have longer, lexically richer, less neutral, and more community-specific titles. A fine-tuned BERT model predicts which of two titles got more engagement with about 74% accuracy, while GPT-4o and time- or view-based baselines stay near random, suggesting the signal is real but subtle. If right, the result means wording alone is a lever for engagement in cross-platform content sharing.","feed_headline":"Longer, richer Reddit titles beat copied YouTube ones","feed_subtitle":"Matched-pair analysis controls for video, time, and subreddit; wording alone shifts engagement.","key_machinery":"The load-bearing object is the multi-phase matched-pair experiment. Posts are paired within the same subreddit: Exact Match pairs share the identical video, Similar Match pairs share videos of comparable view counts (view ratio 0.5–2), and Inverse Match pairs reverse the expected popularity–views relationship. Additional filters restrict pairs to a 0.5-hour posting window, require meaningful title difference via Levenshtein distance thresholds, and demand a doubled score gap. The paper then runs paired t-tests, Wilcoxon signed-rank tests, and McNemar tests on five feature families (structural, lexical, stylistic, readability, sentiment), and validates learnability with a pairwise BERT ranking model.","core_discovery":"The central discovery is that title rewrites measurably improve engagement and that the improvement has a describable text profile. In matched pairs where subreddit, video identity or view similarity, and posting time are controlled, the more popular post typically has a longer title, richer vocabulary as measured by length-adjusted indices, stronger (less neutral) sentiment, and language that resonates with the subreddit's norms. The effect grows in the inverse-match condition, where the more popular post is linked to the less popular video, which the paper reads as evidence that text becomes decisive when video popularity cannot explain the outcome. The paper also shows that a context-aware model trained on these controlled pairs can rank titles by engagement at 74% accuracy, while general-purpose LLM evaluation is near chance, indicating the patterns are learnable but not trivial.","pith_inferences":["The causal read rests on unconfoundedness after matching; titles are user-chosen, not randomized, so hidden differences like poster identity, thumbnail choice, or the earlier post's extra exposure could partially explain the score gaps, and I would expect the true effect size to be smaller than the raw association.","A natural next test is to apply the same pairing design to other platform pairs (e.g., news headlines shared on Twitter/X, TikTok captions) to see whether the same title-feature profile holds or is Reddit-specific.","The large BERT-versus-GPT-4o gap suggests a concrete probe: adversarially edit titles in matched pairs (lengthen, de-neutralize, add community keywords) and check whether the model's predicted ranking flips in the same direction as the observed scores; that would test whether the model learned the paper's stated features or something correlated.","If the finding is robust, one testable extension is an actual A/B test on a platform that permits randomized headline assignment, comparing copied titles against community-matched rewrites with identical timing, video, and account."],"forward_implications":["If wording alone shifts engagement, then low-cost title rewriting is a concrete optimization lever for anyone sharing video content across platforms.","Effective titles are not simply \"simpler\" or \"shorter\": the matched data point to informative, lexically rich, emotionally non-neutral titles, so content strategies should track these features.","Community norms gate the effect: a title that works in r/SquaredCircle or r/kpop is specific to that audience, so global \"best title\" rules will underperform subreddit-aware ones.","Because GPT-4o and simple baselines fail at the pairwise task while BERT succeeds, engagement prediction from titles is a learnable but non-obvious task, which implies automated title-recommendation systems may need fine-tuned models rather than general LLM prompts.","The matched-pair framework provides a template for isolating textual effects in other cross-platform settings, such as Twitter/X or TikTok sharing."],"supporting_citations":[{"why":"Provides the pairwise ranking task and evidence that titles alone can predict Reddit engagement, which the paper's prediction setup builds on.","marker":"[75]"},{"why":"Establishes pairwise ranking that controls for timing and community, the method the paper adapts for its matched-pair comparisons.","marker":"[28]"},{"why":"Shows wording effects on message propagation using topic- and author-controlled natural experiments, motivating the controlled comparison of text.","marker":"[68]"},{"why":"Analyzes the interplay between titles, content, and communities and finds timing effects often overshadow titles, which the Exact and Similar phases directly address.","marker":"[40]"},{"why":"Uses propensity score matching to study headline editing effects on engagement, a causal-inference approach the paper extends to cross-platform video sharing.","marker":"[52]"},{"why":"Supplies BERT, the contextual model that achieves the paper's main prediction results.","marker":"[15]"},{"why":"Supplies VADER, the social-media-tuned lexicon used for continuous sentiment features of titles.","marker":"[34]"}],"fun_headline_variants":["Title rewrites boost Reddit engagement in matched tests","Rewrite beats copy: Reddit titles that engage","BERT ranks title rewrites better than GPT-4o at 74%","Matched-pair test: rewording titles matters on Reddit","Longer, richer titles win on Reddit via rewrites"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"Once subreddit, video identity or similar view counts, and a half-hour posting window are held fixed, the remaining popularity difference within each pair must be driven by the title text alone and not by the earlier post's extra exposure, the poster's identity, the thumbnail, or any other unobserved user choice.","fun_headline_variants_meta":{"raw":{"variants":["Title rewrites boost Reddit engagement in matched tests","Rewrite beats copy: Reddit titles that engage","BERT ranks title rewrites better than GPT-4o at 74%","Matched-pair test: rewording titles matters on Reddit","Longer, richer titles win on Reddit via rewrites"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000824,"raw_usage":{"total_tokens":3588,"prompt_tokens":911,"completion_tokens":2677,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":527,"completion_tokens_details":{"reasoning_tokens":2592}},"tokens_in":527,"tokens_out":2677,"duration_ms":20171,"temperature":1.0,"reasoning_tokens":2592,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T10:03:41.798673+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A matched-pair replication on a fresh sample of Reddit-YouTube posts that applies the same filters but finds no significant difference, or a reversed direction, in the headline title features (length, CTTR/MTLD lexical diversity, neutral-sentiment share) between more and less popular posts would falsify the central claim. More directly, if a randomized field experiment that assigns rewritten versus copied titles to identical videos, subreddits, accounts, and posting times showed no engagement advantage for rewritten titles, the claim would collapse.","supporting_citations":[{"cited_title":"Cats and Captions vs. Creators and the Clock: Comparing Multimodal Content to Context in Predicting Relative Popularity","cited_arxiv_id":"1703.01725","evidence_quote":"Establishes pairwise ranking that controls for timing and community, the method the paper adapts for its matched-pair comparisons."},{"cited_title":"Lakkaraju, Jim Mcauley, and Jure Leskovec","cited_arxiv_id":null,"evidence_quote":"Analyzes the interplay between titles, content, and communities and finds timing effects often overshadow titles, which the Exact and Similar phases directly address."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Uses propensity score matching to study headline editing effects on engagement, a causal-inference approach the paper extends to cross-platform video sharing."}],"review_version":1}