{"id":"c10e0bce-325b-400c-baf8-a1fbc2fa1826","arxiv_id":"1908.08622","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":3,"one_line_summary":"A temporal analysis of Facebook brand pages proposes posting schedules that select high-reaction time buckets, with the best schedule reporting a 7x reactions-per-post lift over the page average in the training data.","lead":"Researchers analyzed 100 Facebook brand pages over five years and proposed six posting schedules based on when posts receive the most comments. The most aggressive schedule is reported to yield seven times more audience reactions than the average, but this gain is measured on the same data used to construct the schedule.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No independent evidence for 7x: schedule and gain are computed from the same 5-year counts, with no holdout, no content-type control, and no denominator floor.","rationale":"The reader's REJECT verdict is supported by the evaluation design. The strongest claim is a 7x reaction gain, but the gain is measured on the same data used to select the schedule, and no temporal holdout, page-level cross-validation, or confounder control is provided. My stress-test adds two technical specifics that make the problem concrete rather than merely a general concern about overfitting. First, because RG is a ratio of bucket-level reactions per post to the global average, buckets with very few posts can have enormous variance; the paper does not report the number of posts in the winning bucket, so a single high-reaction post could drive the headline number. This is an internal statistical robustness problem, not just a generalizability question. Second, the paper's own Table 5 demonstrates that content type is strongly associated with reactions, and the schedule derivation does not condition on content type, page size, or page identity. If posting time is correlated with content type, Eq. (14) does not identify a causal time-of-day effect. The reader's weakest assumption about the absence of a holdout captures the same fundamental issue, and I agree with that framing. The verdict should remain REJECT because the central claim is not supported by the evaluation as presented, even though the descriptive temporal patterns may be useful as exploratory findings. A concrete temporal holdout with content-type stratification would settle whether the 7x figure survives out-of-sample.","tokens_in":14883,"tokens_out":3546,"duration_ms":40899,"concrete_test":"Split the dataset by time: derive each category's SCF R and SWCF R schedules and top-30 buckets from 2011–2013 posts and reactions only (Eqs. 6 and 11), then on the held-out 2014–2015 posts compute the mean reactions per post for posts in recommended buckets versus all other posts, bootstrapping by page and separately within each content type. If the held-out ratio is not robustly above 1, the 7x gain is an in-sample selection artifact.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claim rests on Eq. (14): RG(C_i,k) = δ(C_i,k)/ω(C_i), the ratio of reactions per post in a 15-minute bucket to the global average. The categorized frequent reaction schedule (Eq. 6) ranks buckets by total reactions R_k(C_i), and its top bucket is then reported as having RG near 7. Because the schedule, the bucket ranking, and the gain are all computed from the same aggregated R_k and M_k over the same 2011–2015 data, the reported value is an in-sample maximum with no split, permutation, or confidence interval. Three concrete failure modes are live. First, small-denominator buckets: a time bucket with few posts but one or two unusually successful posts can produce a spuriously high reactions-per-post ratio; the paper gives no minimum post count, variance estimate, or significance test for the top bucket. Second, content confounding: Section 6.5 shows content type strongly affects reactions (links are 78.6% of posts but only 54.16% of reactions, while videos are substantially more engaging per post); if admins post certain content types at certain times, the apparent time effect is not identified as a posting-time effect. Third, page and audience drift: pooling five years of data can make the selected bucket reflect past page behavior rather than a stable property of that time slot. Additionally, Section 2.2 states that only comment timestamps were accessible, so the outcome is comments, not the full set of likes, comments, and shares promised in the abstract. The descriptive daily, weekly, and monthly patterns are plausible, but they do not provide out-of-sample support for the 7x claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper addresses when brand-page administrators should post on Facebook to maximize audience reactions. Using Graph API data from 100 pages (20 each in e-commerce, traffic, telecommunication, hospital, and politics) over 2011–2015, with roughly 0.3 million posts and 10 million comment timestamps, the authors define two aggregate and four category-based 'schedules' (frequent posting, frequent reaction, and weighted variants). Pages are clustered into five categories by k-medoid on reaction-profile similarity using three selected features, and schedules are ranked over 96 fifteen-minute buckets. The headline claim is that the categorized frequent reaction schedule achieves a seven-fold reaction gain over the average reaction rate (Section 6.2); the paper also reports daily, weekly, and monthly audience reaction patterns and shows that content type correlates strongly with reactions (Section 6.5).","tokens_in":15185,"tokens_out":15936,"duration_ms":157271,"significance":"If properly established, the headline result would offer a directly actionable, low-cost tool for brand-page managers, and the large five-year dataset would make the empirical patterns (84% of reactions within 24 hours; diurnal, weekly, and seasonal rhythms; content-type effects) a useful reference for the community. The formal problem statements and explicit schedule equations (Eqs. 2–14) are clear and would be straightforward to re-implement, which is a genuine strength. However, the significance of the central contribution hinges entirely on the validity of the reaction-gain evaluation, and for the reasons in the major comments—the misalignment of reaction-arrival and post-creation buckets, the in-sample nature of the 7x figure, and the absence of confound controls—the paper as it stands does not establish the headline claim; a corrected analysis could plausibly yield a much smaller gain.","major_comments":[{"comment":"The reaction-gain metric is misaligned with the claims it is used to support. R_k(C_i) is defined (Section 5) as reactions whose timestamps fall in bucket k, while M_k(C_i) is the number of posts created in bucket k. Since the paper itself shows that 84% of reactions arrive within 24 hours of posting (Figure 1), most reactions in bucket k are responses to posts created in earlier buckets, not to posts created in the same bucket. The ratio δ(C_i,k)=R_k/M_k therefore does not measure “reactions received when the category posts in time bucket t_k” as claimed in Section 6.1.1; a bucket with high incoming reaction volume but low posting volume will exhibit a large RG even if posting time has no causal effect on engagement. The central 7x figure is consequently not evidence for the posting-time recommendation. The evaluation should instead compute, for each bucket, the average number of reactions received per post created in that bucket (for example within a fixed 24-hour horizon), which the dataset’s post and comment timestamps can support.","section":"§5, §6.1.1, Eqs. (12)–(14)"},{"comment":"The reported 7x gain is an in-sample maximum with no validation. The schedule ranking (Eq. (6) or Eq. (11)) and the reaction gain (Eq. (14)) are both computed from the same cumulative counts R_k(C_i) and M_k(C_i) over the same 100 pages and the same 2011–2015 years. No temporal holdout, page-level cross-validation, bootstrap confidence interval, or permutation test is provided, and no minimum-post-count floor is imposed on δ(C_i,k). Selecting the best of 96 buckets on the basis of its historical per-bucket ratio and then reporting that same ratio as the achievable gain is an upward-biased estimate of the best bucket’s true effect. As written, the paper provides no evidence that the 7x figure would transfer to future posts or to new pages.","section":"§6.2, Eqs. (6), (11), (14)"},{"comment":"Content type is a strong potential confounder for the posting-time effect. Table 5 shows that links constitute 78.6% of posts but yield only 54.16% of reactions, while videos yield more than twice their share of reactions per post. If the content-type mix varies across time buckets (for instance, news links in the morning and videos in the evening), the apparent schedule gains are not identified as posting-time effects. The evaluation contains no control for content type, page size, or page identity, and the abstract’s claim that the posting time is “derived taking other factors into account” is not substantiated by the reported analysis.","section":"§6.5, Table 5"},{"comment":"The outcome variable is comments only, not “likes, comments, shares.” Section 2.2 states that only comment timestamps were accessible and that comments are used as the reaction variable, so the “10 million audience reactions” in Table 1 are 10 million comments. The Abstract and Introduction nevertheless frame the contribution in terms of audience reactions generally (“in the form of likes, comments, shares, etc.”). This mismatch means the headline claims exceed the measured outcome; either the claims should be restricted to comments or additional reaction types must be included.","section":"§2.2 vs. Abstract/§1"},{"comment":"The categorization is validated using the same reaction-profile vectors that define the clustering objective, so the reported within- and across-category correlations largely restate the clustering criterion. K-medoid similarity is defined as the correlation between pages’ cumulative reaction profiles R_k (Section 4.3, cf. Eq. (16)), and the “effectiveness” of the categorization is then measured as the correlation of these same profiles within and across clusters (Tables 3–4). High within-category and low across-category correlation are therefore partly by construction, not independent evidence that the categories are meaningful. An external validation—for instance, association of the learned clusters with the pages’ declared labels on held-out pages, or out-of-sample classification—is needed.","section":"§4.3, §6.3, Tables 3–4"}],"minor_comments":[{"comment":"The phrase “around 10 AM to 12 AM” for the telecommunications category should presumably read “10 AM to 12 PM”; as written it describes a period ending after midnight.","section":"§6.4.1"},{"comment":"The expression “in the tth_k bucket” appears to be a typesetting artifact; it should read “in the t_k-th bucket.”","section":"§5.1, Eq. (2)"},{"comment":"The number of buckets retained per schedule (“top-30”) is arbitrary; the paper should state how a page admin is expected to use the ranked list in practice (how many slots per day, and with what spacing).","section":"§6.2"},{"comment":"The plot should include a clear legend for the six schedules (SAFP, SCFP, SWCFP, SAFR, SCFR, SWCFR), since the text refers to each by acronym.","section":"Figure 3"},{"comment":"The dataset’s geographic location is not stated, although the monthly interpretations (elections, festive sales) are country-specific; the region should be reported for reproducibility.","section":"§2.2"},{"comment":"The statement that “there are no previous baselines on best time to post for Facebook pages” is too strong given the works cited in Section 7 (refs. [16], [35], [45]); at minimum those methods should be discussed as candidate baselines, or the claim should be softened.","section":"§6.2"}],"recommendation":"major_revision","confidential_remarks":"The manuscript header identifies it as the version published in Social Network Analysis and Mining (2018) 8:11, so the substantive concerns above apply to the content as submitted rather than to a new submission. My main worry for the editors is the integrity of the headline 7x claim: the evaluation metric conflates reaction-arrival time with post-creation time, and the gain is an in-sample maximum with no holdout. The descriptive temporal analyses are salvageable and would still be of interest, but the central quantitative claim needs re-analysis (properly aligned per-post outcomes, temporal holdout, confound controls) or retraction. I see a tractable path to a defensible paper, which is why I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The one thing to know: this paper has a genuinely interesting descriptive contribution about when different types of Facebook brand pages get reactions, but the central 7x gain claim is computed from the same data used to select the schedule, so it is an in-sample maximum, not a prediction.\n\nThe new part is the application to brand pages rather than individual users. The authors crawl 100 public pages in five categories (e-commerce, traffic, telecom, hospital, politics) over five years, get 0.3M posts and 10M comments, and derive category-specific posting schedules. The daily, weekly, and monthly reaction patterns are plausible and align with common sense: politics peaks in the evening, traffic during commute hours, e-commerce during office hours, telecom on Sundays. The content-type table (links are 78.6% of posts but only 54.16% of reactions, videos much more engaging per post) is a useful empirical checkpoint. The categorization procedure is standard but carefully done, and the within-vs-across correlation results suggest the clustering is meaningful.\n\nThe soft spot is the evaluation. Equation (14) defines reaction gain as the ratio of reactions per post in a selected bucket to the global average, and the schedule is selected by ranking those same buckets using the same aggregated counts. No temporal holdout, no page-level split, no confidence intervals, no minimum post count for the bucket. A bucket with two posts, one very successful, can give a spuriously high ratio. Section 6.5 shows content type strongly affects engagement, so the schedule may be picking content mix, not posting time. And the paper only has comment timestamps (stated in Section 2.2), yet the abstract and conclusion talk about \"audience reactions\" generally. Those are real problems, and they break the 7x claim as a practical recommendation.\n\nThat said, the descriptive analysis itself is not broken. The patterns are interesting and could be built on. The paper is honest enough to mention the comment-timestamp limitation, even if the abstract then overstates the measure. I would not desk-reject this outright; a serious referee could request a split-sample evaluation, per-page controls, and a denominator floor, and the descriptive sections would still stand. As it is, the headline claim should be treated as an in-sample observation, not a how-to guide.\n\nIf I were editing, I would send it to review but expect major revision. For yourself: worth a skim for the category-level patterns, but don't cite the 7x figure.","headline":"Useful descriptive patterns for Facebook brand pages, but the headline 7x gain is an in-sample artifact and should be read skeptically.","tokens_in":15733,"tokens_out":1582,"would_cite":false,"duration_ms":18052,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Posting at times chosen from a page's audience reaction history can draw seven times more reactions than an average posting time.","keywords":["social media analysis","posting time scheduling","audience reactions","Facebook pages","information spread","temporal analysis","page categorization","reaction gain"],"falsifier":"Take the same pages, randomize new posts between the top reaction bucket and a bucket with average gain, and compare reactions per post after two weeks; if the gap is far below sevenfold, the historical ratio was not predictive. A cheaper version is a temporal holdout: build the schedule on 2011–2013 and evaluate on 2014–2015.","tokens_in":14684,"feed_emoji":"⏰","tokens_out":7417,"duration_ms":70206,"temperature":0.7,"pith_summary":"This paper tries to show that the time of day a brand page posts is a lever large enough to multiply audience reactions by a factor of seven. It builds six posting schedules from five years of Facebook page data (0.3 million posts, 10 million reactions), ranking 15-minute time buckets by where reactions actually arrive. The best schedule, a category-specific schedule ranked by audience reactions, gives a reaction gain of about seven in its top bucket, while schedules based on when admins currently post give a gain below two. If the effect is real, page managers can raise visibility substantially without changing content, only timing.","feed_headline":"Posting at audience-peak times draws seven times more reactions","feed_subtitle":"A reaction-based schedule from 100 Facebook pages and 10 million reactions beats average posting by 7x.","key_machinery":"The carrying object is the reaction profile vector $R_k(P_x)$: for each page, the count of audience reactions received in each 15-minute bucket of the day, summed across years. Rankings derived from these vectors define the schedules, and the reaction gain $RG(C,k) = \\delta(C,k)/\\omega(C)$ evaluates them as reaction per post in a bucket divided by the category's average reaction per post. The same cumulative counts that build the schedule also define its evaluation, so the reported gain is an internal comparison on the training data.","core_discovery":"The central claim is that audience reaction profiles—histograms of how many reactions arrive in each 15-minute bucket of the day—are stable enough to serve as posting schedules, and that posting in the buckets where reactions concentrate produces far more reactions per post than posting at an average time. On its dataset, the categorized frequent reaction schedule has reaction gain 7 in the best bucket, the weighted version 5.4, and all posting-based schedules below 2. The paper also claims that pages cluster into categories with coherent reaction patterns; the top features for categorization are reactions within the first hour, posts per day, and page type.","pith_inferences":["The 7x gain is probably an upper bound: the schedule is selected from the same historical counts that define the gain, so regression toward the mean will shrink the realized effect on new posts.","A temporal holdout (build the schedule on 2011–2013, test on 2014–2015) would tell how much of the gain transfers; the paper reports no such test.","The daily peaks line up with office hours and commutes, so the schedule is likely specific to the city, era, and page types in the dataset; re-running elsewhere would show how portable it is.","A content-and-time policy (post video in the top bucket) is a natural extension the paper stops short of making."],"forward_implications":["A brand page that posts in its category's top reaction bucket should, on the paper's numbers, receive roughly seven times the average reactions per post it would get without a schedule.","A new page with little reaction history can be assigned to a category using reaction-within-first-hour, posts-per-day, and page type, then use that category's schedule.","Reaction-based schedules outperform posting-based schedules by a wide margin, so the paper's result is not just 'post when other admins post.'","Because 84% of reactions arrive within 24 hours, the timing decision is essentially a same-day decision; the top 15-minute bucket is where that day's reaction window concentrates.","Content type matters too: videos and photos earn more reactions per post than links, so timing gains can be combined with content-choice gains."],"supporting_citations":[{"why":"Supplies the data-collection method used to build the 100-page, 0.3-million-post, 10-million-reaction dataset.","marker":"[40]"},{"why":"Supports the premise that social-media content has a short lifespan, motivating the study of posting time as a visibility lever.","marker":"[42]"},{"why":"Documents persistence and decay of social-media trends, used to justify the one-week reaction window and early-reaction analysis.","marker":"[1]"},{"why":"Explains the Facebook News Feed ranking behavior that makes early reactions compound into higher visibility for a post.","marker":"[2]"},{"why":"Supports the use of comments as an implicit measure of audience interest, which is the basis of the reaction timestamp.","marker":"[27]"},{"why":"Also supports the use of comments and reactions as a proxy for post popularity, cited alongside [27] for the reaction measure.","marker":"[37]"},{"why":"Supplies the wrapper-based feature selection used to choose the top reaction-determining features for page categorization.","marker":"[19]"},{"why":"Supplies the entropy-based discretization used to convert continuous page features into the discrete attributes used by the categorizer.","marker":"[12]"}],"fun_headline_variants":["Posting at peak reaction times boosts reactions 7x","Facebook post timing: peak reactions give 7x more","Schedule posts by reaction pattern for 7x audience reactions","Best posting time yields 7x more reactions on Facebook","Reaction-based posting schedule wins 7x more responses"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The 7x result assumes that a 15-minute slot that drew many reactions per post in the past will keep drawing many reactions per post in the future, independently of what is posted, which page posts it, and how audience habits change.","fun_headline_variants_meta":{"raw":{"variants":["Posting at peak reaction times boosts reactions 7x","Facebook post timing: peak reactions give 7x more","Schedule posts by reaction pattern for 7x audience reactions","Best posting time yields 7x more reactions on Facebook","Reaction-based posting schedule wins 7x more responses"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000781,"raw_usage":{"total_tokens":3403,"prompt_tokens":854,"completion_tokens":2549,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":470,"completion_tokens_details":{"reasoning_tokens":2468}},"tokens_in":470,"tokens_out":2549,"duration_ms":15981,"temperature":1.0,"reasoning_tokens":2468,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T11:34:10.163283+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the same pages, randomize new posts between the top reaction bucket and a bucket with average gain, and compare reactions per post after two weeks; if the gap is far below sevenfold, the historical ratio was not predictive. A cheaper version is a temporal holdout: build the schedule on 2011–2013 and evaluate on 2014–2015.","supporting_citations":[{"cited_title":"Semantic Web 4(3), 245– 250 (2013)","cited_arxiv_id":null,"evidence_quote":"Supplies the data-collection method used to build the 100-page, 0.3-million-post, 10-million-reaction dataset."},{"cited_title":"In: Proceedings of the 20th international conference on World wide web","cited_arxiv_id":null,"evidence_quote":"Supports the premise that social-media content has a short lifespan, motivating the study of posting time as a visibility lever."},{"cited_title":"In: Fifth International AAAI Conference on Weblogs and Social Media","cited_arxiv_id":null,"evidence_quote":"Documents persistence and decay of social-media trends, used to justify the one-week reaction window and early-reaction analysis."},{"cited_title":"https://www.facebook.com/ business/news/News-Feed-FYI-A-Window-Into-News-Feed (2013) Maximizing the Visibility of Content in Social Media Brand Pages 21","cited_arxiv_id":null,"evidence_quote":"Explains the Facebook News Feed ranking behavior that makes early reactions compound into higher visibility for a post."},{"cited_title":"In: Proceedings of the 2016 ACM on Multimedia Conference","cited_arxiv_id":null,"evidence_quote":"Supports the use of comments as an implicit measure of audience interest, which is the basis of the reaction timestamp."},{"cited_title":"Social Network Analysis and Mining 4(1), 174 (2014)","cited_arxiv_id":null,"evidence_quote":"Also supports the use of comments and reactions as a proxy for post popularity, cited alongside [27] for the reaction measure."},{"cited_title":"Artiﬁcial intelligence 97(1- 2), 273–324 (1997)","cited_arxiv_id":null,"evidence_quote":"Supplies the wrapper-based feature selection used to choose the top reaction-determining features for page categorization."},{"cited_title":"Machine learning 8(1), 87–102 (1992)","cited_arxiv_id":null,"evidence_quote":"Supplies the entropy-based discretization used to convert continuous page features into the discrete attributes used by the categorizer."}],"review_version":1}