{"id":"063abcfc-818c-434f-a428-1d6200795ada","arxiv_id":"1908.08690","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Category-level dwell time in a smartphone news app differs across eight categories, with politics showing the highest and technology the lowest correlation between dwell time and page length.","lead":"Using one week of browsing data from a Japanese smartphone news app, this paper measures how long readers stay on news pages and compares that time across eight categories. It finds that the link between article length and reading time differs strongly by category, and that short visits often happen when the headline already carries the story.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The politics-vs-technology correlation gap (0.602 vs 0.245) is computed on a top-10% page-view subset with no confidence intervals or filter sensitivity, so the central category-difference claim could rest on sampling or filter artifacts.","rationale":"The paper's contribution is an empirical description, so I looked for the assumption that most directly supports the strongest quantitative claim. The category-specific correlation ranking is the main numeric headline, and it depends critically on the top-10% page-view filter and on Pearson's r, yet no confidence intervals or filter-sensitivity analyses are provided. This is the most load-bearing weakness because if the gap between politics and technology is not robust to reasonable filters or sampling variability, the conclusion that dwell time should be interpreted category by category loses its empirical basis. I partly agree with the reader's weakest_assumption: the same page-view filter is flagged, but the reader also lists category-label validation and the single-author satisfaction annotation, which I see as secondary for the central RQ2 result. The proposed test would settle the concern by quantifying uncertainty and checking whether the finding is an artifact of the filter; if the intervals remain well separated across alternative thresholds, my concern would be resolved. The reader's conditional verdict is appropriate, and my stress-test does not change that verdict.","tokens_in":5620,"tokens_out":4405,"duration_ms":45471,"concrete_test":"Recompute Table I from the full unfiltered set, or at minimum with top-20% and top-50% thresholds, and report bootstrap 95% confidence intervals for each category's Pearson correlation, plus Spearman correlations or log-dwell transforms to check distributional sensitivity. If the politics and technology intervals overlap, or if the ordering changes under a reasonable alternative filter or correlation measure, the claim of large category differences in dwell-time/length correlation should be weakened accordingly.","verdict_should_be":"UNCHANGED","load_bearing_attack":"RQ2's headline contrast, politics r=0.602 versus technology r=0.245, is the quantitative core of the paper. The estimate is computed only on news pages in the top 10% of page views within each category (Section II), using Pearson's correlation (Section III-B), and Table I reports no sample sizes, confidence intervals, or significance tests. The 10% filter truncates the joint distribution of dwell time and page length to popular pages; if popularity correlates with reading behavior, the within-category correlation is a property of popular pages, not of the category as a whole. The paper does not test whether the politics-technology gap survives alternative thresholds or even whether it exceeds sampling error. A difference of 0.357 between correlations can be within noise if per-category sample sizes are moderate, and the table's percentages do not reveal n. Without uncertainty quantification, the central claim that category-level correlation differences are material is not fully secured. I am not claiming the result is wrong; I am claiming the current analysis cannot rule out filter bias or sampling noise as explanations for the observed ordering.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper analyzes one week (July 1-7, 2017) of Gunosy smartphone news browsing data to characterize dwell time across eight news categories: politics, economy, society, international, technology, sports, entertainment, and column. It addresses three research questions: whether dwell-time distributions differ by category (RQ1), whether dwell time correlates with page length by category (RQ2), and what content characterizes short-dwell-time pages (RQ3). Using the median dwell time per news page and restricting to the top 10% of pages by views within each category, the authors report histograms, Pearson correlations (politics 0.602, technology 0.245, whole 0.291), and the results of a single author annotating 321 short-dwell-time pages. They conclude that category-level dwell-time trends differ, that the dwell-time/length correlation differs materially by category, and that short dwell times can still reflect user satisfaction when the title or photo carries the information.","tokens_in":5775,"tokens_out":5771,"duration_ms":58146,"significance":"If the reported category differences are robust, the paper makes a useful practical point: dwell time should be interpreted category by category in mobile news ranking, recommendation, and advertising. The study is a direct empirical description with no fitted model and no circular derivation, and the use of the median rather than the mean is a sensible way to handle app-launch artifacts. The central claims are falsifiable and could inform future work on engagement metrics. However, the current analysis does not yet secure the headline quantitative claims because it lacks uncertainty quantification, provides no sensitivity analysis for the page-view filter, and relies on an unvalidated subjective annotation. The paper's value is conditional on making these supporting analyses available.","major_comments":[{"comment":"The central RQ2 claim that politics has the highest correlation (0.602) and technology the lowest (0.245) is reported without confidence intervals, significance tests, or per-category sample sizes; the '# of news' column reports only percentage shares. A gap of 0.357 between two Pearson correlations can be within sampling error, especially if some categories contribute only a few hundred pages, so the current table does not rule out that the observed ordering is noise. Please report exact counts, confidence intervals (e.g., via Fisher z-transformation or bootstrap), and significance tests for the pairwise differences.","section":"Section III-B, Table I"},{"comment":"The restriction to the top 10% of news pages by page views within each category is a free parameter with no sensitivity analysis. If page popularity is correlated with reading behavior, the within-category distributions and correlations describe only popular pages, not the categories as a whole; the authors should repeat the analysis at alternative thresholds (e.g., 5%, 20%, all pages) or justify the filter empirically.","section":"Section II"},{"comment":"The short-dwell-time conclusion rests on one author's answers to four binary questions for 321 pages, but the paper never defines the 'certain limit' used to select short-dwell-time pages, and there is no inter-rater reliability or validation that the questions measure satisfaction. The abstract's statement that 'a user tends to get sufficient information' overstates what these annotations can support; this should be framed as an exploratory content analysis.","section":"Section III-C"},{"comment":"RQ1 is supported only by visual inspection of histograms whose y-axes differ across categories and whose x-axis uses relative values; no distributional test (e.g., Kolmogorov-Smirnov, quantile comparisons) is reported. The claim of 'different dwell time trends for each category' needs a quantitative test or at least descriptive statistics with confidence intervals.","section":"Section III-A"},{"comment":"The category labels come from Gunosy's proprietary heuristic and machine-learning classifier, but no accuracy or validation of this classifier is reported. Measurement error in category labels could attenuate or distort category-level differences; the paper should report a validation result or discuss the likely impact of label noise.","section":"Section II"}],"minor_comments":[{"comment":"The affiliation 'Toyohashi University of Technorogy' contains a typo; 'Technorogy' should be 'Technology'.","section":"Title page"},{"comment":"The sentence 'The visualization does not include e do not use news pages whose dwell time is over a certain threshold for visualization' is garbled and should be rewritten.","section":"Section III-A"},{"comment":"The 'society' panel in Fig. 2 is mislabeled as 'sciety'.","section":"Fig. 2"},{"comment":"The statement that 'the values of the x-axis and the y-axis are equally spaced' is unclear; please specify the bin widths and what 'relative values' means for the x-axis.","section":"Section III-A"},{"comment":"The page length is said to vary by device, but the paper never defines how page length is measured (e.g., characters, pixels, or scroll height); this should be stated for reproducibility.","section":"Section II"},{"comment":"The Index Terms entry 'click bait' is usually written 'clickbait'.","section":"Index Terms"}],"recommendation":"major_revision","confidential_remarks":"This is a short industry-data paper whose empirical questions are reasonable but whose statistical support is currently thin. The authors need to supply sample sizes, confidence intervals, filter sensitivity, and a more cautious interpretation of the manual annotation. Because the data are proprietary, I would not require full data release, but the analysis should be reproducible in principle from the described pipeline."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things about this paper. It gives category-level dwell-time statistics for a smartphone news app that I have not seen elsewhere, and the category differences in correlations are potentially useful to anyone building engagement signals. But the quantitative core is under-supported: the politics vs. technology correlation gap (0.602 vs. 0.245) is computed on a top-10% page-view subset with no confidence intervals, no significance tests, and no sensitivity analysis. That is the main thing to tell anyone who cites Table I.\n\nWhat is actually new and good: the category-level histograms and the correlation table are a legitimate extension of Yi et al. and Liu et al., and the authors are transparent about using median dwell time to avoid the app-left-open artifact. The RQ3 hypothesis — short dwell times can coexist with user satisfaction when titles and photos carry the content — is worth taking seriously, and their four-case taxonomy is a reasonable starting point. The citation pattern looks right; they cite the relevant dwell-time literature.\n\nThe soft spots are real and they are in the middle of the paper. Table I gives percentages of news pages but not raw n, so the reader cannot judge whether a 0.357 difference in correlations is within sampling noise. The top-10% page-view filter truncates the joint distribution to popular pages; if popularity correlates with reading behavior, the category differences could be an artifact of that filter. No alternative thresholds are tested. That is a load-bearing gap, not a minor one.\n\nRQ3 is weaker still. The threshold for \"short dwell time\" is never disclosed. The four questions were answered by one author on 321 pages, with no second annotator and no inter-rater reliability. So the conclusion that \"users tend to get sufficient information from the title\" is really \"one author believes about two-thirds of these pages would satisfy a user.\" That is an interesting observation, not a measured result. Also, the proprietary category classifier is used without accuracy validation, which is worth a sentence of caution.\n\nWho is this for? Applied researchers in recommender systems and news quality, plus anyone who wants a compact example of why correlation differences need error bars. It deserves a serious referee, not a desk reject — the data are new and the questions are sensible. But I would send it back for major revision: report n and confidence intervals, test the top-10% filter, disclose the dwell-time threshold, and either get a second annotator or soften the satisfaction claim.","headline":"Useful category-level dwell-time data, but the headline correlation gap and the satisfaction claim need more support before I'd trust them.","tokens_in":6309,"tokens_out":1983,"would_cite":false,"duration_ms":21263,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"News dwell time varies by category, and short visits can still satisfy the reader.","keywords":["dwell time","news categories","smartphone news app","user behavior","page length","engagement metrics","clickbait"],"falsifier":"Re-run the dwell-time/length correlation on all news pages, not just the top 10% by views, and have independent judges label categories and rate the 321 short-dwell pages on the four questions; if the politics-technology gap shrinks or the \"yes\" rates fall below a majority, the paper's central claims fail.","tokens_in":5385,"feed_emoji":"📱","tokens_out":7290,"duration_ms":64812,"temperature":0.7,"pith_summary":"This paper tries to establish that dwell time on news pages cannot be treated as one uniform engagement signal: it varies systematically with news category in a smartphone news app. Using eight categories from a Japanese news-curation service, the authors show category-specific dwell-time distributions and report that the correlation between dwell time and page length ranges from r=0.602 (politics) to r=0.245 (technology). They also argue that a short dwell time does not necessarily mean a dissatisfied reader, because a title or an expected photo can carry the information a user came for. The stakes are practical: dwell time is used to rank search results, tune recommender systems, price ads, and judge news quality, so knowing that it is category-dependent changes how those systems should read it.","feed_headline":"News dwell time depends on category—and short visits can satisfy","feed_subtitle":"Politics pages reward longer articles; entertainment pages don't—engagement metrics need a category lens.","key_machinery":"The analysis is carried by three mechanisms. First, per-page dwell time is defined as the median dwell time over all viewers, chosen because a user who leaves the app open produces extreme outliers. Second, the relation to page length is summarized by Pearson's correlation coefficient computed separately within each of the eight categories, and the spread of those coefficients is the evidence for category dependence. Third, the satisfaction claim rests on a four-question manual content review of short-dwell pages, classifying whether the title alone carries enough information and whether photos match the user's expectation. The named object carrying the argument is the category-specific correlation $r$ between dwell time and page length, with politics $r=0.602$ and technology $r=0.245$ as the extremes.","core_discovery":"The paper reports three findings. First, the shape of the dwell-time histogram differs by category: society has fewer very-short-dwell-time pages than other categories, while entertainment has many; economy decays slowly toward long dwell times, and column behaves like society with a gentler decay. Second, the Pearson correlation between per-page dwell time (median across users) and page length is category-dependent: politics, sports, and international sit around 0.53-0.60, whereas technology, entertainment, and column sit around 0.25-0.37, and the overall correlation of 0.291 is dragged by the category mix. Third, in a manual review of the 321 shortest-dwell pages, the author answered \"yes\" to \"does the title have enough information?\" for 253 pages, \"are title and body matched?\" for 246, \"does the title recall photos?\" for 273, and \"does a proper photo appear?\" for 209; the authors take this as evidence that short dwell times often reflect satisfied users who got the content from title and photos rather than from reading the text.","pith_inferences":["If the top-10-percent-by-page-views filter is hiding a different pattern in long-tail pages, then a testable extension is to rerun the analysis on the full page corpus and check whether the politics-technology gap survives.","The photo-recall finding suggests that visual thumbnails, especially face-centered crops, can end a reading session early without dissatisfaction; an A/B test that varies thumbnail cropping while measuring dwell time could turn this observation into a design guideline.","Dwell time may behave differently on desktop or web reading than in a smartphone app, since the title-and-photo satisfaction mechanism is tied to how the app displays lists and thumbnails; replications on other platforms would show how far the category pattern generalizes."],"forward_implications":["Products that score pages by dwell time should calibrate thresholds per category; a time span that flags a technology page as unread may be normal for an entertainment page.","Length-based predictions of engagement will be reliable only for categories like politics and sports, where dwell time and page length correlate around 0.6, and nearly useless for technology and entertainment, where the correlation is below 0.4.","News-quality metrics built on dwell time should treat short visits as inconclusive unless the title and photo content are also assessed, since many short-dwell pages appear to satisfy the user.","The category mix of a traffic sample matters: aggregating pages across categories without controlling for category will understate or obscure the true dwell-time/length relationship."],"supporting_citations":[{"why":"Shows that dwell time in a news recommender varies by category, the prior result this paper extends to a smartphone app.","marker":"[4]"},{"why":"Reports category as the strongest predictor of news clicks, motivating the category-by-category analysis.","marker":"[11]"},{"why":"Documents large category-to-category variation in the parameters of a Weibull dwell-time model, the direct modeling antecedent of the paper's correlation analysis.","marker":"[12]"},{"why":"Predicts dwell time using topic-model-based categories, supporting the link between news category and reading time.","marker":"[13]"}],"fun_headline_variants":["Dwell time by news category: politics length matters, tech doesn't","Short news visits are often satisfied readers, not skimmers","Category reveals dwell time patterns in news apps","News dwell time: social and entertainment differ starkly","Article length correlates with dwell time only for some news"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the pages and category labels used for the analysis represent normal reading behavior; if the top-10-percent popularity filter, the service's automatic category labels, or one person's judgments on 321 short-dwell pages are not representative, the reported category differences could be artifacts.","fun_headline_variants_meta":{"raw":{"variants":["Dwell time by news category: politics length matters, tech doesn't","Short news visits are often satisfied readers, not skimmers","Category reveals dwell time patterns in news apps","News dwell time: social and entertainment differ starkly","Article length correlates with dwell time only for some news"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000265,"raw_usage":{"total_tokens":1604,"prompt_tokens":941,"completion_tokens":663,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":557,"completion_tokens_details":{"reasoning_tokens":584}},"tokens_in":557,"tokens_out":663,"duration_ms":6599,"temperature":1.0,"reasoning_tokens":584,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T11:31:41.033940+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the dwell-time/length correlation on all news pages, not just the top 10% by views, and have independent judges label categories and rate the 321 short-dwell pages on the four questions; if the politics-technology gap shrinks or the \"yes\" rates fall below a majority, the paper's central claims fail.","supporting_citations":[{"cited_title":"Beyond clicks: dwell time for personalization,","cited_arxiv_id":null,"evidence_quote":"Shows that dwell time in a news recommender varies by category, the prior result this paper extends to a smartphone app."},{"cited_title":"The Pulse of News in Social Media: Forecasting Popularity,","cited_arxiv_id":null,"evidence_quote":"Reports category as the strongest predictor of news clicks, motivating the category-by-category analysis."},{"cited_title":"Understanding web browsing behaviors through Weibull analysis of dwell time,","cited_arxiv_id":null,"evidence_quote":"Documents large category-to-category variation in the parameters of a Weibull dwell-time model, the direct modeling antecedent of the paper's correlation analysis."},{"cited_title":"Understanding User Attention and Engage- ment in Online News Reading,","cited_arxiv_id":null,"evidence_quote":"Predicts dwell time using topic-model-based categories, supporting the link between news category and reading time."}],"review_version":1}