{"id":"889e0947-ac84-4d8e-bbd2-5e61706000e7","arxiv_id":"2505.10791","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Using more than 12,000 scanned editions of five Indian newspapers, the study finds corporate ad volume correlates with increased article coverage, while the favorable sentiment link loses significance after adding a popularity control.","lead":"Using image processing and OCR, this paper extracts every advertisement from more than 12,000 editions of five Indian newspapers and maps who advertises, where, when, and how. It reports that companies buying more ad space in a newspaper tend to receive more articles and more positive coverage, though the sentiment link weakens after controlling for company popularity.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Appendix Table 4 contradicts the sentiment headline: after controlling for Google Trends popularity, the ad-ratio coefficient on total sentiment is non-significant and usually negative, yet the abstract and Section 5.2 call the positive relationship 'robust.'","rationale":"The reader's stated weakest assumption is that keyword misclassification is randomly distributed across page numbers and areas, which is a legitimate measurement-error concern. However, the more immediately decisive problem is the paper's own Appendix Table 4, which the reader also flags in the rationale but does not make the central weak point. In Table 2, the sentiment coefficient is positive and highly significant, but in every popularity-controlled specification in Appendix Table 4 the coefficient is non-significant and mostly negative. This is not a question of external consensus or causal language; it is an internal inconsistency between the headline claim and the paper's own auxiliary analysis. The paper itself concedes the point in Section 5.2, yet the abstract and conclusion still describe the sentiment result as robust. The popularity proxy is coarse (Google Trends), so a fully decisive test should first replicate on a common sample and aggregation before concluding that the effect is absent. Because the coverage-volume result remains significant even with popularity controls, the paper still has a defensible empirical contribution; the sentiment claim, however, needs to be reconciled or removed. Keeping the reader's CONDITIONAL verdict is appropriate, but the condition should be made explicit: either extend the popularity-controlled sentiment analysis to the full sample and show the effect survives, or revise the abstract and conclusions to claim only an association with coverage volume, not favorable tone.","tokens_in":16828,"tokens_out":5250,"duration_ms":57911,"concrete_test":"Rerun the Company Total Sentiment regression (Table 2, model 4: newspaper x company + time fixed effects) on the exact 40-company sample of Appendix Table 4, at the same temporal aggregation used in Table 2 (monthly), first without and then with the Google Trends popularity covariate. If adding popularity reduces the ad-ratio coefficient from statistical significance to non-significance, or flips its sign, the sentiment claim in the abstract is not robust to a minimal confound and should be removed or heavily qualified. Also report the analogous 155-company popularity-controlled specification if popularity data can be extended; if it cannot, state the sample restriction explicitly.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's strongest claim is that corporate ad spending is positively associated with both volume and tone of coverage, with Table 2 reporting 0.0189*** for sentiment. The load-bearing weakness is that the only reported specification that adds a popularity control, Appendix E, Table 4, does not reproduce this sentiment result. Across all four models, the Total Ad Page Percent coefficient is -0.0003, 0.0018, -0.0016, and -0.0011, none significant, and three of four are negative. Section 5.2 even states that 'neither ad spending nor popularity shows consistent significance on sentiment across all specifications when company and time effects are included.' This is an internal inconsistency: the abstract's 'robust' favorable-coverage claim rests on a result that the paper's own control analysis fails to confirm. The two tables also use different samples and aggregations (155 entities/72 periods vs. 40/1372), which the paper never reconciles. If the popularity control is valid, the most natural explanation is omitted-variable confounding: large, popular firms both advertise more and attract more positive coverage; once popularity is held constant, ad ratio has no measurable tone effect. That would leave only the coverage-volume result, which is a weaker claim than advertiser influence on news tone.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces an image-processing and OCR pipeline that extracts articles and advertisements from digitized print newspapers, and applies it to five Indian newspapers in three languages over roughly six years, assembling over 12,000 editions and hundreds of thousands of ads. The authors use this dataset to describe the print advertising ecosystem (who advertises, placement, size, timing, and topics) and then run panel regressions relating a weighted ad ratio to the volume and sentiment of news coverage received by government and corporate advertisers. The headline finding is that corporate advertising is positively and robustly associated with both coverage volume and sentiment, while government advertising shows weaker or negative associations.","tokens_in":17063,"tokens_out":2186,"duration_ms":24171,"significance":"If the central claims held, this would be a valuable large-scale contribution to the media-bias and advertiser-influence literature, which has mostly relied on smaller or single-outlet studies. The dataset itself is a substantial asset: the authors make code and data public, cover multiple languages and regions, and document a reproducible extraction pipeline with reported detection and OCR accuracy. The descriptive findings on ad placement, size distributions, and government versus corporate spending are new for the Indian print market and are likely to be useful to other researchers. However, the paper's most consequential claim—that corporate advertising robustly predicts more favorable news tone—is undermined by the paper's own popularity-controlled regressions in Appendix E, Table 4, as detailed below. The coverage-volume result appears more robust and is a meaningful finding in its own right, but it is a weaker claim than advertiser influence on news tone.","major_comments":[{"comment":"The abstract's claim that the ad-sentiment relationship is 'robust over time and across different levels of advertiser popularity' is contradicted by the paper's own Appendix Table 4. In all four specifications that add a Google Trends popularity control, the Total Ad Page Percent coefficient on total sentiment is non-significant (-0.0003, 0.0018, -0.0016, -0.0011), and three of the four coefficients are negative. Section 5.2 even concedes that 'neither ad spending nor popularity shows consistent significance on sentiment across all specifications when company and time effects are included.' This is an internal inconsistency in the central claim. Furthermore, Tables 2 and 4 use different samples and aggregation levels (155 entities / 72 periods versus 40 entities / 1372 periods), and the manuscript does not explain which sample is preferred or how the samples differ. The authors should reconcile these tables, report the popularity-controlled specification as the main specification if it is the more appropriate one, and revise the abstract and Section 5.2 so that the strength of the sentiment claim matches the evidence.","section":"Section 5.2 and Appendix E, Table 4"},{"comment":"The analysis assumes that keyword-based misclassification of articles and ads is randomly distributed across page numbers and page areas. This assumption is load-bearing because the weighted ad ratio is constructed from page-specific scaling factors, and the coverage measures are extracted from the same pages. If misclassification is correlated with page prominence—for example, if front-page articles are more likely to mention large brands, or if front-page advertisements are larger and thus more likely to be correctly detected—then the estimated coefficients linking ad ratio to coverage would be biased. The authors should provide a sensitivity analysis or validation that directly tests whether misclassification rates vary by page number and page area, rather than asserting randomness.","section":"Section 4, Keyword Identification and Section 5.1 regression setup"},{"comment":"The statement that a 1% increase in weighted ad ratio leads to an average increase of 0.0189 units in total sentiment score, and that this is 'substantial' given the -1 to 1 sentiment scale, requires clarification of the dependent variable. If total sentiment is a sum of per-article scores (-1, 0, 1) over a period, then the coefficient is not directly comparable to the per-article scale, and calling it substantial may be misleading. If it is an average, the claim needs a different justification. The manuscript should state clearly how 'total sentiment' is aggregated and provide an effect-size discussion that is consistent with that definition.","section":"Section 5.2, interpretation of effect size"}],"minor_comments":[{"comment":"The text says performance degraded with the larger dataset, yet Table 3 shows that the larger dataset has higher recall (97.8% versus 88.8%) even though mAP and precision are slightly lower. Please clarify whether the degradation refers to mAP alone or to a weighted criterion, and why the first model is preferred.","section":"Appendix A, Table 3"},{"comment":"The single-entity example (Adani) is presented as demonstrating 'the influence of advertising on sentiment,' but a case study of one conglomerate around one scandal cannot establish a general causal relationship. Please either soften the language or add more examples with quantitative support.","section":"Appendix D, Figure 24"},{"comment":"The claim that Tesseract achieves 'an error rate of less than 5%' and that Surya's performance is 'close to 1% error rate' is reported without a citation or evaluation on this dataset. Please provide the evaluation details or reduce the strength of these claims.","section":"Section 3.2, OCR error rates"},{"comment":"There are several minor typographical and formatting issues: 'Corporates' is used inconsistently as a noun, 'cr' appears incomplete in the Introduction ('such as “bribe,” “scam,” “corrupt,” and other relevant terms'), and some references (e.g., [8], [37]) lack access dates or are cited imprecisely. A careful proofread is needed.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The descriptive contribution and dataset are strong and appropriate for cs.CY, but the headline inference is not supported by the paper's own robustness analysis. The authors should be asked to either downgrade the sentiment claim to a conditional or exploratory result, or to provide a principled reconciliation of the main and popularity-controlled specifications. If the sentiment claim is removed or substantially weakened, the remaining coverage-volume result and descriptive analysis would likely be publishable after minor revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nRead this one for the data, not for the headline. The authors built a genuinely new resource: 12,358 editions of five Indian newspapers in three languages, with article and ad segments extracted via a fine-tuned YOLOv8s, OCR, and translation. They release code and data. The descriptive findings—government ads are small, numerous, and spread across pages; corporate ads cluster on premium pages at quarter/half/full-page sizes; print ad volume is stable despite circulation decline—are useful and credible. If you work on media bias or the political economy of news outside the US/Europe, this is a dataset you'll want to cite.\n\nThe regression part is where I'd be careful. The coverage-volume result survives the popularity control: Appendix Table 5 shows ad ratio stays significant for article count in all four specs. That is a real, if modest, finding.\n\nThe sentiment claim does not hold up. Table 2 reports a positive, significant ad-ratio coefficient on total sentiment, and the abstract calls it \"robust over time and across different levels of advertiser popularity.\" But Appendix Table 4—the same analysis adding a Google Trends popularity control—shows the ad coefficient is non-significant in every specification and negative in three of four. Section 5.2 even acknowledges \"neither ad spending nor popularity shows consistent significance on sentiment across all specifications when company and time effects are included.\" That is an internal contradiction with the abstract. The two tables also use different samples and aggregations (155 entities over 72 periods vs. 40 entities over 1,372 periods), and the paper never explains why. If the popularity control is valid, the most natural reading is omitted-variable confounding: big, popular firms both advertise more and get more coverage, with no independent tone effect.\n\nMinor concerns: the keyword-based matching assumption (misclassification random across pages) is stated but not tested, and the causal language is stronger than the panel-correlation design supports. Those are fixable. The inconsistent sentiment claim is not.\n\nWho is this for? Researchers studying print media outside Western contexts, and anyone building similar extraction pipelines. The dataset deserves a serious referee, and the paper should be sent out—but the authors need to either drop the sentiment claim or reconcile Table 2 with Table 4 before publication. As it stands, cite the data, not the conclusion.","headline":"The new print-ad dataset and pipeline are the real contribution; the paper's own appendix undercuts its headline claim that ad spending predicts positive sentiment.","tokens_in":17586,"tokens_out":2161,"would_cite":true,"duration_ms":20706,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Companies that advertise more in a newspaper receive more articles and more positive coverage in that same paper, according to an analysis of 12,358 Indian print editions.","keywords":["Print Media","Advertising","Media Bias","Content Analysis","Information Retrieval","Panel Regression","Sentiment Analysis","Indian Newspapers"],"falsifier":"A company-level event study would settle the claim: find firms whose ad spending in one newspaper fell abruptly for reasons outside that newspaper's control (a budget cut, a boycott, a merger) while their spending in a matched paper stayed flat, and check whether coverage volume and sentiment in the cut paper decline relative to the control. The paper's own anecdote of a conglomerate buying more ads after a scandal, with sentiment gradually recovering, is one instance of the pattern; replicating it systematically across the 155 companies, or exploiting a regulation that forces ad-spend changes, would discriminate the influence channel from coincidence or salience.","tokens_in":16646,"feed_emoji":"📰","tokens_out":19687,"duration_ms":163828,"temperature":0.7,"pith_summary":"This paper asks whether buying newspaper ads buys favorable news coverage, and answers yes for corporate advertisers. The authors built a pipeline that segments and reads the pages of five Indian dailies in three languages, assembling more than 12,000 editions that reach over 100 million readers, and used it to map who advertises, on which pages, at what sizes, and when. Regressing the volume and sentiment of each advertiser's coverage on a page-area-based measure of the advertiser's presence, they find a positive association that survives controls for newspaper, company, and time: a one-percentage-point increase in a company's weighted ad ratio is tied to an average gain of 0.0189 units in sentiment toward the company and to a significant rise in the number of articles mentioning it. The same regressions give a negative or null result for government advertising, which the authors explain by the legal requirement that governments publish ads everywhere, leaving newspapers no incentive to trade favorable coverage for guaranteed revenue. If the result holds, it is large-scale, cross-lingual evidence that print advertising and editorial content are commercially entangled.","feed_headline":"Ads buy coverage: 1% more ad space lifts a firm's coverage tone","feed_subtitle":"A 12,000-edition study of five Indian dailies ties corporate ad share to article count and tone.","key_machinery":"The load-bearing object is the weighted ad ratio, $$\\text{Weighted Ad Ratio} = \\frac{\\text{scaling factor} \\times \\text{ad area}}{\\text{page area}},$$ where the scaling factor is the ratio of a page's actual rate-card price to the paper's base per-square-centimeter rate; it converts every ad into a comparable measure of how much prominence the advertiser bought, which is what makes comparisons across pages, papers, and languages meaningful. The argument then runs through a panel regression of total sentiment and article count on this ratio, with newspaper-by-company fixed effects and time fixed effects that absorb stable differences between papers and advertisers as well as shared shocks. The ratio is computable at scale only because of the extraction pipeline — page segmentation with a fine-tuned object-detection model, OCR (text recognition) in three scripts, and machine translation into English — which the authors validate at 96.8% mean average precision, a standard detection-accuracy score, and with 0.94 F1 for the keyword filters that assign ads and articles to advertisers. Sentiment is scored by a classifier that assigns each article a value of $-1$, $0$, or $1$, and the coverage count is the number of articles per period that match the advertiser's keywords.","core_discovery":"The central claim is that in Indian print newspapers, the more a company advertises in a given paper, the more articles that paper publishes about the company and the more positively those articles are worded. The evidence is a panel regression of monthly coverage on a weighted ad ratio — ad area divided by page area, multiplied by a rate-card scaling factor so that a front-page half-page ad counts for more than an inside-page one — estimated with newspaper-by-company fixed effects and time fixed effects. For corporate advertisers the coefficient on the ad ratio is positive and statistically significant in every specification of the main model: 0.0189 for total sentiment with no fixed effects and 0.0137 with both sets of fixed effects, and between 0.436 and 0.223 for the monthly count of articles. A popularity control built from web-search interest leaves the coverage association significant but weakens the sentiment association to non-significance when company fixed effects are included, a qualification the authors report rather than hide. For government advertisers the same regressions yield negative or null coefficients, which the authors attribute to the legally mandated, non-withdrawable character of most government ads: newspapers have no incentive to return favor for revenue they are guaranteed by law.","pith_inferences":["The correlation is consistent with editors returning favor for ad money, but it is also compatible with a salience story in which newspapers write more about firms that dominate the ad market because those firms are economically important; the design cannot fully separate the two, although the government placebo makes a pure measurement artifact unlikely.","A sharper test the paper's own data enables but does not run: compare articles about a company that appear on the same page as the company's ad with articles about the same company elsewhere in the same issue — if coverage is more positive next to the ad, the bias is partly a page-layout decision rather than a whole-editorial-tone decision.","The mechanism predicts a steeper ad–coverage gradient at newspapers that depend more heavily on a given advertiser's revenue; since the dataset contains ad areas and rate cards, advertiser revenue share per paper is computable and the gradient can be tested directly.","Running the same pipeline on digital news outlets, where ad placement is programmatic and advertisers do not negotiate with the newsroom, would be a natural control — a weaker correlation there would point to negotiated influence rather than general commercial pressure."],"forward_implications":["In the main model, a one-percentage-point increase in a company's weighted ad ratio is associated with a 0.0189-unit gain in total sentiment (on a scale the paper takes as running from $-1$ to $1$) and, depending on the specification, between 0.22 and 0.44 additional articles about the company per time period in the same newspaper.","Corporate advertisers put 27.9% of their ad area and 31.6% of their ad spending on the front, third, and back pages, so any influence they obtain operates on the paper's most visible pages, where readers see ads and coverage together.","The absence of a positive ad–coverage link for legally mandated government advertising supports the authors' reading that the corporate link runs through the advertiser's power to choose and withdraw spending rather than through a property of the measurement itself.","For article volume, the ad ratio stays a significant predictor even when a web-search popularity term and fixed effects are added, while popularity's own coefficient flips sign across models; the authors conclude that ad spending is the more reliable driver of media attention, while noting that the sentiment association weakens under the same controls.","Because the pipeline, code, and dataset are released, the same regressions can be rerun for other countries, languages, and time periods, turning the Indian finding into a testable template for media-influence research elsewhere."],"supporting_citations":[{"why":"Supplies the panel-regression specification with media-by-company and time fixed effects that the study adapts to sentiment and coverage outcomes.","marker":"[2]"},{"why":"Provides the government-advertising and corruption-coverage study whose design motivates the separate government regression and its legal-mandate interpretation.","marker":"[7]"},{"why":"Establishes the ads-influence-editors hypothesis in financial media that the paper tests at scale in Indian print newspapers.","marker":"[29]"},{"why":"Provides the sentiment classifier whose -1, 0, or 1 article scores form the dependent variable in the sentiment regressions.","marker":"[3]"},{"why":"Supplies the object-detection model fine-tuned to segment newspaper pages into articles and ads, the source of every ad-area measurement.","marker":"[14]"},{"why":"Supplies the machine-translation model that converts Hindi and Telugu ad and article text into English for keyword matching.","marker":"[9]"}],"fun_headline_variants":["Corporate ad area in dailies predicts more, friendlier coverage","Govt ads don't buy coverage; corporate ads do","12,000 editions: ad share boosts article volume and tone","Print ad spend correlates with favorable press for firms","In Indian print, corporate ads buy editorial favor"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The analysis stands on the assumption that mistakes in classifying articles and ads by keyword are randomly distributed across page numbers and page positions; if those mistakes cluster on prominent pages, the measured link between advertising and coverage would be biased rather than real.","fun_headline_variants_meta":{"raw":{"variants":["Corporate ad area in dailies predicts more, friendlier coverage","Govt ads don't buy coverage; corporate ads do","12,000 editions: ad share boosts article volume and tone","Print ad spend correlates with favorable press for firms","In Indian print, corporate ads buy editorial favor"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000384,"raw_usage":{"total_tokens":2067,"prompt_tokens":1017,"completion_tokens":1050,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":633,"completion_tokens_details":{"reasoning_tokens":970}},"tokens_in":633,"tokens_out":1050,"duration_ms":10635,"temperature":1.0,"reasoning_tokens":970,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T21:03:01.555983+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A company-level event study would settle the claim: find firms whose ad spending in one newspaper fell abruptly for reasons outside that newspaper's control (a budget cut, a boycott, a merger) while their spending in a matched paper stayed flat, and check whether coverage volume and sentiment in the cut paper decline relative to the control. The paper's own anecdote of a conglomerate buying more ads after a scandal, with sentiment gradually recovering, is one instance of the pattern; replicating it systematically across the 155 companies, or exploiting a regulation that forces ad-spend changes, would discriminate the influence channel from coincidence or salience.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the panel-regression specification with media-by-company and time fixed effects that the study adapts to sentiment and coverage outcomes."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the government-advertising and corruption-coverage study whose design motivates the separate government regression and its legal-mandate interpretation."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes the ads-influence-editors hypothesis in financial media that the paper tests at scale in Indian print newspapers."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the machine-translation model that converts Hindi and Telugu ad and article text into English for keyword matching."}],"review_version":1}