{"id":"78ecc4c5-d367-4e84-a0cd-8f98d12cbac4","arxiv_id":"2607.04042","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"TikTok, Instagram, Mastodon and Facebook largely fail to block verified phishing URLs at posting time; only Twitter blocks a majority.","lead":"Most of the top five Android social media apps allow users to post known malicious phishing URLs with little or no blocking. The finding quantifies a practical gap attackers can exploit through ordinary link-sharing features.","discovery_kind":"incremental","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified beyond the already-flagged post-time sampling limit.","rationale":"The Reader correctly identifies both the strongest claim and its principal limitation. Because the paper never claims more than an exploratory post-time snapshot, and because the tabulated results match that claim without over-reach, the load-bearing concern does not move the verdict. The suggested concrete test simply operationalizes the authors’ own future-work item and would either confirm or mildly qualify the existing numbers; it does not threaten the core finding that four of the five apps blocked almost nothing at the moment of posting. Hence the CONDITIONAL verdict with high confidence remains appropriate.","tokens_in":11123,"tokens_out":402,"duration_ms":4054,"concrete_test":"Re-run the identical 35-URL set (or a fresh PhishTank sample of equal size) on the same five apps, then re-check every successfully posted link after 48 h and after 7 days; if any platform that allowed the original post later removes or flags it, the post-time-only claim would need a temporal qualifier, but the exploratory headline would still hold.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim (most of the five apps lack effective post-time defenses; only 23.8% of verified phishing URLs blocked, almost all by Twitter) rests on a transparent, small-N exploratory measurement that the authors themselves label as pilot work. Tables II–IV and the appendix give per-app, per-date counts that directly support the numbers. The weakest assumption the Reader already isolates—one-time post-time checks of 35 PhishTank URLs, no retroactive monitoring, possible profile-vs-DM discrepancy—is explicitly listed by the authors as future work in §VI and does not invent unsupported strength. No internal inconsistency, circularity, or hidden assumption undermines the modest claim as written.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper reports an exploratory measurement of post-time filtering of known malicious URLs on five Android social apps (TikTok, Instagram, Twitter, Facebook, Mastodon). Over seven dates spanning three months the authors sample 35 PhishTank URLs, dual-validate them with Google Safe Browsing and VirusTotal, create isolated test accounts, and attempt to post each original URL (plus a TinyURL redirection if the original is blocked). Tables II–IV and the appendix show that TikTok, Instagram and Mastodon blocked none of the 35 originals, Facebook blocked four (~10 %), and Twitter blocked the large majority of both originals and redirections (~69 %), yielding an overall block rate of 23.8 % driven almost entirely by Twitter. The authors conclude that most of the examined platforms lack effective defenses against malicious-link posting, discuss security–usability trade-offs, and list several limitations as future work.","tokens_in":11303,"tokens_out":1064,"duration_ms":23069,"significance":"If the modest claim holds, the work supplies concrete, previously sparse evidence that several high-install social platforms perform little or no real-time filtering of known phishing URLs at the moment of posting. That observation is directly relevant to usable-security design, platform policy, and the residual attack surface for link-based social-engineering. Credit is due for the transparent pipeline (Fig. 1), dual external black-list validation, ethical isolation of test accounts, and the fully tabulated per-app/per-date outcomes that let a reader verify the headline numbers. As explicitly labeled pilot work the study usefully surfaces a measurement gap rather than claiming a definitive audit.","major_comments":[{"comment":"§III-C, §IV and Tables II–III: the central claim that most apps ‘did not have an effective defense against the posting and spreading’ of malicious URLs rests solely on one-time post-time checks of 35 URLs. The authors themselves flag in §VI that retroactive removal and profile-versus-DM differences were never measured. Without those data the stronger language of ‘spreading’ and of ‘non-existent’ defenses for three apps is not fully supported; either the experiment must be extended or the conclusions must be substantially qualified to the post-time window actually observed.","section":"§III-C, §IV, Tables II–III"},{"comment":"§III-B: the two-week sampling gap is presented as sufficient for platforms to have ingested the latest blacklist entries, yet no evidence is offered that the five apps consult PhishTank, Safe Browsing or VirusTotal, nor that they update on that cadence. This untested assumption underpins the inference that non-blocking equals inadequate defense rather than delayed or alternative detection; it should be either validated or removed as a load-bearing premise.","section":"§III-B"},{"comment":"Table III and §IV-A: the headline 23.8 % block rate is almost entirely Twitter’s contribution (46 of the blocked instances). Aggregating without stronger per-app emphasis or simple statistical context risks overstating uniformity of the vulnerability across ‘most’ platforms; the discussion should more carefully separate Twitter’s relatively robust behavior from the near-zero rates of the other four apps.","section":"Table III, §IV-A"}],"minor_comments":[{"comment":"Title and abstract: ‘that usability into account’ is missing a verb; several other sentences in the abstract and introduction contain missing articles or run-on constructions.","section":"Abstract"},{"comment":"Section heading ‘METHODOLGY’ is misspelled; Table III header ‘EVLUATIONMATRIX’ is likewise misspelled and inconsistently capitalized.","section":"§III, Table III"},{"comment":"Inconsistent orthography of ‘Redirectional’ / ‘redirectional’ / ‘transformed’ throughout Tables II–IV and the text; pick one term and use it uniformly.","section":"Tables II–IV"},{"comment":"Fig. 1 caption and body text contain several typographic glitches (‘InProcess1&2we’, ‘createtransformedversion’, missing spaces). Clean for camera-ready.","section":"Fig. 1"},{"comment":"Introduction: ‘for personally identifiable informationPIIssuch’ lacks spaces and punctuation; similar concatenation errors appear elsewhere.","section":"§I"}],"recommendation":"major_revision","confidential_remarks":"The manuscript already appeared at USEC 2023; the arXiv posting appears archival. Scope is appropriate for usable-security venues. No citation or novelty red flags beyond the openly exploratory character of the work. The major-revision recommendation is driven by the need to align claim strength with the acknowledged measurement limits rather than by any internal inconsistency."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"This is a straightforward exploratory measurement paper. The new piece is the side-by-side 2022 snapshot: 35 PhishTank URLs (sampled every two weeks, dual-validated with Safe Browsing and VirusTotal) posted as original and TinyURL-shortened versions on TikTok, Instagram, Twitter, Facebook, and Mastodon. Tables II–IV and the appendix give the per-app, per-date counts that back the headline number—only 23.8 % blocked overall, almost all by Twitter. Four apps essentially let everything through; Facebook blocked a handful and was easily bypassed by redirection. That is concrete, actionable data for usable-security people and for anyone writing about platform policy.\n\nWhat the paper does well is transparency and restraint. The pipeline is simple and fully described, ethical controls (private test accounts, immediate deletion) are explicit, and the authors label the work as pilot and list the obvious next steps (profile vs DM, retroactive removal, larger N) in §VI. No circular reasoning, no invented entities, no over-claim of a new attack class. Citations cover the relevant phishing and blacklist literature without padding.\n\nThe soft spots are exactly the ones the authors already flag: small sample, post-time only, possible differences between public posts and DMs, and no longitudinal check for later takedowns. Those limit how far you can generalize, but they do not undermine the modest claim that is actually made. The discussion of security–usability trade-offs (TikTok stripping hyperlinks) is sensible but thin.\n\nWho it is for: anyone measuring social-media phishing defenses or arguing for stronger Play Store policy language. It is not a methods paper and not a systems paper. I would bring it to a reading group if we are talking about empirical baselines or platform accountability; otherwise it is a quick cite for the “most apps do almost nothing” fact. A serious editor should send it to referees—USEC-style venues routinely take clean pilot measurements of this kind—even if the expected outcome is “accept with revisions that expand N or add the retroactive check.” Worth engaging if the topic is on your plate; not a must-read otherwise.","headline":"Clean, small-N pilot that shows four of five popular Android social apps block almost no verified phishing URLs at post time; useful snapshot, not a new attack or defense.","tokens_in":11840,"tokens_out":548,"would_cite":true,"duration_ms":6017,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Most top Android social apps let verified phishing links be posted with little or no blocking.","keywords":["malicious URLs","phishing","social media security","link sharing","Android apps","URL filtering","usable security","PhishTank"],"falsifier":"Re-run the identical sampling and posting protocol on the same five apps today (or on a larger set of URLs) and record whether the block rates for original and redirected links remain as low as the 2022 figures; any large rise in blocking would falsify the claim of systemic weakness.","tokens_in":12051,"feed_emoji":"🔗","tokens_out":838,"duration_ms":7304,"temperature":0.7,"pith_summary":"This exploratory study tests whether the five most popular Android social apps actually stop users from sharing known malicious links. Researchers created throwaway accounts, drew 35 phishing URLs from PhishTank over three months, double-checked them with Google Safe Browsing and VirusTotal, and tried to post both the original URLs and shortened redirect versions. TikTok, Instagram and Mastodon accepted every original link; Facebook blocked only a handful and those were easily bypassed with redirects; only Twitter blocked a substantial share (about 69 percent of its attempts). Overall just 23.8 percent of the harmful links were stopped. The paper therefore claims that everyday link-sharing features on these platforms constitute a real, usable attack surface that attackers or careless users can exploit, and that platform security design has not kept pace with that risk.","feed_headline":"Most top social apps let phishing links through","feed_subtitle":"Only 24 percent of verified malicious URLs were blocked; Twitter did nearly all the work","key_machinery":"A controlled posting pipeline: sample five PhishTank URLs every two weeks for three months, re-verify each with Safe Browsing and VirusTotal, attempt to post the original URL from a private test account, and, only if blocked, re-post a TinyURL redirect of the same destination. Success or failure of each post is the measured outcome.","core_discovery":"When verified phishing URLs drawn from PhishTank are posted on the five highest-ranked Android social apps, the large majority are accepted: only Twitter performs meaningful blocking, and the aggregate block rate across all apps and both original and redirected forms is just 23.8 percent. The study therefore concludes that most of these platforms lack an effective defense against the posting and spreading of malicious links.","pith_inferences":["The same low post-time filtering likely extends to other high-traffic platforms not in the top-five list, widening the practical attack surface.","Retroactive deletion of already-posted links, if it occurs at all, arrives after the window of highest engagement and therefore does little to reduce initial harm.","A public, continuous monitoring dashboard that re-tests the same pipeline weekly could pressure platforms to close the gap without waiting for new regulation."],"forward_implications":["Attackers can currently seed phishing or malware links on TikTok, Instagram and Mastodon with near-certainty of successful posting.","Even the modest blocking present on Facebook is defeated by ordinary URL shorteners, so redirect chains remain an open bypass.","Google Play Store malware-policy language is not strong enough to force apps to stop the spread of known-bad URLs.","Usable-security redesigns that preserve link-sharing while still intercepting known-malicious destinations are still needed."],"fun_headline_variants":["Most top social apps let phishing links through unchecked","Only 24% of verified phishing URLs blocked across top apps","Twitter does nearly all malicious link blocking on social apps","Top Android social apps fail to filter most phishing URLs","Social apps accept bulk of PhishTank malicious links"],"cache_read_input_tokens":128,"weakest_assumption_plain":"A single post-time check of 35 already-flagged phishing URLs is treated as a fair proxy for each platform's real filtering behavior, including any later removal or differences between profile posts and private messages.","fun_headline_variants_meta":{"raw":{"variants":["Most top social apps let phishing links through unchecked","Only 24% of verified phishing URLs blocked across top apps","Twitter does nearly all malicious link blocking on social apps","Top Android social apps fail to filter most phishing URLs","Social apps accept bulk of PhishTank malicious links"]},"model":"grok-4.5","effort":"low","cost_usd":0.006356,"raw_usage":{"total_tokens":1652,"prompt_tokens":793,"num_sources_used":0,"completion_tokens":81,"cost_in_usd_ticks":63560000,"prompt_tokens_details":{"text_tokens":793,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":778,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":793,"tokens_out":81,"duration_ms":6073,"temperature":1.0,"reasoning_tokens":778,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-11T22:04:15.490968+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Re-run the identical sampling and posting protocol on the same five apps today (or on a larger set of URLs) and record whether the block rates for original and redirected links remain as low as the 2022 figures; any large rise in blocking would falsify the claim of systemic weakness.","supporting_citations":[],"review_version":1}