{"id":"d55b4e8d-a50d-48ab-979f-70ea516a1ade","arxiv_id":"2601.17601","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A reader-centered taxonomy of hyperlink-sharing intentions—6 top-level and 26 fine-grained classes—was built from crowdsourced annotations and LLM refinement, and intent labels modestly improved microblog retrieval in a 15-query demo.","lead":"This paper builds a six-category, 26-class taxonomy for why people include hyperlinks in social media posts, using crowd annotations and an LLM-assisted naming step. It reports intention distributions on 1,000 tweets and a small retrieval experiment suggesting intent features can improve BM25 ranking.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The taxonomy's stability and the empirical claims rest on a single, low-agreement AMT labeling run; without independent re-annotation or split-half replication, the 6+26 structure, prevalence figures, and retrieval gains are not yet anchored.","rationale":"The reader's weakest assumption and my concern coincide: the paper treats a single set of five crowd labels with only 'fair' overall agreement as usable ground truth for prevalence and retrieval, and the taxonomy itself is the product of one unreplicated open-coding/grouping exercise. The paper does have real strengths—the expert check on the high-consensus subset, the transparent reporting of low kappa, the release of the taxonomy itself, and a genuinely bottom-up construction path. These justify keeping the current CONDITIONAL verdict rather than moving to REJECT. However, the central claim that this is a stable, reusable reader-centered taxonomy is not yet supported by the evidence, because the only validation of the category structure is a single run with a small group of annotators and no independent replication. My proposed check—an independent re-annotation and rerun of the retrieval experiment—directly tests whether the prevalence and utility numbers survive a change of annotators. If they do, the taxonomy is substantially strengthened; if they do not, the paper's empirical claims need to be revised. The retrieval demonstration on 15 queries also lacks significance testing, but that is secondary to the label-reliability issue and the reader already noted it. I therefore recommend no change to the reader's verdict.","tokens_in":18064,"tokens_out":5807,"duration_ms":72648,"concrete_test":"Release the 1,000 tweet-URL pairs and the five per-item labels, and run an independent replication: have a fresh set of five screened annotators re-label a random 300-tweet subset using the released taxonomy. Then (a) compute majority-label agreement with the original labels (e.g., Cohen's kappa); (b) compare the prevalence distribution from the original and replication labels; (c) re-run the Section 5.2 BM25+intent ranking experiment on the 15 TREC queries using the replication labels instead of the original ones. If majority-label kappa is below 0.6, any top-level prevalence estimate shifts by more than 5 points, or the nDCG@10/MAP gains are not reproduced with the new labels, the central empirical claims should be downgraded from CONDITIONAL to REJECT or UNVERDICTED.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central empirical claims—the 6+26 taxonomy, the prevalence of promote/converse/share, and the retrieval improvement—are all downstream of one small, unreplicated annotation pipeline. The paper's own reliability evidence is weak: Fleiss' kappa is 0.216 over all 1,000 items and only 0.259 even on high-consensus items (Section 4.3). The taxonomy was built from a single open-coding pass by 25 AMT workers on 100 tweets each, then grouped by majority vote of 5 workers (Section 3.2); there is no split-half analysis, no independent re-derivation, and no check that the same categories would emerge from a different sample or different annotators. The expert check (Cohen's kappa 0.793) covers only the 775 high-consensus tweets, so it cannot validate the 22.5% of posts with no high consensus, nor does it establish that the six top-level categories are stable outside that subset. The prevalence distribution in Figure 4 and the BM25+intent gains in Table 5 are computed from majority labels produced by the same five workers. Neither the tweet-URL data nor the per-item labels are released, so the numbers cannot be independently checked from the preprint. This does not disprove the taxonomy, but it makes the core claim of a reusable, empirically grounded reader-centered scheme depend entirely on an unverified single run.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a reader-centered taxonomy of intentions behind URLs in social media posts. It describes a bottom-up crowdsourcing pipeline: 25 screened AMT workers open-code 2,500 tweets into 754 codes, which are cleaned and grouped by 5 AMT workers into 28 groups and then 6 coarse categories, refined with LLM assistance to produce 6 top-level categories and 26 fine-grained intention classes. A second study applies the taxonomy to 1,000 tweets with 5 workers, reporting prevalence (promote, converse, share are most common) and a follow-up context-augmented annotation for low-agreement items. The taxonomy is compared with prior tweet-intent taxonomies, and a retrieval experiment on 15 TREC 2011 queries reports that augmenting BM25 with intent labels improves nDCG@10 from 0.4166 to 0.4374 and MAP from 0.4518 to 0.4757.","tokens_in":18366,"tokens_out":7481,"duration_ms":79860,"significance":"If the taxonomy proves stable, it would be a useful, reusable resource for intent-aware IR and social media analysis. Strengths include the collection of a large tweet-URL dataset, detailed AMT worker screening, public release of the taxonomy, manual curation of authentic examples, a hybrid human-LLM construction described stepwise, and an expert agreement check on a high-consensus subset (Cohen's kappa 0.793). The comparison to prior taxonomies and the Study 2 context-augmentation experiment are also useful. However, the empirical grounding is currently incomplete: the overall annotator agreement is low, the internal validity of the 6+26 structure is not established by independent replication, the pipeline has unexplained numerical transitions, and the retrieval demonstration is too small and lacks controls. The contribution is better viewed as a plausible, well-motivated taxonomy proposal than as a fully validated empirical result.","major_comments":[{"comment":"The pipeline's numerical transitions are unexplained. The text reports 754 codes, then 28 intention groups with 442 fine-grained intentions, then 'another round' yielding 6 coarse classes, but Table 1 lists exactly 26 fine-grained classes. It is not stated how 442 fine-grained intentions were reduced to 26, how the 28 groups were mapped to the 6 top-level categories, or how the LLM refinement changed the earlier crowd groupings. This is not a presentation detail: it is the derivation of the central artifact, and without it the 6+26 structure is not reproducible. Please provide the grouping/merging decision rules or an appendix mapping the 442 codes to the 26 classes, and clarify the role of GPT-5 versus human consensus.","section":"3.2–3.3, Table 1"},{"comment":"The reliability evidence is insufficient for the prevalence claims. Fleiss' kappa = 0.216 over all 1,000 items and 0.259 on the high-consensus subset is low; 'high consensus' is defined as 4/5 agreement, yet 22.5% of items have no high consensus. The expert Cohen's kappa of 0.793 covers only the 775 high-consensus tweets, so it does not validate the taxonomy on the NC-UN subset or on the fine-grained labels outside that subset. The distribution in Figure 4 is based on the same five workers' labels. Because the development and evaluation runs are unreplicated, the stability of the 6+26 structure and the reported prevalence are not established. I would ask for a split-half or independent re-annotation of a random sample, or release of per-item labels so the community can assess agreement.","section":"4.3, Figure 4"},{"comment":"The retrieval demonstration is not yet convincing. The reported gains are from 15 TREC queries with no significance testing; the only comparison is BM25 vs BM25+Intent. The augmentation concatenates inferred query-intent words and tweet-intent labels into the text scored by BM25, so both sides gain the same label vocabulary. A control condition with random labels or with a matched non-intent feature is needed to show that the improvement comes from intent semantics rather than from adding any token to both query and document. Also, the method for inferring query intent from [46] is only referenced, not operationalized, and the tweet-intent annotation in the top-50 is not described (who labeled, with what agreement). Please report per-query results and a significance test.","section":"5.2, Table 5"},{"comment":"The paper releases only the taxonomy (GitHub), not the 2,500 development annotations or the 1,000 labeled tweets used for Figure 4 and Study 2. The central empirical numbers therefore cannot be independently checked. In a submission whose contribution is an 'empirically grounded' scheme, the underlying labels should be made available, or a clear reason given for withholding them. This is especially important given the low overall agreement and the absence of a replication run.","section":"Data availability"}],"minor_comments":[{"comment":"Arithmetic check: 754 codes minus 311 discarded due to similarity yields 443, not 442. If this is not a typo, explain the discrepancy.","section":"3.2"},{"comment":"The context-study results for the 83 NC-UN reply tweets are reported only as the percentage with a majority intention (80% and 78%). Report agreement metrics (e.g., Fleiss' kappa) for that subset as well, since the claim is that context improves annotation reliability.","section":"4.3"},{"comment":"The claim that prior categories are 'fully subsumed' is stronger than the mapping supports; for example, 'Personal Message' from [47] and 'Chat' from [29] are aligned with broader categories. Provide a more detailed mapping with definitions, or soften the claim.","section":"5.1, Table 4"},{"comment":"In the 'thorpe return in 2012 olympics' example, the query intent classification and the linked-content inspection are asserted without evidence. Specify how the query was classified and how the linked content was inspected, or treat this as an illustrative anecdote rather than a demonstration.","section":"5.2"}],"recommendation":"major_revision","confidential_remarks":"I see a publishable qualitative contribution if the authors can close the reliability gap and make the retrieval demonstration credible. The paper would also be strengthened by releasing the labeled data. I would not reject on the basis of disagreement with the taxonomy itself; the problem is the missing evidence between raw annotations and the final structure, and the unvalidated downstream claims."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing to know: the contribution here is the taxonomy itself, not the retrieval experiment. The taxonomy is genuinely new — reader-centered, URL-specific, built bottom-up from crowd codes and then cleaned up with LLM help. I've seen the tweet-level taxonomies they compare against, and the 32.5% 'None/Other' rate in their pilot is a real gap. The authors walk through the construction in enough detail to reproduce the spirit of it, and they've released the taxonomy with examples. That is a solid, citable resource.\n\nThe strongest evidence in the paper is the expert check on the high-consensus subset: Cohen's kappa 0.793 is good. The authors are also transparent about the overall agreement being low — Fleiss' kappa 0.216 is 'fair' by any standard, and they don't pretend otherwise. They investigate the hard cases in Study 2 and show that adding thread context pushes majority agreement to around 80%, which is a useful finding for anyone doing this kind of annotation.\n\nWhere the paper gets wobbly: first, the reliability. The taxonomy was derived from one pass of 25 workers on 100 tweets each, then grouped by majority vote of 5 workers. There is no independent re-derivation, no split-half, no check that the same 6+26 structure would emerge from a different sample. The 0.216 kappa means a fifth of the labels are basically noise at the item level. The prevalence figures (promote, converse, share as top intents) rest on majority labels from those same five workers, so take them as indicative, not definitive.\n\nSecond, the retrieval demo. Fifteen queries, manually annotated intents, no significance test. The nDCG gain of 0.02 could be real, but it's not demonstrated. The authors call it a 'demonstration,' which is honest, but the abstract overstates it slightly.\n\nThird, Section 5.1 overclaims: 'comprehensively covers all previously proposed intent categories' and extends to posts without URLs. The first is plausible but unproven; the second is an assertion with no evidence.\n\nOverall, this is a paper for people who need a starting point for labeling URL intent — they'll get a well-organised, example-backed taxonomy. It deserves a serious referee, but the revision should address the stability of the taxonomy, either by releasing the annotation data or by adding a second annotation run, and by dialing back the retrieval and comprehensiveness claims.","headline":"A genuinely new, well-documented URL-intent taxonomy that deserves a serious look, but the reliability and retrieval claims need a firmer footing.","tokens_in":18899,"tokens_out":3071,"would_cite":true,"duration_ms":32322,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper proposes a reader-centered taxonomy of hyperlink-sharing intent in social posts—six broad categories and 26 fine-grained classes—and shows that adding the inferred intent as a retrieval feature improves microblog search.","keywords":["intent taxonomy","hyperlinks","social media","microblog retrieval","crowdsourcing","large language models","Twitter","reader-centered intent"],"falsifier":"An independent re-annotation of the same 1,000 tweets by a fresh set of annotators that fails to reproduce majority labels at a comparable rate, or a larger-query retrieval experiment in which the intent feature yields no gain over BM25, would undermine the reliability and utility claimed for the taxonomy.","tokens_in":17906,"feed_emoji":"🔗","tokens_out":6619,"duration_ms":61333,"temperature":0.7,"pith_summary":"The paper sets out to answer a question that is hard to observe directly: why do people include a URL in a social post? Rather than surveying authors, it asks readers to interpret the purpose of each link and builds a taxonomy of perceived intentions from open-ended crowd annotations refined with LLM assistance. The resulting scheme has six top-level categories and 26 specific intention classes. Applied to 1,000 real tweets, it finds advertisement, argument, and sharing to be the most common perceived intents. The authors further show that adding this intent signal as a retrieval feature produces a consistent gain in microblog search, suggesting the taxonomy has practical use beyond labeling.","feed_headline":"Link intent in tweets fits 6 categories and lifts search","feed_subtitle":"New taxonomy maps reader-perceived purposes of URLs; intent-aware reranking beats BM25 on TREC-2011 queries.","key_machinery":"The load-bearing object is the intent taxonomy itself: a two-level hierarchy of six categories and 26 fine-grained intention classes. It is produced by a hybrid, three-phase pipeline: (A) open coding—screened crowd workers freely type the perceived purpose of the URL in 2,500 tweets; (B) cleaning—proofreading, standardizing, purging non-intent statements, and deduping reduce 2,500 raw codes to 754; and (C) affinity mapping—two rounds of card-sorting by five workers each group the codes into 28 fine-grained intentions and then 6 coarse categories. An LLM then proposes descriptive names, definitions, and examples, which the research team reviews; the final examples are manually curated from re","core_discovery":"The paper's central claim is that the perceived intention behind a URL in a post can be organized into a compact, reusable taxonomy: six high-level categories—information sharing, entertainment/humor, assistance/information provision, discussion/opinion expression, promotion/advertisement, and request/call for action—with 26 fine-grained intention classes. The taxonomy is derived bottom-up from more than two thousand open codes produced by screened crowd workers, then refined using a large language model to assign descriptive names, definitions, and examples, with all canonical examples manually curated from real posts. The paper also claims empirical support: in 1,000 randomly sampled tweet","pith_inferences":["If the released taxonomy is paired with an automated classifier trained on the 1,000 labeled posts, it could become a general social-media feature for ranking, recommendation, and misinformation filtering; the paper itself does not train such a classifier.","The retrieval gain is demonstrated on only 15 TREC queries, so part of the lift may come from query expansion rather than the semantic content of the intent labels; an ablation that removes the intent term would separate these effects.","The prevalence ranking (advertising, arguing, sharing) is likely specific to this Twitter sample and time window; re-running the same annotation protocol on other platforms or eras would test whether the six-category structure remains stable.","The reader-centered framing suggests that perceived intent, not the author's private motive, is the operationally relevant signal for information retrieval; a small interview study could verify how often perceived intent matches author intent."],"forward_implications":["Intent-aware retrieval: appending hyperlink intent labels to queries and tweets yields higher nDCG@10 (0.4166 to 0.4374) and MAP (0.4518 to 0.4757) on the TREC 2011 microblog queries.","Filtering mismatches: intent labels can remove high-lexical-overlap results whose purpose (e.g., humor) conflicts with a factual query, improving precision.","Crisis response: distinguishing informative, actionable URLs from promotional or deceptive ones can help prioritize useful posts in disaster settings.","Misinformation: categorizing link-sharing intent provides a principled way to characterize and detect deceptive dissemination patterns.","Coverage: the taxonomy subsumes prior tweet-intent schemes and adds an Entertainment/Humor category, making it applicable to both linked and unlinked posts."],"fun_headline_variants":["6 link intents in tweets cut into a new taxonomy","Tweet links decoded: 6 reader-perceived intents","Why link? New taxonomy sorts 6 tweet intents","Taxonomy of link intent improves microblog retrieval","Link purpose in posts: 6 categories aid search"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"That the perceived intentions of a small group of screened crowd workers are a reliable and stable proxy for the true intent behind a URL-sharing post, despite only fair inter-annotator agreement (Fleiss' kappa 0.216) and 22.5% of tweets lacking high consensus.","fun_headline_variants_meta":{"raw":{"variants":["6 link intents in tweets cut into a new taxonomy","Tweet links decoded: 6 reader-perceived intents","Why link? New taxonomy sorts 6 tweet intents","Taxonomy of link intent improves microblog retrieval","Link purpose in posts: 6 categories aid search"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000175,"raw_usage":{"total_tokens":1127,"prompt_tokens":756,"completion_tokens":371,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":500,"completion_tokens_details":{"reasoning_tokens":307}},"tokens_in":500,"tokens_out":371,"duration_ms":4751,"temperature":1.0,"reasoning_tokens":307,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T08:13:54.885562+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"An independent re-annotation of the same 1,000 tweets by a fresh set of annotators that fails to reproduce majority labels at a comparable rate, or a larger-query retrieval experiment in which the intent feature yields no gain over BM25, would undermine the reliability and utility claimed for the taxonomy.","supporting_citations":[],"review_version":1}