{"id":"758b71a5-3034-4e63-9990-51c86e068c09","arxiv_id":"2608.04847","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A new dataset and evaluation show that LLMs define queer slang better when given domain framing and sentential context, but still below a paraphrase-based reference bound.","lead":"The authors built Slang-Q, a curated dataset of 1,024 user-generated English sentences paired with queer slang terms and definitions, and tested four language models on defining these terms under different prompts. They find that models lag behind a reference similarity bound and that adding queer-slang framing and sentence context improves their definitions.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Single-annotator labelling with no reliability check is the load-bearing risk: if the 1,024 sentence–term labels are biased, all model scores and the 'context/framing helps' conclusion inherit that bias.","rationale":"The reader's weakest assumption and my load-bearing concern coincide. The paper's strongest claim is a conjunction: (1) models fall below a human upper bound, (2) withholding framing/context hurts, and (3) framing/context help. All three are measured on a dataset whose labels were produced by one annotator. Unlike the pseudo-human baseline issue (the 'Human' row in Table 5 is actually gold-definition vs. GPT-5.5 paraphrase agreement, not human task performance), which affects only the 'human upper bound' phrasing and can be corrected by rewording, unreliable labels would invalidate the dataset itself and every score built on it. Other issues—the 1,204/1,024 sentence-count discrepancy and the claim that the term inventory was built independently of the corpus while 13 terms were added from the examples—are inconsistencies that suggest documentation gaps, not grounds for overturning the empirical pattern. The manual error analysis on 200 outputs is helpful and partially independent, but it was also performed by the authors on the outputs of only two models, so it does not substitute for label reliability. For these reasons I do not think the paper should be rejected; it should be conditional on a reliability check. The reader's verdict of CONDITIONAL therefore remains unchanged.","tokens_in":16405,"tokens_out":6912,"duration_ms":79533,"concrete_test":"Draw a stratified random sample of 150–200 Slang-Q entries spanning the three taxonomy categories. Have a second annotator with comparable queer-community/online-language expertise independently re-annotate each sentence for (i) whether the matched term is used in its intended queer sense and (ii) whether the content is harmful/vulgar, using the paper's stated guidelines. Compute Cohen's kappa or Krippendorff's alpha for both decisions. If agreement is below 0.7, re-run the four-condition evaluation on the subset of entries where both annotators agree; if the Term-baseline deficit over Context/slang-informed conditions shrinks or disappears in that subset, the central claim is an artifact of annotation bias.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The Slang-Q dataset is the entire evaluation substrate, and every sentence–term pair, including the decision that the matched term is used in its queer sense and that the content is non-harmful, was made by the first author alone (Section 3.3: 'Annotation was carried out by the first author'; only 'fewer ambiguous cases' were discussed with the second author). No inter-annotator agreement, adjudication protocol, or reliability statistic is reported. The central claim has two independent legs—models score below a 'human' bound, and withholding framing/context hurts performance—and both legs are computed against gold definitions attached to these single-annotator pairs. If the annotator systematically retained sentences in which the queer sense is strongly disambiguated by context and discarded genuinely ambiguous ones, the context-benefit effect could be inflated or even manufactured: the dataset would contain few cases where context is needed, yet the remaining context would be maximally helpful. Similarly, if gold definitions are drawn from external glossaries that do not match the sense actually intended in the sentence (Section 3.1.1), ROUGE-L/BERTScore comparisons are unreliable independent of model behavior. This is not a cosmetic limitation: without reliability evidence, the dataset cannot support the paper's headline conclusions.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces Slang-Q, a dataset of 1,024 naturally occurring English sentences paired with 118 queer slang terms and reference definitions drawn from external lexicons, together with a taxonomy of queer slang. Using this resource, the authors evaluate four LLMs (Claude Sonnet 4.6, LLaMA 3.3 70B, LLaMA 4 Scout, Qwen3 32B) on a definition-generation task under four prompt conditions that vary whether the prompt is slang-informed and whether an example sentence is provided. Automatic evaluation with ROUGE-L and BERTScore is supplemented by a manual error analysis on two models. The paper concludes that models fall below a human upper bound, that withholding domain framing and sentential context hurts performance, and that both slang-informed prompting and contextual grounding help models converge on the intended queer-specific meaning.","tokens_in":16623,"tokens_out":4561,"duration_ms":49627,"significance":"If the results are reliable, this is a useful contribution to an underserved area: queer slang is underrepresented in NLP, and Slang-Q could become a reference resource for evaluating LLM understanding of community-specific language. The authors are transparent about data sources, make the repository available, include a data-contamination check, and provide a manual error analysis that partially corroborates the automatic findings. However, the current evidence is exploratory: the automatic metric differences are small, the 'human' bound is not a true human-performance measurement, and the entire dataset rests on single-annotator judgments without reliability checks. These issues are central to the paper's headline claims rather than cosmetic.","major_comments":[{"comment":"The annotation was carried out by the first author alone, with only 'fewer ambiguous cases' discussed with the second author, and no inter-annotator agreement or reliability statistic is reported. Because every sentence-term pairing, every judgment that a term is used in its intended queer sense, and every harmfulness decision is single-annotator, the two headline findings--that models score below a human bound and that context/framing improve performance--inherit any systematic bias in these labels. If the annotator retained sentences in which the queer sense is strongly disambiguated by context and discarded genuinely ambiguous ones, the context-benefit effect could be inflated or even manufactured. The manuscript should report a reliability study (e.g., a second annotator on a stratified sample with Cohen's kappa or equivalent) and an adjudication protocol, or the conclusions must be substantially softened.","section":"Section 3.3"},{"comment":"The row labeled 'Human' is not a measure of human performance; it is the mean similarity of the gold reference definition to the two GPT-5.5-generated alternative definitions. The conclusion in Section 6 that 'models fall below the human upper bound' is therefore not supported by the data as presented. Either collect and score actual human-written definitions for a sample of terms, or relabel this quantity as a reference-agreement ceiling and revise the wording of the finding accordingly.","section":"Section 5, Table 5"},{"comment":"The reported condition differences are very small relative to the standard deviations (e.g., ROUGE-L 0.16 vs 0.18, BERTScore 0.85 vs 0.86, with SDs around 0.05-0.07), and no significance tests are reported. The claim in Section 6 that withholding domain framing and sentential context 'consistently hurts performance' requires paired significance testing across terms (e.g., Wilcoxon signed-rank or bootstrap) and effect sizes. The manual evaluation is suggestive but covers only two models and two conditions, so it cannot by itself establish the cross-model claim.","section":"Section 5, Table 5"},{"comment":"Claude Sonnet 4.6 shows a high verbatim contamination range (57.30-62.00%) with the source corpus, and the context condition supplies the exact example sentences from that corpus. Although the authors argue that memorization of a sentence does not entail ability to solve the definition task, memorized examples could plausibly inflate performance specifically in the context conditions. Please provide a per-model breakdown of the context benefit and discuss how contamination affects the interpretation, or explicitly control for it in the analysis.","section":"Section 4.4 and Section 5"}],"minor_comments":[{"comment":"The discard counts (1,833 semantically unrelated + 327 harmful/vulgar = 2,160) leave 1,024 of 3,184 sentences; stating the retained percentage explicitly (about 32%) would prevent reader confusion.","section":"Section 3.3"},{"comment":"There is a typo in 'showes' (should be 'shows'), and the condition naming is inconsistent between 'Term baseline' in the text and 'Term (Base)' in the first result paragraph.","section":"Section 5"},{"comment":"The 'Human' row would benefit from a footnote stating explicitly that it measures agreement among the gold and paraphrased reference definitions, not human-written outputs.","section":"Table 5"},{"comment":"The category panels aggregate across all models and conditions, which can mask important interactions; adding error bars or significance markers would make the visual claims more interpretable.","section":"Figures 1 and 2"},{"comment":"The limitations paragraph should mention the single-annotator reliability issue and the lack of significance tests, since these are the main threats to the stated conclusions.","section":"Section 6"}],"recommendation":"major_revision","confidential_remarks":"The paper is framed as a first exploratory evaluation, which is appropriate for a workshop venue, but the strength of the claims in Section 6 exceeds what the evidence supports. The single-annotator issue is the most serious concern; I would require at least a reliability check on the dataset before publication, along with a re-labelling of the 'Human' row and significance testing for the condition comparisons."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. This paper introduces Slang-Q, a new dataset of 1,024 sentences with queer slang terms and definitions, plus a 118-term taxonomy. That's a real resource, and the paper is the first I know to ask whether LLMs can define queer slang rather than just measuring bias or harm. The experimental design is clean: four conditions cross context and domain framing, four models, automatic and manual evaluation.\n\nThe good news: the manual error analysis is the most convincing part. On 200 outputs, Claude goes from 70% correct in the term-only condition to 100% with slang-informed context; Llama 4 goes from 64% to 90%. That's a genuine effect, and the qualitative examples are instructive. The taxonomy is thoughtfully organized and should be useful to others.\n\nThe soft spots are real. Automatic metric differences are tiny—ROUGE-L 0.16 to 0.18, BERTScore 0.85 to 0.86—with no significance tests. The 'human upper bound' is not human performance; it's the similarity between the gold definition and two GPT-5.5 paraphrases, so it sets a ceiling but doesn't tell us what a person would score. The abstract says 1,204 sentences, the body says 1,024. The inventory is described as independent of the corpus, yet 13 terms were added after seeing the examples. Most importantly, the whole dataset rests on single-annotator labeling with no inter-annotator agreement. If the annotator kept only sentences where the queer sense is strongly disambiguated, the context benefit could be inflated. That's the load-bearing risk, and the stress-test note is right about it.\n\nNone of these are fatal alone. The central direction—framing and context help—is plausible and supported by the manual sample. But the quantitative core is weaker than the prose suggests. It reads more like a dataset paper with a preliminary evaluation than a definitive finding about model capability.\n\nWho is this for? People working on slang, queer NLP, and LLM evaluation. The dataset is a starting point, the taxonomy is citable. With revisions—IAA on a subset, fixing the count, reframing the human baseline, and either running significance tests or clearly labeling the numbers as descriptive—this would be a solid workshop or short-paper contribution.\n\nMy recommendation: send it to peer review. It deserves a serious referee. It is not ready as is, but it is a legitimate first step.","headline":"A genuinely useful new dataset and a plausible first result on LLMs and queer slang, but the quantitative core is thin and the single-annotator labeling is the load-bearing risk.","tokens_in":17177,"tokens_out":4697,"would_cite":true,"duration_ms":44157,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Large language models can define queer slang reliably when given an example sentence or a domain cue, but without either they fall back on the word's everyday meaning.","keywords":["queer slang","LGBTQIA+ language","user-generated content","large language models","definition generation","prompt design","slang taxonomy","polysemy"],"falsifier":"Independently re-annotate a random sample of the 1,024 sentences with several annotators who are familiar with online queer slang, and measure inter-annotator agreement; if a substantial share of the queer-sense and harmfulness labels are not reproduced, the gold standard and the prompt-condition scores built on it collapse.","tokens_in":16187,"feed_emoji":"🌈","tokens_out":9862,"duration_ms":98130,"temperature":0.7,"pith_summary":"This paper asks whether large language models actually understand queer slang as it appears in real user-generated text, and it answers with a new evaluation resource. The authors build Slang-Q, a manually curated set of 1,024 English sentences from an online slang dictionary, paired with 118 queer-related terms, definitions, and a two-level taxonomy. They then test four instruction-tuned models on a definition-generation task under four prompt conditions. The consistent finding is that models score below a human reference bound, and performance drops most when the model receives only the bare term with no indication that it is queer slang; adding a domain cue, an example sentence, or both pushes definitions toward the intended queer-specific sense. Manual review shows fully incorrect answers drop sharply in the context-plus-framing condition, which matters because it means simple prompting choices can change whether an LLM gives accurate information about community-specific language.","feed_headline":"One example sentence lifts AI definitions of queer slang","feed_subtitle":"Models given only the term default to everyday senses of words like bear and tea; framing or an example restores the queer sense.","key_machinery":"The load-bearing object is Slang-Q: a manually curated dataset of 1,024 user-generated English sentences, each paired with one of 118 queer-related terms and a reference definition, organized by a two-level taxonomy with broad categories (Identity, Slang, Intersectional) and optional subcategories (Reclaimed, Shorthand, Idiomatic expression, Pronoun, Spelling variation). The evaluation machinery is the crossed design of two binary prompt dimensions—domain framing (generic language expert versus queer internet slang expert) and context (term only versus term plus example sentence)—which isolates the effect of each on definition generation. Scores are computed with a lexical-overlap metric and a semantic-similarity metric against three reference definitions per term, then aggregated at the term level so that frequent terms do not dominate the results.","core_discovery":"On its own terms, the central discovery is that what an LLM knows about queer slang is not a fixed quantity: it is strongly conditioned by whether the prompt identifies the domain and whether an example sentence is present. Across all four evaluated models, the bare-term baseline scores lowest on both automatic metrics, while slang-informed framing alone and contextual grounding alone each improve scores, and the two together perform best. Manual inspection of 200 outputs from the two conditions with the strongest contrast supports this: in the bare-term condition, roughly a quarter to a third of definitions miss the intended sense, while in the slang-informed context condition the share of fully incorrect definitions drops to zero in one model and to a small fraction in the other, though some definitions still omit sociocultural nuance. The paper also finds that all models fall below the human reference bound, that no model dominates, and that intersectional terms—those shared with other communities such as AAVE or fandom—are the hardest category. A data-contamination probe finds that one proprietary model appears to have memorized far more of the source sentences than the open-weight models, yet this does not translate into better definitional performance.","pith_inferences":["Because the model with the highest estimated memorization of the source corpus does not outperform the others, the paper's data point implicitly suggests that exposure to slang-heavy text during training is not the limiting factor for community-specific understanding; prompt and context design may be more actionable.","The taxonomy's 'Intersectional' label acknowledges that queer slang overlaps with AAVE and fandom usage; a fairer scoring protocol might compare model definitions against multiple gold definitions, one per community, since a single reference definition will penalize valid senses.","A direct testable extension is to apply the same two-by-two prompt design to other dynamic sociolects, such as disability community language or regional slang, to see whether the context-and-framing effect generalizes beyond queer slang."],"forward_implications":["A model asked to define a slang term in isolation will make noticeably more errors than the same model given a single example sentence, so user-facing systems that explain community slang should avoid bare-term queries.","Explicitly telling the model that it is dealing with queer slang is enough by itself to shift some definitions toward the queer-specific sense, and combining that framing with an example sentence yields the strongest correctness in the manual evaluation.","Intersectional terms that circulate in multiple communities, such as AAVE or fandom, are the hardest to define, meaning queer-slang evaluation should treat multi-community usage as a first-class difficulty rather than noise.","Since all four tested models cluster within a narrow performance band, the results point to broadly similar queer-slang knowledge across model families and sizes, with none reaching the human reference bound.","The Slang-Q dataset provides a reusable evaluation set for future work on definition generation, prompt design, and slang handling in large language models."],"supporting_citations":[{"why":"Supplies the source corpus of naturally occurring user-generated slang sentences that the extraction pipeline starts from.","marker":"[5]"},{"why":"Provides the seed term inventory of 46 terms and prior evidence that LLMs respond differently to queer slang, motivating the evaluation.","marker":"[13]"},{"why":"Contributes 24 queer-community terms to the inventory and frames the question of coverage of queer minorities in lexical resources.","marker":"[23]"},{"why":"Supplies the taxonomy inspiration and many of the reference definitions used as gold standards for evaluation.","marker":"[16]"},{"why":"Provides the black-box contamination quiz used to estimate how much of the source corpus each model may have memorized.","marker":"[30]"},{"why":"Provides the lexical-overlap metric used to score generated definitions against the gold references.","marker":"[28]"},{"why":"Provides the semantic-similarity metric used alongside the lexical metric to capture meaning beyond surface form.","marker":"[29]"}],"fun_headline_variants":["Queer slang knowledge in LLMs hinges on prompt context","Prompt framing changes how LLMs define queer slang","LLMs flounder on queer slang without context clues","Context boosts LLM definitions of queer slang terms","Intersectional queer slang stumps language models"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole evaluation rests on one author's manual judgment, without measured inter-annotator agreement, of which sentences use each term in its intended queer sense and which content is harmful; if those labels are wrong or inconsistent, the gold definitions and all model scores built on them shift.","fun_headline_variants_meta":{"raw":{"variants":["Queer slang knowledge in LLMs hinges on prompt context","Prompt framing changes how LLMs define queer slang","LLMs flounder on queer slang without context clues","Context boosts LLM definitions of queer slang terms","Intersectional queer slang stumps language models"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000155,"raw_usage":{"total_tokens":1178,"prompt_tokens":874,"completion_tokens":304,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":490,"completion_tokens_details":{"reasoning_tokens":230}},"tokens_in":490,"tokens_out":304,"duration_ms":3722,"temperature":1.0,"reasoning_tokens":230,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T15:00:43.468336+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Independently re-annotate a random sample of the 1,024 sentences with several annotators who are familiar with online queer slang, and measure inter-annotator agreement; if a substantial share of the queer-sense and harmfulness labels are not reproduced, the gold standard and the prompt-condition scores built on it collapse.","supporting_citations":[{"cited_title":"Tint, Guardrails, not guidance: Understanding responses to LGBTQ+ language in large language models, in: A","cited_arxiv_id":null,"evidence_quote":"Provides the seed term inventory of 46 terms and prior evidence that LLMs respond differently to queer slang, motivating the evaluation."},{"cited_title":"URL: https://lexicon.library.lgbt/","cited_arxiv_id":null,"evidence_quote":"Supplies the taxonomy inspiration and many of the reference definitions used as gold standards for evaluation."},{"cited_title":"Lin, ROUGE: A package for automatic evaluation of summaries, in: Text Summarization Branches Out, Association for Computational Linguistics, Barcelona, Spain, 2004, pp","cited_arxiv_id":null,"evidence_quote":"Provides the lexical-overlap metric used to score generated definitions against the gold references."}],"review_version":1}