{"id":"5caf7988-8bed-4dff-aa9b-0a17050aba93","arxiv_id":"1908.07336","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Applying TCAV interpretability to two toxicity models reveals that a debiased model version changes how LGBTIQ+ identity terms are associated with toxicity in internal representations.","lead":"This paper introduces Project Respect, a crowdsourcing effort to collect positive statements from LGBTIQ+ communities, and applies TCAV to two Perspective API toxicity models to measure how internal representations encode social identity concepts. The authors find that the two model versions encode different associations between LGBTIQ+ identity terms and toxicity, suggesting TCAV can serve as a normative auditing lens for ML models.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central claim depends on TCAV concept sets isolating 'respect', but Project Respect statements differ from baseline examples in sentiment and register; without matched controls, TCAV scores may reflect toxicity/sentiment rather than normative values.","rationale":"The reader's verdict is CONDITIONAL with medium correctness risk, and the identified weakest assumption concerns confounds between the two model versions. I agree that model-version confounds are real, but the more load-bearing gap lies one level earlier: TCAV concept sets do not isolate the normative concept 'respect' from general sentiment, register, or identity-term presence. Even if the two model versions were otherwise identical, the central claim that TCAV provides a 'normative lens' would still require evidence that the learned concept directions correspond to the intended normative values. The paper provides no such evidence: no TCAV accuracy values, no matched control concepts, no permutation baselines, and no statistical tests beyond the displayed intervals. This concern is consistent with the reader's conditional verdict rather than a reason to reject outright; the paper is an early demonstration and the missing validation is addressable with additional experiments. I therefore recommend keeping the verdict unchanged, with the condition that the TCAV concept validity and model-version confounds be addressed before the claim is accepted as robust.","tokens_in":2088,"tokens_out":2628,"duration_ms":29420,"concrete_test":"Run a matched-control TCAV analysis on both toxicity@1 and toxicity@6. Construct a control concept set of positive, respectful statements about non-marginalized, neutral topics (e.g., parks, books, science) matched to Project Respect statements for sentiment, length, and register, and a second control set in which LGBTIQ+ identity terms in Project Respect statements are replaced with neutral nouns. Compare TCAV scores for these controls against the Project Respect concept scores. If the controls produce statistically indistinguishable TCAV scores, the reported differences reflect sentiment or register rather than social respect; if controls score near zero while Project Respect scores remain non-zero, the normative-lens interpretation is supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"TCAV estimates the importance of a concept by learning a direction in activation space that separates concept examples from random examples, then measuring how often the model's prediction changes when moved along that direction. For the claim that 'the internal representations of the two neural networks are encoding different kinds of information about the interaction between social identities and normative values' to hold, each learned direction must correspond to the intended concept. The four example sets differ on many confounded dimensions at once: toxic vs non-toxic content, presence vs absence of LGBTIQ+ identity terms, positive vs neutral register, and different data sources (web comments vs Project Respect submissions). A direction separating Project Respect positive statements from random non-toxic examples can be driven by sentiment, formality, topic, or even sentence length rather than by 'respect' as a normative value. In particular, toxicity@6 was trained with bias mitigation likely to reduce identity-term/toxicity correlations; a positive-statement direction correlated with non-toxicity would then show lower TCAV scores mechanically, without implying a change in normative encoding. The paper reports only point estimates and confidence intervals in Figure 1, with no TCAV accuracy per concept, no overlap/randomization checks, no matched control concepts, and no statistical test of the difference. Without such validation, the interpretive step from TCAV scores to a 'normative lens' is not established.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a normative lens for interpreting machine learning models by combining TCAV (Testing with Concept Activation Vectors) with a new crowdsourced dataset, Project Respect, containing positive statements from marginalized communities. It applies TCAV to two versions of the Perspective API toxicity model and reports mean TCAV scores with 90% confidence intervals for four concept sets: toxic web comments about LGBTIQ+ identities, neutral web comments about those identities, and two Project Respect positive-statement sets. The authors claim the differing TCAV scores between toxicity@1 and toxicity@6 show that the two models encode different information about the interaction between social identities and normative values, and they argue this demonstrates a usable normative lens for auditing model internals.","tokens_in":2358,"tokens_out":3476,"duration_ms":36627,"significance":"If the empirical result were validated, the paper would make a useful contribution by extending TCAV from visual and generic concepts to normative values in language models, and by introducing Project Respect as a resource for value-sensitive auditing. The paper's strengths are that it uses a well-established interpretability method (TCAV), a concrete artifact (Project Respect), and a clearly stated normative motivation connecting model internals to disparate impact. These are meaningful first steps in an important direction. However, the central empirical claim currently rests on a single qualitative figure, and the missing controls and statistical validation are substantial. The proposed resource and framing are likely to be of interest to the fairness, accountability, and transparency community, even though the evidence in this version is preliminary.","major_comments":[{"comment":"The comparison between toxicity@1 and toxicity@6 is confounded. The paper states only that version 6 had bias mitigation 'similar to' the techniques in [1], without reporting whether the architecture, training data, training procedure, or other hyperparameters also changed. As a result, any TCAV difference between the two models cannot be attributed specifically to bias mitigation or to a change in normative encoding. The authors should provide the exact model versions and their training details, or use controlled model pairs that differ only in the bias mitigation technique.","section":"Normative model insights, Figure 1"},{"comment":"The four concept sets differ simultaneously along sentiment, presence of identity terms, register, source, and likely sentence length. The TCAV direction for the Project Respect examples may therefore capture general positivity, formality, non-toxicity, or topic rather than 'respect' as a normative value. In particular, since the Project Respect statements are positive by design, their TCAV direction is likely correlated with non-toxicity, so lower TCAV scores in toxicity@6 could follow mechanically from bias mitigation reducing the identity-toxicity correlation, without implying any change in normative encoding. The authors should add matched control concepts (e.g., positive statements about non-marginalized topics, sentiment-matched non-toxic comments from the same source), report TCAV accuracy per concept and concept set sizes, and include randomization or overlap checks to demonstrate that each learned direction corresponds to the intended concept.","section":"Normative model insights, concept sets"},{"comment":"No statistical test is reported for the difference between toxicity@1 and toxicity@6, so the claim of 'very different results' is not quantitatively established. The figure lists point estimates and 90% confidence intervals, but without a significance test, an effect size, or a report of whether the intervals overlap, the reader cannot assess whether the differences are meaningful. Additionally, the paper omits key TCAV parameters such as the target layer, number of trials, sizes of random baseline sets, and whether the identical example sets were used across both models; these details are necessary for reproducibility and for judging whether the results are robust.","section":"Figure 1"},{"comment":"The conclusion that the two neural networks are 'encoding different kinds of information about the interaction between social identities and normative values' is stronger than the evidence supports. TCAV measures the sensitivity of a model's prediction to movement along a learned concept direction; it does not directly characterize the model's internal representation. The claim should be softened to state that the models' predictions show different sensitivity to the concept directions, or additional probing experiments should be provided to substantiate the representational claim.","section":"Discussion and Conclusion"}],"minor_comments":[{"comment":"Use 'TCAV' consistently; the text alternates between 'tcav' and 'TCAV'.","section":"Throughout"},{"comment":"The figure would be easier to interpret with a table of numeric values and explicit labels for the two Project Respect conditions; the current grayscale rendering makes it difficult to distinguish the concepts.","section":"Figure 1"},{"comment":"Provide the number of submissions, the distribution of identity terms, and the total size of the four concept sets, since the TCAV estimates depend on these quantities.","section":"Project Respect"},{"comment":"Reference [1] is a GitHub README; cite a stable or peer-reviewed description of the Perspective API and its model versions, and clarify the exact relationship between [1] and the bias mitigation applied to toxicity@6.","section":"References"},{"comment":"The phrase 'normative lens' is used informally; a short formal definition of how TCAV scores operationalize normative values would clarify the proposal.","section":"Normative model insights"}],"recommendation":"major_revision","confidential_remarks":"The paper is a short application of TCAV to a normative setting, and the empirical core is not yet convincing. With additional experiments, matched controls, and fuller reporting of the TCAV setup, the central claim could become defensible, but in its current form the load-bearing evidence is a single unvalidated figure. The authors may also consider whether a workshop or short-paper venue is more appropriate than a full archival journal for a preliminary result of this kind."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know this is a short workshop-style paper with two contributions: the Project Respect corpus of positive statements from LGBTIQ+ communities, and the first application of TCAV to language models. The dataset is the more solid piece—a real, hard-to-collect resource that could be reused. The TCAV experiments are the weaker half, and the paper's headline claim that they reveal a 'normative lens' on model internals is not supported by the evidence as presented.\n\nThe conceptual framing is genuinely useful: distinguishing descriptive from normative ML and asking what values we want systems to encode is the right question to raise. The application to two versions of the Perspective API is a reasonable proof-of-concept. That said, the empirical comparison has a load-bearing confound. The four concept sets differ simultaneously in toxicity/sentiment, register, source, and identity-term distribution. A TCAV direction learned by separating Project Respect positive statements from random non-toxic examples could just as easily be a sentiment or politeness direction as a respect direction. Since toxicity@6 was trained with bias mitigation intended to weaken correlations between identity terms and toxicity, a sentiment-correlated concept direction would show lower TCAV scores mechanically, without implying that the model's normative encoding changed. The paper offers no matched negative controls, no concept-set accuracy or randomization checks, no significance tests on the differences in Figure 1, and no details on TCAV parameters. The stress-test note is not hyperbole; it lands directly on the central interpretation.\n\nI would still give the authors credit for asking the right question and for collecting data that the field needs. The paper is honest about the data collection but silent on the confounding problem, which is a real omission. For whom is this paper? For a workshops audience or for someone working on interpretability and fairness, it is a useful prompt and a potential data reference. For a serious archival venue, the empirical section needs to be redone with matched concept sets, permutation tests, and a much more careful discussion of what the TCAV scores can and cannot mean.\n\nRecommendation: send it to peer review, but expect heavy revision. The idea and the dataset deserve referee time; the current experimental evidence does not support the strong conclusion.","headline":"A genuinely useful dataset and a first-of-its-kind TCAV application, but the central empirical claim is confounded and the paper overreads its own evidence.","tokens_in":2829,"tokens_out":1630,"would_cite":true,"duration_ms":18491,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that TCAV, applied to two versions of a toxicity classifier with concept examples built from LGBTIQ+ identity terms and positive self-statements, reveals that the models' internal representations encode different…","keywords":["TCAV","normative values","toxicity classification","social identity","bias mitigation","model interpretability","LGBTIQ+","concept activation vectors"],"falsifier":"Hold architecture, training data, and training procedure fixed, and rerun the TCAV analysis on a model with and without the specific bias-mitigation step. If the two versions produce indistinguishable TCAV scores for the same concept sets, the paper's attribution of the encoding difference to bias mitigation would be falsified.","tokens_in":1924,"feed_emoji":"🤖","tokens_out":5690,"duration_ms":49117,"temperature":0.7,"pith_summary":"The paper sets out to show that machine learning models are not value-neutral: the patterns they learn carry normative implications, and those implications can be inspected. It presents the first application of TCAV to language models, probing two versions of a toxicity classifier with concept sets built from LGBTIQ+ identity terms, toxic and neutral web comments, and positive self-descriptive statements from the Project Respect crowdsourcing program. The results show the two model versions encode different information about the interaction between social identities and normative values. A sympathetic reader would care because this offers a way to see what values a model has internalized before it produces disparate impacts.","feed_headline":"Two toxicity models encode LGBTIQ+ identity differently","feed_subtitle":"A TCAV probe of Perspective API versions 1 and 6 shows bias mitigation changes how identity terms relate to toxicity.","key_machinery":"The central machinery is TCAV (Testing with Concept Activation Vectors), a method that represents a human-interpretable concept as a direction in the activation space of a neural network. The direction is learned from a set of example inputs; a TCAV score then measures how strongly moving an input's internal representation along that direction changes the model's prediction. Here the concept directions are learned from LGBTIQ+ identity terms, toxic comments, neutral comments, and Project Respect's positive statements, and the score measures each concept's association with the toxicity prediction. This machinery is what lets the paper translate model internals into a normative claim.","core_discovery":"On the paper's own terms, the discovery is that TCAV scores expose a normative difference between model versions. In toxicity@1, both toxic and neutral comments containing LGBTIQ+ identity terms receive high TCAV scores for toxicity, suggesting the model strongly associates identity terms with toxicity regardless of comment sentiment. In toxicity@6, toxic and neutral comments receive very different TCAV scores, indicating the bias-mitigated model no longer conflates identity with toxicity in the same way. The Project Respect positive statements also score differently across versions, and the paper reads this as evidence that the internal representations encode different kinds of information about how social identities relate to normative values.","pith_inferences":["If the result generalizes, TCAV scores could be tracked over model versions as a regression test for value alignment, not just bias.","The observed differences are correlational; a stronger test would intervene on training data distribution and observe shifts in concept directions.","The same normative lens could be applied to generative models, asking whether identity terms are internally associated with harm during generation."],"forward_implications":["Bias mitigation can change not only the model's outputs but the internal geometry linking identity concepts to toxicity.","TCAV with normative concept sets can serve as a pre-deployment audit check for whether a model encodes identity in conflation with harm.","The Project Respect positive-statement data supplies a concept set for positive value encoding, not just toxicity, extending the probe beyond negative categories.","The technique is model- and domain-agnostic, so the same concept sets can probe other pre-trained language models and other social categories."],"supporting_citations":[{"why":"Supplies the two toxicity-model versions whose internal representations are compared.","marker":"[1]"},{"why":"Supplies the crowdsourced positive statements that define the normative concept set.","marker":"[2]"},{"why":"Provides evidence of abusive identity-term usage online and the bias-mitigation approach behind version 6.","marker":"[3]"},{"why":"Provides the TCAV method used to derive concept vectors and scores.","marker":"[4]"}],"fun_headline_variants":["TCAV exposes model shift on LGBTIQ+ toxicity association","Bias mitigation decouples identity from toxicity in Perspective API","Identity terms and toxicity: TCAV contrasts two model versions","Model versions encode LGBTIQ+ identity differently per TCAV"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central comparison assumes that the only meaningful difference between toxicity@1 and toxicity@6 is the bias-mitigation technique, so that differing TCAV scores can be attributed to that technique rather than to differences in architecture, training data, or other changes between the versions.","fun_headline_variants_meta":{"raw":{"variants":["TCAV exposes model shift on LGBTIQ+ toxicity association","Bias mitigation decouples identity from toxicity in Perspective API","Identity terms and toxicity: TCAV contrasts two model versions","Model versions encode LGBTIQ+ identity differently per TCAV"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000239,"raw_usage":{"total_tokens":1431,"prompt_tokens":781,"completion_tokens":650,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":397,"completion_tokens_details":{"reasoning_tokens":575}},"tokens_in":397,"tokens_out":650,"duration_ms":6115,"temperature":1.0,"reasoning_tokens":575,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T15:38:55.758119+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Hold architecture, training data, and training procedure fixed, and rerun the TCAV analysis on a model with and without the specific bias-mitigation step. If the two versions produce indistinguishable TCAV scores for the same concept sets, the paper's attribution of the encoding difference to bias mitigation would be falsified.","supporting_citations":[{"cited_title":"https://github.com/ conversationai/perspectiveapi/blob/ master/api_reference.md#models","cited_arxiv_id":null,"evidence_quote":"Supplies the two toxicity-model versions whose internal representations are compared."},{"cited_title":"https://projectrespect","cited_arxiv_id":null,"evidence_quote":"Supplies the crowdsourced positive statements that define the normative concept set."},{"cited_title":"Dixon, J","cited_arxiv_id":null,"evidence_quote":"Provides evidence of abusive identity-term usage online and the bias-mitigation approach behind version 6."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the TCAV method used to derive concept vectors and scores."}],"review_version":1}