{"id":"fb5282a9-f052-4ef3-a4ec-a277bb0742c2","arxiv_id":"2504.13486","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"In a preliminary comparison, ChatBlackGPT provided more culturally specific resources and historical context than ChatGPT for Black travel inquiries, suggesting culturally tailored assistants add value.","lead":"This paper compares a culturally tailored AI assistant, ChatBlackGPT, with ChatGPT on travel questions from Black communities. It finds ChatBlackGPT gives more historical context, concrete resources, and safety guidance, and proposes a research agenda for designing AI that reflects Black lived experience.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Distinguishing-feature claims rely on four hand-selected prompts with no selection criteria, codebook, or saturation check; the conclusion generalizes beyond the evidence.","rationale":"I read the paper in good faith: it is an extended abstract that explicitly frames itself as preliminary, includes a positionality statement, and acknowledges that a larger qualitative sample could provide additional insights. Those are real strengths. However, the central empirical claim in the conclusion is not the exploratory agenda but the comparative finding about ChatBlackGPT's distinguishing features, and the only evidence for those features is the four-prompt thematic analysis. The quantitative metrics are orthogonal to the distinguishing features, so they do not corroborate RQ2. The weakest link is the unstated prompt-selection procedure and the absence of a saturation or reliability check. A random or exhaustive coding pass with a pre-specified codebook would settle whether the features are representative or an artifact of hand-picked examples. This is the same load-bearing assumption the reader identified, and the CONDITIONAL verdict remains appropriate: the direction is plausible and socially meaningful, but the generalization is not yet established.","tokens_in":12435,"tokens_out":4398,"duration_ms":43779,"concrete_test":"Draw a random sample of at least 50 prompts from the 271 CultureBank prompts used in Section 3.2 (or code all 271), and have two or more coders who are blind to the research hypothesis apply an a priori codebook with operational definitions of the three Section 4.2 features, including 'concrete resources' as a named book, location, or business. Compute the prevalence of each feature in ChatBlackGPT versus ChatGPT responses and report inter-rater reliability (e.g., Cohen's kappa). If the features are not present in a majority of ChatBlackGPT responses, or if kappa is below 0.6, the Section 4.2 generalization and the Section 6 conclusion are not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing step is the move from Section 4.2's thematic analysis to the conclusion in Section 6. The three distinguishing features (contextualizing Black history, providing concrete resources, tone) are generated from four hand-selected prompts out of the 271-prompt CultureBank set. The paper states no selection criterion, provides no codebook, reports no inter-rater reliability, and offers no saturation check. Because the choice of prompts is not described, the observed features could be artifacts of which four examples were chosen rather than systematic differences between ChatBlackGPT and ChatGPT. The quantitative analyses in Section 4.1.1 measure only readability and sentiment, so they do not independently measure the features that carry RQ2. The conclusion that 'ChatBlackGPT offers concrete advice and culturally relevant resources when prompted about Black travel related inquiries' therefore rests on an unstated, non-random sample and is broader than the evidence presented.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents a preliminary comparative study of two text-based AI assistants, ChatGPT and ChatBlackGPT, focused on travel-related information-seeking prompts drawn from the CultureBank dataset (271 evaluation questions about Black and African diaspora culture). The authors apply readability metrics (Gunning Fog, Flesch Reading Ease, Flesch-Kincaid), a sentiment classifier, and an inductive thematic analysis of four hand-selected sample prompts. For RQ1 (similarities), they report that both assistants use a chronological before/during/after structure and a positive/neutral tone. For RQ2 (distinguishing features), they claim ChatBlackGPT contextualizes Black history, provides concrete resources (e.g., Black-owned business recommendations), and adopts a more personal, empathetic tone. The paper concludes that ChatBlackGPT offers concrete advice and culturally relevant resources for Black travel inquiries and proposes future work with surveys, interviews, and workshops.","tokens_in":12741,"tokens_out":3412,"duration_ms":31785,"significance":"If the findings hold, this work would empirically document differences between a culturally tailored and a general-purpose AI assistant, providing a concrete example of how 'for us, by us' design can manifest in output quality. The topic is timely for HCI and the authors' positionality statement and use of an existing public dataset (CultureBank) are strengths, as is the provision of a GitHub repository with supplementary outputs that supports reproducibility of the data-collection portion. However, the evidence base is narrow: the qualitative claims rest on a non-random sample of four prompts with no reliability checks, and the quantitative results lack inferential statistics. The paper is therefore best viewed as a proposal for a research agenda rather than a fully supported comparative evaluation.","major_comments":[{"comment":"The qualitative analysis that answers RQ2 and supports the central conclusion in §6 is based on only four hand-selected prompts from the 271-prompt CultureBank set. The paper states no selection criteria, provides no codebook, reports no inter-rater reliability, and offers no saturation check. The authors' own limitation statement (§3.3.3) acknowledges that a larger sample could provide additional insights. The features listed in §4.2.1–4.2.3 (historical context, concrete resources, tone) are then generalized in §6 to 'ChatBlackGPT offers concrete advice and culturally relevant resources.' Because the four prompts were not shown to be representative of the dataset, this conclusion overstates what the evidence can support. The authors should either (a) provide selection criteria and inter-rater agreement and ideally cross-validate the features on a larger sample, or (b) explicitly restrict the concluding claims to the four examined prompts and frame the broader statements as hypotheses.","section":"§3.3.3, §4.2, §6"},{"comment":"The quantitative results in §4.1.1 are reported descriptively: grade-level distributions 'ranged from 8 to 14' with ChatBlackGPT 'mostly at 12–14' and ChatGPT 'clustering at 8–10'; sentiment is described as 'uniformly neutral or positive' with 'no negative sentiments.' No standard deviations, confidence intervals, effect sizes, or significance tests are reported for any of the three readability metrics or the sentiment scores. With 271 paired responses per assistant, simple paired tests (e.g., Wilcoxon signed-rank for readability scores) are feasible and would substantially strengthen the RQ1 similarity claim. As written, the reader cannot distinguish systematic differences from noise, and the 'slightly more readable' claim is not quantified.","section":"§4.1.1, Figure 1"},{"comment":"The comparison is not reproducible because the specific versions of ChatGPT and ChatBlackGPT are not disclosed. ChatGPT has multiple model versions with materially different output styles, and the paper does not state which model (e.g., GPT-3.5, GPT-4) was used, the access date, or any temperature/settings. If ChatBlackGPT is built on an underlying ChatGPT model (a common architecture for such tools, as suggested by reference [23] on 'tailored ChatGPTs'), the reported differences may reflect a system prompt rather than a fundamentally different model family. The authors should describe the exact systems, versions, and data-collection dates in §3.1, and ideally provide the raw prompts/outputs in the repository for independent verification.","section":"§3.1"},{"comment":"The claim that ChatBlackGPT 'consistently suggested supporting Black-owned businesses' and extended this 'universally, including to non-Black users' is based on four prompts, at least one of which explicitly involved identity. Such universal claims require full-dataset evidence or a quantitative count; otherwise they read as overgeneralizations of a small sample. The same issue appears in §5.1, where the 'microaggressions' observation (ChatGPT not mentioning 'Black' or 'woman') is described as 'exemplified' by a single response type. The authors should either quantify the frequency of these patterns across all 271 prompts or carefully qualify them as observations from the illustrative sample.","section":"§4.2.2"}],"minor_comments":[{"comment":"Figure 1 would be more informative as boxplots or violin plots with individual data points; a simple bar chart of central tendency obscures the spread of readability scores across the 271 prompts.","section":"Figure 1"},{"comment":"The paper states that filtering resulted in 130 Reddit and 141 TikTok evaluation questions. It would clarify whether all 271 were used for the quantitative analysis, and if any were excluded (e.g., duplicate or malformed prompts).","section":"§3.2"},{"comment":"The example comparing 'engaging with the locals' to explicit identity mention is illustrative but not tied to a specific prompt; providing a short verbatim snippet or identifying the prompt number would make the qualitative evidence more transparent.","section":"§4.2.3"},{"comment":"The phrase 'passive sentiment observed in ChatGPT's response' is ambiguous; if it refers to the sentiment classifier's neutral/positive labels, say so directly.","section":"§5.1"},{"comment":"Reference [2] is a news article about Black Twitter; the citation context in §2.2 is clear, but the reference should include access date and publication outlet details consistently with other citations.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"This manuscript is an extended abstract (CHI EA) and is positioned as a 'preliminary' study, which mitigates some concerns about scope. However, for a full journal paper, the gap between the qualitative sample (four prompts) and the general conclusion (§6) is too large, and the lack of model-version disclosure is a reproducibility problem. The core idea is valuable and the authors' transparency about positionality is commendable. I would encourage the editor to consider whether a paper with such a thin empirical base fits the journal's standards, even as a 'research agenda' contribution, unless the authors can expand the qualitative analysis or substantially temper the claims."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nQuick take: this is a genuinely first comparative look at ChatBlackGPT against ChatGPT for culturally loaded travel queries, and the authors are appropriately honest that it is exploratory. The quantitative work (readability, sentiment on 271 CultureBank prompts) is a reasonable start, and the qualitative coding, while small, points to plausible distinguishing features: historical context, concrete resources, and tone. The positionality statement is a plus, and the paper is clearly written.\n\nWhere it gets soft is exactly where the stress-test lands. The three distinguishing features in Section 4.2 come from coding only four hand-selected prompts out of 271, with no selection criteria, no codebook, no inter-rater reliability, and no saturation check. The paper acknowledges the small sample as a limitation, but the conclusion in Section 6 (“ChatBlackGPT offers concrete advice and culturally relevant resources…”) is stated as a finding, not a hypothesis. That is a load-bearing gap. The quantitative results also lack any error bars or significance tests, so the readability and sentiment differences are descriptive at best. The supplementary repository is mentioned, but I could not verify what is actually shipped; if the prompts and outputs are not there, reproducibility takes a further hit.\n\nThat said, this is an extended abstract, and on that scale the ambition is right. The topic matters, the comparison is new, and the authors treat the work as a foundation for a survey and interviews. I do not think the paper is circular or self-serving; the CultureBank prompts are external, and the coding is inductive. The central problem is simply that the evidence does not yet support the breadth of the conclusion.\n\nWho is this for? HCI and fairness researchers who want a first data point on culturally tailored assistants, and anyone planning user studies with Black communities around AI. It is not a definitive empirical result, but it is a legitimate research agenda worth a serious referee — though the referee should push for a much larger, randomly sampled qualitative set and a pre-registered analysis if this becomes a full paper.\n\nRecommendation: send it to peer review as an exploratory study, but with clear expectations that the conclusions be trimmed to match the evidence or the analysis be scaled up.","headline":"A timely but preliminary comparison of ChatBlackGPT vs. ChatGPT whose qualitative claims rest on four hand-picked prompts; the direction is promising, the evidence is thin.","tokens_in":13090,"tokens_out":1335,"would_cite":true,"duration_ms":13428,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Culturally informed AI built for Black users gives travel answers that differ from ChatGPT's: it adds Black history, names concrete resources and Black-owned businesses, and speaks to the user's identity, where ChatGPT stays generic.","keywords":["culturally informed AI","ChatBlackGPT","ChatGPT","Black travel","comparative analysis","thematic analysis","AI assistants","cultural relevance"],"falsifier":"Code a random sample of 50 prompt pairs from the full 271, using the paper's own feature definitions for 'historical context' and 'concrete resources,' and count how often each assistant supplies them; if ChatGPT matches or exceeds ChatBlackGPT's frequency on either feature, the paper's central distinction fails.","tokens_in":12275,"feed_emoji":"🧳","tokens_out":6741,"duration_ms":56765,"temperature":0.7,"pith_summary":"The paper is trying to establish that a culturally informed AI assistant built by and for Black communities, ChatBlackGPT, answers Black travel-related questions in a way that a general-purpose assistant does not: it layers in Black historical context, points to named resources and Black-owned businesses, and keeps a calm, identity-aware tone instead of ChatGPT's detached, overtly cheerful one. The authors contend that this difference is not cosmetic; it is the kind of output that can affect whether a Black traveler gets safety-relevant information, such as awareness of sundown towns, and whether the user feels the tool was built with them in mind. The paper positions the finding as preliminary evidence for a larger research agenda on when and how Black communities engage with culturally tailored AI assistants. A reader should care because the result grounds the 'for us, by us' design philosophy in a concrete, observed output difference rather than in principle alone.","feed_headline":"Black-built AI gives more concrete travel answers than ChatGPT","feed_subtitle":"On the same 271 travel prompts, ChatBlackGPT adds history, named resources, and Black-owned businesses.","key_machinery":"The load-bearing comparison is a paired-prompt design. The authors took 271 travel-related evaluation questions from CultureBank, a community-sourced dataset of cultural knowledge derived from Reddit and TikTok, ran the identical prompts through ChatGPT and ChatBlackGPT, and then compared the outputs three ways: a sentiment model for emotional tone, three readability metrics for grade level, and an inductive thematic analysis of four selected prompt pairs. The thematic coding is what carries the argument: it is the mechanism that surfaces the distinguishing features, historical context, concrete resources, and identity-aware tone, that the quantitative metrics alone could not separate, since both assistants scored similarly on sentiment and readability.","core_discovery":"On 271 travel-related evaluation questions drawn from the CultureBank dataset, ChatGPT and ChatBlackGPT produced outputs that looked alike in broad structure—both offered before/during/after advice in a positive or neutral tone, with similar formatting. The paper's central discovery is in what separates them. ChatBlackGPT's responses contextualized the history behind places and experiences, such as the origins and cultural weight of HBCUs; attached concrete, findable resources including named literature; and recommended supporting Black-owned businesses to all users, not only Black ones. ChatGPT, by contrast, kept to broad advice, did not repeat the user's identity even when the prompt named 'Black woman,' and referred to Black communities as 'the locals.' The authors conclude that ChatBlackGPT offers concrete advice and culturally relevant resources for Black travel inquiries, and argue that such tools can serve as a meaningful resource for culturally nuanced information seeking.","pith_inferences":["A testable extension follows: if the three distinguishing features are real, they should be measurable at scale by counting named entities such as businesses, authors, and organizations, and historical references across all 271 response pairs; a simple density metric would let other community-built assistants be benchmarked against general-purpose baselines.","The paper's contrast between ChatGPT dropping 'Black woman' from its reply and ChatBlackGPT preserving it points to an epistemic difference, not just a stylistic one; a user study measuring perceived trust and safety could test whether identity omission is what drives alienation.","The same comparative design could shift domains: health and education are where Black users already report chatbot use, and the concrete-resources feature would plausibly matter even more there, where a named clinic or vetted organization can change an outcome.","The finding that ChatBlackGPT recommends Black-owned businesses to all users, including non-Black ones, suggests a design principle worth testing in other tools: resource recommendations can express a community's values without being conditional on the user's identity."],"forward_implications":["If the finding holds, Black travelers seeking information through an AI assistant can get safety-relevant details that a general-purpose tool omits, such as the historical exclusion of sundown towns.","The observed difference gives designers a concrete target: culturally tailored assistants can be evaluated by whether they name resources, ground advice in history, and honor stated identity, rather than by tone alone.","The paper's proposed research agenda, a survey, interviews, and demo workshops with Black adults, becomes the natural next test of whether these output differences translate into differences in trust, use, and decision-making.","The tendency of ChatGPT to drop or distance the user's stated identity suggests that general-purpose assistants may continue to produce the detached responses that prior work links to user alienation, reinforcing the case for community-built alternatives."],"supporting_citations":[{"why":"Supplies the 271 travel-related evaluation questions used as identical prompts for both assistants.","marker":"[64]"},{"why":"Defines ChatBlackGPT, the culturally informed assistant whose outputs are under study.","marker":"[62]"},{"why":"Defines ChatGPT, the general-purpose baseline assistant used for comparison.","marker":"[55]"},{"why":"Provides the inductive thematic analysis approach used to identify the distinguishing features.","marker":"[20]"},{"why":"Documents ChatBlackGPT's origin and design intent as a tool created by and for Black users.","marker":"[23]"},{"why":"Establishes Black travelers' use of online spaces for culturally tailored information, motivating the travel domain.","marker":"[25]"},{"why":"Supplies the 'for us, by us' design philosophy the paper claims ChatBlackGPT instantiates.","marker":"[26]"}],"fun_headline_variants":["ChatBlackGPT adds history and resources ChatGPT misses","Culturally informed AI gives Black travelers concrete advice","ChatBlackGPT: richer travel answers, not just generic tips","Study: Black-focused AI offers more contextual travel guidance","ChatBlackGPT vs ChatGPT: cultural context wins for travel"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The three distinguishing features come from qualitative coding of only four hand-picked prompts out of the 271 total, with no selection criteria or saturation check, so the generalization from those four to the full dataset is assumed.","fun_headline_variants_meta":{"raw":{"variants":["ChatBlackGPT adds history and resources ChatGPT misses","Culturally informed AI gives Black travelers concrete advice","ChatBlackGPT: richer travel answers, not just generic tips","Study: Black-focused AI offers more contextual travel guidance","ChatBlackGPT vs ChatGPT: cultural context wins for travel"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000638,"raw_usage":{"total_tokens":2896,"prompt_tokens":861,"completion_tokens":2035,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":477,"completion_tokens_details":{"reasoning_tokens":1957}},"tokens_in":477,"tokens_out":2035,"duration_ms":12085,"temperature":1.0,"reasoning_tokens":1957,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T12:06:28.798786+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Code a random sample of 50 prompt pairs from the full 271, using the paper's own feature definitions for 'historical context' and 'concrete resources,' and count how often each assistant supplies them; if ChatGPT matches or exceeds ChatBlackGPT's frequency on either feature, the paper's central distinction fails.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines ChatBlackGPT, the culturally informed assistant whose outputs are under study."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines ChatGPT, the general-purpose baseline assistant used for comparison."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Documents ChatBlackGPT's origin and design intent as a tool created by and for Black users."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes Black travelers' use of online spaces for culturally tailored information, motivating the travel domain."},{"cited_title":"For Us By Us","cited_arxiv_id":null,"evidence_quote":"Supplies the 'for us, by us' design philosophy the paper claims ChatBlackGPT instantiates."}],"review_version":1}