{"id":"0dbad149-1e27-43e4-881a-565f606dabb8","arxiv_id":"2504.16671","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"LLMCode uses IoU and Modified Hausdorff Distance to measure human-LLM coding alignment; two studies with 26 designers show LLMs capture surface coding patterns but not the designer's deeper interpretive lens.","lead":"This paper introduces an open-source tool, LLMCode, that uses two overlap and distance metrics to compare how human designers and large language models code qualitative text. In two small studies with 26 designers, the model matched surface coding patterns but not the designers' deeper interpretive lens, and users and model influenced each other during iterative refinement.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central negative result—'deeper interpretive lens' is limited—is inferred from IoU/MHD similarity scores with no independent measure of depth and no human-human baseline, so the claim is not yet anchored.","rationale":"The reader's CONDITIONAL verdict is well calibrated. My stress-test identifies the same central vulnerability: the quantitative story is built on IoU and MHD, whose construct validity for 'deeper interpretive lens' is assumed, not shown. I would go slightly further than the reader by emphasizing the missing human-human baseline and the uncalibrated meaning of 'limited'; without a comparison level, even a valid similarity metric does not tell us whether the LLM is worse than a reasonable human interpreter. The paper's own Section 6.3 limitation (no held-out test set in Study 2) and Section 5.1.1 caveat (MHD gains may be codebook convergence) are acknowledged, which is honest but does not remove the need for construct validation. The deductive-coding claim in the abstract is also not directly tested. These concerns do not warrant rejection: the tool is open-source, the metrics are clearly defined, and the qualitative thematic analysis gives some independent, if partial, support for the conclusion that LLM codes can feel vague or decontextualized. The right outcome remains conditional acceptance pending the proposed validation.","tokens_in":20192,"tokens_out":7554,"duration_ms":74562,"concrete_test":"Conduct a construct-validation substudy on a random subset of the Study 1/2 texts: for each designer, have a second human coder independently code the same texts and have the original designer rate each LLM-generated annotation on 'how well it captures my intended interpretive lens' (1-7). Compute the same character-level IoU and MHD for human-human and human-LLM pairs, plus the Spearman correlation between the designer ratings and IoU/MHD. If human-LLM scores fall within the human-human distribution, or if IoU/MHD do not correlate with the designer's own depth ratings, the central claim is not supported; if human-LLM scores are significantly lower and the correlations are strong, the claim stands.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract and Section 7 claim that gpt-4o can align with a designer's coding at a surface level but cannot emulate the designer's 'deeper interpretive lens.' The only quantitative support is IoU and MHD (Section 3.1). These metrics measure character-level overlap of highlights and embedding distance between code labels; they are similarity measures, not measures of interpretive depth. The paper itself concedes in Section 5.1.1 that MHD improvement may reflect codebook convergence rather than deeper emulation, and Section 6.3 admits Study 2's alignment improvements are measured on the validation set with no held-out test set, while participants' own annotations could shift during iteration. More fundamentally, no human-human or other baseline is reported, so 'limited' is uncalibrated: a low or medium IoU/MHD could reflect the inherent subjectivity of reflexive coding rather than an LLM-specific deficit. The qualitative findings (vague codes, missing meme context) are suggestive, but the abstract's binary claim—good deductive, limited deeper lens—also goes beyond the evidence: the deductive-coding claim is inferred from MHD behavior when codes already appear in examples, not from a direct deductive-coding experiment.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces LLMCode, an open-source tool for LLM-assisted qualitative coding that integrates two alignment metrics, character-level Intersection over Union (IoU) and embedding-based Modified Hausdorff Distance (MHD), to compare human and LLM annotations. The central research questions concern how to measure researcher-AI alignment, whether an LLM can emulate an individual designer's interpretive perspective through in-context learning, and how LLM assistance shapes design research insights. Two studies are reported: Study 1 with 19 design students who manually coded forum data and then used LLMCode in a notebook environment, and Study 2 with 7 designers who used a graphical interface to iteratively refine few-shot examples and prompts. The paper reports that IoU and MHD improve with more few-shot examples, that the model's performance degrades when texts are dissimilar to the examples, and that iterative human refinement yields improvements over a random baseline, especially for IoU. Qualitative analysis suggests participants both adapted to and were influenced by the model's outputs. The authors conclude that while LLMs can capture surface patterns and perform well in deductive coding, they are limited in emulating the designer's deeper interpretive lens, and they discuss overfitting and the need for held-out test sets.","tokens_in":20435,"tokens_out":3228,"duration_ms":30676,"significance":"The paper addresses a timely and important problem: how to evaluate and support appropriate reliance on LLM-assisted qualitative analysis in research-for-design contexts, where interpretive depth is central. The main contributions are the design and open-source release of a tool with two concrete, computable alignment metrics, and two empirical studies that provide qualitative insights into how designers interact with such a tool, including reciprocal influence between human and AI coding. The qualitative findings, particularly participants' mental models of few-shot example selection and their willingness to adopt AI-suggested codes, are valuable for the HCI and CSCW communities. If the central claims were fully supported, the work would provide a practical framework for calibrating trust in LLM-driven qualitative coding. However, the quantitative evidence is thin and, as the paper itself acknowledges, Study 2 has no held-out test set; the 'limited deeper interpretive lens' conclusion is inferred from surface and semantic similarity metrics without a human-human baseline or an independent measure of interpretive depth.","major_comments":[{"comment":"The claim that human oversight improves alignment is not fully established because the reported improvements are computed on the validation set used to iteratively select examples and refine prompts. The paper explicitly acknowledges in Section 6.3 that no separate test set was employed and that participants' own annotations may have shifted during iteration, yet Section 5.2 presents the improvement over the random baseline as a finding. This is a validation-loop circularity: the examples were chosen using the same IoU and MHD scores that are then reported as the outcome, so part of the observed gain may reflect optimization toward the evaluation metric rather than genuine generalization. To support the claim, the authors should either re-analyze the data on a held-out test set (even a small one, coded after the iteration phase) or explicitly reframe Section 5.2 as demonstrating validation-set alignment, with the generalization question left open.","section":"Section 5.2 and Section 6.3"},{"comment":"The central negative claim—that the LLM cannot emulate a designer's deeper interpretive lens—is derived from IoU and MHD, which measure character-level highlight overlap and embedding distance between code labels. These are surface and semantic-similarity measures, not direct measures of interpretive depth. The paper itself concedes in Section 5.1.1 that improvement in MHD may reflect codebook convergence rather than deeper emulation of the researcher's perspective. Moreover, no human-human baseline is reported, so the observed divergence is uncalibrated: in a reflexive, constructivist coding task, disagreement between two human coders could plausibly be of a similar magnitude. The authors should either add a human-human comparison (e.g., have two human coders code the same texts and compute the same metrics) or restrict the conclusion to 'the model's outputs are not highly similar to this specific designer's annotations,' removing the unsupported inference about 'deeper interpretive lens.'","section":"Section 3.1, Section 5.1, Section 6.1"},{"comment":"The abstract and conclusion state that 'the model performs well with deductive coding,' but no direct deductive coding experiment is reported. No study condition uses a fixed, predefined codebook applied to new texts; the claim is instead inferred from the MHD behavior when codes already appear in the few-shot examples (Section 5.1.1). This is an overgeneralization from an indirect observation. The authors should either add a direct evaluation of deductive coding (e.g., supplying a fixed codebook and measuring assignment accuracy) or soften the claim to something like 'the model can apply seen codes to related texts,' which is what the current data support.","section":"Abstract, Section 7, Section 5.1.1"},{"comment":"The correlation evidence in Figure 5 is reported as a Pearson coefficient of 0.53 with no confidence interval, no significance test, and no accounting for the non-independence of multiple texts from the same participant. With N=8 participants, this is thin quantitative support for the statement that 'the model performs better on texts that are more similar to its example set.' Additionally, the number of K-Means clusters (five) is presented without justification, and the analysis depends on this free parameter. A mixed-effects model with participant as a random effect, or at minimum bootstrap confidence intervals for the correlation, would materially strengthen the claim.","section":"Section 5.1.2, Figure 5"}],"minor_comments":[{"comment":"There is a duplicated word in the sentence 'allowing themes to to be determined organically'; it should read 'allowing themes to be determined organically.'","section":"Section 5.3"},{"comment":"The caption uses 'Hausdorff distance' while the paper's metric is the Modified Hausdorff Distance (MHD). Please use the full term consistently, and clarify in the caption that MHD is an average of embedding-based cosine distances, not the original point-cloud distance, since the interpretation of the values depends on this.","section":"Table 1 caption"},{"comment":"The description of example-set selection differs between the interface description (manual selection) and the Study 1 analysis (automatic chronological selection for the learning-curve experiment). Consider adding a sentence in Section 5.1.1 clarifying that the automatic selection was for the controlled comparison and that manual selection is the intended use, to avoid confusion about how the results relate to the tool's workflow.","section":"Section 3.2, Section 4.1"}],"recommendation":"major_revision","confidential_remarks":"The manuscript has a useful contribution and is honest about its limitations, but the abstract and conclusion currently overstate the strength of the quantitative evidence. The major revision should require either (a) a held-out test set for Study 2 and a human-human baseline for the central comparison, or (b) a substantial softening of the claims to match what the current data can support. The qualitative findings and the open-source tool are likely to be of interest to the CHI/CSCW community, and the paper is close to publishable if the framing is aligned with the evidence."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague—this one is worth your time if you work on LLM-assisted qualitative analysis, but the paper's own headline claim is ahead of its evidence.\n\nWhat's actually new: LLMCode, an open-source toolkit that wraps IoU and Modified Hausdorff Distance around an in-context learning workflow for qualitative coding. The metric ingredients aren't novel—IoU is Jaccard, MHD is a standard vision distance, and embedding-based codebook comparison appears in Dai et al. and Zhao et al.—but the interactive workflow and the two studies are new. The empirical observation that in-context learning improves IoU while MHD gains coincide with codebook saturation is a useful, non-obvious result. Study 2's qualitative finding that participants both refine the model and adopt its codes is a real contribution.\n\nWhere it gets soft. The central claim—good deductive, limited deeper lens—is not directly tested. MHD measures embedding distance between labels, not interpretive depth. The paper itself notes that MHD improvement may reflect codebook convergence rather than deeper emulation (Section 5.1.1), and the 'deductive' claim is inferred from MHD behavior when codes already appear in examples, not from a controlled deductive-coding experiment. There is no human-human baseline, so 'limited' is uncalibrated: a low IoU/MHD might reflect inherent subjectivity of reflexive coding rather than an LLM-specific deficit. Study 2 has no held-out test set (admitted in Section 6.3), so part of the reported improvement is optimization toward the validation metric. N is 7–11, and there are no significance tests or confidence intervals. These are all real, but they are also mostly acknowledged in the text—the paper is honest about its limits, which counts for something.\n\nBottom line: it's a solid, well-scoped HCI submission with an open-source artifact and an honest limitations section. The stress-test's complaint that the abstract's binary claim overreaches is fair. A serious referee should ask for a human-human baseline or an independent depth proxy, a held-out test set in Study 2, and more cautious wording. I'd send it to peer review; it deserves referee time and will likely come back substantially revised.","headline":"An honest, well-scoped HCI paper with a useful open-source tool and a real empirical observation, but the abstract's headline claim overreaches the evidence.","tokens_in":20951,"tokens_out":1979,"would_cite":true,"duration_ms":18865,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"An open-source tool with two alignment metrics shows that LLM-assisted qualitative coding replicates a designer's surface patterns from few examples but not the deeper interpretive lens behind them.","keywords":["qualitative coding","large language models","researcher-AI alignment","in-context learning","Intersection over Union","Modified Hausdorff Distance","research for design","human-AI collaboration"],"falsifier":"Give designers a text that contains a genuinely new theme, ask them to code it, have the LLM code it after few-shot training on a different theme, and check whether the model's codes, when shown to an independent panel, are judged to reflect the designer's reasoning rather than merely reuse surface vocabulary. If the model's codes are repeatedly judged as sharing the designer's interpretive stance even when IoU and MHD are low, the paper's central limitation claim would be refuted. Alternatively, measure whether designers' own codes change systematically toward the model's embeddings when no explicit alignment instruction is given; if they do not, the reciprocal-influence claim weakens.","tokens_in":19998,"feed_emoji":"🤖","tokens_out":4943,"duration_ms":43410,"temperature":0.7,"pith_summary":"The paper introduces LLMCode, an open-source tool that scores how closely an LLM's qualitative coding matches a designer's own coding, using two metrics: character-level Intersection over Union for what text gets highlighted, and Modified Hausdorff Distance between code-label embeddings for how the highlighted text is labelled. Across two studies with 26 designers, the paper argues that a state-of-the-art LLM can learn to reproduce the surface patterns of a designer's coding from a few examples, but its ability to extrapolate the designer's broader interpretive lens to new, unfamiliar material is limited. The result matters because research for design relies on each designer's situated, reflexive perspective, and tools that only imitate surface patterns could flatten the diversity of insights while appearing trustworthy. The paper also documents a reciprocal dynamic: designers refine the model's outputs, but also revise their own codes in response to the model's suggestions. The authors position the two metrics as a practical way to make that alignment visible and to guide iterative example selection.","feed_headline":"LLMs mimic coders' style but miss their interpretive lens","feed_subtitle":"New IoU and embedding-distance metrics show AI coding aligns on the surface, not in deeper meaning.","key_machinery":"The carrying mechanism is the pair of alignment metrics inside LLMCode. IoU (Intersection over Union, character-level) measures overlap between the text segments a human highlights and the segments the model highlights, capturing agreement about what content is salient. Modified Hausdorff Distance (MHD) maps the human's and the model's code labels into a shared embedding space and measures the average distance between the two point clouds, capturing whether the labels are semantically related even when wording differs. Together the two metrics make the 'lens' partially visible: IoU tracks focus, MHD tracks conceptual association, and their joint trend across example-set size is the evidence base for the paper's claim about surface versus deep alignment.","core_discovery":"The central claim is that LLM-assisted qualitative coding can be evaluated and steered by comparing highlights and code labels against a single designer's annotations, and that under this comparison the model's apparent learning is mostly surface-level. In Study 1, increasing the number of few-shot examples raised IoU and lowered MHD, but the improvement tracked how representative the examples were of the test texts; when test texts' codes were dissimilar from the example set, the model's output diverged, with a Pearson correlation of 0.53 between that dissimilarity and the model's error. The paper reads this as evidence that the model reuses codes it has seen rather than inductively adopting the researcher's evolving viewpoint. In Study 2, designers who iteratively refined their examples improved IoU beyond a random baseline but the MHD trend was less conclusive, and think-aloud data showed participants adopting some of the model's codes as their own. The paper's conclusion is that in-context learning supports deductive-style application of established codes, but genuine alignment with a designer's emergent interpretation remains unmet and requires human oversight.","pith_inferences":["The two metrics could be borrowed by other qualitative tools as a cheap proxy for 'who is driving the interpretation,' but character overlap and embedding distance will miss structural reinterpretations where a designer re-segments text or renames codes in ways embeddings do not capture.","A testable extension is to run the same few-shot scaling analysis with other LLMs and embedding models to see whether the apparent plateau in inductive alignment is model-specific or a general property of in-context learning.","The reciprocal-influence finding implies that reported alignment gains may be inflated because the human half of the pair is drifting toward the model; using a pre-registered, separately-coded test set would quantify that drift directly.","A concrete design principle follows from the paper's observations: future tools should let users show corrections ('you misunderstood this text') rather than only positive examples, aligning the interaction with the chat mental models users already have."],"forward_implications":["Researchers can use IoU and MHD to locate texts where the model's interpretation diverges, sort by disagreement, and decide where to intervene.","Scaling manual coding to larger corpora is most reliable when the codebook is already established and new texts stay thematically close to the annotated examples.","When novel themes appear in new data, the model should not be expected to keep pace without human recoding or new examples.","Designers will change their own codes in response to the model's output, so the baseline for alignment shifts during iteration and alignment scores must be read with that in mind.","Separating validation and test sets during example iteration is necessary for trustworthy alignment claims, and the paper spells out this recommendation."],"supporting_citations":[{"why":"Defines the in-context learning paradigm that the few-shot coding workflow relies on.","marker":"[6]"},{"why":"Demonstrates that LLMs can perform thematic coding via in-context learning with human-annotated examples, the baseline the paper extends.","marker":"[8]"},{"why":"Shows LLM coding with in-context learning in HCI contexts, which LLMCode builds on.","marker":"[17]"},{"why":"Provides the Modified Hausdorff Distance used to compare code-label point clouds.","marker":"[11]"},{"why":"Supplies the Jaccard Index on which the paper's character-level IoU is based.","marker":"[37]"},{"why":"Provides the embedding model whose vector space grounds MHD comparisons between human and LLM code labels.","marker":"[34]"},{"why":"Specifies the frontier LLM whose few-shot behavior the studies measure.","marker":"[36]"},{"why":"Frames warranted trust as requiring evaluation of AI outputs, motivating the alignment metrics.","marker":"[20]"},{"why":"Establishes appropriate reliance as the design goal that the alignment metrics serve.","marker":"[24]"}],"fun_headline_variants":["Surface-fit only: LLM coding lacks deep interpretation","LLMs handle deductive coding, struggle with interpretive nuance","IoU and MHD show AI coding is shallow, not deep","AI alignment in coding: style over substance","LLMCode metrics: surface mimicry, limited depth"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that character-level highlight overlap (IoU) and embedding distance between code labels (MHD) adequately capture what it means for an AI to share a designer's interpretive lens, so that the conclusion about 'surface-level' emulation rests on these two measures rather than on an independent test of interpretive depth.","fun_headline_variants_meta":{"raw":{"variants":["Surface-fit only: LLM coding lacks deep interpretation","LLMs handle deductive coding, struggle with interpretive nuance","IoU and MHD show AI coding is shallow, not deep","AI alignment in coding: style over substance","LLMCode metrics: surface mimicry, limited depth"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000214,"raw_usage":{"total_tokens":1421,"prompt_tokens":940,"completion_tokens":481,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":556,"completion_tokens_details":{"reasoning_tokens":402}},"tokens_in":556,"tokens_out":481,"duration_ms":5405,"temperature":1.0,"reasoning_tokens":402,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T10:57:56.768331+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Give designers a text that contains a genuinely new theme, ask them to code it, have the LLM code it after few-shot training on a different theme, and check whether the model's codes, when shown to an independent panel, are judged to reflect the designer's reasoning rather than merely reuse surface vocabulary. If the model's codes are repeatedly judged as sharing the designer's interpretive stance even when IoU and MHD are low, the paper's central limitation claim would be refuted. Alternatively, measure whether designers' own codes change systematically toward the model's embeddings when no explicit alignment instruction is given; if they do not, the reciprocal-influence claim weakens.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the Modified Hausdorff Distance used to compare code-label point clouds."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the Jaccard Index on which the paper's character-level IoU is based."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the embedding model whose vector space grounds MHD comparisons between human and LLM code labels."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Specifies the frontier LLM whose few-shot behavior the studies measure."}],"review_version":1}