{"id":"3cabd128-3772-4bfa-8898-474a4fbc1c9c","arxiv_id":"2507.13881","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Zero-shot LLMs identify several construct-relevant features in open-response situational judgment test answers with moderate agreement with human raters, below human-level reliability for most features.","lead":"This paper tests whether large language models can automatically detect the same behavioral features, such as empathy and problem solving, that human raters look for in written answers to a situational judgment test. The best models agree with human raters on some features but fall short on most, so automated scoring of these tests is feasible in part but not yet ready.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The prompt-engineering improvement is measured on the same 162 responses used to design the level descriptions, so the reported delta-kappa is an overfit estimate and cannot support the conclusion that adding level details improves LLM extraction.","rationale":"The paper is a transparent pilot: the zero-shot multi-model comparison, human annotation of 162 responses, and the worked example are useful, and the authors do not overclaim predictive validity. The central feasibility claim, that zero-shot LLMs can assign these seven feature labels with low-to-moderate agreement with trained raters, is supported by Table 3 and does not depend on the enriched-prompt experiment. What the paper additionally claims is that adding level descriptions produces reliable gains (delta-kappa of 0.08-0.206). That claim requires the enriched prompts to generalize, but the design cannot establish this because the same responses were used both to motivate the prompt edits (Table 4) and to evaluate them (Table 3, last column). This is a classic selection-on-evaluation problem, not a question of author intent. The low human kappa values make the problem worse: tuning to noisy labels on the same sample can chase rater-specific patterns. The proposed holdout or fresh-batch check would settle it. I therefore keep the reader's CONDITIONAL verdict: feasibility is plausible and worth a preregistered replication, but the prompt-engineering result is not yet established. This is closely related to the reader's concern about noisy human labels and unstable ground truth, but the most load-bearing issue is the evaluation leakage, so agreement is partial.","tokens_in":853,"tokens_out":1616,"duration_ms":100061,"concrete_test":"Hold out one-third of the 162 responses (n=54) before writing any enriched prompts. Use only the remaining 108 responses to inspect o4-mini's label distributions and to draft feature-level inclusion and exclusion criteria; freeze the prompts, then apply them to the held-out 54 responses and compute per-feature kappa against the same two human raters. Repeat in a three-fold cycle, redrafting prompts on the development folds and evaluating on the held-out fold, and report the average held-out delta-kappa with a bootstrap confidence interval. If the held-out delta-kappa is near zero or no longer consistently positive across features, the paper's prompt-engineering conclusion should be downgraded to exploratory. A simpler variant: annotate a fresh batch of Casper responses and apply the finalized enriched prompts without further edits.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 5.2 presents the paper's strongest evidence for the headline conclusion that LLMs can be instructed to extract construct-relevant features more accurately. But the enriched prompts were not independent of the evaluation data. After inspecting Table 4, which shows o4-mini and human classification distributions on the same 162 responses, the authors wrote new level descriptions and then computed kappa on those same responses. Any prompt edit that moves o4-mini's labels toward the observed human labels, including rater-specific quirks and sampling noise, will inflate the reported delta-kappa values of 0.08-0.206. The zero-shot comparison in Section 5.1 is not contaminated by this step, so the basic feasibility claim survives; however, the quantitative claim that level descriptions improve agreement, and the resulting guidance to build a future scoring system on this approach, is not identifiable. With human-human kappa as low as 0.356 (VAGUE) and 0.510 (CREAT), tuning to human labels on the same sample is especially dangerous: the prompts can fit one rater's idiosyncrasies rather than a stable construct. The paper acknowledges small sample size in Section 6 but does not acknowledge this evaluation leakage.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper investigates whether zero-shot LLMs can identify seven construct-relevant features in 162 open-response Casper SJT responses. Study 1 compares five LLMs (GPT-4o-mini, DeepSeek-R1, Llama 4 Maverick, o4-mini, Claude Sonnet 4) against two human raters, reporting Cohen's kappa for each feature and finding that Claude Sonnet 4 and o4-mini generally perform best, though mostly below human-rater agreement. Study 2 uses o4-mini with prompts augmented by level descriptions, reporting improved kappa values (Δκ ≈ 0.08–0.206). The paper concludes that LLM-based feature extraction is promising and that prompt details can improve agreement, laying groundwork for automated SJT scoring.","tokens_in":9336,"tokens_out":3078,"duration_ms":37649,"significance":"If the zero-shot results are taken on their own, the paper provides a useful, honestly reported feasibility demonstration: multiple LLMs were evaluated with a reproducible prompt, and the authors disclose low human-human agreement for several features. The study also ships a minimal reproducible example and makes no exaggerated claim of human-level performance for most features. However, the headline improvement claim in Study 2 is compromised by evaluation leakage, because the level-description prompts were designed after inspecting model and human classifications on the same 162 responses. As a result, the quantitative evidence for the paper's central methodological recommendation is not identifiable from the reported data. The work is still publishable after a revision that validates the prompt-engineering step on a holdout set and substantially adds uncertainty-aware reporting.","major_comments":[{"comment":"The improvement reported in the last column of Table 3 is not a valid estimate of the benefit of level descriptions, because the enriched prompts were developed after inspecting o4-mini's and the human raters' distributions on the same 162 responses (as stated in the Table 4 discussion: \"We used these results to motivate our prompt engineering strategy\"). Any prompt edit that shifts o4-mini's labels toward the observed human labels, including rater idiosyncrasies and sampling noise, will inflate Δκ. The paper acknowledges the small sample but does not acknowledge this evaluation leakage. To support the claim that adding level details improves agreement, the authors must validate on a held-out set of responses (or use cross-validation) and report the corresponding kappa.","section":"§5.2, Table 3, Table 4"},{"comment":"For VAGUE, human-human Cohen's κ is 0.356, and for CREAT it is 0.510. With such noisy reference labels, the absolute LLM-human kappa values are hard to interpret; a low LLM-human kappa could reflect rater unreliability rather than LLM deficiency, and a high kappa could be partially an artifact of fitting one rater's pattern. The paper should report LLM-human κ separately for each human rater (or the range across raters) and discuss the ceiling imposed by rater disagreement. As it stands, statements such as Claude Sonnet 4 \"generally outperforms\" on features with κ differences of 0.03–0.07 (e.g., INT 0.404 vs o4-mini 0.343) are not supported without uncertainty quantification.","section":"Table 1, §5.1"},{"comment":"No confidence intervals, standard errors, or significance tests are reported for any kappa values. With n=162 and many comparisons across models and features, small numerical differences (e.g., GPT-4o-mini 0.658 vs DeepSeek-R1 0.603 for LACKINF, or Claude Sonnet 4 0.277 vs o4-mini 0.054 for CREAT) cannot be interpreted as meaningful without error bars. The phrase \"super-human agreement\" for GPT-4o-mini on LACKINF (0.658 vs human-human 0.640) is particularly misleading because human-human agreement is not a hard ceiling and the difference is within likely sampling error. At minimum, the authors should provide bootstrap confidence intervals for all reported kappa values.","section":"§5.1, Table 3"}],"minor_comments":[{"comment":"The feature name \"VAGUE\" is typeset with a space as \"V AGUE\" and the model \"Llama 4 Maverick\" appears as \"Lllama 4 Maverick\"; these typos should be corrected.","section":"Table 1 and Table 3"},{"comment":"The system prompt labels the scenario as an \"ethical dilemma,\" but not all Casper scenarios are ethical dilemmas. Using neutral wording such as \"situation\" would avoid biasing the model and better match the assessment design.","section":"§4, prompt example"},{"comment":"The paper omits two of the nine features from Iqbal et al. because they are scenario-specific, but does not describe what those features are. A sentence in §2 or §3 listing the omitted features would help readers judge the generalizability of the approach.","section":"§2 and §3"},{"comment":"The paper collects \"reasoning\" outputs but never analyzes them. Since the authors explicitly cite Casabianca et al. on the validity value of reasoning traces, a brief explanation of why the reasoning outputs were not inspected (or a plan to inspect them in future work) would strengthen the validity discussion.","section":"§5.2"},{"comment":"The abstract calls the approach \"novel,\" but prior work on LLM-based feature extraction in essay scoring is cited in §1. The claimed novelty should be scoped to open-response SJTs specifically, or the wording should be softened.","section":"Abstract and §6"}],"recommendation":"major_revision","confidential_remarks":"The authors are employees of Acuity Insights, the company that administers the Casper SJT. The manuscript does not include a competing-interests statement, which is a significant oversight for a journal publication and should be flagged to the editor. Additionally, the evaluation leakage in Study 2 is the key technical issue; if the authors can produce a holdout validation, the paper could be acceptable, but without it the improvement claim is not supported."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Cole, What you need to know: the zero-shot feasibility claim in this paper is plausible and honestly reported, but the headline result about prompt engineering improving agreement is not valid as estimated, because the improved prompts were designed after looking at the same 162 responses they were then evaluated on. The second study (Section 5.2) should be treated as a hypothesis-generating exercise, not as evidence of a real gain. What's actually new: this is the first application of zero-shot LLM classification to extracting construct-relevant features from open-response SJT responses, using a feature taxonomy from the authors' prior mixed-methods work. They test five LLMs, report per-feature kappas, and include a minimal reproducible prompt. They also honestly report human-human kappas, some of which are low (VAGUE 0.356, CREAT 0.510), and they don't overclaim: they say most features fall well short of human agreement in the zero-shot condition. That is a useful pilot for anyone working on automated scoring of situational judgment tests. Soft spots. The main one is the leakage in Study 2. After inspecting Table 4, which shows o4-mini and human label distributions on the same responses, the authors wrote new level descriptions and computed kappa on that same sample. Any prompt edit that moves the model's labels toward the observed human labels, including rater quirks and sampling noise, will inflate the delta-kappa. The reported improvements of 0.08 to 0.206 are not identifiable. The paper notes the small sample but never mentions this holdout problem. Second, the human labels themselves are noisy for several features; comparing LLMs to a shaky benchmark caps what we can conclude. Third, there are no confidence intervals and no simple baseline (keyword search or a bag-of-words classifier), so we do not know whether the moderate kappas on features like LACKINF require an LLM at all. The feature framework comes from the authors' own prior work, which is fine given the zero-shot setup, but the construct validity of the whole chain still rests on that prior study. These concerns are real but mostly constrain the generality claims, not the basic feasibility demonstration. Take-home: this is a solid pilot that deserves a peer-review round. The prompt-engineering conclusion should be reframed as exploratory or re-run with a holdout sample. For a reader working on automated assessment of open responses, the zero-shot agreement table is a useful reference point. I would send it to review with major revision.","headline":"Zero-shot LLM feature extraction for SJTs is a plausible, honestly reported pilot, but the Study 2 prompt-engineering gain is inflated by evaluating on the same responses used to tune the prompts.","tokens_in":9851,"tokens_out":3403,"would_cite":false,"duration_ms":32615,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Zero-shot LLMs can extract construct-relevant features from open-ended situational judgment test responses, with agreement that improves when prompts include level descriptions.","keywords":["situational judgment test","large language models","feature extraction","construct-relevant features","zero-shot classification","Cohen's kappa","Casper SJT","automated scoring"],"falsifier":"A direct test would be to extract the same features from a larger set of responses scored by a larger panel of raters and check whether the reported $\\kappa$ improvements hold against the more reliable reference; a simpler check is whether the extracted features predict Casper's holistic human scores, a correlation the paper does not report.","tokens_in":8912,"feed_emoji":"🤖","tokens_out":6288,"duration_ms":55911,"temperature":0.7,"pith_summary":"This paper asks whether large language models can identify the features that trained human raters use when scoring open-response situational judgment tests (SJTs). Using 162 responses from the Casper SJT and seven construct-relevant features, the authors show that zero-shot LLMs reach moderate agreement with human raters on some features and near-human agreement on one feature, LACKINF. Adding level descriptions to the prompt improves agreement for every feature tested, most sharply for DISRES. The authors position this as a foundation for automated scoring of personal and professional skills, not yet as a replacement for human raters.","feed_headline":"LLMs spot skills features in open-ended test answers","feed_subtitle":"Zero-shot prompts reach moderate human agreement on Casper SJT responses, and level details push agreement higher.","key_machinery":"The load-bearing object is the set of seven construct-relevant features (INT, LACKINF, JUST, VAGUE, PERSP, DISRES, CREAT) taken from a prior mixed-methods study of what drives Casper raters' scores. Each feature is defined by levels on an ordinal or binary scale, and the task is a per-feature classification of each response. The method is zero-shot prompting: one LLM at a time receives the scenario context, the questions, and a feature description, and returns a JSON decision plus reasoning. Agreement is measured with Cohen's $\\kappa$ with quadratic weighting against each of two human raters, averaged. The level descriptions added in the second study are the mechanism that reduces threshold misalignment, which the authors diagnose by comparing the proportions of levels selected by humans versus the LLM.","core_discovery":"The paper's central claim is that construct-relevant features of SJT responses can be extracted automatically by prompting LLMs to classify one feature at a time, and that richer level descriptions in the prompt move LLM classifications closer to human raters. With a shared zero-shot prompt, the best model achieved the highest average Cohen's $\\kappa$ on four of seven features and near-human agreement on LACKINF; no model approached human-level agreement on most features, with gaps between 0.209 and 0.352 in $\\kappa$. Prompt engineering with inclusion and exclusion criteria for each feature level improved agreement on all six features tested, with the largest gain on DISRES ($\\Delta\\kappa = 0.206$). The authors conclude this is a feasible route toward automated scoring, while noting that the dataset is small and human-human agreement is itself imperfect.","pith_inferences":["If feature extraction becomes reliable, the same pipeline could extend beyond SJTs to personal essays and reference letters, where generative AI has made authenticity harder to judge.","The 'reasoning' field the models return is currently unused; using it for validity evidence or respondent feedback is a natural next step the paper names but does not test.","The level-description gain suggests that calibrating thresholds from human annotations could close more of the remaining gap than prompt text alone.","A testable extension is to measure whether these features predict holistic Casper scores, connecting feature extraction to construct validity."],"forward_implications":["Zero-shot LLM feature extraction is feasible for open-response SJTs and can support the development of automated scoring.","Classifying one feature per response, rather than all features at once, is a workable prompting strategy for nuanced constructs.","Adding level descriptions and inclusion or exclusion criteria to prompts materially improves LLM-human agreement.","A production system may use different LLMs for different features, or an ensemble voting scheme, since each model excelled on at least one feature.","Features like LACKINF that track specific wording may be extractable with simpler keyword or semantic methods."],"supporting_citations":[{"why":"Supplies the seven construct-relevant features, their level definitions, and the original human-classified dataset this study expands.","marker":"Iqbal et al., 2025"},{"why":"Supplies the one-aspect-at-a-time and level-description prompting strategies used for classification.","marker":"Lee et al., 2024"},{"why":"Defines Casper, the open-response SJT whose responses are used as the test bed.","marker":"Dore et al., 2017"},{"why":"Establishes that open-response SJTs relate to personal and professional skills, motivating the feature approach.","marker":"McDaniel et al., 2007"},{"why":"Motivates including a reasoning field in the prompt as future validity evidence for LLM-based scoring.","marker":"Casabianca et al., 2025"},{"why":"Shows LLMs scoring hard-to-automate constructs, the precedent for applying LLMs to SJT features.","marker":"Organisciak et al., 2023"},{"why":"Provides the ensemble idea the authors suggest for combining multiple LLMs.","marker":"Dietterich, 2000"}],"fun_headline_variants":["LLMs extract skill features from open-ended test answers","Zero-shot LLMs match humans on skill feature detection","Level details push LLM skill ratings closer to humans","Automated scoring for open-ended skills tests via LLMs","LLM prompts with level details improve SJT scoring agreement"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole comparison assumes the seven features are the right ones for the construct and that the two human raters' classifications are a stable, valid reference; with human-human kappas as low as 0.356 for VAGUE, the benchmark itself is noisy.","fun_headline_variants_meta":{"raw":{"variants":["LLMs extract skill features from open-ended test answers","Zero-shot LLMs match humans on skill feature detection","Level details push LLM skill ratings closer to humans","Automated scoring for open-ended skills tests via LLMs","LLM prompts with level details improve SJT scoring agreement"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000497,"raw_usage":{"total_tokens":2397,"prompt_tokens":869,"completion_tokens":1528,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":485,"completion_tokens_details":{"reasoning_tokens":1450}},"tokens_in":485,"tokens_out":1528,"duration_ms":13515,"temperature":1.0,"reasoning_tokens":1450,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T16:14:23.860726+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A direct test would be to extract the same features from a larger set of responses scored by a larger panel of raters and check whether the reported $\\kappa$ improvements hold against the more reliable reference; a simpler check is whether the extracted features predict Casper's holistic human scores, a correlation the paper does not report.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the seven construct-relevant features, their level definitions, and the original human-classified dataset this study expands."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the one-aspect-at-a-time and level-description prompting strategies used for classification."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines Casper, the open-response SJT whose responses are used as the test bed."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes that open-response SJTs relate to personal and professional skills, motivating the feature approach."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Shows LLMs scoring hard-to-automate constructs, the precedent for applying LLMs to SJT features."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the ensemble idea the authors suggest for combining multiple LLMs."}],"review_version":1}