{"id":"70a030fd-b5fc-465b-8ee4-a4bbd16fbc72","arxiv_id":"2508.13240","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Using LLM-annotated operator notes from 17 hackers, the authors report a borderline negative correlation between general risk propensity and persistence technique use, but the overall model does not reach significance.","lead":"This paper uses GPT-4o to turn handwritten attacker notes from a 17-person hacking exercise into action sequences, then looks for a link between persistence tactics and psychological test scores. The link is weak and mostly not statistically significant, so the study is a proof of concept rather than a demonstrated result.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central inference from persistence counts to loss aversion is not established: the only significant predictor is a generic risk scale, while the loss-framing scales fail, and the proxy itself is untested.","rationale":"The reader's weakest_assumption correctly identifies the unvalidated proxy linking persistence behavior to loss aversion. My stress-test agrees and sharpens the concern: the paper's own data provide a direct test of that proxy, because ADMC RC2 specifically measures loss-framing susceptibility and fails to predict persistence, while GRiPS, the only significant predictor, is a generic risk scale. This makes the conceptual gap more fundamental than the LLM-annotation issue, since even perfect annotation would not turn a correlation with general risk propensity into evidence about loss aversion. I also considered the statistical fragility of the GRiPS coefficient in a non-significant overall model (F(4,14)=2.11, p=0.133, n=17); that is real but secondary, and the authors acknowledge the need for replication. The reader's CONDITIONAL verdict is therefore appropriate and should remain unchanged. The proposed within-subject trigger analysis would settle whether persistence behavior actually responds to threats of loss, which is the core of the loss-aversion claim.","tokens_in":8616,"tokens_out":4115,"duration_ms":43245,"concrete_test":"Compare persistence behavior before versus after access-threatening operational triggers (e.g., simulated maintenance alerts) within each participant. For each participant, split annotated action sequences into pre-trigger and post-trigger windows matched for duration, and test whether persistence technique counts or uniqueness increase after the threat of losing access (paired test). If no within-subject increase is found, the assertion that persistence counts measure loss aversion is unsupported; if an increase is found, it provides direct behavioral evidence for the proxy, independent of GRiPS/ADMC correlations.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing step is the construct-level claim, stated in the Cognitive Bias Hypothesis section, that 'an adversary exhibiting loss aversion will show a heightened focus on maintaining access through the use of persistence mechanisms.' This premise is asserted, not validated, and the psychometric evidence in the paper undermines rather than supports it. ADMC RC2—the scale designed to measure resistance to loss framing, the mechanism most proximal to loss aversion—is non-significant both as a correlation (r = -0.20, p = 0.415) and as a regression predictor (β = -2.45, p = 0.290). The only nominally significant result is the GRiPS coefficient (β = -4.42, p = 0.045), but GRiPS is a measure of general risk propensity, not loss aversion; equating low GRiPS with high loss aversion is an additional unvalidated assumption. This matters because the title and conclusion make claims about loss aversion specifically. Even if the LLM annotation pipeline were perfectly accurate, the correlation between persistence counts and generic risk propensity would not establish that persistence behavior is driven by loss aversion rather than by expertise, engagement, or other factors. The proxy gap is therefore more load-bearing than the pipeline-validation gap.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper describes a proof-of-concept study in which GPT-4o is used to parse operational notes (OPNOTES) written by 17 penetration testers during a controlled red-team exercise and to label actions against MITRE ATT&CK persistence techniques. The authors then compute counts of persistence techniques per participant and relate these counts to three psychometric measures: GRiPS (general risk propensity) and two ADMC Resistance to Framing subscales (RC1 for gains, RC2 for losses). The results show a moderate negative correlation between GRiPS and persistence count (r = -0.43, p = 0.065) and a nominally significant negative regression coefficient for GRiPS (β = -4.42, p = 0.045); the ADMC subscales are non-significant, and the overall regression model is non-significant (F(4,14) = 2.11, p = 0.133). The authors conclude that the approach demonstrates that LLMs can reliably extract temporal action sequences, identify MITRE persistence techniques, and associate these actions with cognitive traits such as loss aversion.","tokens_in":8815,"tokens_out":5683,"duration_ms":54094,"significance":"If the central claim were established, the work would offer a scalable method for extracting behavioral indicators of cognitive bias from off-the-shelf text artifacts, with potential applications in active cyber defense. The paper has several strengths: the use of externally validated psychometric instruments (GRiPS, ADMC) and the transparency of the statistical reporting (confidence intervals, model fit statistics). The attempt to ground a cognitive construct in observable attack behaviors is a valuable direction for the ReSCIND program. However, the evidence presented does not support the specific claim of loss aversion: the scale designed to measure loss-framing resistance (ADMC RC2) is non-significant, the only significant predictor is a generic risk propensity measure, and the proxy linking persistence counts to loss aversion is asserted rather than validated. As presently stated, the title and conclusion outrun the data.","major_comments":[{"comment":"The central premise that 'an adversary exhibiting loss aversion will show a heightened focus on maintaining access through the use of persistence mechanisms' is asserted without supporting evidence or a literature citation. No data are provided to show that persistence behavior is driven by loss aversion rather than by mission requirements, expertise, or habits. This is load-bearing because the dependent variable in the analysis is persistence count; without construct validation, the statistical results cannot be interpreted as evidence about loss aversion. The paper should provide external validation (e.g., expert ratings of whether each persistence action was motivated by loss aversion) or visibly restrict the claims to 'risk aversion' or 'persistence behavior'.","section":"Cognitive Bias Hypothesis"},{"comment":"The paper claims that LLMs can 'reliably extract temporal action sequences' and that classifications were 'verified during manual review', but no reliability metrics are reported. In particular, the reader cannot assess precision or recall of the LLM's MITRE persistence technique annotations against human expert annotations, inter-annotator agreement, or the consistency of the segmentation. Since the persistence counts that form the dependent variable come entirely from this pipeline, a systematic bias in the LLM's labeling could create the observed correlations. Report agreement statistics on a held-out subset or a comparison against a human-labeled baseline.","section":"Data Processing & Annotation Pipeline"},{"comment":"The regression evidence is too weak to support the paper's conclusion. The overall model is non-significant (F(4,14) = 2.11, p = 0.133), the adjusted R² is only 0.198, and the single nominally significant coefficient (GRiPS, p = 0.045) would not survive a conservative multiple-comparison correction across the three psychometric predictors plus division. With n = 17, the estimate is fragile. The paper should present the results as exploratory, report effect sizes with confidence intervals, and avoid causal or strong confirmatory language in the abstract and conclusion.","section":"Results and Findings / Table 2"},{"comment":"The interpretation of GRiPS as a surrogate for loss aversion is not defended. GRiPS measures general risk propensity, and the loss-framing subscale RC2, which is the operationalization most directly tied to loss aversion, is non-significant. To claim loss aversion specifically, the authors need a theoretical or empirical justification for why low GRiPS should be equated with high loss aversion in this population, or they should reframe the paper to be about risk aversion rather than loss aversion. Without this, the title and abstract overstate the construct.","section":"Analysis Approach"}],"minor_comments":[{"comment":"The sentence 'We process the hacker generated notes using LLMs using it to segment the various actions' is grammatically awkward and should be reworded.","section":"Abstract"},{"comment":"In the Related Work section, 'take palace' should be 'take place'.","section":"Related Work"},{"comment":"The author biography contains 'mulitmodal', which should be 'multimodal'.","section":"About the Authors"},{"comment":"The 'LA_' prefix in the column headers is undefined; please define the abbreviation or rename the columns for clarity.","section":"Table 1"},{"comment":"The axes of Figure 1 are not labeled; the y-axis presumably denotes occurrence counts but this is not identified.","section":"Figure 1"},{"comment":"The reference 'Kahneman, D., & Tversky, A. (2013)' is a reprint; consider citing the original 1979 Econometrica article or clearly noting that it is a reprint, and check consistency of author name ordering elsewhere (e.g., 'Nir, D.' versus 'Daniel et al.' in the text).","section":"References"},{"comment":"The paper does not mention institutional review board approval or ethical approval for the human-subjects experiment, despite recruiting human participants as penetration testers; this should be stated.","section":"Research Methods / Experimental Setup"}],"recommendation":"major_revision","confidential_remarks":"This is a short proof-of-concept conference paper. The gaps are substantial, and I would urge the editor to require a serious revision that either adds validation of the annotation pipeline and the persistence-loss-aversion construct or substantially narrows the paper's claims. Reframing from 'loss aversion' to 'risk aversion' with explicit limitations would make the contribution defensible as an exploratory study, but the current wording overstates the evidence."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's the quick version: novel dataset and a genuinely interesting pipeline, but the paper's own statistics don't support the loss-aversion claim, and the abstract oversells what the experiment can show. The most honest reading is that it is a proof-of-concept for LLM-based annotation of adversary notes, not a measurement of loss aversion.\n\nWhat's new: they take real OPNOTES from a 17-person red-team exercise, run them through GPT-4o to extract MITRE persistence techniques, then correlate counts with three psychometric scales. As far as I know, no one has done exactly this combination before. The pipeline is described carefully and the authors are transparent about the small sample, the non-significant overall model, and the low adjusted R². That transparency deserves credit.\n\nWhere it gets soft: the construct gap is load-bearing. The hypothesis section simply asserts that persistence mechanisms are a behavioral manifestation of loss aversion. Nothing validates that mapping. The direct loss-framing measure, ADMC RC2, fails to predict persistence (r = -0.20, p = 0.41; β = -2.45, p = 0.29). GRiPS—a generic risk-propensity scale—does show a hint of an effect, but that is not loss aversion. Equating low GRiPS with high loss aversion is an extra step the authors just assume. So the central inference is not established.\n\nThe statistics also get somewhat over-interpreted. The GRiPS regression coefficient is significant at p = 0.045, but the overall model is not (F(4,14) = 2.11, p = 0.133), and with n = 17 and four predictors, that is a single weak signal among several nulls. The correlation is borderline (p = 0.065). A more careful paper would present this as exploratory and avoid claiming evidence of loss aversion.\n\nThere is also no validation of the LLM annotation itself. The authors mention 'manual review' but report no inter-annotator agreement, no human baseline, no error analysis. Given that the whole empirical chain rests on those annotations, that is a real gap, though a fixable one.\n\nWho is it for: anyone working on cyberpsychology, LLM behavior annotation, or IARPA-style ReSCIND research. It is a useful illustration of how to set up a pipeline and how to overreach when the construct measurement is weak.\n\nMy recommendation: don't desk-reject it, but don't accept it as is either. Send it to peer review with the expectation that the authors will need to validate the annotation pipeline, add a human-coded ground truth, and either reframe the claims around risk propensity or provide direct evidence linking persistence to loss aversion.","headline":"A transparent proof-of-concept for LLM-based annotation of hacker notes, but the paper's own statistics do not support the loss-aversion claim it leads with.","tokens_in":9384,"tokens_out":2392,"would_cite":false,"duration_ms":23227,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that large language models can extract behavioral evidence of loss aversion from hackers' operational notes by counting persistence techniques, with lower self-reported risk propensity predicting more persistence.","keywords":["loss aversion","large language models","MITRE ATT&CK","persistence techniques","operational notes","cyberpsychology","cognitive bias","red team experiment"],"falsifier":"A reader could settle this by having two expert human coders independently label the same 17 sets of operational notes for persistence techniques and for stated motivations, then comparing their tallies to the LLM's counts; if the LLM counts diverge substantially, or if a larger preregistered sample fails to reproduce the negative GRiPS coefficient, the central claim would be falsified.","tokens_in":8363,"feed_emoji":"🛡️","tokens_out":13845,"duration_ms":129757,"temperature":0.7,"pith_summary":"This paper tries to show that large language models can turn free-form notes written by hackers during a simulated attack into quantitative evidence of a cognitive bias: loss aversion, the tendency to weigh potential losses more heavily than equivalent gains. Its central claim is that attackers with lower general risk propensity use more persistence techniques—technical moves for keeping access to a compromised system—and that this behavioral pattern is a signature of loss aversion. The authors build a multi-stage LLM annotation pipeline, apply it to operational notes from 17 red-team participants, and report a significant negative coefficient for risk-propensity scores ($\\beta = -4.42$, $p = 0.045$) alongside a non-significant overall model ($F(4,14) = 2.11$, $p = 0.133$). If the claim holds, defenders could infer an adversary's psychological state from text logs and adapt defenses in near real time, moving beyond static fortification. The paper frames the result as a proof of concept with a small sample.","feed_headline":"Risk-averse hackers leave more persistence techniques","feed_subtitle":"AI analysis of attacker notes links lower risk tolerance to more persistence moves in a red-team study.","key_machinery":"The central object is the two-stage LLM annotation pipeline built on GPT-4o. First, the model segments each participant's operational notes into discrete timestamped actions; second, it classifies each action against the MITRE ATT&CK framework's persistence techniques, producing explicit reasoning before assigning a label. The output is a per-participant count of persistence techniques, which the paper treats as a behavioral proxy for loss aversion and feeds into Pearson correlations and a multivariate linear regression with GRiPS, ADMC RC1, ADMC RC2, and participant division as predictors. This machinery is what converts unstructured self-report text into a testable quantitative signal.","core_discovery":"The paper reports that LLM-based analysis of operational notes from a controlled red-team exercise can extract temporal action sequences, tag actions with MITRE ATT&CK persistence techniques, and link those tags to psychometric indicators of loss aversion. The flagship quantitative result is a negative relationship between General Risk Propensity Scale (GRiPS) scores and persistence counts: participants with lower self-reported risk propensity used more persistence techniques ($\\beta = -4.42$, $p = 0.045$; $r = -0.43$, $p = 0.065$). The two Adult Decision-Making Competence (ADMC) resistance-to-framing subscales, RC1 and RC2, did not significantly predict persistence usage ($p = 0.121$ and $p = 0.290$), and division membership was only marginal ($p = 0.109$). The model explained 37.6 percent of the variance (adjusted $R^2 = 0.198$) but was not significant overall, and the authors interpret the GRiPS finding as consistent with loss aversion while acknowledging the need for replication with larger datasets.","pith_inferences":["A validation study comparing LLM-extracted persistence counts against independent expert coders and against a direct behavioral loss-aversion task (e.g., mixed gambles) would settle whether the counts measure loss aversion rather than routine attacker tradecraft.","The gap between raw and adjusted R-squared (0.376 vs 0.198) and the sample size of 17 suggest the effect size is likely optimistic; a preregistered replication would clarify how much of the association is real.","If the proxy is validated, the same pipeline could run on machine-generated network logs instead of self-reported notes, allowing near-real-time cognitive inference in live intrusions.","The combination of a significant GRiPS effect and null framing effects hints that what the paper calls loss aversion may actually be general risk avoidance rather than framing-specific loss aversion; a targeted experiment with loss-framed scenarios could resolve this."],"forward_implications":["If the central claim holds, defender systems can process attacker-authored notes and treat persistence-heavy behavior as a real-time indicator of loss aversion, enabling adaptive countermeasures.","The significant GRiPS coefficient implies that a standard psychometric risk-propensity questionnaire, combined with LLM-extracted behavioral counts, can anticipate how much effort an attacker will invest in maintaining access.","The null ADMC framing results imply that susceptibility to gain/loss framing is a weaker signal than general risk propensity in operational cyber settings, steering future measurement toward risk-propensity instruments.","Because the annotation pipeline is modular, the same segmentation-and-classification approach can be applied to other MITRE tactics and to other cognitive biases, such as the sunk-cost fallacy or confirmation bias."],"supporting_citations":[{"why":"Supplies the survey of MITRE ATT&CK usage that grounds the choice of persistence techniques as the annotation taxonomy.","marker":"Al-Sada et al. (2024)"},{"why":"Shows ATT&CK can categorise post-compromise attacker techniques, the precedent for mapping actions to persistence.","marker":"Oosthoek & Doerr (2019)"},{"why":"Provides experimental evidence that red-team adversaries exhibit cognitive biases, the prior result this study extends.","marker":"Ferguson-Walter et al. (2018)"},{"why":"Defines loss aversion in prospect theory, the psychological construct the study operationalises.","marker":"Kahneman & Tversky (2013)"},{"why":"The GPT-4o model used for action segmentation and persistence classification, the pipeline's core tool.","marker":"OpenAI (2024)"},{"why":"Shows LLMs can infer psychological dispositions from text, supporting the use of LLMs for cognitive inference.","marker":"Peters & Matz (2024)"},{"why":"Validates LLMs as behavioural text annotators, supporting the reliability of the annotation outputs.","marker":"Bunt et al. (2025)"}],"fun_headline_variants":["Loss aversion predicts hacker persistence moves","Low risk tolerance, more persistence in attackers","Hackers who fear loss plant more persistence","Risk-averse cyber attackers favor persistence","LLM study links risk aversion to persistence"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the count of persistence techniques an LLM extracts from a hacker's notes is a valid behavioral stand-in for loss aversion, with no independent check that those actions reflect fear of losing access rather than routine tradecraft or model artifacts.","fun_headline_variants_meta":{"raw":{"variants":["Loss aversion predicts hacker persistence moves","Low risk tolerance, more persistence in attackers","Hackers who fear loss plant more persistence","Risk-averse cyber attackers favor persistence","LLM study links risk aversion to persistence"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000618,"raw_usage":{"total_tokens":2882,"prompt_tokens":976,"completion_tokens":1906,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":592,"completion_tokens_details":{"reasoning_tokens":1842}},"tokens_in":592,"tokens_out":1906,"duration_ms":15451,"temperature":1.0,"reasoning_tokens":1842,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T17:19:48.660912+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A reader could settle this by having two expert human coders independently label the same 17 sets of operational notes for persistence techniques and for stated motivations, then comparing their tallies to the LLM's counts; if the LLM counts diverge substantially, or if a larger preregistered sample fails to reproduce the negative GRiPS coefficient, the central claim would be falsified.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the survey of MITRE ATT&CK usage that grounds the choice of persistence techniques as the annotation taxonomy."}],"review_version":2}