{"id":"08be7cd7-e9b4-40e8-b608-ea17f2d43ded","arxiv_id":"2506.12189","paper_version":2,"verdict":"REJECT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":3,"one_line_summary":"LLMs asked to rank critical events in articles show different, consistent event-selection styles that the authors label as personality traits using an LLM judge.","lead":"This paper introduces the Supernova Event Dataset, in which LLMs pick and rank the five most important events from biographies, news stories, and scientific articles. A second LLM then labels each model's style as a personality, and the authors report stable differences, such as Orca 2 being emotional and Qwen 2.5 being strategic.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central personality claim rests entirely on unvalidated LLM-as-judge labels; a single judge (Qwen 2.5) that also evaluates its own outputs cannot establish distinct traits until the measurement is independently validated.","rationale":"The reader's weakest assumption identifies the same load-bearing point: the entire personality inference depends on the validity of LLM-as-judge measurements, and this validity is never established. My read does not change the reader's REJECT verdict. The paper is self-aware and discloses the lack of human validation and the risk of stylistic judge bias (Sec. 8), and the dataset/code release is a useful resource. However, those acknowledgments do not supply the missing measurement evidence. The central claim — that models exhibit stable, distinct personality traits without personality prompting — would be true only if judge labels reliably track something about the target models rather than the judge or the prompt. Since the only judge for the main analysis is qwen2.5:14b, and since it evaluates its own outputs, the observed separation in Fig. 3 could be an artifact of self-preference or category wording. Similarly, in the scientific-discovery subanalysis, o3 both generates the reflective labels and assigns them to the codebook, which is circular for the claim that o3 is causality-centric. The proposed cross-judge replication test would directly settle whether the profiles are stable across independent measurement instruments. Until such a test is run, the central claim remains unsupported, and the paper should not be accepted as establishing LLM personalities.","tokens_in":23247,"tokens_out":5700,"duration_ms":70644,"concrete_test":"Take the complete set of target-model responses used in Sec. 5.1 and Sec. 6.3. Have three independent judge LLMs that are not among the target models (e.g., GPT-4o, Claude 3.7, Llama-3.1-70B) apply the same Box 3 prompt and the same three-way codebook, with judge identity blinded, and compute pairwise agreement (e.g., Cohen's kappa) and the resulting per-model personality distributions. If the profiles do not replicate across judges — or if qwen2.5 is labeled differently by judges other than qwen2.5 itself — the distinct-personality conclusion is a judge artifact rather than a property of the target models.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The core result in Sec. 6 — that phi4, orca2, and qwen2.5 occupy distinct personality profiles, and that Claude/Gemini/o3 are synthesis/enablement/causality-centric — is inferred solely from labels produced by an LLM judge. In Sec. 5.1, qwen2.5:14b is the only judge for all three target models, including qwen2.5 itself; the judge prompt (Appendix A.1, Box 3) supplies the seven category names and asks for a single-line classification. No human labels, inter-rater reliability, or cross-judge agreement are reported. LLM judges are known to have stylistic and self-preference biases (the paper cites Cao 2024 and Krumdick et al. 2025 but does not control for them). Consequently, the reported trait distributions in Fig. 3a/3b and the 'emotional vs strategic' contrast could reflect the judge's response tendencies or the particular category wording rather than stable properties of the target models. The scientific-discovery analysis compounds this: the three-way causality/enablement/synthesis codebook is applied by o3 (Sec. 6.3), and o3 is one of the three models being characterized, so its own labels and their assignment are not independent of the trait being measured. Because every headline personality claim passes through this unvalidated measurement step, the central claim is unsupported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces the Supernova Event Dataset, a collection of Wikipedia biographies, historical and news articles, and scientific-discovery narratives, and proposes a task in which an LLM extracts and ranks the five most critical events from a document. The authors then use a second LLM as a judge to classify the target model's personality from its event selections and rankings. They report that Phi-4 is strategic/achievement-oriented, Orca 2 is emotional, Qwen 2.5 is balanced/strategic, and that among stronger models Claude 3.7 is synthesis-centric, Gemini 2.5 is enablement-centric, and o3 is causality-centric. A movie-script ablation is presented as confirming these profiles. The paper frames this as a prompt-agnostic behavioral probe for LLM personality and releases the dataset and code.","tokens_in":23706,"tokens_out":3337,"duration_ms":43043,"significance":"If the personality measurements were valid, the Supernova dataset and critical-event-ranking task would be a useful addition to subjective and long-context LLM evaluation, and the observation that personality-like patterns emerge without explicit personality prompting would be of broad interest. The authors provide a new dataset, a concrete task formulation, and public code and data. However, the central empirical claim rests on a single LLM judge that is itself one of the evaluated models, with no human validation, no inter-rater reliability, and no cross-judge agreement. The scientific-discovery analysis similarly uses o3 both as a target and as the label assigner. Because every headline trait attribution passes through this unvalidated measurement step, the current results should be treated as exploratory hypotheses rather than established findings. The paper acknowledges these limitations in Section 8, but the presentation throughout Sections 6 and 7 states the personality traits as conclusions, not as provisional observations.","major_comments":[{"comment":"The personality classifications underlying all headline claims are produced solely by Qwen 2.5 14B, which is also one of the three evaluated targets. Section 5.1 states that qwen 2.5:14b evaluates phi4, orca2:13b, and qwen 2.5:14b, and the judge prompt in Appendix A.1 supplies the seven category names and asks for a one-line classification. No human labels, inter-rater reliability, or agreement with an independent judge are reported. The paper cites known LLM-judge stylistic and self-preference biases (Cao, 2024; Krumdick et al., 2025) but does not control for them. Consequently, the trait distributions in Fig. 3a and the semantic separation in Fig. 3b may reflect the judge's response tendencies or the supplied category wording rather than stable properties of the target models. This is load-bearing: the central claim that the models have distinct personalities is unsupported without an independent, validated measurement.","section":"Sec. 5.1, Fig. 3"},{"comment":"The scientific-discovery reasoning profiles are assigned using o3 with the 'finalized three-way codebook', while o3 is one of the three models being characterized. The codebook itself is derived post hoc through keyword counting and open coding, and the assignment of every label to a category is performed by o3. This is not an independent measurement: o3's own labels and their assignment are entangled with the trait being measured. The claim that o3 is causality-centric, Gemini is enablement-centric, and Claude is synthesis-centric therefore needs either human annotation with reported agreement or an independent judge that is not one of the target models. The current figure and table do not provide that evidence.","section":"Sec. 6.3, Fig. 2"},{"comment":"Section 5.1 concedes that the personality categories are 'empirically derived rather than grounded in established psychological frameworks' and that the judge approach 'introduces potential biases and lacks human validation.' Yet Section 6.1 and the abstract present the resulting attributions ('Orca 2 demonstrates emotional reasoning', 'Qwen 2.5 displays a more strategic, analytical style') as findings. Given that the categories are post hoc and the judge is a single model, the paper should either substantially temper these claims or provide the missing validation. As written, the results section overstates the evidential status of the measurements.","section":"Sec. 5.1, Sec. 6.1"},{"comment":"No sample size, confidence interval, or statistical comparison is reported for the personality-category distributions. It is therefore unclear whether the differences between phi4, orca2, and qwen2.5 in Fig. 3a are stable across documents or within the range of judge noise. Reporting the number of judged responses, per-model counts, and a measure of judge consistency would be necessary to support the claim that the models 'occupy distinct regions in the personality space.'","section":"Sec. 6.1, Fig. 3a"}],"minor_comments":[{"comment":"The judge prompt contains a typo: 'Idealogical' should be 'Ideological', and the category list is inconsistently capitalized relative to the categories used in Figure 3a (e.g., 'Public Influence' vs. 'Influencer').","section":"Appendix A.1, Box 3"},{"comment":"The figure contains stray text elements ('1% 21. 7%') that appear to be rendering artifacts; please clean the figure so the axis labels and category percentages are legible.","section":"Fig. 3a"},{"comment":"The description of the retrieval pipeline is underspecified: 'MultiQueryRetriever' and 'two-stage prompting' are mentioned but the actual query reformulation behavior is not described beyond the prompt in Box 1. Please clarify how many queries are generated and how retrieval quality was checked.","section":"Sec. 4"},{"comment":"The sentence 'We do not apply any post-processing and verify for hallucination before saving our articles' is self-contradictory; hallucination verification is a form of post-processing. Please rephrase to describe the verification procedure.","section":"Sec. 3.3"},{"comment":"The citation to Cao (2024) is titled 'Writing Style Matters: An Examination of Bias and Fairness in Information Retrieval Systems'; this does not appear to be the intended reference for LLM-judge stylistic bias. Please verify and correct the citation.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper is well positioned as an exploratory workshop contribution, but for a journal venue the measurement validity problem is severe: all personality attributions flow through a single LLM judge that is also a target, and the scientific-discovery analysis uses o3 as both target and labeler. This is fixable with human validation, independent judges, and more cautious framing, but the current manuscript does not support its central claims as stated. I would not recommend rejection outright because the dataset and task could be valuable if the measurement is properly validated."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick summary: the dataset and the task are real assets, and the paper is honestly written about its own limits. But the core personality claim does not survive contact with the measurement setup.\n\nWhat is genuinely new and good: the Supernova dataset (bio, history, news, scientific discoveries, plus the movie-script ablation) is a reasonable, reusable resource for long-context and event-salience work. Framing critical-event ranking as an implicit behavioral probe, rather than asking models to self-report personality, is a legitimate extension of the existing LLM-personality literature, and the paper cites that literature properly. The authors also disclose the major weaknesses in Section 8: they know LLM judges carry stylistic bias and that there is no human validation. That matters; this is not a paper that hides its problems.\n\nWhere it falls apart: every personality label in Figure 3 and the scientific discovery analysis comes from a single judge (Qwen 2.5 14B) that is itself one of the evaluated models (Section 5.1), and the judge prompt in Appendix A.1 supplies the seven category names in advance. The judge cannot report anything outside those labels, and since Qwen 2.5 judges Qwen 2.5, the \"distinct\" profiles for phi4, orca2, and qwen2.5 may just be the judge's own stylistic tendencies. The scientific discovery analysis is worse: o3 generates the labels, then o3 is evaluated against a codebook it helped produce (Section 6.3, Table 6), and the open coding that \"converges\" on the three-way scheme is not shown as an independent process with reliability metrics. The paper acknowledges these issues, but acknowledgment does not rescue the inference. The observed event-ranking differences themselves are real and sometimes interesting (e.g., Qwen ranking systemic causes first in the financial crisis example), but calling them personality traits requires a validated measurement step that is entirely absent. There are also free parameters everywhere: the seven personality categories, the three-way codebook, the dataset thresholds, none of which are given stability tests.\n\nBottom line: as a dataset-plus-task contribution, this deserves attention and could be useful to people working on long-context evaluation or event salience. As a personality-interpretation result, it is exploratory at best. A serious referee should send it back for major revision, with human labels or a multi-judge validation as a prerequisite for any claim about stable model traits. I would not cite the personality findings, but I would point people to the dataset.","headline":"Useful dataset and a genuinely new task framing, but the personality inference rests entirely on an unvalidated self-referential LLM judge, so the headline claim is unsupported as stated.","tokens_in":24102,"tokens_out":633,"would_cite":false,"duration_ms":10258,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Language models, without any personality prompt, exhibit consistent decision-making styles—emotional, strategic, causal—when ranking critical events, and those styles can be read as personality traits by a judge model.","keywords":["LLM personality","critical event ranking","LLM-as-a-judge","event salience","model interpretability","subjective benchmarks","long-context reasoning"],"falsifier":"A decisive check would be to run the identical event-ranking outputs through several different LLM judges, or through human annotators who are blind to model identity, and see whether the personality assignments—emotional for Orca 2, strategic for Qwen 2.5, causality-centric for o3—reproduce; if judges disagree, or if the target's label flips when the judge prompt is reworded without the seven pre-supplied categories, then the measured personality belongs to the judge, not the target.","tokens_in":23049,"feed_emoji":"🎭","tokens_out":5886,"duration_ms":56578,"temperature":0.7,"pith_summary":"This paper tries to establish that a large language model's personality can be read from how it picks and ranks critical events in narratives, without telling the model to adopt any role. The authors build the Supernova Event Dataset—long Wikipedia biographies, news, historical events, and scientific-discovery articles—and ask six models to select and order the five most decisive events, using counterfactual tests ('Would the narrative have unfolded differently?') as selection criteria. A second LLM (Qwen 2.5) then classifies each model's rankings into personality categories. The reported result is stable differentiation: Orca 2 favors emotional and interpersonal moments, Phi-4 and Qwen 2.5 prefer strategic, achievement-oriented events, and among stronger models Claude Sonnet 3.7 frames concepts, Gemini 2.5 Pro prioritizes validation and enabling methods, and o3 follows causal chains. If right, critical-event ranking becomes a behavioral probe for model personality and value alignment, usable where factual benchmarks fall short.","feed_headline":"LLMs reveal fixed personalities in how they rank key events","feed_subtitle":"Without a role-play prompt, Orca 2 reads emotional, Qwen 2.5 strategic, o3 causal—a judge reads their event picks.","key_machinery":"The load-bearing mechanism is the critical-event sampling-and-ranking task combined with an LLM judge. Each target model receives one long article (via retrieval-augmented generation over chunked text) and is prompted to extract exactly five critical events, rank them from most to least critical, and explain why the top event is decisive, using counterfactual tests ('Would the narrative have unfolded differently?') as the selection criterion. The judge model (Qwen 2.5 14B) then reads the target's ranked list and classifies it into one of seven supplied categories—Ideological, Emotional, Strategic, Creative, Observational, Public Influence, Community Support—or, for the scientific-discovery runs, into causality-centric, enablement-centric, and synthesis-centric categories. The mechanism works by converting an open, subjective judgment into a forced-choice classification of the target's choices, which is what lets the authors call the differences 'personality.'","core_discovery":"The central discovery the paper defends is that personality-like behavioral patterns emerge in LLMs without explicit personality prompting, and these patterns can be recovered from a subjective task: identifying and ranking the five critical events that most decisively changed a narrative's trajectory. Using its proposed Supernova Event Dataset, the paper reports that Phi-4 exhibits a strategic-achiever orientation, Orca 2 an emotional orientation centered on relationships, and Qwen 2.5 a strategic, systemic style; on scientific-discovery articles, o3 favors step-by-step causality, Gemini 2.5 Pro emphasizes empirical validation and enabling methods, and Claude Sonnet 3.7 favors conceptual framing. These labels are produced by an LLM judge that inspects the target model's ranked event lists, motivated by evidence that models' self-explanations misrepresent their reasoning. The paper treats 'personality' as a metaphor for consistent behavioral patterns, not consciousness or emotion.","pith_inferences":["An implication the paper leaves implicit is that the inferred personality is best read as a property of the model–judge pair: because the judge's own stylistic tendencies and the seven supplied category names shape the label, the same target model might type differently under a different judge, and a testable extension is to hold the target outputs fixed and sweep across judge models.","If the reported differences are real, a natural next question is what causes them—training data composition, post-training alignment, or decoding strategy; the paper does not address this, but the dataset could be adapted to compare checkpoints of the same base model before and after alignment to localize the origin.","The three-way codebook for scientific discovery (causality, enablement, synthesis) could be validated by having human scientists label the same event lists, which would show whether the categories capture recognizable reasoning styles or are artifacts of the open-coding procedure."],"forward_implications":["Event ranking can serve as a prompt-free behavioral probe: any LLM's stable decision-making style can be profiled without role-play instructions.","Model selection becomes more informed: users could choose a model whose inferred priorities (e.g., relational/emotional vs. strategic/causal) match the needs of a task.","The Supernova Event Dataset supports additional research on long-context reasoning, causal-chain modeling, and counterfactual reasoning beyond personality labeling.","The framework shifts evaluation away from factual accuracy toward subjective judgment and value alignment, which matters for high-stakes deployments in healthcare, law, and finance."],"supporting_citations":[{"why":"Motivates the indirect judge-based assessment by showing that models' self-explanations misrepresent their reasoning, which is why the paper avoids self-report.","marker":"Lindsey et al., 2025"},{"why":"Shows LLMs can express personality when prompted; the present work extends this by claiming the traits appear without prompting.","marker":"Jiang et al., 2023"},{"why":"Demonstrates that emulated personality traits affect decision-making, the link the paper uses to infer traits from event-ranking decisions.","marker":"Wang et al., 2025"},{"why":"Shows LLMs simulate Big Five traits when prompted; the paper contrasts its prompt-agnostic setup with this line of work.","marker":"Sorokovikova et al., 2024"},{"why":"Documents the limitation of LLM-as-judge without human grounding, which the paper cites as a known caveat for its inference.","marker":"Krumdick et al., 2025"},{"why":"Provides evidence that LLM judges exhibit stylistic biases that could shape trait inference, a limitation the paper acknowledges.","marker":"Cao, 2024"}],"fun_headline_variants":["LLMs show distinct traits in event ranking, no prompts needed","O3 causal, Orca 2 emotional: LLM personality via event picks","Event selection exposes LLM personalities without role-play","Judge LLM reads model personality from critical event choices","Claude, Gemini, o3 differ in event analysis: new dataset"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the judge LLM's classification of the target model's event rankings is a valid measurement of the target's personality, rather than a reflection of the judge's own stylistic biases or of the seven category labels the prompt supplies.","fun_headline_variants_meta":{"raw":{"variants":["LLMs show distinct traits in event ranking, no prompts needed","O3 causal, Orca 2 emotional: LLM personality via event picks","Event selection exposes LLM personalities without role-play","Judge LLM reads model personality from critical event choices","Claude, Gemini, o3 differ in event analysis: new dataset"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000189,"raw_usage":{"total_tokens":1356,"prompt_tokens":988,"completion_tokens":368,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":604,"completion_tokens_details":{"reasoning_tokens":281}},"tokens_in":604,"tokens_out":368,"duration_ms":99438,"temperature":1.0,"reasoning_tokens":281,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T00:55:40.643604+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A decisive check would be to run the identical event-ranking outputs through several different LLM judges, or through human annotators who are blind to model identity, and see whether the personality assignments—emotional for Orca 2, strategic for Qwen 2.5, causality-centric for o3—reproduce; if judges disagree, or if the target's label flips when the judge prompt is reworded without the seven pre-supplied categories, then the measured personality belongs to the judge, not the target.","supporting_citations":[],"review_version":1}