{"id":"22878d52-9ce1-44f8-8d27-e957219bb1ef","arxiv_id":"2504.14706","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"LLMs given numerical arousal and valence coordinates produce text that sentiment analysis places in the same region of Russell's circumplex, demonstrating limited but real control over emotional tone.","lead":"The authors asked nine large language models to role-play emotional states defined by two numerical dials, arousal and valence, and then used a sentiment classifier to see if the responses matched. Most models steered their text toward the requested state, with GPT-4 and Llama3-70B being the most controllable. The work is a capability check for emotion-tunable AI assistants, not a claim that models actually feel emotions.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Unvalidated Table 3 label-to-Russell mapping is the load-bearing link in every cosine similarity; independent mapping or human ratings are needed before the central claim can be accepted.","rationale":"The reader's weakest-assumption analysis identified exactly the same load-bearing point: the entire quantitative evaluation depends on the researcher-authored mapping in Table 3 from GoEmotions labels to Russell's arousal-valence space. My review confirms this is the right concern and that it is more serious than a minor methodological caveat. The mapping is used twice, once to convert predicted labels into vectors for generation scores and once in Appendix A to validate the sentiment model, so the Appendix cannot serve as an independent check. The claim about model superiority, e.g., GPT-4, GPT-4 turbo, and Llama3-70B performing best, is a ranking claim that could shift under a different mapping. The paper is transparent about limitations and provides useful checks, such as the word-count correlation in Appendix C and the word-specified comparison in Appendix D, but those do not address the mapping validity. A conditional verdict remains appropriate: the qualitative direction of the finding is plausible, and the generated examples show some emotional expressiveness, but the quantitative support is not secure until an independent mapping or human judgment is applied. No change to the reader's verdict is needed.","tokens_in":15317,"tokens_out":4392,"duration_ms":47610,"concrete_test":"Have the 27 non-neutral GoEmotions labels rated by three or more human annotators on continuous arousal and valence scales, or locate each label's coordinates in an established dimensional norm such as Warriner et al., 2013, and recompute Appendix A's mean cosine, Table 2's totals, and the model ranking using this independent mapping. If the recomputed totals remain clearly above the heuristic baseline and the top-three model set is unchanged, the mapping concern is resolved; if not, the central claim of controlled emotional expression is not established by the current evaluation.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that LLMs can control emotional expression is established solely by cosine similarities between the specified arousal-valence vector and the vector of the GoEmotions label predicted by a sentiment classifier. Both vectors are placed on the same circumplex through Table 3, a one-to-two mapping authored by the researchers with no human validation or independent norming. The same mapping is reused in Appendix A to validate the sentiment model by comparing correct and predicted labels, so Appendix A cannot provide independent support: a wrong mapping would make both the 'evaluator limit' and the generation scores shift in the same direction. There are entries that are plainly questionable, e.g., 'confusion' mapped to 'alarmed', 'optimism' to 'at ease', 'disgust' to 'frustrated', and 'desire' to 'excited, aroused'. Because only 27 discrete label vectors are available, every per-question mean similarity reflects how close a label lies to the target direction. The paper's own Limitation acknowledges that the discrete classifier may influence metrics, but it does not test the mapping's validity. A different mapping could lower Table 2 totals and alter the reported model ranking, including the claim that GPT-4, GPT-4 turbo and Llama3-70B are superior. Hence the quantitative central claim is conditional on an unvalidated assumption.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a framework for exploring whether current LLMs can express specified emotional states in generated text. Using Russell's Circumplex Model, the authors specify 12 emotional states as (valence, arousal) unit vectors and prompt nine LLMs (GPT-3.5/4/4-turbo/4o, Gemini 1.5 Flash/Pro, Llama3-8B/70B, Command R+) to answer ten questions while role-playing an agent with each specified state. The outputs are evaluated with a BERT-based sentiment classifier trained on GoEmotions; classifier predictions are mapped through a hand-crafted correspondence table (Table 3) to unit vectors in the same arousal–valence space, and cosine similarity between specified and evaluated vectors is the main metric. The authors report positive mean similarities for most model–question combinations, with GPT-4, GPT-4 turbo, and Llama3-70B Instruct showing the highest totals, and conclude that LLMs can control their outputs with specified emotional states. The paper also includes a baseline (0.061), an appendix validating the sentiment model via the same mapping, a check that text length does not correlate with similarity, and a word-conditioned control experiment.","tokens_in":15584,"tokens_out":5754,"duration_ms":50231,"significance":"If the central claim is adequately supported, the paper would provide a simple and transparent protocol for continuous emotional control of LLM-based agents, with plausible applications in advisory systems, creative generation, and human–AI interaction. The experimental design is commendably explicit: all models, prompts, and questions are listed, a heuristic baseline is used, and the authors checked for text-length confounds. However, the quantitative conclusion rests entirely on the authors' unvalidated mapping from GoEmotions labels to Russell coordinates, and the reported numbers lack any measure of uncertainty. Without external validation of that mapping and significance testing against the baseline, the model ranking and the 'capability for emotional expression' claim are conditional. With such validation, the paper would be a useful empirical contribution.","major_comments":[{"comment":"Every cosine similarity reported in Table 2 and Figure 4 is computed from the unit vectors assigned in Table 3, yet that table is the product of the authors' own judgment and is not validated against human ratings, established affective norms, or an independent source. Entries such as 'confusion' mapped to 'alarmed', 'optimism' to 'at ease', 'disgust' to 'frustrated', and 'desire' to 'excited, aroused' illustrate that plausible alternative mappings exist, and a different mapping would change the numerical results and potentially the model ranking, including the claimed superiority of GPT-4, GPT-4 turbo, and Llama3-70B Instruct. This mapping is load-bearing for the central claim, so the authors should validate it externally and report a robustness analysis over alternative mappings.","section":"Section 3.2, Table 3 (Appendix B)"},{"comment":"Appendix A is presented as evidence that the sentiment analysis model can estimate emotional states in Russell space, but it uses the same Table 3 mapping to place both correct and predicted labels. This makes the reported mean cosine similarity of 0.680 a consistency check of the mapping, not an independent validation: a systematic error in Table 3 would shift the Appendix A results and the generation results in Table 2 in the same direction. The paper therefore cannot use Appendix A as support for the validity of the evaluation, and the central claim remains conditional on the unvalidated mapping.","section":"Appendix A"},{"comment":"Table 2 reports only means over the ten questions, with no standard deviations, confidence intervals, or significance tests, and its caption states that 'positive significance of all values' confirms the capability even though two cells are negative (GPT-3.5 Q2: -0.048; Gemini 1.5 Flash Q9: -0.049) and several cells fall below the 0.061 baseline. The discussion claims that 'LLMs can control their outputs with specified emotional states within a certain range,' but no statistical test quantifies this range or compares the observed aggregate similarities against the baseline. Please add dispersion measures and appropriate inferential tests, and correct the caption to reflect the actual values.","section":"Section 4, Table 2, Figure 4"}],"minor_comments":[{"comment":"In the reference to Khan (2022), 'Febrary' should be 'February'.","section":"References"},{"comment":"Figure 1 appears as garbled or obfuscated text in the manuscript; if the published figure is likewise unreadable, it should be replaced with the actual prompt text.","section":"Figure 1"},{"comment":"Table 5 (Appendix D) uses inconsistent decimal formatting, e.g., '0.66' in the GPT-3.5 Q4 column among three-decimal entries.","section":"Table 5 (Appendix D)"},{"comment":"The phrase 'the positive significance of all values' in the Table 2 caption contradicts the negative entries; reword to 'the majority of values are positive' unless a significance test is added.","section":"Table 2 caption"}],"recommendation":"major_revision","confidential_remarks":"The key issue is the unvalidated mapping from GoEmotions labels to Russell coordinates. I believe this is fixable with a norming study or use of existing affective norms, so I recommend major revision rather than rejection. The paper is within scope but currently overclaims given the dependency on that mapping."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague—this one is worth a look if you work on affective computing, but keep the right skepticism. The paper does something simple and useful: it asks nine LLMs to answer ten questions while role-playing an emotional state given as numeric arousal and valence coordinates, then checks whether a sentiment classifier maps the outputs back to the instructed region of Russell's circumplex. The numeric-prompting setup is the actual contribution; prior work mostly conditions on emotion words or scenarios. The comparison across open and closed models is also refreshingly straightforward, and the authors are transparent about limitations, including a check that text length doesn't bias the scores.\n\nThe soft spot is exactly where the stress-test put it. Table 3, the mapping from GoEmotions labels to Russell coordinates, is hand-built, with entries like confusion→alarmed and optimism→at ease that any affective scientist will blink at. Every cosine similarity in Table 2 is computed through that mapping, and Appendix A's validation of the sentiment model uses the same mapping, so it cannot independently support the evaluation. A different mapping could shift the numbers and the model ranking. That is a load-bearing assumption, not a minor detail. On top of that, Table 2 reports means with no standard deviations or significance tests, and the caption's claim that all values are positively significant is simply false—there are negative cells (GPT-3.5 Q2, Gemini Flash Q9). The baseline of 0.061 is also near the average cosine for random label pairs, so values like 0.147 for GPT-3.5 need variance to be interpretable.\n\nNone of this destroys the qualitative result that LLMs can coarsely steer emotional tone. The examples in Figure 3 look plausible, and the word-conditioning comparison in Appendix D shows the numeric approach is competitive. But the precise \"superior performance\" claims for GPT-4 and Llama3-70B are not established.\n\nWho benefits? Anyone building emotion-controlled agents or assessing LLM affective behavior. It's a decent empirical starting point, not a final benchmark. I'd send it to peer review—a serious referee can ask for human norming of the label mapping, error bars and significance tests, and a sensitivity analysis. Without that, the quantitative conclusions remain conditional.","headline":"Simple, honest experiment on numeric emotion prompting; the unvalidated GoEmotions-to-Russell label mapping makes the quantitative model ranking conditional.","tokens_in":16068,"tokens_out":3010,"would_cite":true,"duration_ms":28508,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Told to feel a specific arousal-valence state, most large language models reply in that emotional register.","keywords":["emotional expression","large language models","Russell's circumplex model","arousal","valence","GoEmotions","sentiment analysis","role-play prompting"],"falsifier":"Re-run the protocol with the Table 3 mapping replaced by coordinates from human raters or an independently published norming study, and compare mean cosine similarities. If the values fall to the 0.061 baseline or below across models, the reported emotional control was an artifact of the mapping rather than a property of the LLMs.","tokens_in":15116,"feed_emoji":"🎭","tokens_out":5829,"duration_ms":49257,"temperature":0.7,"pith_summary":"This study asks whether current large language models can make their written answers carry a specified emotion, not just a specified topic. The authors define each target emotion numerically, using the two axes of Russell's circumplex model (arousal and valence), feed the numbers to nine large language models in a role-play prompt, and score the replies with an independent sentiment classifier trained on GoEmotions. Across 12 evenly spaced emotional states and 10 questions, mean cosine similarity between the specified state and the state inferred from the text is positive for most model-question pairs and exceeds the heuristic baseline of 0.061. The authors conclude that LLMs can control their output emotional state within a range, which matters because emotion-typed agents become usable as advisors, consultants, or creative partners.","feed_headline":"Prompted with a two-number emotion dial, LLMs follow the tone","feed_subtitle":"Role-playing prompts made nine models' inferred emotions track their specified arousal and valence.","key_machinery":"The load-bearing object is Russell's circumplex model of affect, which places any emotional state at a point on a circle with axes arousal (sleepy-activated) and valence (pleasure-displeasure). The experiment fixes each target as a unit vector at one of 12 equally spaced angles, feeds the two coordinates into a role-play prompt, and measures the output by cosine similarity between the target vector and the vector of the sentiment label assigned by the GoEmotions-trained classifier. The chain lives or dies with the manual Table 3 mapping that converts 27 GoEmotions labels into Russell coordinates, because that mapping decides whether a response matches its instructed emotion.","core_discovery":"The central claim is that LLMs can translate a numerically specified emotional state into text whose emotional content matches the specification. The experiment uses the Russell circumplex: target states are unit vectors at 12 evenly spaced angles in the arousal-valence plane, and each prompt tells the model to role-play an agent experiencing that state while answering one of ten open questions. A BERT-based sentiment classifier trained on GoEmotions labels each generated text, and the label is mapped through the authors' Table 3 correspondence to a point in the same plane. Positive mean cosine similarities, generally above the 0.061 chance-level baseline, lead the authors to conclude that the models control their expressed emotional state within a certain range; GPT-4, GPT-4 turbo, and Llama3 70B Instruct show the steadiest performance, while GPT-3.5 turbo lags.","pith_inferences":["The mapping in Table 3 could be doing much of the work: if a different but equally plausible label-to-coordinate mapping lowered all similarities, the models' apparent emotional expression would shrink toward a property of the evaluator rather than the generator.","Because the sentiment evaluator was trained on Reddit comments, its label geometry reflects one online register; an English-only, Reddit-derived evaluation may understate how well LLMs express emotion in other cultural or conversational contexts.","Showing that a model can imitate a specified emotional register is not evidence that the model experiences emotion; the data support instruction-following in expressive style, not inner states.","A natural next test is human-participant evaluation: if human raters agree with the classifier's label assignments on the same generated texts, the reported alignment would be robust to the choice of evaluator; if not, the effect is partly measurement-specific."],"forward_implications":["If the claim is right, the emotional tone of an LLM agent can be selected at inference time from two continuous numbers, without retraining or fine-tuning.","Because the same control appeared across ten unrelated questions, the effect is not tied to one topic: tone-setting generalizes to a range of conversations.","Continuous arousal and valence coordinates open the door to scripting emotional dynamics, such as an agent that grows calmer or more agitated over the course of a dialogue.","Emotion-typed agents become feasible for applications where a deliberate emotional stance matters, such as advisors that can disagree with a user while staying in a chosen register.","The observed differences across models suggest that emotional control is not simply a function of parameter count; training data and alignment choices matter."],"supporting_citations":[{"why":"Supplies the circumplex model of affect whose two axes (arousal and valence) define every target emotional state in the experiment.","marker":"Russell, 1980"},{"why":"Provides the GoEmotions dataset and its 28 fine-grained emotional labels, the label set on which the sentiment classifier was trained.","marker":"Demszky et al., 2020"},{"why":"The HuggingFace sentiment model (sentiment-model-sample-27go-emotion) used to evaluate the emotion inferred from each generated text.","marker":"Khan, 2022"},{"why":"The BERT architecture underlying the sentiment model, establishing that the evaluator is independent of the GPT-family generators.","marker":"Devlin et al., 2019"},{"why":"Cited to support the claim that the GoEmotions-trained classifier outperforms general LLMs on emotion classification, justifying its use as an evaluator.","marker":"Koco´n et al., 2023"}],"fun_headline_variants":["Two-number emotion dial steers LLM text tone","LLMs follow emotion dial: tone matches numbers","Emotion dial works: LLM text matches specified mood","LLMs translate emotion numbers into text sentiment","Two numbers set LLM emotional tone"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole measurement rests on the authors' hand-built mapping from GoEmotions labels to coordinates on Russell's circle, which was not validated against human judgments or an independent reference.","fun_headline_variants_meta":{"raw":{"variants":["Two-number emotion dial steers LLM text tone","LLMs follow emotion dial: tone matches numbers","Emotion dial works: LLM text matches specified mood","LLMs translate emotion numbers into text sentiment","Two numbers set LLM emotional tone"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000276,"raw_usage":{"total_tokens":1638,"prompt_tokens":927,"completion_tokens":711,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":543,"completion_tokens_details":{"reasoning_tokens":640}},"tokens_in":543,"tokens_out":711,"duration_ms":6786,"temperature":1.0,"reasoning_tokens":640,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T11:41:45.277998+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the protocol with the Table 3 mapping replaced by coordinates from human raters or an independently published norming study, and compare mean cosine similarities. If the values fall to the 0.061 baseline or below across models, the reported emotional control was an artifact of the mapping rather than a property of the LLMs.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the circumplex model of affect whose two axes (arousal and valence) define every target emotional state in the experiment."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The HuggingFace sentiment model (sentiment-model-sample-27go-emotion) used to evaluate the emotion inferred from each generated text."}],"review_version":1}