{"id":"f9ed25b6-e6c6-4b5c-9b00-13be12cc5532","arxiv_id":"2501.00581","paper_version":2,"verdict":"REJECT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":6,"one_line_summary":"A dependency graph of 17 values learned from two LLMs predicts side effects of role and SAE steering, but the causal and human-alignment claims are unsupported.","lead":"Researchers mined questionnaire responses from two LLMs to build graphs of how value dimensions change together, then used those graphs to steer the models with role-playing prompts and sparse autoencoder features. The authors claim this exposes a latent causal structure that differs from human values, but no human data are collected and the causal interpretation rests on unverified assumptions.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"PC-based causal discovery conflates value-level causality with the effects of role and SAE steering: because role and feature are unobserved common causes of many value dimensions, the recovered graph is not identified as a latent causal value graph.","rationale":"The reader's weakest assumption identifies the load-bearing flaw: PC on orientation vectors from role and SAE interventions cannot separate true causal edges among values from correlations induced by the interventions acting as common causes. I agree, and would sharpen the point: this is not merely an untested assumption but a violation of PC's causal-sufficiency requirement, because the manipulated role and SAE feature are unobserved variables external to the 17 value nodes that jointly affect many of them. The held-out role/SAE split is a predictive check for the same intervention family, not a test of causal structure. The paper's SAE steering results and thought/answer consistency checks are useful engineering evidence, but they do not establish the central causal claim. The explicit-value-instruction experiment in Section 3.3 could have served as a validation of individual edges, yet it is not used for that purpose. My reading therefore leaves the reader's REJECT verdict unchanged.","tokens_in":17253,"tokens_out":5263,"duration_ms":59829,"concrete_test":"Use the explicit value instruction prompts from Appendix A.2 as genuine single-node interventions: for each value v, apply the positive and negative instruction templates on held-out roles while holding the role prompt fixed, and measure the full 17-dimensional orientation vector. For every directed edge v->w in the discovered graph, test whether w changes significantly more often than a random non-successor node, using a paired permutation test across roles. If the graph's edges are causal, this must replicate; if successors do not move under explicit single-value steering, the PC edges are artifacts of the role/SAE common causes rather than causal value relations.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3.2 applies PC to 2,525 orientation vectors generated by varying 101 role prompts and 25 SAE features across 17 value dimensions. PC is only guaranteed to recover a DAG under causal sufficiency, Markov, faithfulness, and i.i.d. sampling from a fixed distribution. Here each row is produced by deliberately applying a role r and an SAE feature f, neither of which is among the 17 value nodes. Those two manipulated variables are therefore unobserved common causes of many value dimensions at once: a role like 'Energy Manager, ESTP' or a +100 activation on a 'your values' feature can shift several values simultaneously. Conditional dependencies among value orientations are then explainable by variation in r and f rather than by causal edges among values. The held-out role/SAE split in Section 3.2.1 does not repair this, because the same confounding mechanism operates on the test rows. Consequently, the graph may be a useful summary of steering co-variation, but it is not evidence for a latent causal value graph. The human-comparison claim is also under-supported, since the reference graph is generated by GPT-4o rather than from human questionnaire data.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper argues that a latent causal value graph underlies the value dimensions of LLMs, that this structure remains largely unchanged by alignment training, and that it differs from human value systems as described by Schwartz's theory. The graph is learned by applying the PC algorithm to 2,525 orientation vectors obtained from 101 role prompts and 25 SAE steering features across 17 value dimensions (Section 3.2). The learned graph is then used to predict which value dimensions should change when another value is steered, comparing against reference graphs generated by GPT-4o, by the LLMs themselves, and by ValueBench's hierarchical structure (Sections 2.3 and 3.2.1). The paper also claims that SAE steering is more targeted than role-based prompting, based on the number of source nodes affected and on steering consistency metrics (Section 3.3).","tokens_in":17515,"tokens_out":4288,"duration_ms":45745,"significance":"If the causal interpretation were justified, the paper would offer a useful tool for explainable and controllable value steering in LLMs, and its comparison with human value theories would be an important empirical finding. The authors are to be commended for using held-out roles and questions in the evaluation, for testing multiple reference graphs, and for providing detailed tables of steering effects. However, the manuscript's central claim of a latent causal value graph is not supported by the methods: the data generation process violates the assumptions of the causal discovery algorithm, the evaluation metric is not an external validity check, and the 'human value system' comparison is based on a language model's output rather than human data. The practical contribution of a steering co-variation graph remains plausible, but the paper as written does not establish its causal or human-comparison conclusions.","major_comments":[{"comment":"The PC algorithm is applied to data whose rows are generated by deliberately varying role prompts and SAE features, which are not included among the 17 value nodes. Under PC's assumption of causal sufficiency, the absence of these two manipulated variables means that they act as unobserved common causes of many value dimensions simultaneously. The resulting graph is therefore not identified as a causal graph among values; it can at best be interpreted as a dependency summary of steering-induced co-variation. The held-out role split in Section 3.2.1 does not repair this, because the same confounding mechanism operates on the test rows.","section":"Section 3.2, PC application on 2,525 orientation vectors"},{"comment":"The evaluation metric is not an external validity check. The 'expected' successors are defined as successors in the learned graph, and c(v', v) measures the frequency with which v' changes when v changes under the same kind of role and SAE steering interventions used to build the graph. Hence the metric is essentially a train/test evaluation of dependency co-variation, not a test of whether the graph captures causal effects. The higher prediction numbers compared with the GPT-4o reference graph are expected, because the learned graph is fit to the same type of data that defines the evaluation.","section":"Section 2.3, Eq. (1) and Section 3.2.1"},{"comment":"The claim that the LLM value structure differs from human value systems is not supported by the evidence. The 'reference graph' representing human values is generated by GPT-4o from a prompt based on Schwartz's theory, rather than from human questionnaire data. The abstract and Remark 1 assert a difference from human value systems; the data only support a difference from one language model's paraphrase of Schwartz's theory. No human subjects data are collected or analyzed anywhere in the manuscript.","section":"Section 3.2 and Appendix A.3"},{"comment":"The description of the procedure as 'passive causal discovery' is misleading and the required assumptions are unmet. PC requires i.i.d. samples from a fixed distribution satisfying Markov, faithfulness, and causal sufficiency. Here the rows are produced by active interventions (role changes and SAE feature amplifications at a chosen layer and strength), so the samples are not i.i.d. from a fixed distribution. The manuscript neither verifies nor discusses any of these assumptions, nor does it report sensitivity of the recovered graph to the PC significance level alpha=0.05, the choice of SAE layer, the amplification magnitude, or the number of selected features. Without such evidence, the causal claim is not identifiable from these data.","section":"Section 2.1 and Section 3.2"}],"minor_comments":[{"comment":"The caption refers to 'Gemma-3B-IT' while the text and Table 1 refer to 'Gemma-2B-IT'; this inconsistency should be corrected.","section":"Section 3.2.1, Figure 4 caption"},{"comment":"The notation sr0[v] and the baseline role r0 are introduced informally; the definition should state explicitly that r0 is a fixed baseline role and that sr0[v] is its orientation on value v.","section":"Section 2.3"},{"comment":"The text and figure captions repeatedly use 'casual graph' where 'causal graph' is intended.","section":"Figures 4, 6, and 7"},{"comment":"The explanation of the data sampling is confusing: it says 40% of training data are randomly sampled for each role-SAE dyad, but the relationship between this sampling and the 70%/30% question split is not made clear.","section":"Section 3.1"},{"comment":"The Limitations section mentions the breadth of ValueBench and the scale of models, but it does not acknowledge the fundamental limitations of the causal discovery step, namely the unobserved confounders introduced by role and SAE steering and the absence of human data for the value comparison.","section":"Limitations"}],"recommendation":"reject","confidential_remarks":"The central contribution is the causal value graph, but the identification argument is not there. The paper could perhaps be revised into a study of steering-induced dependency structure, but that would be a substantially different contribution and would require removing or heavily qualifying the causal and human-comparison claims. I would advise the editor that the current causal framing is likely to mislead readers and that the paper's useful practical observations about SAE steering precision should not be used to mask the unsupported core claim."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe useful part is the steering recipe; the causal graph claim does not hold up. Mining a graph from questionnaire orientation vectors and using it to coordinate role-prompt and SAE steering is genuinely new as far as I know. The held-out role split is a reasonable check, and the SAE steering consistency numbers in Table 1 look decent. If all they claimed was \"a graph of steering co-variation helps predict side effects,\" I would call this a solid empirical paper.\n\nThe problem is the \"latent causal value graph\" framing. PC is run on 2,525 orientation vectors produced by varying 101 roles and 25 SAE features. Role and SAE feature are unobserved common causes of many value dimensions at once. PC requires causal sufficiency; with unobserved common causes, the edges are not identified as causal. The graph may be a useful dependency summary, but it is not evidence of a latent causal structure among values. The held-out split does not repair this, because the same confounding mechanism operates on the test rows.\n\nThe \"different from human values\" claim is also under-supported. The reference graph is generated by GPT-4o, not from human questionnaire responses. That is a semantic comparison, not an empirical one. There are no human data in the paper.\n\nLesser issues: no code or data shipped, no error bars, and the prediction metric in Section 2.3 defines expected changes as successors in the same graph, so it is not an external validity check. The authors acknowledge limitations, but they do not address the confounding.\n\nThe paper deserves a serious referee because the application is novel and the practical steering results may be useful. But the causal language needs to be softened, and the human comparison needs actual human data. Single-value interventions (e.g., explicit value instruction for one value and then measuring others) would give a cleaner test. As it stands, the central causal claim is not established.\n\nRecommendation: send to peer review, but expect major revision. It is not a desk reject.","headline":"A useful recipe for predicting steering side-effects wrapped in an unverified causal claim; the graph is a dependency summary, not evidence of latent value causality.","tokens_in":18048,"tokens_out":2446,"would_cite":false,"duration_ms":24214,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims a latent causal value graph organizes LLM responses, that alignment training leaves it structurally different from human value theory, and that the graph predicts and limits steering side effects.","keywords":["causal value graph","value alignment","sparse autoencoder steering","role-based prompting","causal discovery","Peter-Clark algorithm","LLM value structure","Schwartz value theory"],"falsifier":"Take a value pair the graph marks with an edge and one it marks as non-adjacent; steer only the cause node, using a single SAE feature at a single token position whose training profile shows it moves that one value, and record on new roles whether the predicted successor flips while the non-successor stays fixed. If the prediction advantage over the reference graphs disappears when interventions are restricted to single nodes, the discovered edges are artifacts of the multi-node interventions.","tokens_in":17061,"feed_emoji":"🧭","tokens_out":9867,"duration_ms":87744,"temperature":0.7,"pith_summary":"This paper tries to establish that the value-related behavior of a large language model is organized by a latent causal graph — directed links showing which value dimensions, such as achievement, social cynicism, or uncertainty avoidance, push which others. The authors mine the graph by collecting value-orientation vectors from two kinds of intervention, assigning the model 101 different personas and amplifying 25 sparse-autoencoder features, and feeding the resulting 2,525 vectors to the Peter–Clark causal-discovery algorithm. They find this recovered structure predicts the side effects of steering better than reference graphs built from Schwartz's theory of human values, on held-out roles and questions, for both Gemma-2B-IT and Llama3-8B-IT. From that, they conclude that alignment training has not made LLM value systems match human ones, and that knowing the model's own graph makes value control more predictable and precise.","feed_headline":"Causal maps reveal LLM value wiring diverges from human theory","feed_subtitle":"From 2,525 role and SAE interventions, the graphs beat Schwartz-style references at predicting steer side effects.","key_machinery":"The load-bearing object is the value causal graph $G = (V, E)$ over 17 ValueBench value dimensions, constructed with the Peter–Clark (PC) algorithm, a standard passive causal-discovery procedure, run at significance level 0.05 on 2,525 orientation vectors (101 roles times 25 SAE features). Each vector records the model's average ternary score on the value questionnaire under that particular persona or feature-amplification setting. The graph carries the whole argument: its successor and non-successor relations define the prediction metrics, the comparison against Schwartz-style reference graphs, and the claimed precision advantage of SAE steering, which is quantified as the number of source nodes each intervention family touches.","core_discovery":"The central claim is that LLMs possess a latent causal value graph, a directed structure over value dimensions in which a change to one value propagates along edges to others, and that this graph remains significantly different from the structure of human values described by theories such as Schwartz's quasi-circumplex model. The evidence is predictive: when a role-prompt or SAE intervention changes the model's orientation on a target value, dimensions the learned graph marks as successors change more often than dimensions marked as successors by the reference graphs (0.57–0.74 versus 0.43–0.51 across the two models), while non-successor dimensions change less often. The authors further claim that SAE steering is the more surgical instrument, because it activates only about four source nodes in the graph on average, compared with 7.7–14.6 for persona prompts, which is why it produces fewer unexpected side effects. The reason to accept a graph at all, on this account, is that it predicts what a tester can actually observe after steering.","pith_inferences":["Because the PC step treats the role and SAE interventions as passive observations, its edges could in principle reflect the steering methods acting as common causes on several values at once; a decisive check the paper does not report is whether graphs refit from role-prompt data alone and from SAE data alone agree on edge directions.","If the graph is stable across models of different sizes, it could serve as a transferable wiring diagram for predicting fine-tuning side effects or for ordering alignment curricula along the model's own causal directions instead of a human ordering.","The ternary response coding discards the size of each orientation change, so graded scores might reveal weaker causal edges that the binarized metric misses and could narrow the measured gap between LLM and human value structures.","Since SAE features are layer-specific, steering the same value through different layers and checking whether successor activation patterns change would test whether the causal graph is a property of the whole model or of the readout layer."],"forward_implications":["An operator can consult the graph before steering: it enumerates which value dimensions will shift as side effects, so a target value can be pursued while minimizing collateral change.","SAE feature steering becomes a precision tool for value control, since it perturbs far fewer causal source nodes than persona-based prompting and therefore triggers fewer unexpected changes.","Alignment that targets coarse values such as helpfulness, harmlessness, and honesty leaves an internal value structure that differs from human value theory, so effective fine-grained alignment should be guided by the model's own discovered graph rather than by human taxonomies alone.","The prediction metrics give a quantitative way to compare an LLM's value organization against alternatives — human theory, the model's own self-report, or ValueBench's hierarchy — and in every tested comparison the discovered graph predicts observed steering effects better."],"supporting_citations":[{"why":"Supplies the Peter–Clark algorithm that turns the 2,525 orientation vectors into the causal value graph.","marker":"Spirtes et al., 2001"},{"why":"Provides ValueBench, the questionnaire and 17 value dimensions from which all orientation scores are computed.","marker":"Ren et al., 2024"},{"why":"Supplies the human value structure that guides the reference graphs the learned graph is compared against.","marker":"Schwartz and Boehnke, 2004"},{"why":"Provides the pretrained sparse autoencoders whose selected features are amplified to steer values.","marker":"Bloom and Chanin, 2024"},{"why":"The Gemma-2B-IT model on which half of the steering experiments and graph mining are run.","marker":"Team et al., 2024"},{"why":"The Llama3-8B-IT model that provides the second, independent test bed for the graphs.","marker":"Dubey et al., 2024"},{"why":"The RLHF paradigm whose coarse-grained value coverage motivates the need for fine-grained causal value control.","marker":"Ouyang et al., 2022"},{"why":"Grounds the manipulation of individual sparse-autoencoder features as a mechanism for steering model behavior.","marker":"Cunningham et al., 2023"}],"fun_headline_variants":["Causal graphs show LLM values diverge from humans","LLM value graphs don't match human theory","Causal value graph reveals LLM-human value gap","SAE steering highlights LLM-human value divergence","LLM values structurally differ from humans, causal study finds"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the Peter–Clark algorithm applied to the 2,525 role-and-SAE orientation vectors recovers genuine causal edges among value dimensions rather than correlations created by the intervention families acting as common causes on several values at once; if that fails, the graph describes the steering machinery, not the model's values.","fun_headline_variants_meta":{"raw":{"variants":["Causal graphs show LLM values diverge from humans","LLM value graphs don't match human theory","Causal value graph reveals LLM-human value gap","SAE steering highlights LLM-human value divergence","LLM values structurally differ from humans, causal study finds"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001384,"raw_usage":{"total_tokens":5586,"prompt_tokens":913,"completion_tokens":4673,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":529,"completion_tokens_details":{"reasoning_tokens":4597}},"tokens_in":529,"tokens_out":4673,"duration_ms":30940,"temperature":1.0,"reasoning_tokens":4597,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T22:47:08.208776+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a value pair the graph marks with an edge and one it marks as non-adjacent; steer only the cause node, using a single SAE feature at a single token position whose training profile shows it moves that one value, and record on new roles whether the predicted successor flips while the non-successor stays fixed. If the prediction advantage over the reference graphs disappears when interventions are restricted to single nodes, the discovered edges are artifacts of the multi-node interventions.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the human value structure that guides the reference graphs the learned graph is compared against."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the pretrained sparse autoencoders whose selected features are amplified to steer values."}],"review_version":1}