REVIEW 4 major objections 5 minor 32 references
CALMA: A Process for Deriving Context-aligned Axes for Language Model Alignment
T0 review · 4 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read A new four-stage process derives language-model alignment values from community interaction rather than researcher labels.
desk verdict CALMA is a genuinely new participatory process for deriving alignment axes, but the pilot evidence doesn't support the paper's headline claims about novel benchmarks and reduced researcher bias. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery is the four-stage CALMA cycle. Familiarize orients participants to the deployment context and to grounded-theory coding using examples from an unrelated context to avoid priming. Interact lets participants hold open-ended conversations with the model, generating the interaction space to be annotated. Reflect applies two coding passes to those interactions: initial coding that labels observed values, and focused coding that groups labels into attribute clusters. Discuss brings participants together to define, refine, and rank the final set of axes, using artifacts like an embedding-space map of their labels. The load-bearing idea is that participant-generated interactions plus open coding plus structured dialogue can replace researcher-defined labels with locally grounded axes.
What would settle it
A controlled replication would run CALMA twice with matched groups from the same community under identical training and facilitation; if the final consensus axes diverge substantially across runs, the process is not reliable. A second decisive test is to run the full process and a variant that skips the open-ended interaction stage, with the same training and discussion format; if the axes are nearly identical, then the interaction step contributes nothing beyond what participants already know.
Extended reading notes
Core claim
In the paper's own terms, the central discovery is that context-specific alignment axes can be elicited from observed human interactions through a grounded, participatory, non-prescriptive process, instead of being imposed by researchers. The pilot showed that different participants interpret the same model response very differently, and that group dialogue can consolidate these interpretations into defensible attribute clusters. The derived axes reflected each community's particular historical and cultural position, such as the Indian group's emphasis on colonial and religious bias and the US student group's emphasis on cultural framing and empathy. The paper presents this as evidence that open-ended, use-case-driven evaluation surfaces priorities that standard benchmarks omit.
Load-bearing premise
The load-bearing premise is that participants can perform grounded-theory coding competently enough for their labels to be meaningful; the paper itself notes that initially inadequately trained participants produced generic labels and no inter-coder reliability is reported.
Editorial extensions
If this is right
- The axes and example interactions produced by CALMA can be turned into training and preference datasets for context-specific alignment techniques such as RLHF and SteerLM.
- CALMA outputs can supply annotation guidelines for evaluating model outputs within a specific deployment context, and can support context-specific LLM-as-judge setups.
- Models aligned on CALMA-derived axes can be A/B tested within the same community against models trained on traditional datasets to measure whether the grounded axes change user satisfaction.
- For scalable deployment, the paper suggests training lightweight classifiers on participant labels and using few-shot prompts with user definitions, enabling semi-automated dataset expansion.
Reading between the lines
- A natural extension not explored in the paper: if CALMA is run with the same community but different facilitation styles (for example, more or less moderator involvement in Discuss), the stability of the final axes could be tested, since the paper's own pilot found group size and format affected consensus.
- The paper's distinction between which axes matter and where along an axis a community falls implies that future pluralistic alignment pipelines may need to separate axis discovery from preference aggregation, treating the two as different participatory tasks.
- If CALMA is scaled by training classifiers on participant labels, one risk is that the classifier inherits the same researcher bias the process avoids; a testable variant would compare classifier-clustered axes with full dialogue-derived axes to see if clustering preserves the novel priorities.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces CALMA (Context-aligned Axes for Language Model Alignment), a four-stage participatory methodology (Familiarize, Interact, Reflect, Discuss) grounded in grounded theory, intended to elicit context-specific axes for LLM evaluation and alignment from observed human interactions. The authors report a pilot with two groups (MIT students and Indian professionals) for a history-educational-assistant use case, producing axes such as Cultural Context, Source, Fact/Power, Complexity, and Completeness. The paper claims that CALMA 'surfaced novel priorities that are absent from standard benchmarks' (Abstract) and that it 'reduces researcher bias' (Section 7), and it draws qualitative lessons about open-ended interaction, dialogic consensus-building, and context-specific definitions.
Significance. If the headline claims were supported, CALMA would be a valuable contribution to participatory and pluralistic alignment: it is transparent, open-ended, and grounds axis derivation in participant interactions rather than in researcher-defined labels. The process description is detailed, the pilot materials are reported, and the qualitative excerpts illustrate meaningful pluralism in annotation. The authors also deserve credit for explicitly acknowledging limitations (Section 6, Appendix A.4). However, the current evidence is a small, non-representative pilot with no benchmark comparison and no measurement of researcher bias, so the paper's strong empirical claims outrun what the data can support. The contribution is best understood as a promising process proposal with an illustrative demonstration, not as a validated method for surfacing novel, community-representative values.
major comments (4)
- [Abstract and Section 7] The claim that CALMA 'surfaced novel priorities that are absent from standard benchmarks' is not supported by the data. The pilot produces axes such as Cultural Context, Source, Fact/Power, Complexity, and Completeness (Tables 2-4), but the paper never compares these axes against any existing benchmark taxonomy. Several of these axes plausibly overlap with established evaluation categories (e.g., truthfulness, citation quality, helpfulness, and empathy), so the novelty claim is untested. The authors should either provide an explicit mapping to one or more standard benchmarks or taxonomies, or revise the claim to say that the process surfaces priorities that are not named in the specific pilot's predefined prompt.
- [Section 7 and Appendix A.2-A.3] The statement that CALMA 'reduces researcher bias' is asserted without any comparison condition or measurement. Researcher decisions remain embedded in the system prompt (Appendix A.2), the training materials, the researcher-constructed embedding space used in the Discuss stage (Appendix A.3, Segment Two), and group facilitation choices. To support the claim, the paper would need a bias measurement (e.g., comparing axes derived with versus without specific researcher choices) or a more modest formulation such as 'CALMA is designed to reduce researcher bias,' accompanied by qualitative evidence for how each design choice mitigates a specific bias pathway.
- [Appendix A.4] Appendix A.4 states that 'the groups were not sampled to be representative or act as a set of experts representing a given community, so the final attributes cannot be ascribed as indicative of any community's preferences.' This disclaimer directly conflicts with the abstract's implication that CALMA elicits context-specific community values. The paper should frame the study explicitly as an illustrative pilot and temper the abstract and conclusion accordingly, rather than presenting the derived axes as community-relevant outputs.
- [Section 6 and Appendix A.4] The Reflect stage's output is load-bearing for all downstream claims, but the paper reports no inter-coder reliability or other validation of participant coding quality. Section 6 concedes that inadequate training initially produced generic labels ('summary', 'introduction', 'conclusion') and that the training had to be updated, and Appendix A.4 notes over-representation of male-identifying participants and format differences between groups. Because the final axes inherit any unreliability in participant coding, the paper should either report coding reliability statistics or further soften the central claims, presenting the axes as illustrative of the process rather than as validated outputs.
minor comments (5)
- [Section 3] The phrase 'Figure 1 and details in Appendix' is incomplete; it should reference a specific appendix subsection.
- [Appendix A.2 and A.3] There is a missing space in 'the design of theInteract phase', and the heading 'Group Discussion Design: Interact' should presumably be 'Discuss' to match the stage name used in Section 3.
- [Table 4] The row labeled 'Cultural Context' in Table 4 has a definition that matches 'Schools of Thought' in Table 2; this appears to be a labeling error that should be corrected.
- [References and Section 6] The project name is spelled inconsistently as 'WikiBench' in Section 2.2 and 'Wikibench' in Section 6; the paper should standardize the spelling.
- [Section 4.1] The sentence 'they curated the remainder of their data through unique, self-directed interactions' could be clarified to specify whether participants selected or generated these interactions.
Circularity Check
No significant circularity: the pilot's axes are direct participant outputs, and the headline claims are under-supported but not circular.
full rationale
CALMA is a qualitative, participatory process paper; it contains no fitted parameters, no formal derivation, and no equation whose output is an input by construction. The axes in Tables 2-4 are direct products of participant coding and group discussion, not quantities predicted from a model fit to the same data. The paper's two headline claims, namely that CALMA surfaced novel priorities absent from standard benchmarks and that CALMA reduces researcher bias, are empirical generalizations that the pilot does not adequately support: no benchmark mapping is provided and no bias baseline is measured. However, unsupported empirical claims are a correctness and evidence concern, not circularity. The only self-citations (Casper et al. 2025 and Price 2022) are background and tool provenance, and neither is load-bearing for the derivation. The appendix explicitly disclaims representativeness in Section A.4, stating that the final attributes cannot be ascribed as indicative of any community's preferences, which undercuts the abstract's community-values framing but does not make the process circular. The Limitations section also concedes that inadequate training initially produced generic labels, but this again is a reliability threat rather than a circular step. The method's outputs are what they claim to be: participant-derived labels. The weakness is lack of external validation, not circular reasoning.
Assumptions & free parameters
assumptions (5)
- domain assumption Grounded Theory coding, as taught in a short Familiarize stage, produces valid and reliable value labels from lay participants.
- domain assumption Self-directed open-ended interactions with the LLM are representative of real deployment use of the educational assistant.
- ad hoc to paper The researcher-built embedding space and word cloud shown in the Discuss stage do not materially prime participants' final axes.
- domain assumption Group consensus among 6 and 12 participants is sufficient to stand for context-specific priorities of a broader community.
- domain assumption Group discussions can surface plural views without re-imposing societal power hierarchies.
Cite this review
Pith. "Pith review of CALMA: A Process for Deriving Context-aligned Axes for Language Model Alignment." pith.science (2026). https://pith.science/paper/QPFXRXYY
@misc{pith2026250709060,
author = {Pith},
title = {Pith review of: CALMA: A Process for Deriving Context-aligned Axes for Language Model Alignment},
year = {2026},
howpublished = {\url{https://pith.science/paper/QPFXRXYY}},
note = {Machine review of arXiv:2507.09060}
}
read the original abstract
Datasets play a central role in AI governance by enabling both evaluation (measuring capabilities) and alignment (enforcing values) along axes such as helpfulness, harmlessness, toxicity, quality, and more. However, most alignment and evaluation datasets depend on researcher-defined or developer-defined axes curated from non-representative samples. As a result, developers typically benchmark models against broad (often Western-centric) values that overlook the varied contexts of their real-world deployment. Consequently, models trained on such proxies can fail to meet the needs and expectations of diverse user communities within these deployment contexts. To bridge this gap, we introduce CALMA (Context-aligned Axes for Language Model Alignment), a grounded, participatory methodology for eliciting context-relevant axes for evaluation and alignment. In a pilot with two distinct communities, CALMA surfaced novel priorities that are absent from standard benchmarks. Our findings demonstrate the value of evaluation practices based on open-ended and use-case-driven processes. Our work advances the development of pluralistic, transparent, and context-sensitive alignment pipelines.
Figures
Reference graph
Works this paper leans on
-
[1]
Constitutional ai: Harmlessness from ai feedback
Bai, Y., Kadavath, S., Kundu, S., Askell, A., Kernion, J., Jones, A., Chen, A., Goldie, A., Mirhoseini, A., McKinnon, C., et al. Constitutional ai: Harmlessness from ai feedback. arXiv preprint arXiv:2212.08073, 2022
arXiv 2022
-
[2]
Fine-tuning language models to find agreement among humans with diverse preferences
Bakker, M., Chadwick, M., Sheahan, H., Tessler, M., Campbell-Gillingham, L., Balaguer, J., McAleese, N., Glaese, A., Aslanides, J., Botvinick, M., et al. Fine-tuning language models to find agreement among humans with diverse preferences. Advances in Neural Information Processing Systems, 35: 0 38176--38189, 2022
work page 2022
-
[3]
Stela: a community-centred approach to norm elicitation for ai alignment
Bergman, S., Marchal, N., Mellor, J., Mohamed, S., Gabriel, I., and Isaac, W. Stela: a community-centred approach to norm elicitation for ai alignment. Scientific Reports, 14 0 (1): 0 6616, 2024
work page 2024
-
[4]
Pitfalls of evidence-based ai policy
Casper, S., Krueger, D., and Hadfield-Menell, D. Pitfalls of evidence-based ai policy. arXiv preprint arXiv:2502.09618, 2025
arXiv 2025
-
[5]
Chang, J. C., Amershi, S., and Kamar, E. Revolt: Collaborative crowdsourcing for labeling machine learning datasets. In Proceedings of the 2017 CHI conference on human factors in computing systems, pp.\ 2334--2346, 2017
work page 2017
-
[6]
Constructionism and the grounded theory method
Charmaz, K. Constructionism and the grounded theory method. Handbook of constructionist research, 1 0 (1): 0 397--412, 2008
work page 2008
-
[7]
Chen, Q., Bragg, J., Chilton, L. B., and Weld, D. S. Cicero: Multi-turn, contextual argumentation for accurate crowdsourcing. In Proceedings of the 2019 chi conference on human factors in computing systems, pp.\ 1--14, 2019
work page 2019
-
[8]
Dong, Y., Wang, Z., Sreedhar, M. N., Wu, X., and Kuchaiev, O. Steerlm: Attribute conditioned sft as an (user-steerable) alternative to rlhf. arXiv preprint arXiv:2310.05344, 2023
arXiv 2023
Show all 32 references
-
[9]
Kto: Model alignment as prospect theoretic optimization
Ethayarajh, K., Xu, W., Muennighoff, N., Jurafsky, D., and Kiela, D. Kto: Model alignment as prospect theoretic optimization. arXiv preprint arXiv:2402.01306, 2024
2024 arXiv
-
[10]
o lz, P., Parkes, D. C., Procaccia, A. D., Rusak, G., Shapira, I., and W \
Fish, S., G \"o lz, P., Parkes, D. C., Procaccia, A. D., Rusak, G., Shapira, I., and W \"u thrich, M. Generative social choice. arXiv preprint arXiv:2309.01291, 2023
2023 arXiv
-
[11]
Collective constitutional ai: Aligning a language model with public input, Oct 2023
Ganguli, D., Huang, S., Lovitt, L., Siddharth, D., Liao, T., and Durmus, E. Collective constitutional ai: Aligning a language model with public input, Oct 2023. URL https://www.anthropic.com/news/collective-constitutional-ai-aligning-a-language-model-with-public-input
2023
-
[12]
Improving alignment of dialogue agents via targeted human judgements
Glaese, A., McAleese, N., Tr e bacz, M., Aslanides, J., Firoiu, V., Ewalds, T., Rauh, M., Weidinger, L., Chadwick, M., Thacker, P., et al. Improving alignment of dialogue agents via targeted human judgements. arXiv preprint arXiv:2209.14375, 2022
2022 arXiv
-
[13]
W., and On, K.-W
Jung, S., Han, G., Nam, D. W., and On, K.-W. Binary classifier optimization for large language model alignment. arXiv preprint arXiv:2404.04656, 2024
2024 arXiv
-
[14]
M., Kirk, H
Khandelwal, K., Tonneau, M., Bean, A. M., Kirk, H. R., and Hale, S. A. Indian-bhed: A dataset for measuring india-centric biases in large language models. In Proceedings of the 2024 International Conference on Information Technology for Social Good, pp.\ 231--239, 2024
2024
-
[15]
M., and Seo, M
Kim, S., Bae, S., Shin, J., Kang, S., Kwak, D., Yoo, K. M., and Seo, M. Aligning large language models through synthetic feedback. arXiv preprint arXiv:2305.13735, 2023
2023 arXiv
-
[16]
R., Whitefield, A., R \"o ttger, P., Bean, A., Margatina, K., Ciro, J., Mosquera, R., Bartolo, M., Williams, A., He, H., et al
Kirk, H. R., Whitefield, A., R \"o ttger, P., Bean, A., Margatina, K., Ciro, J., Mosquera, R., Bartolo, M., Williams, A., He, H., et al. The prism alignment project: What participatory, representative and individualised human feedback reveals about the subjective and multicult...
2024 arXiv
-
[17]
Wikibench: Community-driven data curation for ai evaluation on wikipedia
Kuo, T.-S., Halfaker, A., Cheng, Z., Kim, J., Wu, M.-H., Wu, T., Holstein, K., and Zhu, H. Wikibench: Community-driven data curation for ai evaluation on wikipedia. arXiv preprint arXiv:2402.14147, 2024
2024 arXiv
-
[18]
Second thoughts are best: Learning to re-align with human values from text edits
Liu, R., Jia, C., Zhang, G., Zhuang, Z., Liu, T., and Vosoughi, S. Second thoughts are best: Learning to re-align with human values from text edits. Advances in Neural Information Processing Systems, 35: 0 181--196, 2022 a
2022
-
[19]
Aligning generative language models with human values
Liu, R., Zhang, G., Feng, X., and Vosoughi, S. Aligning generative language models with human values. In Findings of the Association for Computational Linguistics: NAACL 2022, pp.\ 241--252, 2022 b
2022
-
[20]
I., Lau, N
Meadows, G. I., Lau, N. W. L., Susanto, E. A., Yu, C. L., and Paul, A. Localvaluebench: A collaboratively built and extensible benchmark for evaluating localized value alignment and ethical safety in large language models. arXiv preprint arXiv:2408.01460, 2024
2024 arXiv
-
[21]
M., and Diaz, M
Prabhakaran, V., Davani, A. M., and Diaz, M. On releasing annotator-level labels and information in datasets. arXiv preprint arXiv:2110.05699, 2021
2021 arXiv
-
[22]
Open Coding for Machine Learning
Price, M. Open Coding for Machine Learning. PhD thesis, Massachusetts Institute of Technology, 2022
2022
-
[23]
Mapping affinities: visualizing academic practice through collaboration
Rodighiero, D. Mapping affinities: visualizing academic practice through collaboration. Technical report, EPFL, 2018
2018
-
[24]
Large language model alignment: A survey
Shen, T., Jin, R., Huang, Y., Liu, C., Dong, W., Guo, Z., Wu, X., Liu, Y., and Xiong, D. Large language model alignment: A survey. arXiv preprint arXiv:2309.15025, 2023
2023 arXiv
-
[25]
Polis: Scaling deliberation by mapping high dimensional opinion spaces
Small, C., Bjorkegren, M., Erkkil \"a , T., Shaw, L., and Megill, C. Polis: Scaling deliberation by mapping high dimensional opinion spaces. Recerca: revista de pensament i an \`a lisi , 26 0 (2), 2021
2021
-
[26]
and Dennison, C
Solaiman, I. and Dennison, C. Process for adapting language models to society (palms) with values-targeted datasets. Advances in Neural Information Processing Systems, 34: 0 5861--5873, 2021
2021
-
[27]
M., Ye, A., Jiang, L., Lu, X., Dziri, N., et al
Sorensen, T., Moore, J., Fisher, J., Gordon, M., Mireshghallah, N., Rytting, C. M., Ye, A., Jiang, L., Lu, X., Dziri, N., et al. A roadmap to pluralistic alignment. arXiv preprint arXiv:2402.05070, 2024
2024 arXiv
-
[28]
Arithmetic control of llms for diverse user preferences: Directional preference alignment with multi-objective rewards
Wang, H., Lin, Y., Xiong, W., Yang, R., Diao, S., Qiu, S., Zhao, H., and Zhang, T. Arithmetic control of llms for diverse user preferences: Directional preference alignment with multi-objective rewards. arXiv preprint arXiv:2402.18571, 2024
2024 arXiv
-
[29]
R., Everett, R., Huang, S., Zhu, T
Weidinger, L., McKee, K. R., Everett, R., Huang, S., Zhu, T. O., Chadwick, M. J., Summerfield, C., and Gabriel, I. Using the veil of ignorance to align ai systems with principles of justice. Proceedings of the National Academy of Sciences, 120 0 (18): 0 e2213709120, 2023
2023
-
[30]
and Klein, D
Yang, K. and Klein, D. Fudge: Controlled text generation with future discriminators. arXiv preprint arXiv:2104.05218, 2021
2021 arXiv
-
[31]
P., Zhang, H., Gonzalez, J
Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E. P., Zhang, H., Gonzalez, J. E., and Stoica, I. Judging llm-as-a-judge with mt-bench and chatbot arena, 2023. URL https://arxiv.org/abs/2306.05685
2023 arXiv
-
[32]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.