Pith. sign in

REVIEW 4 major objections 5 minor 32 references

CALMA: A Process for Deriving Context-aligned Axes for Language Model Alignment

T0 review · 4 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read A new four-stage process derives language-model alignment values from community interaction rather than researcher labels.

desk verdict CALMA is a genuinely new participatory process for deriving alignment axes, but the pilot evidence doesn't support the paper's headline claims about novel benchmarks and reduced researcher bias. read the letter →

arxiv 2507.09060 v2 pith:QPFXRXYY submitted 2025-07-11 cs.CY

classification cs.CY
keywords CALMAalignmentaxesparticipatorygroundedtheorylanguagemodelevaluationcontext-specificvaluesvaluepluralismhumanfeedback
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Alignment and evaluation of language models currently rely on axes such as helpfulness and harmlessness that researchers define in advance, often from non-representative samples. This paper claims that these researcher-defined axes miss the values of the communities where models are actually deployed, and proposes a four-stage process called CALMA to elicit axes directly from a community's own interaction with the model. Participants first get oriented, then freely converse with the model, then code their own conversations for implicit values using grounded-theory-style open coding, and finally discuss as a group to define a ranked set of axes. In a pilot with two different communities, the process produced axes such as 'Schools of Thought', 'Fact/Power', and 'Localization/Geographic Breadth' that are absent from standard benchmarks. The authors argue this reduces researcher bias and offers a path toward pluralistic, context-sensitive alignment pipelines.

What carries the argument

The machinery is the four-stage CALMA cycle. Familiarize orients participants to the deployment context and to grounded-theory coding using examples from an unrelated context to avoid priming. Interact lets participants hold open-ended conversations with the model, generating the interaction space to be annotated. Reflect applies two coding passes to those interactions: initial coding that labels observed values, and focused coding that groups labels into attribute clusters. Discuss brings participants together to define, refine, and rank the final set of axes, using artifacts like an embedding-space map of their labels. The load-bearing idea is that participant-generated interactions plus open coding plus structured dialogue can replace researcher-defined labels with locally grounded axes.

What would settle it

A controlled replication would run CALMA twice with matched groups from the same community under identical training and facilitation; if the final consensus axes diverge substantially across runs, the process is not reliable. A second decisive test is to run the full process and a variant that skips the open-ended interaction stage, with the same training and discussion format; if the axes are nearly identical, then the interaction step contributes nothing beyond what participants already know.

Watch

Extended reading notes

Core claim

In the paper's own terms, the central discovery is that context-specific alignment axes can be elicited from observed human interactions through a grounded, participatory, non-prescriptive process, instead of being imposed by researchers. The pilot showed that different participants interpret the same model response very differently, and that group dialogue can consolidate these interpretations into defensible attribute clusters. The derived axes reflected each community's particular historical and cultural position, such as the Indian group's emphasis on colonial and religious bias and the US student group's emphasis on cultural framing and empathy. The paper presents this as evidence that open-ended, use-case-driven evaluation surfaces priorities that standard benchmarks omit.

Load-bearing premise

The load-bearing premise is that participants can perform grounded-theory coding competently enough for their labels to be meaningful; the paper itself notes that initially inadequately trained participants produced generic labels and no inter-coder reliability is reported.

Editorial extensions

If this is right

  • The axes and example interactions produced by CALMA can be turned into training and preference datasets for context-specific alignment techniques such as RLHF and SteerLM.
  • CALMA outputs can supply annotation guidelines for evaluating model outputs within a specific deployment context, and can support context-specific LLM-as-judge setups.
  • Models aligned on CALMA-derived axes can be A/B tested within the same community against models trained on traditional datasets to measure whether the grounded axes change user satisfaction.
  • For scalable deployment, the paper suggests training lightweight classifiers on participant labels and using few-shot prompts with user definitions, enabling semi-automated dataset expansion.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A natural extension not explored in the paper: if CALMA is run with the same community but different facilitation styles (for example, more or less moderator involvement in Discuss), the stability of the final axes could be tested, since the paper's own pilot found group size and format affected consensus.
  • The paper's distinction between which axes matter and where along an axis a community falls implies that future pluralistic alignment pipelines may need to separate axis discovery from preference aggregation, treating the two as different participatory tasks.
  • If CALMA is scaled by training classifiers on participant labels, one risk is that the classifier inherits the same researcher bias the process avoids; a testable variant would compare classifier-clustered axes with full dialogue-derived axes to see if clustering preserves the novel priorities.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper introduces CALMA (Context-aligned Axes for Language Model Alignment), a four-stage participatory methodology (Familiarize, Interact, Reflect, Discuss) grounded in grounded theory, intended to elicit context-specific axes for LLM evaluation and alignment from observed human interactions. The authors report a pilot with two groups (MIT students and Indian professionals) for a history-educational-assistant use case, producing axes such as Cultural Context, Source, Fact/Power, Complexity, and Completeness. The paper claims that CALMA 'surfaced novel priorities that are absent from standard benchmarks' (Abstract) and that it 'reduces researcher bias' (Section 7), and it draws qualitative lessons about open-ended interaction, dialogic consensus-building, and context-specific definitions.

Significance. If the headline claims were supported, CALMA would be a valuable contribution to participatory and pluralistic alignment: it is transparent, open-ended, and grounds axis derivation in participant interactions rather than in researcher-defined labels. The process description is detailed, the pilot materials are reported, and the qualitative excerpts illustrate meaningful pluralism in annotation. The authors also deserve credit for explicitly acknowledging limitations (Section 6, Appendix A.4). However, the current evidence is a small, non-representative pilot with no benchmark comparison and no measurement of researcher bias, so the paper's strong empirical claims outrun what the data can support. The contribution is best understood as a promising process proposal with an illustrative demonstration, not as a validated method for surfacing novel, community-representative values.

major comments (4)
  1. [Abstract and Section 7] The claim that CALMA 'surfaced novel priorities that are absent from standard benchmarks' is not supported by the data. The pilot produces axes such as Cultural Context, Source, Fact/Power, Complexity, and Completeness (Tables 2-4), but the paper never compares these axes against any existing benchmark taxonomy. Several of these axes plausibly overlap with established evaluation categories (e.g., truthfulness, citation quality, helpfulness, and empathy), so the novelty claim is untested. The authors should either provide an explicit mapping to one or more standard benchmarks or taxonomies, or revise the claim to say that the process surfaces priorities that are not named in the specific pilot's predefined prompt.
  2. [Section 7 and Appendix A.2-A.3] The statement that CALMA 'reduces researcher bias' is asserted without any comparison condition or measurement. Researcher decisions remain embedded in the system prompt (Appendix A.2), the training materials, the researcher-constructed embedding space used in the Discuss stage (Appendix A.3, Segment Two), and group facilitation choices. To support the claim, the paper would need a bias measurement (e.g., comparing axes derived with versus without specific researcher choices) or a more modest formulation such as 'CALMA is designed to reduce researcher bias,' accompanied by qualitative evidence for how each design choice mitigates a specific bias pathway.
  3. [Appendix A.4] Appendix A.4 states that 'the groups were not sampled to be representative or act as a set of experts representing a given community, so the final attributes cannot be ascribed as indicative of any community's preferences.' This disclaimer directly conflicts with the abstract's implication that CALMA elicits context-specific community values. The paper should frame the study explicitly as an illustrative pilot and temper the abstract and conclusion accordingly, rather than presenting the derived axes as community-relevant outputs.
  4. [Section 6 and Appendix A.4] The Reflect stage's output is load-bearing for all downstream claims, but the paper reports no inter-coder reliability or other validation of participant coding quality. Section 6 concedes that inadequate training initially produced generic labels ('summary', 'introduction', 'conclusion') and that the training had to be updated, and Appendix A.4 notes over-representation of male-identifying participants and format differences between groups. Because the final axes inherit any unreliability in participant coding, the paper should either report coding reliability statistics or further soften the central claims, presenting the axes as illustrative of the process rather than as validated outputs.
minor comments (5)
  1. [Section 3] The phrase 'Figure 1 and details in Appendix' is incomplete; it should reference a specific appendix subsection.
  2. [Appendix A.2 and A.3] There is a missing space in 'the design of theInteract phase', and the heading 'Group Discussion Design: Interact' should presumably be 'Discuss' to match the stage name used in Section 3.
  3. [Table 4] The row labeled 'Cultural Context' in Table 4 has a definition that matches 'Schools of Thought' in Table 2; this appears to be a labeling error that should be corrected.
  4. [References and Section 6] The project name is spelled inconsistently as 'WikiBench' in Section 2.2 and 'Wikibench' in Section 6; the paper should standardize the spelling.
  5. [Section 4.1] The sentence 'they curated the remainder of their data through unique, self-directed interactions' could be clarified to specify whether participants selected or generated these interactions.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the pilot's axes are direct participant outputs, and the headline claims are under-supported but not circular.

full rationale

CALMA is a qualitative, participatory process paper; it contains no fitted parameters, no formal derivation, and no equation whose output is an input by construction. The axes in Tables 2-4 are direct products of participant coding and group discussion, not quantities predicted from a model fit to the same data. The paper's two headline claims, namely that CALMA surfaced novel priorities absent from standard benchmarks and that CALMA reduces researcher bias, are empirical generalizations that the pilot does not adequately support: no benchmark mapping is provided and no bias baseline is measured. However, unsupported empirical claims are a correctness and evidence concern, not circularity. The only self-citations (Casper et al. 2025 and Price 2022) are background and tool provenance, and neither is load-bearing for the derivation. The appendix explicitly disclaims representativeness in Section A.4, stating that the final attributes cannot be ascribed as indicative of any community's preferences, which undercuts the abstract's community-values framing but does not make the process circular. The Limitations section also concedes that inadequate training initially produced generic labels, but this again is a reliability threat rather than a circular step. The method's outputs are what they claim to be: participant-derived labels. The weakness is lack of external validation, not circular reasoning.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

No numerical free parameters are fit to data; design choices such as group size, model choice, system prompt, and discussion format are methodological decisions rather than fitted quantities. The axioms listed are the load-bearing assumptions about participant competence, representativeness, and researcher neutrality that the central claims depend on. No invented physical or formal entities are introduced; the alignment axes are participant-generated labels, not falsifiable entities with independent evidence.

assumptions (5)
  • domain assumption Grounded Theory coding, as taught in a short Familiarize stage, produces valid and reliable value labels from lay participants.
    The Reflect stage depends on participants applying initial and focused coding correctly. Section 6 admits that inadequate training initially produced generic labels ('summary', 'introduction', 'conclusion'), and no inter-coder reliability check is reported.
  • domain assumption Self-directed open-ended interactions with the LLM are representative of real deployment use of the educational assistant.
    Section 4.1 claims a 'realistic and diverse interaction space', but participants curate their own prompts with no constraint, and no deployment or log data is used for grounding.
  • ad hoc to paper The researcher-built embedding space and word cloud shown in the Discuss stage do not materially prime participants' final axes.
    Appendix A.3 describes researchers constructing the embedding space 'to present the group's collective interaction data' and a word cloud; the influence of these artifacts on consensus is unmeasured.
  • domain assumption Group consensus among 6 and 12 participants is sufficient to stand for context-specific priorities of a broader community.
    Appendix A.4 states the groups were not sampled to be representative and 'the final attributes cannot be ascribed as indicative of any community's preferences', which weakens any generalization from the pilot.
  • domain assumption Group discussions can surface plural views without re-imposing societal power hierarchies.
    Section 6 acknowledges that 'group settings can reproduce societal power dynamics, risking the exclusion of subaltern perspectives', yet the pilot's findings are interpreted as community-derived without measuring inclusion.

how reviews work

0 comments
Cite this review

Pith. "Pith review of CALMA: A Process for Deriving Context-aligned Axes for Language Model Alignment." pith.science (2026). https://pith.science/paper/QPFXRXYY

@misc{pith2026250709060,
  author       = {Pith},
  title        = {Pith review of: CALMA: A Process for Deriving Context-aligned Axes for Language Model Alignment},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/QPFXRXYY}},
  note         = {Machine review of arXiv:2507.09060}
}
read the original abstract

Datasets play a central role in AI governance by enabling both evaluation (measuring capabilities) and alignment (enforcing values) along axes such as helpfulness, harmlessness, toxicity, quality, and more. However, most alignment and evaluation datasets depend on researcher-defined or developer-defined axes curated from non-representative samples. As a result, developers typically benchmark models against broad (often Western-centric) values that overlook the varied contexts of their real-world deployment. Consequently, models trained on such proxies can fail to meet the needs and expectations of diverse user communities within these deployment contexts. To bridge this gap, we introduce CALMA (Context-aligned Axes for Language Model Alignment), a grounded, participatory methodology for eliciting context-relevant axes for evaluation and alignment. In a pilot with two distinct communities, CALMA surfaced novel priorities that are absent from standard benchmarks. Our findings demonstrate the value of evaluation practices based on open-ended and use-case-driven processes. Our work advances the development of pluralistic, transparent, and context-sensitive alignment pipelines.

Figures

Figures reproduced from arXiv: 2507.09060 by the authors.

Figure 1
Figure 1. The CALMA process A.2. CALMA Tool Details The CALMA methodology was tested by using a platform adapted from the Open Coding for Machine Learning tool1 (Price, 2022). The language model interface for the Interact step was built using Llama-70B & Mistral-8x7B, where the following system prompt was used to contextualize the model responses: “You are an education assistant for high school students studying history in In… view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

32 extracted references · 16 canonical work pages

  1. [1]

    Constitutional ai: Harmlessness from ai feedback

    Bai, Y., Kadavath, S., Kundu, S., Askell, A., Kernion, J., Jones, A., Chen, A., Goldie, A., Mirhoseini, A., McKinnon, C., et al. Constitutional ai: Harmlessness from ai feedback. arXiv preprint arXiv:2212.08073, 2022

  2. [2]

    Fine-tuning language models to find agreement among humans with diverse preferences

    Bakker, M., Chadwick, M., Sheahan, H., Tessler, M., Campbell-Gillingham, L., Balaguer, J., McAleese, N., Glaese, A., Aslanides, J., Botvinick, M., et al. Fine-tuning language models to find agreement among humans with diverse preferences. Advances in Neural Information Processing Systems, 35: 0 38176--38189, 2022

  3. [3]

    Stela: a community-centred approach to norm elicitation for ai alignment

    Bergman, S., Marchal, N., Mellor, J., Mohamed, S., Gabriel, I., and Isaac, W. Stela: a community-centred approach to norm elicitation for ai alignment. Scientific Reports, 14 0 (1): 0 6616, 2024

  4. [4]

    Pitfalls of evidence-based ai policy

    Casper, S., Krueger, D., and Hadfield-Menell, D. Pitfalls of evidence-based ai policy. arXiv preprint arXiv:2502.09618, 2025

  5. [5]

    C., Amershi, S., and Kamar, E

    Chang, J. C., Amershi, S., and Kamar, E. Revolt: Collaborative crowdsourcing for labeling machine learning datasets. In Proceedings of the 2017 CHI conference on human factors in computing systems, pp.\ 2334--2346, 2017

  6. [6]

    Constructionism and the grounded theory method

    Charmaz, K. Constructionism and the grounded theory method. Handbook of constructionist research, 1 0 (1): 0 397--412, 2008

  7. [7]

    B., and Weld, D

    Chen, Q., Bragg, J., Chilton, L. B., and Weld, D. S. Cicero: Multi-turn, contextual argumentation for accurate crowdsourcing. In Proceedings of the 2019 chi conference on human factors in computing systems, pp.\ 1--14, 2019

  8. [8]

    N., Wu, X., and Kuchaiev, O

    Dong, Y., Wang, Z., Sreedhar, M. N., Wu, X., and Kuchaiev, O. Steerlm: Attribute conditioned sft as an (user-steerable) alternative to rlhf. arXiv preprint arXiv:2310.05344, 2023

Show all 32 references
  1. [9]

    Kto: Model alignment as prospect theoretic optimization

    Ethayarajh, K., Xu, W., Muennighoff, N., Jurafsky, D., and Kiela, D. Kto: Model alignment as prospect theoretic optimization. arXiv preprint arXiv:2402.01306, 2024

  2. [10]

    o lz, P., Parkes, D. C., Procaccia, A. D., Rusak, G., Shapira, I., and W \

    Fish, S., G \"o lz, P., Parkes, D. C., Procaccia, A. D., Rusak, G., Shapira, I., and W \"u thrich, M. Generative social choice. arXiv preprint arXiv:2309.01291, 2023

  3. [11]

    Collective constitutional ai: Aligning a language model with public input, Oct 2023

    Ganguli, D., Huang, S., Lovitt, L., Siddharth, D., Liao, T., and Durmus, E. Collective constitutional ai: Aligning a language model with public input, Oct 2023. URL https://www.anthropic.com/news/collective-constitutional-ai-aligning-a-language-model-with-public-input

  4. [12]

    Improving alignment of dialogue agents via targeted human judgements

    Glaese, A., McAleese, N., Tr e bacz, M., Aslanides, J., Firoiu, V., Ewalds, T., Rauh, M., Weidinger, L., Chadwick, M., Thacker, P., et al. Improving alignment of dialogue agents via targeted human judgements. arXiv preprint arXiv:2209.14375, 2022

  5. [13]

    W., and On, K.-W

    Jung, S., Han, G., Nam, D. W., and On, K.-W. Binary classifier optimization for large language model alignment. arXiv preprint arXiv:2404.04656, 2024

  6. [14]

    M., Kirk, H

    Khandelwal, K., Tonneau, M., Bean, A. M., Kirk, H. R., and Hale, S. A. Indian-bhed: A dataset for measuring india-centric biases in large language models. In Proceedings of the 2024 International Conference on Information Technology for Social Good, pp.\ 231--239, 2024

  7. [15]

    M., and Seo, M

    Kim, S., Bae, S., Shin, J., Kang, S., Kwak, D., Yoo, K. M., and Seo, M. Aligning large language models through synthetic feedback. arXiv preprint arXiv:2305.13735, 2023

  8. [16]

    R., Whitefield, A., R \"o ttger, P., Bean, A., Margatina, K., Ciro, J., Mosquera, R., Bartolo, M., Williams, A., He, H., et al

    Kirk, H. R., Whitefield, A., R \"o ttger, P., Bean, A., Margatina, K., Ciro, J., Mosquera, R., Bartolo, M., Williams, A., He, H., et al. The prism alignment project: What participatory, representative and individualised human feedback reveals about the subjective and multicult...

  9. [17]

    Wikibench: Community-driven data curation for ai evaluation on wikipedia

    Kuo, T.-S., Halfaker, A., Cheng, Z., Kim, J., Wu, M.-H., Wu, T., Holstein, K., and Zhu, H. Wikibench: Community-driven data curation for ai evaluation on wikipedia. arXiv preprint arXiv:2402.14147, 2024

  10. [18]

    Second thoughts are best: Learning to re-align with human values from text edits

    Liu, R., Jia, C., Zhang, G., Zhuang, Z., Liu, T., and Vosoughi, S. Second thoughts are best: Learning to re-align with human values from text edits. Advances in Neural Information Processing Systems, 35: 0 181--196, 2022 a

  11. [19]

    Aligning generative language models with human values

    Liu, R., Zhang, G., Feng, X., and Vosoughi, S. Aligning generative language models with human values. In Findings of the Association for Computational Linguistics: NAACL 2022, pp.\ 241--252, 2022 b

  12. [20]

    I., Lau, N

    Meadows, G. I., Lau, N. W. L., Susanto, E. A., Yu, C. L., and Paul, A. Localvaluebench: A collaboratively built and extensible benchmark for evaluating localized value alignment and ethical safety in large language models. arXiv preprint arXiv:2408.01460, 2024

  13. [21]

    M., and Diaz, M

    Prabhakaran, V., Davani, A. M., and Diaz, M. On releasing annotator-level labels and information in datasets. arXiv preprint arXiv:2110.05699, 2021

  14. [22]

    Open Coding for Machine Learning

    Price, M. Open Coding for Machine Learning. PhD thesis, Massachusetts Institute of Technology, 2022

  15. [23]

    Mapping affinities: visualizing academic practice through collaboration

    Rodighiero, D. Mapping affinities: visualizing academic practice through collaboration. Technical report, EPFL, 2018

  16. [24]

    Large language model alignment: A survey

    Shen, T., Jin, R., Huang, Y., Liu, C., Dong, W., Guo, Z., Wu, X., Liu, Y., and Xiong, D. Large language model alignment: A survey. arXiv preprint arXiv:2309.15025, 2023

  17. [25]

    Polis: Scaling deliberation by mapping high dimensional opinion spaces

    Small, C., Bjorkegren, M., Erkkil \"a , T., Shaw, L., and Megill, C. Polis: Scaling deliberation by mapping high dimensional opinion spaces. Recerca: revista de pensament i an \`a lisi , 26 0 (2), 2021

  18. [26]

    and Dennison, C

    Solaiman, I. and Dennison, C. Process for adapting language models to society (palms) with values-targeted datasets. Advances in Neural Information Processing Systems, 34: 0 5861--5873, 2021

  19. [27]

    M., Ye, A., Jiang, L., Lu, X., Dziri, N., et al

    Sorensen, T., Moore, J., Fisher, J., Gordon, M., Mireshghallah, N., Rytting, C. M., Ye, A., Jiang, L., Lu, X., Dziri, N., et al. A roadmap to pluralistic alignment. arXiv preprint arXiv:2402.05070, 2024

  20. [28]

    Arithmetic control of llms for diverse user preferences: Directional preference alignment with multi-objective rewards

    Wang, H., Lin, Y., Xiong, W., Yang, R., Diao, S., Qiu, S., Zhao, H., and Zhang, T. Arithmetic control of llms for diverse user preferences: Directional preference alignment with multi-objective rewards. arXiv preprint arXiv:2402.18571, 2024

  21. [29]

    R., Everett, R., Huang, S., Zhu, T

    Weidinger, L., McKee, K. R., Everett, R., Huang, S., Zhu, T. O., Chadwick, M. J., Summerfield, C., and Gabriel, I. Using the veil of ignorance to align ai systems with principles of justice. Proceedings of the National Academy of Sciences, 120 0 (18): 0 e2213709120, 2023

  22. [30]

    and Klein, D

    Yang, K. and Klein, D. Fudge: Controlled text generation with future discriminators. arXiv preprint arXiv:2104.05218, 2021

  23. [31]

    P., Zhang, H., Gonzalez, J

    Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E. P., Zhang, H., Gonzalez, J. E., and Stoica, I. Judging llm-as-a-judge with mt-bench and chatbot arena, 2023. URL https://arxiv.org/abs/2306.05685

  24. [32]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.