Pith. sign in

REVIEW 3 major objections 5 minor 26 references

A gentle push funziona benissimo: making instructed models in Italian via contrastive activation steering

T0 review · 3 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read A 30-prompt activation push matches fine-tuning for steering an LLM into Italian.

desk verdict Useful application of steering to Italian, but the headline comparison to fine-tuning is undercut by a narrow answer-extraction regex that appears to penalize the fine-tuned baseline. read the letter →

arxiv 2411.18247 v1 pith:LZGN5IEI submitted 2024-11-27 cs.CL cs.LG

classification cs.CLcs.LG
keywords activationsteeringcontrastiveadditionItalianlanguageadaptationinstructiontuningalternativeinference-timeinterventionbenchmarkingconsistencycatastrophicforgetting
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper claims that activation steering, adding a vector to a model's internal activations during generation, can adapt an English-instructed LLM to Italian as effectively as instruction fine-tuning, but with fewer than 100 prompts and no parameter updates. The authors test Italian steering on Llama 3 8B, Phi 3 3.8B, and Llama 2 7B, comparing against fine-tuned Italian models. They report that the steered Llama 3 matches or beats the fine-tuned ANITA on Italian MMLU, HellaSwag, and ARC, while generating more consistently Italian text. If this holds, steering is a dramatically cheaper route to language adaptation, especially when only machine-translated training data is available.

What carries the argument

The machinery is the Italian steering vector, defined as the difference between averaged activations over Italian responses and averaged activations over English responses, collected from the last token of each attention head across K = 30 Alpaca-style prompts. A variant called ITA uses English questions with Italian answers, aiming to capture the language-switch direction. At generation time, the vector is added to every layer-head activation with a multiplier alpha that starts at 1.5 and linearly decays to 0 over the generated tokens, so the push is strongest early and fades as the model settles into Italian.

What would settle it

Re-score the fine-tuned ANITA baseline by response likelihood instead of the paper's regex extractor: if ANITA's Italian benchmark scores rise to or above the steered models' scores, the claim of comparable-or-better performance would be undermined.

Watch

Extended reading notes

Core claim

The central claim is that a gentle push, a contrastive activation steering vector computed from 30 instruction prompts, is enough to make an English-instructed LLM answer Italian benchmarks in Italian, at a level comparable to or better than a model fine-tuned on roughly 240,000 Italian instruction examples. On Llama 3 8B, the ITA steering variant scores 55.95 on Italian MMLU versus 55.01 for ANITA; 50.00 versus 42.49 on HellaSwag; and 71.38 versus 72.54 on ARC, with Italian language detection 0.996 versus 0.715. The authors emphasize that steering pushes the model's language direction without teaching it new facts, so it preserves most of the original model's correct answers, unlike fine-tuning, which loses some. The paper does not present this as a full replacement for fine-tuning where new knowledge must be injected, but as a cheaper and often better alternative when the goal is purely language adaptation.

Load-bearing premise

The load-bearing premise is that a direction computed from 30 contrastive prompts isolates a general, transferable Italianness in the model's activation space, and that the base model already contains enough latent Italian knowledge for the push to make a difference.

Editorial extensions

If this is right

  • Steering with under 100 contrastive prompts can substitute for instruction fine-tuning when the goal is language adaptation, cutting data requirements by more than 99 percent.
  • Because steering does not update weights, a steered model keeps the original model's correct answers while gaining Italian fluency, avoiding the capability loss observed in fine-tuned ANITA.
  • The method generalizes across model families (Llama 3, Phi 3, Llama 2), so it is a model-agnostic recipe rather than a one-off hack.
  • When the only available fine-tuning data is machine-translated, steering is the more effective use of that data, since it only needs a handful of prompts rather than hundreds of thousands.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because steering only reweights existing latent behavior, the method's success should depend on how much Italian the base model saw in pre-training; for languages with negligible pre-training exposure, instruction fine-tuning or additional data would remain necessary.
  • The extraction uses 30 fixed prompts; varying the prompt set and measuring the variance of downstream scores would test how robust the Italian direction is, a diagnostic the paper does not run.
  • If the steering vector captures a general language direction, the same recipe might transfer to other languages or to dialect, register, or style control, and possibly be composed with other steering vectors, though the paper only tests Italian.
  • The regex-based answer extraction likely penalizes ANITA more than the steered models, so a likelihood-based re-scoring could change the ranking on individual benchmarks.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The manuscript investigates contrastive activation steering as a low-resource alternative to instruction tuning for adapting English-centric LLMs to Italian. The authors extract a steering vector from the mean activation difference between English and Italian (or Italian-answer-to-English-question) Alpaca prompts, then add this direction to all attention-head outputs during generation, with a diminishing intensity. They evaluate on Llama 3 8B Instruct, Phi 3 mini Instruct, and Llama 2 7B Instruct on Italian MMLU, HellaSwag, ARC, and a language-detection metric, comparing against the fine-tuned Italian models ANITA and LLaMAntino 2. Their headline finding is that steering with 30 prompts and no training performs comparably to or better than fine-tuning on several benchmarks; for example, ITA steering on Llama 3 scores 55.95 vs 55.01 (ANITA) on MMLU, 50.00 vs 42.49 on HellaSwag, and 71.38 vs 72.54 on ARC, while also improving Italian-language consistency (0.996 vs 0.715). On Llama 2 ARC, steering ITA-full raises accuracy from 32.84 to 41.06, exceeding the fine-tuned model's 34.98.

Significance. The paper addresses a practical problem—language adaptation without expensive fine-tuning—and the proposed method is clearly described and easy to reproduce from the text, including the contrastive dataset construction, K=30, the steering intensity, and the evaluation protocol. The reported gain on Llama 2 ARC is striking and, if real, would be an important demonstration that activation steering can transfer to unseen reasoning benchmarks. At the same time, the central comparison against fine-tuning is weakened by two issues: a regex-based answer extractor that appears to penalize ANITA's response format, and a complete absence of uncertainty estimates. Substantial revision is therefore needed to determine whether the headline claim survives a format-agnostic evaluation.

major comments (3)
  1. [§4.2, Table 2, Appendix B] The regex evaluation in Appendix B requires a separator (':' or 'e’') immediately before the answer letter. Table 2 contains ANITA answers marked correct that would not match this pattern: 'A\n(mixed Thai...', 'B \n(Ela explicação)...', 'C. Gli atomi rimangono gli stessi.', and 'D (un prato erboso...'. Since the steered models in the same table tend to produce answers like 'La risposta corretta è (B)...', which do match, the scoring is asymmetric and may under-report ANITA's accuracy on HellaSwag and other benchmarks. Because the central claim of comparability with fine-tuning rests on these numbers, please re-evaluate all models with a format-independent answer extractor (e.g., first-letter match after normalizing punctuation) or report the fraction of responses that fail the regex for each model, and update the tables accordingly.
  2. [§4.2, Tables 1 and 3] All reported scores are point estimates from single deterministic runs with no error bars, confidence intervals, or significance tests. For ARC, which contains roughly 1,000 items, the difference between ITA (71.38) and ANITA (72.54) is well within one standard error of the estimate; the MMLU gap of 0.94 points is similarly not interpretable without variance information. To support the claim that steering is 'comparable to, or even better than' fine-tuning, the authors should provide bootstrap confidence intervals over the benchmark items (or an equivalent uncertainty quantification).
  3. [§3, Eq. (3), 'Steering vector extraction'] The method relies on two user-set hyperparameters, K (number of contrastive prompts) and the steering intensity valmax, without any sensitivity analysis. Since the paper's main premise is that a very small number of examples (30) suffices to produce a transferable language direction, the authors should demonstrate that the reported outcomes do not sharply depend on these choices. I request a sensitivity study that varies K (e.g., 10, 30, 60) and valmax (e.g., 0.5–3.0) for at least one model/benchmark pair and reports the performance curve.
minor comments (5)
  1. [Appendix A] The section heading 'Promtps' contains a typo; it should be 'Prompts'.
  2. [Appendix B] The two regex patterns are presented without a clear statement of how they are combined (e.g., alternatives or sequential), and without a fallback rule for responses that match neither; please specify the exact scoring function, including what is scored when no match is found.
  3. [Section 5 (and Abstract)] The claim that steering uses 'less than 0.5%' of the fine-tuning data is imprecise: 30 examples is 0.0125% of 240K; giving the exact count and percentage would be more informative.
  4. [§4.2, Generation quality] Tables 7 and 8 contain only a few hand-picked generations; the abstract's claim of 'higher quality' generations is not supported by a systematic human evaluation or a quantitative metric beyond lang-detect. Please either soften the claim or provide stronger evidence.
  5. [Reproducibility] Please state whether the code, the extracted steering vectors, and the exact benchmark prompts will be made publicly available; this is especially relevant for a method whose value lies in being cheap and fast.

Circularity Check

1 steps flagged · score 2.0 of 10

No load-bearing circularity: benchmark gains are independent of the steering-vector construction; only the lang-detect result is a by-construction sanity check.

  1. self definitional [Section 3 (Steering vector extraction) and Section 4.2 (Generation quality, Table 1)]
    "The aim is to emphasize the language switch task, pushing the model to respond in Italian even to an English prompt. ... Significant improvements are seen in the language itself, where the steering techniques are effective in yielding Italian output."

    The steering vector is defined as the averaged activation difference between Italian-answer and English-answer versions of the same Alpaca prompts (Eq. 1; Delta_ITA = a_ITA - a_ENG), i.e., it is literally a 'become more Italian' direction. The lang-detect column measures the probability that generated text is Italian. Reporting that steered models increase on lang-detect is therefore verifying the construction rather than testing an external prediction. The benchmark scores (MMLU, HellaSwag, ARC) are not part of the contrast set, so the headline 'comparable to fine-tuning' does not reduce to this step.

full rationale

The central derivation chain is not circular. The steering vector is extracted from 30 contrastive English/Italian Alpaca prompts, while the headline comparisons (MMLU, HellaSwag, ARC) are held-out benchmarks that play no role in the construction; no parameter is fitted to those benchmarks, and the alpha=1.5 diminishing schedule is inherited from prior work [4] rather than tuned on the test data. The self-citation [4] is methodological and not load-bearing for the claim that steering transfers to Italian benchmarks. One secondary step is by-construction: the lang-detect result measures exactly the 'Italianness' property that Delta_ITA = a_ITA - a_ENG is defined to inject, so it is a manipulation check rather than an independent prediction; it does not affect the benchmark conclusions. The Appendix B regex-based scorer and the lack of error bars are possible fairness and robustness limitations, but a scoring artifact is not a circularity because the benchmark numbers are still external measurements, not functions of the steering construction. Overall, the paper's main claim survives as an independent empirical comparison.

Assumptions & free parameters 4 free parameters · 5 assumptions · 0 invented entities

The central claim rests on four free parameters (valmax = 1.5, K = 30, unreported M, and the post-hoc ITA versus ITA-full choice), five domain assumptions about how language is represented and how the intervention behaves, and no invented entities. The largest unknown is whether the 30-prompt contrast extracts a transferable language direction, which the paper does not test with variance or ablations.

free parameters (4)
  • valmax (steering intensity) = 1.5
    Multiplier for the added steering vector, taken from prior work [4]; it regulates how strongly the Italian direction is injected and is not fit to the target benchmarks.
  • K (number of contrastive prompts) = 30
    Activations are averaged over K = 30 prompts (Section 3); the choice is made by hand and its sensitivity is not tested.
  • M (maximum generated tokens) = not reported
    Appears in Eq. 3 as the schedule length for the linearly decaying steering factor, but its value is never stated, so the decay schedule is not fully specified.
  • Steering variant selection (ITA vs ITA-full) = ITA chosen as primary
    The paper reports that ITA 'generally proves to be more effective' and uses it as the headline variant; choosing the better of two variants on the evaluation benchmarks is a post-hoc model selection.
assumptions (5)
  • domain assumption Linear representation hypothesis: high-level concepts such as language are directions in activation space
    Section 2, cited to Park et al. [10]; the entire steering vector extraction relies on this being true for the 'Italian' concept.
  • domain assumption The base models already contain sufficient latent Italian competence for steering to unlock it
    Section 3, first paragraph: 'the model already sees a small amount of the target language'; if false, steering cannot create Italian ability it merely amplifies.
  • domain assumption Contrastive activation addition is a safe intervention that shifts language without destroying task competence
    Section 3, injection equation (2); inherited from prior work [4], [5], [11] and assumed to hold for the benchmark tasks.
  • domain assumption Machine-translated Alpaca suffices to extract a clean Italian direction
    Appendix A discusses translation imperfections yet uses the translated Alpaca as the contrastive source; the assumption is that translation quality is adequate for direction extraction even if inadequate for fine-tuning.
  • domain assumption Greedy decoding with regex answer extraction measures correctness fairly across all compared models
    Appendix B replaces likelihood-based evaluation with a regex and greedy decoding; the fairness of this metric across differently formatted models is assumed.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A gentle push funziona benissimo: making instructed models in Italian via contrastive activation steering." pith.science (2026). https://pith.science/paper/LZGN5IEI

@misc{pith2026241118247,
  author       = {Pith},
  title        = {Pith review of: A gentle push funziona benissimo: making instructed models in Italian via contrastive activation steering},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/LZGN5IEI}},
  note         = {Machine review of arXiv:2411.18247}
}
read the original abstract

Adapting models to a language that was only partially present in the pre-training data requires fine-tuning, which is expensive in terms of both data and computational resources. As an alternative to fine-tuning, we explore the potential of activation steering-based techniques to enhance model performance on Italian tasks. Through our experiments we show that Italian steering (i) can be successfully applied to different models, (ii) achieves performances comparable to, or even better than, fine-tuned models for Italian, and (iii) yields higher quality and consistency in Italian generations. We also discuss the utility of steering and fine-tuning in the contemporary LLM landscape where models are anyway getting high Italian performances even if not explicitly trained in this language.

Figures

Figures reproduced from arXiv: 2411.18247 by the authors.

Figure 1
Figure 1. Graphical representation of all the correct answer combinations given by models on the ARC challenge. Each column shows a different combination of correct answers between all the different approaches with their respective cardinality8 (e.g. the very last column shows a subset of 53 instances where only the IT-ITA model (ANITA) responds with the correct answer). The steered and the IT-ITA models have limited overlap … view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

26 extracted references · 15 canonical work pages

  1. [4]

    Scalena, G

    D. Scalena, G. Sarti, M. Nissim, Multi-property steering of large language models with dynamic activation composition, in: Y. Belinkov, N. Kim, J. Jumelet, H. Mohebbi, A. Mueller, H. Chen (Eds.), Proceedings of the 7th BlackboxNLP Workshop: Analyzing and Interpreting Neural Networks for NLP, Association for Computational Linguistics, Miami, Florida, US, 2...

  2. [1]

    Taori, I

    R. Taori, I. Gulrajani, T. Zhang, Y. Dubois, X. Li, C. Guestrin, P. Liang, T. B. Hashimoto, Stanford al- paca: An instruction-following llama model, https: //github.com/tatsu-lab/stanford_alpaca, 2023

  3. [2]

    Santilli, E

    A. Santilli, E. Rodolà, Camoscio: an italian instruction-tuned llama, 2023. arXiv:2307.16456

  4. [3]

    Basile, E

    P. Basile, E. Musacchio, M. Polignano, L. Siciliani, G. Fiameni, G. Semeraro, Llamantino: Llama 2 mod- els for effective text generation in italian language,

  5. [5]

    Panickssery, N

    N. Panickssery, N. Gabrieli, J. Schulz, M. Tong, E. Hubinger, A. M. Turner, Steering llama 2 via contrastive activation addition, 2024. URL: https: //arxiv.org/abs/2312.06681. arXiv:2312.06681

  6. [6]

    Team, Introducing meta llama 3: The most ca- pable openly available llm to date, https://ai.meta

    M. Team, Introducing meta llama 3: The most ca- pable openly available llm to date, https://ai.meta. com/blog/meta-llama-3/, 2024

  7. [7]

    Team, Phi-3 technical report: A highly ca- pable language model locally on your phone,

    M. Team, Phi-3 technical report: A highly ca- pable language model locally on your phone,

  8. [8]

    Ahuja, H

    K. Ahuja, H. Diddee, R. Hada, M. Ochieng, K. Ramesh, P. Jain, A. Nambi, T. Ganu, S. Segal, M. Ahmed, K. Bali, S. Sitaram, MEGA: Multilin- gual evaluation of generative AI, in: H. Bouamor, J. Pino, K. Bali (Eds.), Proceedings of the 2023 Con- ference on Empirical Methods in Natural Language Processing, Association for Computational Linguis- tics, Singapore...

Show all 26 references
  1. [9]

    Polignano, P

    M. Polignano, P. Basile, G. Semeraro, Advanced natural-based interaction for the italian language: Llamantino-3-anita, 2024. arXiv:2405.07101

  2. [10]

    K. Park, Y. J. Choe, V. Veitch, The linear representa- tion hypothesis and the geometry of large language models, 2023. URL: https://arxiv.org/abs/2311.03658. arXiv:2311.03658

  3. [11]

    A. M. Turner, L. Thiergart, G. Leech, D. Udell, J. J. Vazquez, U. Mini, M. MacDiarmid, Activation addi- tion: Steering language models without optimiza- tion, 2024. URL: https://arxiv.org/abs/2308.10248. arXiv:2308.10248

  4. [12]

    Ilharco, M

    G. Ilharco, M. T. Ribeiro, M. Wortsman, S. Guru- rangan, L. Schmidt, H. Hajishirzi, A. Farhadi, Edit- ing models with task arithmetic, 2023. URL: https: //arxiv.org/abs/2212.04089. arXiv:2212.04089

  5. [13]

    Mikolov, K

    T. Mikolov, K. Chen, G. S. Corrado, J. Dean, Efficient estimation of word representations in vector space, in: International Conference on Learning Represen- tations, 2013. URL: https://api.semanticscholar.org/ CorpusID:5959482

  6. [14]

    NLP, Minerva llms, https://nlp.uniroma1.it/ minerva/, 2024

    S. NLP, Minerva llms, https://nlp.uniroma1.it/ minerva/, 2024

  7. [15]

    Hendrycks, C

    D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, J. Steinhardt, Measuring mas- sive multitask language understanding, Proceed- ings of the International Conference on Learning Representations (ICLR) (2021)

  8. [16]

    Zellers, A

    R. Zellers, A. Holtzman, Y. Bisk, A. Farhadi, Y. Choi, HellaSwag: Can a machine really finish your sen- tence?, in: A. Korhonen, D. Traum, L. Màrquez (Eds.), Proceedings of the 57th Annual Meeting of the Association for Computational Linguis- tics, Association for Computationa...

  9. [17]

    Clark, I

    P. Clark, I. Cowhey, O. Etzioni, T. Khot, A. Sab- harwal, C. Schoenick, O. Tafjord, Think you have solved question answering? try arc, the ai2 reason- ing challenge, 2018. URL: https://arxiv.org/abs/1803. 05457. arXiv:1803.05457

  10. [18]

    Biderman, H

    S. Biderman, H. Schoelkopf, L. Sutawika, L. Gao, J. Tow, B. Abbasi, A. F. Aji, P. S. Ammanamanchi, S. Black, J. Clive, A. DiPofi, J. Etxaniz, B. Fattori, J. Z. Forde, C. Foster, J. Hsu, M. Jaiswal, W. Y. Lee, H. Li, C. Lovering, N. Muennighoff, E. Pavlick, J. Phang, A. Skowron...

  11. [19]

    Kirkpatrick, R

    J. Kirkpatrick, R. Pascanu, N. Rabinowitz, J. Ve- ness, G. Desjardins, A. A. Rusu, K. Milan, J. Quan, T. Ramalho, A. Grabska-Barwinska, D. Hassabis, C. Clopath, D. Kumaran, R. Hadsell, Overcom- ing catastrophic forgetting in neural networks, Proceedings of the National Academy...

  12. [20]

    ( R i s p o s t a : | r i s p o s t a e ’ ) \ s ∗ \ ( ? ( [ ABCD ] ) \ ) ?

    Y. Du, D. Nguyen, Measuring the instability of fine- tuning, 2023. URL: https://arxiv.org/abs/2302.07778. arXiv:2302.07778. A. Promtps and instructions When extracting the behavior from the models, we em- ploy different versions of Alpaca. Examples of the three versions listed...

  13. [22]

    Ciò può aumentare il traffico sulla strada

    **Attrazione di traffico**: la creazione di una corsia in più può attirare più veicoli sulla strada, poiché gli automobilisti potrebbero essere più propensi a utilizzare la strada con più corsie. Ciò può aumentare il traffico sulla strada

  14. [23]

    Ecco alcune possibili motivazioni:

    **Cambiamenti di comportamento**: la creazione di una corsia in più può cambiare il comportamento degli automobilisti, che potrebbero essere più propensi a utilizzare la strada per raggiungere i loro obiettivi più rapidamente, aumentando il traffico.[...] ANITA [9] In generale...

  15. [24]

    La costruzione di una corsia in più in una strada non è sufficiente per ridurre il traffico in modo significativo

    **Attrazione di nuovo traffico**: la creazione di una corsia in più può attirare nuovi veicoli sulla strada, poiché la capacità di attraversamento aumenta, il che potrebbe annullare l’effetto di riduzione del traffico previsto.[...] ITA-full La risposta è no, non necessariamen...

  16. [25]

    La corsia in più non è necessariamente una misura per ridurre il traffico. La costruzione di una corsia in più può attirare più traffico nella zona (domanda indotta), poiché gli automobilisti coglieranno l’opportunità di poter usare la nuova corsia[...] ITA La risposta è no, n...

  17. [26]

    Original Il testo descrive il Monte Bianco, la montagna più alta delle Alpi, d’Italia, di Francia e d’Europa, con un’altitudine di 4805,59 metri. È noto anche come

    **Riduzione del traffico**: se la corsia in più è destinata ai mezzi pubblici o ai pedoni e ciclisti, può ridurre il traffico generale sulla strada,[...] System Rispondi alle domande dell’utente Prompt Riassumi il seguente testo: "Il Monte Bianco (Mont Blanc in francese e in a...

  18. [2024]

    arXiv:2404.14219

    URL: https://arxiv.org/abs/2404.14219. arXiv:2404.14219

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.