{"total":12,"items":[{"citing_arxiv_id":"2607.07916","ref_index":18,"ref_count":1,"confidence":0.88,"is_internal_anchor":false,"paper_title":"Persona Cartography: Charting Language Model Personality Traits in Weight Space","primary_cat":"cs.AI","submitted_at":"2026-07-08T21:00:44+00:00","verdict":"CONDITIONAL","verdict_confidence":"HIGH","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Composable LoRA adapters can amplify or suppress OCEAN traits in LLMs, combine approximately additively, preserve moderate-scale capability, and move safety-relevant behaviours.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.25380","ref_index":101,"ref_count":1,"confidence":0.88,"is_internal_anchor":false,"paper_title":"A Survey of Toxicity Detection and Mitigation Strategies for Multilingual Language Models","primary_cat":"cs.CL","submitted_at":"2026-06-24T04:24:30+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":1.0,"formal_verification":"none","one_line_summary":"A survey that catalogs threat models, detection approaches, and mitigation strategies for toxicity in multilingual LLMs while identifying challenges such as uneven language coverage and culturally variable harm definitions.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.11316","ref_index":36,"ref_count":1,"confidence":0.88,"is_internal_anchor":false,"paper_title":"Sch\\\"utzen: Evaluating LLM Safety in Bulgarian and German Contexts","primary_cat":"cs.CL","submitted_at":"2026-06-09T18:01:19+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Schützen is a German-Bulgarian LLM safety dataset showing pronounced cross-language differences in model safety behavior.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.30169","ref_index":38,"ref_count":1,"confidence":0.88,"is_internal_anchor":false,"paper_title":"Dissociative Identity: Language Model Agents Lack Grounding for Reputation Mechanisms","primary_cat":"cs.CY","submitted_at":"2026-05-28T16:20:19+00:00","verdict":"CONDITIONAL","verdict_confidence":"MODERATE","novelty_score":7.0,"formal_verification":"none","one_line_summary":"VLMs preserve linearly separable visual magnitudes and can compare them, yet collapse at symbolic mapping because visual and textual number spaces remain fractured and disjoint.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.27025","ref_index":5,"ref_count":1,"confidence":0.88,"is_internal_anchor":false,"paper_title":"Attribute-Based Diagnosis of LLM Alignment with Hate Speech Annotations","primary_cat":"cs.CL","submitted_at":"2026-05-26T13:44:48+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"LLMs show split alignment with human hate speech annotations (strong on explicit attributes, inverted on evaluative ones), and attribute-based ridge regression reconstructs continuous scores with R² up to 0.71.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.23069","ref_index":15,"ref_count":1,"confidence":0.88,"is_internal_anchor":false,"paper_title":"DFKI-MLT at SemEval-2026 TASK 7: Steering Multilingual Models Towards Cultural Knowledge","primary_cat":"cs.CL","submitted_at":"2026-05-21T21:58:20+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":3.0,"formal_verification":"none","one_line_summary":"Activation steering with FLORES-derived language vectors produces modest, layer-sensitive and language-dependent gains on cultural awareness tasks, with some settings degrading performance and strong interaction with prompt design.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.17694","ref_index":92,"ref_count":1,"confidence":0.88,"is_internal_anchor":false,"paper_title":"Do LLM Agents Mirror Socio-Cognitive Effects in Power-Asymmetric Conversations?","primary_cat":"cs.CL","submitted_at":"2026-05-17T23:23:45+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"LLMs assigned high or low status personas in multi-turn dialogues exhibit socio-cognitive effects including language coordination, pronoun patterns, persuasion success, and compliance with unsafe requests.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.11789","ref_index":4,"ref_count":1,"confidence":0.88,"is_internal_anchor":false,"paper_title":"Beyond Inefficiency: Systemic Costs of Incivility in Multi-Agent Monte Carlo Simulations","primary_cat":"cs.AI","submitted_at":"2026-05-12T08:54:03+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"Monte Carlo simulations of LLM agents confirm that toxic debates take 25% longer to converge, with larger delays in smaller models, and show a first-mover advantage independent of toxicity.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.10843","ref_index":10,"ref_count":1,"confidence":0.88,"is_internal_anchor":false,"paper_title":"Training-Free Cultural Alignment of Large Language Models via Persona Disagreement","primary_cat":"cs.CL","submitted_at":"2026-05-11T16:55:16+00:00","verdict":"CONDITIONAL","verdict_confidence":"MODERATE","novelty_score":6.0,"formal_verification":"none","one_line_summary":"DISCA converts within-country disagreement among World Values Survey personas into a bounded logit correction that reduces cultural misalignment by 10-24% on MultiTP for models 3.8B and larger across 20 countries, without any weight updates.","context_count":1,"top_context_role":"background","top_context_polarity":"background","context_text":"CAReDiO: Cultural alignment of LLM via represen- tativeness and distinctiveness guided data optimization.arXiv preprint arXiv:2504.08820, 2025. doi: 10.48550/arXiv.2504.08820. URLhttps://arxiv.org/abs/2504.08820. 12 A. Zewail, A. Figueroa, J. Graham, and M. Atari. Moral stereotyping in large language models. Proceedings of the National Academy of Sciences, 123(10):e2519941123, 2026. doi: 10.1073/ pnas.2519941123. URLhttps://www.pnas.org/doi/10.1073/pnas.2519941123. B. Zhang, X. Zhao, J. Li, H. Chen, and Z. Chen. Mind the gap in cultural alignment: Task-aware culture management for large language models.arXiv preprint arXiv:2602.22475, 2026. A. Zou, L. Phan, S. Chen, J. Campbell, P. Guo, et al. Representation engineering: A top-down"},{"citing_arxiv_id":"2605.07462","ref_index":109,"ref_count":1,"confidence":0.88,"is_internal_anchor":false,"paper_title":"The Moltbook Files: A Harmless Slopocalypse or Humanity's Last Experiment","primary_cat":"cs.CL","submitted_at":"2026-05-08T09:10:17+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"An AI-agent social platform generated mostly neutral content whose use in fine-tuning reduced model truthfulness comparably to human Reddit data, suggesting limited unique harm but flagging tail risks like secret leaks.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2604.02585","ref_index":6,"ref_count":1,"confidence":0.88,"is_internal_anchor":false,"paper_title":"Mitigating LLM biases toward spurious social contexts using direct preference optimization","primary_cat":"cs.AI","submitted_at":"2026-04-02T23:42:20+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Debiasing-DPO reduces bias to spurious social contexts by 84% and improves predictive accuracy by 52% on average for LLMs evaluating U.S. classroom transcripts.","context_count":1,"top_context_role":"background","top_context_polarity":"background","context_text":"Evaluating large language models in analysing classroom dialogue.npj Science of Learning, 9(1), October 2024. ISSN 2056-7936. doi: 10.1038/s41539-024-00273-3. URLhttp://dx.doi.org/10.1038/s41539-024-00273-3. 12 Preprint. Pedro Henrique Luz de Araujo and Benjamin Roth. Helpful assistant or fruitful facilitator? investigating how personas affect language model behavior.PLOS One, 20(6):e0325664, June 2025. ISSN 1932-6203. doi: 10.1371/journal.pone.0325664. URL http://dx.doi.org/ 10.1371/journal.pone.0325664. Lars Malmqvist. Sycophancy in large language models: Causes and mitigations, 2024. URL https://arxiv.org/abs/2411.15287. Baptiste Moreau-Pernet, Yu Tian, Sandra Sawaya, Peter Foltz, Jie Cao, Brent Milne, and Thomas Christie."},{"citing_arxiv_id":"2408.12935","ref_index":174,"ref_count":1,"confidence":0.88,"is_internal_anchor":false,"paper_title":"AI Safety Landscape for Large Language Models: Taxonomy, State-of-the-art, and Future Directions","primary_cat":"cs.AI","submitted_at":"2024-08-23T09:33:48+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":4.0,"formal_verification":"none","one_line_summary":"The paper introduces a taxonomy of AI safety for LLMs organized into Trustworthy AI, Responsible AI, and Safe AI perspectives, accompanied by a review of state-of-the-art methods, challenges, and future directions.","context_count":1,"top_context_role":"background","top_context_polarity":"background","context_text":"an encoder-decoder design, where the encoder and decoder modules consist of a stack of transformer blocks, each ACM Comput. Surv. AI Safety Landscape for Large Language Models: Taxonomy, State-of-the-art, and Future Directions 5 comprising a Multi-Head Attention layer and a feedforward layer, connected by layer normalization [31] and residual connection modules [296]. In practice, the architectures are implemented to be encoder-only [174, 297, 439], decoder- only [585, 586] and encoder-decoder [ 391, 588] models, depending on their use cases. These Transformer-based Pre-trained Language Models (PLMs) have been applied in a wide range of downstream tasks, such as information retrieval [111, 112, 784], question answering [479, 804], and text generation [124, 316], often achieving state-of-the-art"}],"limit":50,"offset":0}