{"total":17,"items":[{"citing_arxiv_id":"2606.16723","ref_index":12,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"AgentFairBench: Do LLM Agents Discriminate When They Act?","primary_cat":"cs.AI","submitted_at":"2026-06-15T13:50:26+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"AgentFairBench is a multi-domain benchmark for demographic disparity in LLM agent actions, with a pilot showing no significant effect for Claude Haiku 4.5 after arity-matched noise correction.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.00334","ref_index":43,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Isolating LLM Lexical Bias: A Curation-Free Triangulated Metric for Preference-Stage Learning","primary_cat":"cs.CL","submitted_at":"2026-05-29T20:19:49+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Introduces a triangulation-based metric to quantify lexical shifts attributable to preference tuning without requiring manual curation of examples.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.00019","ref_index":9,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"LLMs in the Real World: Evaluating \"AI\" in Emergency Contexts","primary_cat":"cs.CY","submitted_at":"2026-05-29T19:27:25+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":2.0,"formal_verification":"none","one_line_summary":"AI researchers should take greater responsibility for publicly explaining the limitations of their technologies to prevent misuse in high-stakes applications such as emergency translation services.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.21919","ref_index":7,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"SDGBiasBench: Benchmarking and Mitigating Vision--Language Models' Biases in Sustainable Development Goals","primary_cat":"cs.CV","submitted_at":"2026-05-21T02:44:54+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"SDGBiasBench reveals intrinsic SDG biases in VLMs driven by priors rather than evidence, and CADE mitigates them with up to 25% accuracy gains and 12-point MAE reductions.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.10442","ref_index":16,"ref_count":2,"confidence":0.9,"is_internal_anchor":false,"paper_title":"StereoTales: A Multilingual Framework for Open-Ended Stereotype Discovery in LLMs","primary_cat":"cs.CY","submitted_at":"2026-05-11T12:12:28+00:00","verdict":"ACCEPT","verdict_confidence":"MODERATE","novelty_score":7.0,"formal_verification":"none","one_line_summary":"StereoTales shows that all tested LLMs emit harmful stereotypes in open-ended stories, with associations adapting to prompt language and targeting locally salient groups rather than transferring uniformly across languages.","context_count":1,"top_context_role":"background","top_context_polarity":"background","context_text":"The control of the false discovery rate in multiple testing under dependency.The Annals of Statistics, 29(4):1165-1188, 2001. ISSN 00905364, 21688966. URLhttp://www.jstor.org/stable/2674075. [15] W. Bergsma. A bias-correction for cramér's v and tschuprow's t.Journal of the Korean Statistical Society, 42(3):323-328, 2013. ISSN 1226-3192. doi: https://doi.org/10.1016/j.jkss. 2012.10.002. [16] S. L. Blodgett, S. Barocas, H. Daum'e, and H. M. Wallach. Language (technology) is power: A critical survey of \"bias\" in nlp.ArXiv, abs/2005.14050, 2020. URL https: //api.semanticscholar.org/CorpusID:218971825. [17] S. L. Blodgett, G. Lopez, A. Olteanu, R. Sim, and H. Wallach. Stereotyping norwegian salmon: An inventory of pitfalls in fairness benchmark datasets."},{"citing_arxiv_id":"2605.02640","ref_index":30,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Trustworthy AI Suffers from Invariance Conflicts and Causality is The Solution","primary_cat":"cs.AI","submitted_at":"2026-05-04T14:26:28+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":4.0,"formal_verification":"none","one_line_summary":"Causality resolves trade-offs in trustworthy AI by treating them as invariance conflicts under different data-generating process changes.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2604.18761","ref_index":2,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"From Tokens to Ties: Network and Discourse Analysis of Web3 Ecosystems","primary_cat":"cs.SI","submitted_at":"2026-04-20T19:11:06+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"Network and discourse analysis of NFT collections shows holding behavior builds dense, socially embedded Web3 communities with ongoing participation, unlike fragmented transactional networks from trading and speculation.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2604.05483","ref_index":3,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Can We Trust a Black-box LLM? LLM Untrustworthy Boundary Detection via Bias-Diffusion and Multi-Agent Reinforcement Learning","primary_cat":"cs.AI","submitted_at":"2026-04-07T06:24:01+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"GMRL-BD detects untrustworthy topic boundaries for black-box LLMs by combining bias-diffusion on a Wikipedia KG with multi-agent RL, supported by a released dataset labeling biases in models like Llama2 and Qwen2.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2604.04735","ref_index":1,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Lighting Up or Dimming Down? Exploring Dark Patterns of LLMs in Co-Creativity","primary_cat":"cs.CL","submitted_at":"2026-04-06T15:03:26+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"Sycophancy appears in 91.7% of LLM responses during co-creative writing tasks, especially on sensitive topics, while anchoring varies by literary form and is most common in folktales.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2506.06816","ref_index":15,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"How do datasets, developers, and models affect biases in a low-resourced language?: The Case of the Bengali Language","primary_cat":"cs.CL","submitted_at":"2025-06-07T14:46:35+00:00","verdict":"CONDITIONAL","verdict_confidence":"MODERATE","novelty_score":5.0,"formal_verification":"none","one_line_summary":"Bengali sentiment analysis models exhibit persistent identity-based biases across datasets and developer backgrounds despite similar semantic content.","context_count":1,"top_context_role":"background","top_context_polarity":"background","context_text":"gorithmic extrapolations, interpolations, and decisions [ 37, 48]. For example, studies have found that NLP tools are often unable to un- derstand racial, ethnic, and religious minorities' dialects [81] or clas- sify their linguistic practices as negative and abusive [ 37, 42, 109]. Researchers previously examined the biases of computational sys- tems across different social identity dimensions [ 15, 87], such as gender [72], race [ 109], nationality [ 134], religion [ 12], caste [ 6], age [ 44], occupation [ 132], disability [ 135], and political affilia- tions [1]. Such biases can be put into three categories [ 58]: preex- isting, technical, and emergent. Preexisting bias has its roots in social institutions, practices, and prejudicial attitudes, which can be reinforced in sociotechnical sys-"},{"citing_arxiv_id":"2408.09049","ref_index":7,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Inertia in Moral and Value Judgments of Large Language Models","primary_cat":"cs.CL","submitted_at":"2024-08-16T23:24:10+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":4.0,"formal_verification":"none","one_line_summary":"LLMs exhibit persistent inertia in value orientations, with harm avoidance and fairness remaining skewed across persona prompts.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2305.10403","ref_index":14,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"PaLM 2 Technical Report","primary_cat":"cs.CL","submitted_at":"2023-05-17T17:46:53+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"PaLM 2 reports state-of-the-art results on language, reasoning, and multilingual tasks with improved efficiency over PaLM.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2303.08774","ref_index":40,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"GPT-4 Technical Report","primary_cat":"cs.CL","submitted_at":"2023-03-15T17:15:04+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"GPT-4 is a scaled Transformer model with post-training alignment that reaches human-level performance on academic and professional benchmarks via infrastructure enabling performance prediction from much smaller models.","context_count":1,"top_context_role":"method","top_context_polarity":"use_method","context_text":"This report focuses on the capabilities, limitations, and safety properties of GPT-4. GPT-4 is a Transformer-style model [39] pre-trained to predict the next token in a document, using both publicly available data (such as internet data) and data licensed from third-party providers. The model was then fine-tuned using Reinforcement Learning from Human Feedback (RLHF) [ 40]. Given both the competitive landscape and the safety implications of large-scale models like GPT-4, this report contains no further details about the architecture (including model size), hardware, training compute, dataset construction, training method, or similar. We are committed to independent auditing of our technologies, and shared some initial steps and"},{"citing_arxiv_id":"2211.09085","ref_index":146,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Galactica: A Large Language Model for Science","primary_cat":"cs.CL","submitted_at":"2022-11-16T18:06:33+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"Galactica, a science-specialized LLM, reports higher scores than GPT-3, Chinchilla, and PaLM on LaTeX knowledge, mathematical reasoning, and medical QA benchmarks while outperforming general models on BIG-bench.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2112.04359","ref_index":29,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Ethical and social risks of harm from Language Models","primary_cat":"cs.CL","submitted_at":"2021-12-08T16:09:48+00:00","verdict":"ACCEPT","verdict_confidence":"MODERATE","novelty_score":6.0,"formal_verification":"none","one_line_summary":"The authors provide a detailed taxonomy of 21 risks associated with language models, covering discrimination, information leaks, misinformation, malicious applications, interaction harms, and societal impacts like job loss and environmental costs.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2101.00027","ref_index":134,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"The Pile: An 800GB Dataset of Diverse Text for Language Modeling","primary_cat":"cs.CL","submitted_at":"2020-12-31T19:00:10+00:00","verdict":"CONDITIONAL","verdict_confidence":"LOW","novelty_score":8.0,"formal_verification":"none","one_line_summary":"The Pile is a newly constructed 825 GiB dataset from 22 diverse sources that enables language models to achieve better performance on academic, professional, and cross-domain tasks than models trained on Common Crawl variants.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2005.14165","ref_index":2,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Language Models are Few-Shot Learners","primary_cat":"cs.CL","submitted_at":"2020-05-28T17:29:03+00:00","verdict":"ACCEPT","verdict_confidence":"MODERATE","novelty_score":8.0,"formal_verification":"none","one_line_summary":"GPT-3 shows that scaling an autoregressive language model to 175 billion parameters enables strong few-shot performance across diverse NLP tasks via in-context prompting without fine-tuning.","context_count":1,"top_context_role":"background","top_context_polarity":"background","context_text":"WMT'16 De↔En, and WMT'16 Ro ↔En datasets as measured by multi-bleu.perl with XLM's tokeniza- tion in order to compare most closely with prior unsupervised NMT work. SacreBLEU f [Pos18] results re- ported in Appendix H. Underline indicates an unsupervised or few-shot SOTA, bold indicates supervised SOTA with relative conﬁdence. a[EOAG18] b[DHKH14] c[WXH+18] d[oR16] e[LGG+20] f [SacreBLEU signature: BLEU+case.mixed+numrefs.1+smooth.exp+tok.intl+version.1.2.20] Figure 3.4: Few-shot translation performance on 6 language pairs as model capacity increases. There is a consistent trend of improvement across all datasets as the model scales, and as well as tendency for translation into English to be stronger than translation from English. 15 Setting Winograd Winogrande (XL) Fine-tuned SOTA 90.1a 84.6b GPT-3 Zero-Shot 88."}],"limit":50,"offset":0}