{"total":13,"items":[{"citing_arxiv_id":"2606.31213","ref_index":31,"ref_count":1,"confidence":0.88,"is_internal_anchor":false,"paper_title":"Can LLMs Imagine Moral Alternatives Beyond Binary Dilemmas?","primary_cat":"cs.CL","submitted_at":"2026-06-30T06:49:06+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"LLMs often prefer compromise alternatives over binary moral options and generate alternatives rated higher than human-authored ones on structural and ethical criteria.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.23595","ref_index":65,"ref_count":1,"confidence":0.88,"is_internal_anchor":false,"paper_title":"SPIRAL: Learning to Search and Aggregate","primary_cat":"cs.AI","submitted_at":"2026-06-22T17:02:09+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"SPIRAL is a reinforcement learning framework that jointly optimizes sequential reasoning, parallel trace generation, and aggregation in language models for improved test-time performance.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.07951","ref_index":100,"ref_count":1,"confidence":0.88,"is_internal_anchor":false,"paper_title":"From `May' to `Is': Certainty Distortion in Language Model Rewriting","primary_cat":"cs.CL","submitted_at":"2026-06-06T02:53:31+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"LMs systematically inflate expressed certainty during rewriting, affecting up to 75% of outputs with a 1.5-2x bias toward increasing rather than decreasing certainty, and the effect compounds over iterations.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.05384","ref_index":63,"ref_count":1,"confidence":0.88,"is_internal_anchor":false,"paper_title":"Stability vs. Manipulability: Evaluating Robustness Under Post-Decision Interaction in LLM Judges","primary_cat":"cs.AI","submitted_at":"2026-06-03T19:37:23+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"LLM judges exhibit high stability under neutral re-evaluation but substantial reversibility under targeted post-decision challenges, quantified via a new Evaluation Robustness Score (ERS).","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.01914","ref_index":46,"ref_count":1,"confidence":0.88,"is_internal_anchor":false,"paper_title":"Mechanistic Diagnostics of Spatial Lexical Bias in Multimodal Large Language Model Spatial Reasoning","primary_cat":"cs.CL","submitted_at":"2026-06-01T08:49:47+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"MLLMs exhibit spatial lexical bias on multiple-choice spatial questions, traced via mechanistic tools to language-side channels rather than vision, and largely mitigated by LLM-only DPO on synthetic data.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.22785","ref_index":40,"ref_count":1,"confidence":0.88,"is_internal_anchor":false,"paper_title":"Evaluating Commercial AI Chatbots as News Intermediaries","primary_cat":"cs.CL","submitted_at":"2026-05-21T17:42:07+00:00","verdict":"CONDITIONAL","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"Commercial AI chatbots reach over 90% multiple-choice accuracy on recent news facts but lose 11-17% in free response and drop to 19-70% on subtle false-premise questions, with retrieval failures causing most errors and clear Anglophone bias.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.20128","ref_index":46,"ref_count":1,"confidence":0.88,"is_internal_anchor":false,"paper_title":"MixRea: Benchmarking Explicit-Implicit Reasoning in Large Language Models","primary_cat":"cs.CL","submitted_at":"2026-05-19T17:15:08+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"MixRea benchmark reveals LLMs achieve at most 42.8% consistency on explicit-implicit reasoning tasks, with PRCP prompting proposed to recover overlooked relations.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.18890","ref_index":47,"ref_count":1,"confidence":0.88,"is_internal_anchor":false,"paper_title":"Stop Drawing Scientific Claims from LLM Social Simulations Without Robustness Audits","primary_cat":"physics.soc-ph","submitted_at":"2026-05-17T00:21:53+00:00","verdict":"ACCEPT","verdict_confidence":"MODERATE","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Minor perturbations in persona format, instruction framing, and network structure shift cooperation by up to 76 percentage points and polarization metrics consistently, showing that LLM social simulations require per-claim robustness audits via the new TRAILS taxonomy.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.15393","ref_index":51,"ref_count":1,"confidence":0.88,"is_internal_anchor":false,"paper_title":"LPDS: Evaluating LLM Robustness Through Logic-Preserving Difficulty Scaling","primary_cat":"cs.LG","submitted_at":"2026-05-14T20:26:59+00:00","verdict":"CONDITIONAL","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"LPDS quantifies difficulty of logic-preserving problem variations and searches for the hardest ones, producing up to 5x larger performance drops than random sampling and better robustness gains from fine-tuning on difficult examples.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.01846","ref_index":20,"ref_count":1,"confidence":0.88,"is_internal_anchor":false,"paper_title":"Do Large Language Models Plan Answer Positions? Position Bias in Multiple-Choice Question Generation","primary_cat":"cs.CL","submitted_at":"2026-05-03T12:29:41+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"LLMs implicitly plan answer positions during MCQ generation, as shown by predictive signals in hidden representations and controllable shifts via activation steering.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.00907","ref_index":32,"ref_count":1,"confidence":0.88,"is_internal_anchor":false,"paper_title":"TRIP-Evaluate: An Open Multimodal Benchmark for Evaluating Large Models in Transportation","primary_cat":"cs.CV","submitted_at":"2026-04-29T04:29:48+00:00","verdict":"ACCEPT","verdict_confidence":"MODERATE","novelty_score":7.0,"formal_verification":"none","one_line_summary":"TRIP-Evaluate is a new open multimodal benchmark with 837 text, image, and point-cloud items organized by a role-task-knowledge taxonomy to evaluate large models on transportation workflows.","context_count":1,"top_context_role":"background","top_context_polarity":"unclear","context_text":", and C. Garcia-Garcia. Differential length and overlap with the stem in multiple-choice item options: A pilot experiment. Educational Psychology, 2018, 25: 43-48. doi:10.5093/psed2018a20. [31] Zheng, C., H. Zhou, F. Meng, et al. Large language models are not robust multiple-choice selectors. CoRR, 2023, abs/2309.03882. doi:10.48550/arXiv.2309.03882. [32] Pezeshkpour, P., and E. Hruschka. Large language models sensitivity to the order of options in multiple- choice questions. In Findings of the Association for Computational Linguistics: NAACL 2024, 2024, pp. 2006-2017. doi:10.18653/v1/2024.findings-naacl.130. 18 A preprint - May 5, 2026 [33] Chen, X., H. Ma, J. Wan, et al. Multi-view 3D object detection network for autonomous driving."},{"citing_arxiv_id":"2605.06672","ref_index":16,"ref_count":1,"confidence":0.88,"is_internal_anchor":false,"paper_title":"More Thinking, More Bias: Length-Driven Position Bias in Reasoning Models","primary_cat":"cs.AI","submitted_at":"2026-04-21T04:14:13+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"Position bias scales positively with reasoning trajectory length in CoT models, shown by partial correlations and truncation interventions across multiple benchmarks and model scales.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2509.23542","ref_index":28,"ref_count":1,"confidence":0.88,"is_internal_anchor":false,"paper_title":"On the Shelf Life of Fine-Tuned LLM-Judges: Future-Proofing, Backward-Compatibility, and Question Generalization","primary_cat":"cs.CL","submitted_at":"2025-09-28T00:43:52+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Fine-tuned LLM judges struggle with future-proofing to newer generators but maintain backward-compatibility more easily; DPO training and continual learning improve adaptation while all models degrade on unseen questions.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null}],"limit":50,"offset":0}