LLMs display a consistent pattern of elevated form-meaning divergence and uniform rhetorical device use in argumentative texts compared to humans, quantified by new metrics FMD, GPR, and RDDE.
Title resolution pending
6 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
years
2026 6verdicts
UNVERDICTED 6roles
background 1polarities
background 1representative citing papers
Causal mediation analysis shows harmful LLM outputs arise in late layers from MLP failures and gating neurons, with early layers handling harm context detection and signal propagation.
MLLMs show self-preference bias and family-level mutual bias when judging captions; Philautia-Eval quantifies it and Pomms ensemble reduces it.
Introduces FARO, a scalable quadratic optimization approach for fairness-aware top-k retrieval in RAG that mitigates generation bias via controlled reranking and position-aware propagation modeling.
A three-phase ML-assisted curation creates a Cardiology Interface Terminology (CIT) from SNOMED and EHR data that highlights details in cardiology notes with 74.21% coverage, 98.2% average completeness, and 84.2% average conciseness on test data.
All five tested LLMs deviated from US race-stratified disease distributions in synthetic case generation, while retrieval-based agentic workflows improved mean p-value by 0.0348, median p-value by 0.1166, and mean difference by 0.0949 for DeepSeek V3 in diagnosis ranking.
citing papers explorer
-
Saying More Than They Know: A Framework for Quantifying Epistemic-Rhetorical Miscalibration in Large Language Models
LLMs display a consistent pattern of elevated form-meaning divergence and uniform rhetorical device use in argumentative texts compared to humans, quantified by new metrics FMD, GPR, and RDDE.
-
Why Do Large Language Models Generate Harmful Content?
Causal mediation analysis shows harmful LLM outputs arise in late layers from MLP failures and gating neurons, with early layers handling harm context detection and signal propagation.
-
MLLM-as-a-Judge Exhibits Model Preference Bias
MLLMs show self-preference bias and family-level mutual bias when judging captions; Philautia-Eval quantifies it and Pomms ensemble reduces it.
-
Fairness-Aware Retrieval Optimization for Retrieval-Augmented Generation
Introduces FARO, a scalable quadratic optimization approach for fairness-aware top-k retrieval in RAG that mitigates generation bias via controlled reranking and position-aware propagation modeling.
-
Curation of a Cardiology Interface Terminology for Highlighting Electronic Health Records using Machine Learning
A three-phase ML-assisted curation creates a Cardiology Interface Terminology (CIT) from SNOMED and EHR data that highlights details in cardiology notes with 74.21% coverage, 98.2% average completeness, and 84.2% average conciseness on test data.
-
First, Do No Harm (With LLMs): Mitigating Racial Bias via Agentic Workflows
All five tested LLMs deviated from US race-stratified disease distributions in synthetic case generation, while retrieval-based agentic workflows improved mean p-value by 0.0348, median p-value by 0.1166, and mean difference by 0.0949 for DeepSeek V3 in diagnosis ranking.