LLMs show significant pro-female bias on Japanese resumes across five models; name removal nearly eliminates the effect while prompt instructions do not, and privacy filters trigger high refusal rates on GPT-4o.
Robustly improving llm fairness in realistic settings via interpretability
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it