REVIEW 3 cited by
JBBQ: Japanese Bias Benchmark for Analyzing Social Biases in Large Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
With the development of large language models (LLMs), social biases in these LLMs have become a pressing issue. Although there are various benchmarks for social biases across languages, the extent to which Japanese LLMs exhibit social biases has not been fully investigated. In this study, we construct the Japanese Bias Benchmark dataset for Question Answering (JBBQ) based on the English bias benchmark BBQ, with analysis of social biases in Japanese LLMs. The results show that while current open Japanese LLMs with more parameters show improved accuracies on JBBQ, their bias scores increase. In addition, prompts with a warning about social biases and chain-of-thought prompting reduce the effect of biases in model outputs, but there is room for improvement in extracting the correct evidence from contexts in Japanese. Our dataset is available at https://github.com/ynklab/JBBQ_data.
Forward citations
Cited by 3 Pith papers
-
McBE: A Multi-task Chinese Bias Evaluation Benchmark for Large Language Models
A new Chinese bias benchmark with 4,077 instances and five tasks indicates larger language models are less biased than smaller ones when bias is measured through understanding tasks.
-
Intersectional Bias in Japanese Large Language Models from a Contextualized Perspective
Using the new inter-JBBQ benchmark, biased responses of Japanese LLMs vary with scenario context even when similar social attribute combinations are tested.
-
GG-BBQ: German Gender Bias Benchmark for Question Answering
A manually corrected German translation of the BBQ gender bias QA dataset shows that all evaluated German LLMs exhibit measurable gender bias.
Discussion (0). Sign in to comment.