Pith. sign in

Paper Citation Record · LEDGER

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs

As of 8 August 2026, this Paper Citation Record lists 84 of 84 outbound references and 0 inbound Pith citation observations for arXiv:2505.17131.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.17131 v1

Coverage vector

measured 84 of 84 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:12:26.426274Z

measured 84 of 84 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

84 of 84 outbound references displayed

  • verified exact2
  • verified fuzzy28
  • unresolved54
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e4bdad71-1fe1-4c64-bfdc-823656e5396c · outbound

This paper cites https://aws.amazon.com/bedrock, 2024.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs https://aws.amazon.com/bedrock, 2024

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:12:17.646921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:12:17.646921Z digest=sha256:c097ea754b557675b1996d264394383225bcf68c7aa7b5aa317e0d214ef75d9e

Observation fe12084b-f7d6-471a-ae69-fc55aedc4e25 · outbound

This paper cites https://www.deepseek.com, 2024.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs https://www.deepseek.com, 2024

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:12:17.746789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:12:17.746789Z digest=sha256:e62712b569c4292f69340419747995189063cc532560dd7963240b63f83b387b

Observation d3263a64-e21f-4c0a-b2a8-9466450e3b1b · outbound

This paper cites https://aistudio.google.com, 2024.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs https://aistudio.google.com, 2024

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T15:12:17.883888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:12:17.883888Z digest=sha256:fd8851f4549eed716768a99cc0527733d90e9a71f5cb7471354e02c5e11cd70e

Observation 27091ff0-102d-4124-8d26-6e80a6136dfe · outbound

This paper cites https://www.meta.ai, 2024.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs https://www.meta.ai, 2024

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:12:18.019509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:12:18.019509Z digest=sha256:b9bf202e69fd1658fcff1f98a628b70cc6a312b78f5fce8daf9e8fbe83712f75

Observation cf38449e-a6ba-49f6-b810-513208dd96c1 · outbound

This paper cites Tukey’s honestly significant difference (hsd) test.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs Tukey’s honestly significant difference (hsd) test

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:33.024376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:12:18.140041Z digest=sha256:a1db3c02777818b8ecfa07d6e4d1352b0d1fd177bf995870d48f0537eff62a5f

Observation 5fb07169-7a57-4fd0-a9a8-ce1ea80e5927 · outbound

This paper cites Mitigating Language-Dependent Ethnic Bias in BERT.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs Mitigating Language-Dependent Ethnic Bias in BERT

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:12:18.275983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:12:18.275983Z digest=sha256:494082c73446ac2978ee8a02efe5bd339d83fc653ffd73b7f1c28493a2ce3e6b

Observation f4578404-f9f2-4b19-8eb3-86f2bbf2f202 · outbound

This paper cites Amazon bedrock guardrails.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs Amazon bedrock guardrails

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:32.863742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:12:18.427567Z digest=sha256:c4bdb90d8246379d84bd438d5d45fd71d8d3b2ddc593545112275653f66dc72a

Observation 11d2797a-da77-43e9-8bac-7afda67c3f9b · outbound

This paper cites A Human-AI Comparative Analysis of Prompt Sensitivity in LLM-Based Relevance Judgment.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs A Human-AI Comparative Analysis of Prompt Sensitivity in LLM-Based Relevance Judgment

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T15:12:18.548069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:12:18.548069Z digest=sha256:c57452f460c7757f0972d6fd5e8b446e264ce6e4cf37ea72faf6e5399e1b2e81

Observation 02e014cc-a2ce-4e43-808d-8de825d45fef · outbound

This paper cites Non-Determinism of "Deterministic" LLM Settings.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs Non-Determinism of "Deterministic" LLM Settings

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T15:12:18.716951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:12:18.716951Z digest=sha256:5c6d7afc1c424dbd457246388e545359a764a6fdb8e9fd0b4fb2d5584ec310be

Observation fe4888e9-46c0-4c98-bb3c-0c640548b716 · outbound

This paper cites Language models are few-shot learners.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs Language models are few-shot learners

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T15:12:18.815946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:12:18.815946Z digest=sha256:654d6de1ab23bb1193f9ffece1fb9b9b76648dc9ac29ba513c0f2e5a9b391835

Observation bb3a2d54-1d8a-4e21-99cc-d06ef3418dd8 · outbound

This paper cites FairFil: Contrastive Neural Debiasing Method for Pretrained Text Encoders.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs FairFil: Contrastive Neural Debiasing Method for Pretrained Text Encoders

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:12:27.164329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:12:18.982402Z digest=sha256:353f291d5605866c0f6db8d916285615b4c3a20bf387ae38978466dd6633561a

Observation 0758b737-4732-4884-9924-d4b18d25fad9 · outbound

This paper cites OR-Bench: An Over-Refusal Benchmark for Large Language Models.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs OR-Bench: An Over-Refusal Benchmark for Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T15:12:19.066466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:12:19.066466Z digest=sha256:51a1944f4f0cc8bfdc8c417124bf4d4b90cb92930eda7d635daabafc0e39100e

Observation c338b927-9b1c-4f4d-9ffc-207a5ac2d65c · outbound

This paper cites Bert: Pre-training of deep bidi- rectional transformers for language understanding.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs Bert: Pre-training of deep bidi- rectional transformers for language understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:12:19.173065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:12:19.173065Z digest=sha256:2e242acdc25a4da70162c1c68ee47bec7c190bbcba59437d1e7ca66b4a830fb0

Observation 6b5ca39a-ffbc-4f02-b6d7-08215b5c9787 · outbound

This paper cites Bold: Dataset and metrics for measuring biases in open-ended language generation.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs Bold: Dataset and metrics for measuring biases in open-ended language generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:12:19.297539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:12:19.297539Z digest=sha256:f1ab6e6a86025fd06dd073400d9f996608937fc9483eb197929e907718937287

Observation 8157e21b-4b7c-4a7f-a39b-ebe9562fb848 · outbound

This paper cites Disclosure and Mitigation of Gender Bias in LLMs.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs Disclosure and Mitigation of Gender Bias in LLMs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:12:19.415481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:12:19.415481Z digest=sha256:7fb6db31ec1e63a875f4fda55fb9423daa096fa92d266a7382ab27756db5a446

Observation 5825adc7-a251-4d94-8bbf-d85ae8576f63 · outbound

This paper cites SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:12:19.560729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:12:19.560729Z digest=sha256:d1f2bb23679432fc0607ed58b80f2377146515361b3cfca8c50251e523a7f003

Observation 82026c9e-a093-442e-a7e2-0a7c6444b669 · outbound

This paper cites Alpacafarm: A simulation framework for methods that learn from human feedback.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs Alpacafarm: A simulation framework for methods that learn from human feedback

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:32.722693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:12:19.703634Z digest=sha256:e810ae04866c37df34cb05ecef36b9cc0b2f235c3fb6d8e9dc6a0a7ba91ca76f

Observation 0a9afb5e-8da9-4087-a629-60bb14512d12 · outbound

This paper cites Meta ai refusing to answer questions related to politicians and par- ties ahead of elections in india, 2024.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs Meta ai refusing to answer questions related to politicians and par- ties ahead of elections in india, 2024

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:32.563582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:12:19.856933Z digest=sha256:a360a35fd3f6fcf971ee3adedd85f230c7f15dd04915652bd21b05eb4b2b5991

Observation b2492a96-992d-453f-8d21-e86786a60a30 · outbound

This paper cites Toy Models of Superposition.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs Toy Models of Superposition

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:12:20.004814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:12:20.004814Z digest=sha256:18090c74c549e2b68efe3a0a7c9570340716ff43470a30525d5d8d0f564cf84f

Observation fc80d2f6-491b-453f-af53-37fcca23bb84 · outbound

This paper cites ROBBIE: Robust Bias Evaluation of Large Generative Language Models.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs ROBBIE: Robust Bias Evaluation of Large Generative Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:12:20.139639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:12:20.139639Z digest=sha256:fbf0217d0f0da79f59e021c48026072a83af9ec9156ab68e522d52c6963e738b

Observation 4bc9522b-1911-46b0-8868-8bf5a10ab785 · outbound

This paper cites Bias and fairness in large language models: A survey.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs Bias and fairness in large language models: A survey

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:12:20.291546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:12:20.291546Z digest=sha256:e93e335c1d0fa24ab69aa88efdc876339e2846de2d6c9a14a8df369c8e905134

Observation 13c4e8a8-9369-4abf-9209-9364879cd4e4 · outbound

This paper cites Pairwise multiple comparison procedures with unequal n’s and/or variances: a monte carlo study.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs Pairwise multiple comparison procedures with unequal n’s and/or variances: a monte carlo study

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:32.413947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:12:20.426690Z digest=sha256:22486b53be7f33d09648c803b8d2c94ef2cc80743960d3097dd222a4002f36cd

Observation cddec1c9-c2d5-4615-bddb-26582ad667b3 · outbound

This paper cites RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:12:20.521002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:12:20.521002Z digest=sha256:f62a1ce0f8e0a6cbc64cb4bb76ec054d7cc4c1e8d4beed3ef04139767d10c261

Observation bd52ac49-2568-4496-83b5-0452e8d5603f · outbound

This paper cites Debiasing pre-trained language models via efficient fine-tuning.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs Debiasing pre-trained language models via efficient fine-tuning

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:32.190325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:12:20.612760Z digest=sha256:a37e6ea1e27ef8456b90291199d464fa73d8c132fe988ce6710f6ec40a7aa10e

Observation 81337885-9f20-44bf-84ba-34c00a739afa · outbound

This paper cites A Survey on LLM-as-a-Judge.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs A Survey on LLM-as-a-Judge

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:12:20.754742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:12:20.754742Z digest=sha256:f5c46c8ccdd1bd14a74edcc529f1c2670741603c575740f7b2c8ff60b44af74f

Observation bc5e5836-6e71-49db-a11c-68f0ec9b3020 · outbound

This paper cites We tried out deepseek.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs We tried out deepseek

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:31.997050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:12:20.827485Z digest=sha256:9d8bd4daf669084238cc971b93fd72a247c753b1ab0d2c44b156ad541e27cec2

Observation b9f7f6dd-7bd2-4dc5-bbcb-34757dbca244 · outbound

This paper cites Auto-debias: Debiasing masked language models with automated biased prompts.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs Auto-debias: Debiasing masked language models with automated biased prompts

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:31.826061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:12:20.874761Z digest=sha256:779df43b053acd526bac786d9580ad878880eec2ba7236d05069cebf5878a8c1

Observation 5cf0360d-fe30-4af0-801e-713682834ef7 · outbound

This paper cites Does Prompt Formatting Have Any Impact on LLM Performance?.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs Does Prompt Formatting Have Any Impact on LLM Performance?

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T15:12:20.912641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:12:20.912641Z digest=sha256:46684ff3e2318f9f124a929d3fc6f0eb2e2d62fee777104b628dc2867be6bd1b

Observation 4875d47e-8d33-40d7-a52a-bfac4fa55d63 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs Measuring Massive Multitask Language Understanding

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T15:12:21.025106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:12:21.025106Z digest=sha256:72dcd6a64c39ce7ae02b06dd19328080d42f40526b1d5eca72dc79a1057de14d

Observation 9288de9f-697d-41e5-8006-36910969edb9 · outbound

This paper cites Reducing Sentiment Bias in Language Models via Counterfactual Evaluation.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs Reducing Sentiment Bias in Language Models via Counterfactual Evaluation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T15:12:21.086207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:12:21.086207Z digest=sha256:c5dcfddfe90a8ba082f9877f20698156133915c5dd08ac5668a75317c75bf480

Observation 94768ba5-bedc-4d62-9e37-9b389fae0686 · outbound

This paper cites Perspective api, 2025.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs Perspective api, 2025

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:31.572271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:12:21.163584Z digest=sha256:bbce2fbc462af111334f9a8699eecca61343fedd59f40a4dd3f4330812d24d3f

Observation 4c37d7f6-31c5-40eb-88bd-3596d7a7721b · outbound

This paper cites Debiasing Pre-trained Contextualised Embeddings.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs Debiasing Pre-trained Contextualised Embeddings

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T15:12:21.217936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:12:21.217936Z digest=sha256:40abccf420581850db0cbf6685e7911ecad6d19925023bd4a0e49320fddedccf

Observation f776040b-4aa1-490c-98c5-ba51bdd90fc0 · outbound

This paper cites Pretraining language models with human preferences.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs Pretraining language models with human preferences

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T15:12:21.293155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:12:21.293155Z digest=sha256:875b3c4aaa40a076b2f7ab002288cc982c7c278f37673e71fd29a6cc16a17433

Observation 04241856-2d8a-48a5-90dc-7c056fc46297 · outbound

This paper cites Measuring Bias in Contextualized Word Representations.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs Measuring Bias in Contextualized Word Representations

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T15:12:21.340647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:12:21.340647Z digest=sha256:c2bf87e5d468ca9b214b73aa6e6291f3e1b8b4b634734ebf333e43cbb8ab4214

Observation 13f4ea0b-56d7-465c-8718-3f306199ce9b · outbound

This paper cites Equivalence tests: A practical primer for t tests, correlations, and meta-analyses.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs Equivalence tests: A practical primer for t tests, correlations, and meta-analyses

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:31.339605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:12:21.373901Z digest=sha256:bc65c7f19bc14f6b5c3b8dd21cfa9e71755cabb58373ae9dab3268f378f9be0b

Observation 900fb795-066d-4dff-ae43-98f629135183 · outbound

This paper cites Fairness Testing of Large Language Models in Role-Playing.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs Fairness Testing of Large Language Models in Role-Playing

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T15:12:21.650019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:12:21.650019Z digest=sha256:189be7c04ae03e4b9818805b5ba8c17862994f4bb61aac29af98aabdae3ddf67

Observation d05f91ee-fda1-4647-b3ad-ebb11077e1e8 · outbound

This paper cites Alpacaeval: An automatic evaluator of instruction-following models, 2023.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs Alpacaeval: An automatic evaluator of instruction-following models, 2023

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T15:12:21.727008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:12:21.727008Z digest=sha256:d6dc260e4280f25e784483b6e206169007c709469267ce33dafa9963e3de1e93

Observation ecde6a30-1e10-4d50-aea4-2ca2f4ab2083 · outbound

This paper cites Towards Debiasing Sentence Representations.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs Towards Debiasing Sentence Representations

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T15:12:21.833828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:12:21.833828Z digest=sha256:8f9e57549f73b37f1b477c963584808014ad29ffbfffb5c4b41f7515247c19ff

Observation 94d18f79-d22a-4a1d-8332-1afc7a36329e · outbound

This paper cites Towards understanding and mitigating social biases in language models.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs Towards understanding and mitigating social biases in language models

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:31.100134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:12:21.919579Z digest=sha256:ef51b15ce0f55dc9974705b548e3cf2b58cd58aaf796b9854acfe3920211dc44

Observation 9e817d53-75a9-4742-972f-1035cb756a88 · outbound

This paper cites Holistic Evaluation of Language Models.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs Holistic Evaluation of Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T15:12:22.022328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:12:22.022328Z digest=sha256:fceca38c7c0c2aef7f7e85a2cd4b0272594635067d089eeb6f1b740284a6ef0e

Observation 5be1b9cd-4210-480c-846b-3d16adc3b2f5 · outbound

This paper cites Investigating Bias in LLM-Based Bias Detection: Disparities between LLMs and Human Perception.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs Investigating Bias in LLM-Based Bias Detection: Disparities between LLMs and Human Perception

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T15:12:22.173007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:12:22.173007Z digest=sha256:5c2d31907ad6b861bcdebbec414a4d81a483b3c9b3d400abe51e44855d67c71e

Observation bed8add1-4382-4d3a-a30b-00692b898d15 · outbound

This paper cites Does Gender Matter? Towards Fairness in Dialogue Systems.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs Does Gender Matter? Towards Fairness in Dialogue Systems

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T15:12:22.311528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:12:22.311528Z digest=sha256:f9a895ff8264be29e54f5fbb5808c770e7d3bc015229e8a9a205f11f9603e948

Observation d5d6003c-7c16-44c6-85d0-3af142aefd78 · outbound

This paper cites Lmarena: Open platform for crowdsourced ai benchmarking.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs Lmarena: Open platform for crowdsourced ai benchmarking

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:30.817269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:12:22.483302Z digest=sha256:f887cb2d5fa094f6dc54472c277a8dbaa1906bad7d61266d92cc11e12f429cea

Observation 7f818f25-51e8-4eb6-9bee-290dcef3de71 · outbound

This paper cites Azure openai service content filtering.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs Azure openai service content filtering

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:30.573600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:12:22.598191Z digest=sha256:80b1a1c508ed58093a95fe567497bfe5527cfc3bb79839b5cba4e74cd2e69610

Observation debc0755-9045-45b1-9b60-0bd2280fead5 · outbound

This paper cites CrowS-Pairs: A Challenge Dataset for Measuring Social Biases in Masked Language Models.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs CrowS-Pairs: A Challenge Dataset for Measuring Social Biases in Masked Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T15:12:22.712392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:12:22.712392Z digest=sha256:edc4654b70c92c016451fe2e3373ff9544ef3cb69e4dcedd99c5a6bf5423da60

Observation 82e35708-b7a9-4509-960f-b2fcb2468cf6 · outbound

This paper cites Sentence-T5: Scalable Sentence Encoders from Pre-trained Text-to-Text Models.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs Sentence-T5: Scalable Sentence Encoders from Pre-trained Text-to-Text Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T15:12:22.808661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:12:22.808661Z digest=sha256:1fcc322c23b6f4f624d79903e07cdaf77eff16ec382a35971c66e1a968b736df

Observation bc9135a4-6626-4692-8217-3ef9f92ed405 · outbound

This paper cites Honest: Measuring hurtful sentence completion in language models.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs Honest: Measuring hurtful sentence completion in language models

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:30.369802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:12:22.902161Z digest=sha256:4031f49cf49ba5d04fd41ca024b9ce2a8c21366cc250430595f4b70debeb3432

Observation 7f840faa-bef0-42ee-b0dc-83a1ca6f1fa7 · outbound

This paper cites Large Language Model (LLM) Bias Index -- LLMBI.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs Large Language Model (LLM) Bias Index -- LLMBI

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T15:12:22.978844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:12:22.978844Z digest=sha256:82a4e3e23fffb01927c370986f9ffc32226ee181d007f476b0e1b008256a19a3

Observation db77cd14-10cc-4dad-942e-f3e128a84168 · outbound

This paper cites GPT-4 Technical Report.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs GPT-4 Technical Report

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T15:12:23.078776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:12:23.078776Z digest=sha256:b18f066da6b2e2a032a7b0761fae7f6390997d91c664e2601024ce9937d6e24b

Observation 7b5cda11-1dbf-44bf-adf9-4569df9bbcfc · outbound

This paper cites The perils and promises of fact-checking with large language models.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs The perils and promises of fact-checking with large language models

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:30.102080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:12:23.154522Z digest=sha256:e1a1b7ecf4f521951a3ea9437520f1a94f9ffc7637ec6a4820c0ca7284843ad7

Observation 4aa33908-6d17-4353-a76f-5c31de903cc5 · outbound

This paper cites MBIAS: Mitigating Bias in Large Language Models While Retaining Context.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs MBIAS: Mitigating Bias in Large Language Models While Retaining Context

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T15:12:23.243382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:12:23.243382Z digest=sha256:3770ab1973bb1b6a987cda222bddd1208063b656be3a8170ffef5b6ae813d70b

Observation 717f6f6f-6ed1-4ab2-8aac-de77b9e10e61 · outbound

This paper cites NeMo Guardrails: A Toolkit for Controllable and Safe LLM Applications with Programmable Rails.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs NeMo Guardrails: A Toolkit for Controllable and Safe LLM Applications with Programmable Rails

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T15:12:23.329400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:12:23.329400Z digest=sha256:4d1e359d1b1914adaf8638f79d6b61d3f26f9a5739f66e6e26524e6c78b2c158

Observation 023e3dd2-a012-45f6-8ce4-d8c286ca2e5b · outbound

This paper cites Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T15:12:23.428995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:12:23.428995Z digest=sha256:6fd559ae74533760b7762eefbe1278f9d8d4e085be40a94ed63aa365bffa1a80

Observation d7384e92-8db2-4a4b-82cd-52ed3d1c5261 · outbound

This paper cites Does deepseek censor its answers? we asked 5 questions on sensitive china top- ics.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs Does deepseek censor its answers? we asked 5 questions on sensitive china top- ics

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:29.864986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:12:23.541843Z digest=sha256:ace3005c2f5b39870262ae2f4ac43a69e61994c877d6bf357e0993b829d54f39

Observation 5dad7343-2525-4cc9-a388-3578f07946de · outbound

This paper cites Introduction to probability models.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs Introduction to probability models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T15:12:23.660114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:12:23.660114Z digest=sha256:5da8ce2e9ce0e7cf54849bfe58dcb53f99416e8c8ef3c29fe6ec90640f675066

Observation c786a44b-f57f-4d3f-9d2d-5d2398be8f55 · outbound

This paper cites Self-diagnosis and self-debiasing: A proposal for reducing corpus-based bias in nlp.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs Self-diagnosis and self-debiasing: A proposal for reducing corpus-based bias in nlp

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:29.641418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:12:23.743605Z digest=sha256:44223b34776e60a024853dd32657c627fe0ee8f91500915905f6c2b89209af59

Observation 237f47cf-82c3-43bb-9f8b-520b225f8c36 · outbound

This paper cites A comparison of the two one-sided tests procedure and the power approach for assessing the equivalence of average bioavailability.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs A comparison of the two one-sided tests procedure and the power approach for assessing the equivalence of average bioavailability

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:29.430369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:12:23.964865Z digest=sha256:e6fc053f244d8df8ed88547b478f399614f44a1a505ec8ddfdd2ac9f0a7f1e21

Observation 5e59c711-700f-44cf-b5de-bb0b683911ad · outbound

This paper cites Large Language Model Alignment: A Survey.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs Large Language Model Alignment: A Survey

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T15:12:24.027242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:12:24.027242Z digest=sha256:39cf748fdb9a6305ddda9ab25a1c725bae43cf0cd2e6b1a12f092128f9ed9d6e

Observation 99c86017-6745-4a4d-935a-2c35e05ab2d2 · outbound

This paper cites Prompting GPT-3 To Be Reliable.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs Prompting GPT-3 To Be Reliable

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T15:12:24.224762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:12:24.224762Z digest=sha256:e0b732f5a72844f0e80e18136c16ab01384b60ddb2e479a74a1dd731a2825042

Observation 3ea5021e-aa06-4c4f-a1cb-bce3c7a11016 · outbound

This paper cites The Good, The Bad, and The Greedy: Evaluation of LLMs Should Not Ignore Non-Determinism.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs The Good, The Bad, and The Greedy: Evaluation of LLMs Should Not Ignore Non-Determinism

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T15:12:24.340787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:12:24.340787Z digest=sha256:1f363af00f3860fd5baabc6a7f3eb7e772d3de21a7b918d744195291cfd3788a

Observation 8d7614bc-8ed8-49dc-a1c9-73f71e6fe84a · outbound

This paper cites Analysis of variance (anova).Chemometrics and intelligent laboratory systems, 6(4):259–272, 1989.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs Analysis of variance (anova).Chemometrics and intelligent laboratory systems, 6(4):259–272, 1989

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:29.270644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:12:24.408924Z digest=sha256:23926c6d2c9ef1f4b5628eccaabc81cdf4647d8046bade16aeda1d3d4ef5bdec

Observation 118dff2d-0753-4dd6-9bc8-c24c37b20be7 · outbound

This paper cites One Embedder, Any Task: Instruction-Finetuned Text Embeddings.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs One Embedder, Any Task: Instruction-Finetuned Text Embeddings

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T15:12:24.472276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:12:24.472276Z digest=sha256:e8db421bafe6681ee07e80d447475dbf0ca69c6042568a82239f0edd4b01043f

Observation d94475bd-d37f-40e5-b28d-2b83a782add4 · outbound

This paper cites Grok 3 appears to have briefly censored unflattering mentions of trump and musk, 2025.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs Grok 3 appears to have briefly censored unflattering mentions of trump and musk, 2025

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:29.089474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:12:24.566817Z digest=sha256:38e6f6dbeecc80bd0982709028bceeef619c18a5ce8e59ccd7740089f1443e39

Observation 2554893e-4bef-45d0-a450-5f78479ec2bc · outbound

This paper cites A Robust Bias Mitigation Procedure Based on the Stereotype Content Model.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs A Robust Bias Mitigation Procedure Based on the Stereotype Content Model

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:12:26.644599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:12:24.658029Z digest=sha256:109f081c741efca0c26082e2de0a47d8c80b7eaeadd4ca12e41fb59303420b56

Observation ea0bf318-8eca-4227-be09-d081ae39eca8 · outbound

This paper cites Llm leaderboard, 2025.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs Llm leaderboard, 2025

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:28.824739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:12:24.738891Z digest=sha256:4a8def2f108bc0b00baa05711c8b4a0be05a0b5b658357ed6acb618e0ae71870

Observation 0fed6980-7565-461d-85dd-9683beba6d21 · outbound

This paper cites The dark side of generative artificial intelligence: A critical analysis of controversies and risks of chatgpt.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs The dark side of generative artificial intelligence: A critical analysis of controversies and risks of chatgpt

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:28.629921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:12:24.829066Z digest=sha256:92f20a1981d0d4d51d5ce4abf58f117cf01a4bdeede53e9fa9cfb722e5b268dd

Observation 9f96eb18-a34c-4bf0-a191-435305c35cf0 · outbound

This paper cites Large Language Models are not Fair Evaluators.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs Large Language Models are not Fair Evaluators

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T15:12:24.922647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:12:24.922647Z digest=sha256:95bfaccbad5b9faba50568ed5d63af16aab7fb36f644de5ee63fd3924b430975

Observation d4827d86-f6f6-4b39-b673-e5803bfd6fe1 · outbound

This paper cites Self-Instruct: Aligning Language Models with Self-Generated Instructions.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs Self-Instruct: Aligning Language Models with Self-Generated Instructions

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T15:12:24.998012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:12:24.998012Z digest=sha256:411fd9e1bebac5ed164d12346e942ca45a8b55c1f5c050f194b1ac88ffd239e5

Observation 4c3efecd-3656-42c7-801a-bbec0a5ce01b · outbound

This paper cites Mmlu-pro: A more robust and challenging multi-task language understanding benchmark.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs Mmlu-pro: A more robust and challenging multi-task language understanding benchmark

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T15:12:25.092944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:12:25.092944Z digest=sha256:df9d39161d86383232b161f1b9a8e9593aebe7fd61b6f78967b2f456bae9f057

Observation d37038b2-8922-475e-98fc-2cf000c107c7 · outbound

This paper cites All of statistics: a concise course in statistical inference.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs All of statistics: a concise course in statistical inference

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T15:12:25.172063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:12:25.172063Z digest=sha256:f2aea8efb1f7916af73f1710ecc0fa84da86b99ccf4f95f9af64137b00fd8993

Observation 91b9fee3-05ff-4342-92ce-8b9e21dd9b7b · outbound

This paper cites Measuring and Reducing Gendered Correlations in Pre-trained Models.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs Measuring and Reducing Gendered Correlations in Pre-trained Models

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T15:12:25.231885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:12:25.231885Z digest=sha256:0bf7e72840d83da52626a60b0d49ee165987799da208df02d1cb5b0a5e8c9c12

Observation a60db2fb-3c48-4f80-9cb9-f55dad0982d9 · outbound

This paper cites The generalization of ‘student’s’problem when several different population varlances are involved.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs The generalization of ‘student’s’problem when several different population varlances are involved

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T15:12:25.322067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:12:25.322067Z digest=sha256:ac80dad0a0d82c3fef2a07ff8730eea876dfaab3566ca373fe7068c249f28c4b

Observation fdbc3afe-5b05-4fca-8ae5-f5e313faf5b8 · outbound

This paper cites Livebench: A challenging, contamination-free LLM benchmark.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs Livebench: A challenging, contamination-free LLM benchmark

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:28.344544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:12:25.414011Z digest=sha256:97a17e109677600afdcd9294e738e88d0e583ebfb56881521c2abea85772611f

Observation 8ea27ce6-d086-4b6b-a7ff-3aa3a6784b2c · outbound

This paper cites This powerful new chatbot works great—unless you ask about china, 2025.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs This powerful new chatbot works great—unless you ask about china, 2025

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:28.133989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:12:25.526743Z digest=sha256:8ab675a104748e1938b24f42c55c3a114e638d2354dd49ddae40777c5e79f24a

Observation 8a853731-12f4-440a-ad5f-9c07a3867568 · outbound

This paper cites Compensatory debiasing for gender imbalances in language models.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs Compensatory debiasing for gender imbalances in language models

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:27.962499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:12:25.608617Z digest=sha256:1c1bf03a99ecb8cb53c5a6f1fdfea2d151770bda73ec81072e1feadde209d95f

Observation febd100a-0219-41c8-9fd7-2d4b3b64b0cb · outbound

This paper cites Order Matters in Hallucination: Reasoning Order as Benchmark and Reflexive Prompting for Large-Language-Models.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs Order Matters in Hallucination: Reasoning Order as Benchmark and Reflexive Prompting for Large-Language-Models

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T15:12:25.691449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:12:25.691449Z digest=sha256:76c33491b8aacdc6b1e9af067a90822dea6e8a9b98d9b7a0156d099087441099

Observation c813e29d-1692-4e8f-b78b-f080f31fb17d · outbound

This paper cites Large language model as attributed training data generator: A tale of diversity and bias.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs Large language model as attributed training data generator: A tale of diversity and bias

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:27.777884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:12:25.797224Z digest=sha256:80e36160903e2e2587feefc5c2bdce4ceb618da6e373fd20f47e3d0dfe4a4226

Observation b497a7b1-6cbe-42cb-8f8f-6b4ef4ab3a57 · outbound

This paper cites Wider and Deeper LLM Networks are Fairer LLM Evaluators.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs Wider and Deeper LLM Networks are Fairer LLM Evaluators

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T15:12:25.897572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:12:25.897572Z digest=sha256:cdce23d4f7b34a373a4012748faa5ae4a7e78a194b41d7ba5ebcbf7fb681d042

Observation a85f9f47-6d90-443d-ada6-a810ed66b41e · outbound

This paper cites Explainability for large language models: A survey.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs Explainability for large language models: A survey

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T15:12:25.991115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:12:25.991115Z digest=sha256:19830165d113d94d5a476465aa2f31108821799c89bf6caf604516efeb0c9bfe

Observation 0186f663-106e-4d56-b52e-bc8829b8046b · outbound

This paper cites GPTBIAS: A Comprehensive Framework for Evaluating Bias in Large Language Models.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs GPTBIAS: A Comprehensive Framework for Evaluating Bias in Large Language Models

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T15:12:26.108646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:12:26.108646Z digest=sha256:dfd142ce97c7740f20ed6b505631783f10e2e1120e77a3d64e0dcf5b38b0e220

Observation 36bb37d0-4df9-483f-8d44-fdea6d64eb38 · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs Judging llm-as-a-judge with mt-bench and chatbot arena

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T15:12:26.214959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:12:26.214959Z digest=sha256:10f39b1385f5175e386b943e45d9a5792295773ca50005a340fd64fe61a68912

Observation 3483dc30-69ef-424c-8e25-217f447deeca · outbound

This paper cites Lima: Less is more for alignment.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs Lima: Less is more for alignment

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T15:12:26.258317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:12:26.258317Z digest=sha256:19865c55a843d2f2dff594457d8efba3bd494a889b823656988e4928b77d2fe9

Observation 13ab7a7b-870e-4cd0-a527-687eedc76b63 · outbound

This paper cites {question}.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs {question}

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:27.596932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:12:26.342749Z digest=sha256:b45c086c292135e0bfaac60b463c2819d57bd466fe67152c5310981ae6e616e3

Observation 5ba7887e-9682-456e-966d-43e54c02743d · outbound

This paper cites A.3 Bias Criteria Table Score Description 1 Fully neutral and factual.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs A.3 Bias Criteria Table Score Description 1 Fully neutral and factual

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:27.434147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:12:26.426274Z digest=sha256:88b0f21ed7f36539bdde51a4b68a206ce240f6a93550d7b2c0569f334306fdc8

Pith citing papers

No inbound Pith citation observations are available.