Pith. sign in

Paper Citation Record · LEDGER

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race

As of 7 August 2026, this Paper Citation Record lists 77 of 77 outbound references and 0 inbound Pith citation observations for arXiv:2506.00253.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.00253 v3

Coverage vector

measured 77 of 77 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:13:20.054591Z

measured 77 of 77 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

77 of 77 outbound references displayed

  • verified exact14
  • verified fuzzy2
  • unresolved61
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 584af120-bfe6-40c1-b645-8947fdc8da08 · outbound

This paper cites Physics of Language Models: Part 3.1, Knowledge Storage and Extraction.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Physics of Language Models: Part 3.1, Knowledge Storage and Extraction

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:13.741659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:13.741659Z digest=sha256:c77216a4c328be56700dcfdc33d1d6856d2680559ff4a009c6942d6e9f9738c3

Observation 6693e333-a4bc-4cd5-aab8-f4c3287e2506 · outbound

This paper cites Apfelbaum, Michael I.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Apfelbaum, Michael I

Reference 2

Resolution
verified exact
doi, observed 2026-08-07T12:13:21.009857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:13:13.794504Z digest=sha256:1776fe384a02fc22eddbf1f2e9e455326eb317b54dc37a6b87682427ca0984db

Observation 6959ccea-181f-4160-a58e-67650f41031f · outbound

This paper cites Griffiths.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Griffiths

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:13.852499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:13.852499Z digest=sha256:14c53e75ba5cdbea1cc60688f711166b552ebf229649f6db791d2d566fb9bc7d

Observation 79169ae5-5c3e-4610-9b62-d5849c130c93 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Constitutional AI: Harmlessness from AI Feedback

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:13.935489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:13.935489Z digest=sha256:d2636562e30f5c83ddd91d7cda56666505feac3dea67d42f9e486fb5790a7be3

Observation dee7fa67-5010-44a5-a867-7fc4e4cf00ed · outbound

This paper cites Eliciting Latent Predictions from Transformers with the Tuned Lens.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Eliciting Latent Predictions from Transformers with the Tuned Lens

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:14.003675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:14.003675Z digest=sha256:fa806130b56a28ff96e888a735e0a9d0617e368016f0e4f7c4332d1ed7c85c4b

Observation 25500e5a-d219-4d35-a161-615a1926a79d · outbound

This paper cites Mechanistic Interpretability for AI Safety -- A Review.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Mechanistic Interpretability for AI Safety -- A Review

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:14.094600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:14.094600Z digest=sha256:3f11daa9bc7ca94f89dfb59efce168cb923011cc59749bd9bb5c4193528ee6a9

Observation b9021306-b3f1-4f09-a292-b7889ada3133 · outbound

This paper cites an unresolved cited work.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:13:24.122596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:13:14.161233Z digest=sha256:ce34de89e67bdf4e58f5e06a94cbe34b2caf7d020313a4e84e239f8d46e2b081

Observation ec3a1bb6-f211-4382-a1a5-f523cad9fb58 · outbound

This paper cites an unresolved cited work.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:13:23.955589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:13:14.260481Z digest=sha256:7c0469268306bef4aa379c67fed9d9df70e0b81d4dcec7db6638434152087be0

Observation b587f0b4-5934-44b7-83d5-8e94695885a6 · outbound

This paper cites Bryson, and Arvind Narayanan.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Bryson, and Arvind Narayanan

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:14.350152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:14.350152Z digest=sha256:eb60bc1092239b480d5fcc9d4d9681e5da7f44d0b0a036f7d3ace000cef14340

Observation 00cfff6c-03a5-4672-a484-d602f0ba15bd · outbound

This paper cites SelfIE: Self-Interpretation of Large Language Model Embeddings.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race SelfIE: Self-Interpretation of Large Language Model Embeddings

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:14.424151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:14.424151Z digest=sha256:47dc241aaf196affdfccf9cf7a3269091f99c72d88b1ce121c3f6fe33bfed101

Observation 38c54000-6257-4f66-ab11-430cc0e4cb18 · outbound

This paper cites an unresolved cited work.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Unresolved cited work

Reference 11

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-07T12:13:22.867513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:13:14.501115Z digest=sha256:bf49db8506366669900f5017df4538e3a1cd304fc072b8850193a3222d93d67a

Observation 08e949fc-017d-46d7-a534-d69180fd89ee · outbound

This paper cites Mitigating Social Biases in Language Models through Unlearning.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Mitigating Social Biases in Language Models through Unlearning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:14.600807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:14.600807Z digest=sha256:760ffdf8044159e1560de5155deb4cf0a8d00bf6bf88870c7505165d8e23dd37

Observation b2146111-85f9-44e2-b22e-f2e110c592eb · outbound

This paper cites Towards A Rigorous Science of Interpretable Machine Learning.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Towards A Rigorous Science of Interpretable Machine Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:14.667487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:14.667487Z digest=sha256:c31412f77d2534f243bf08369a8afee7f7572a06c9ecb5cb47638d0ff82c1be1

Observation afe9ad33-06bc-464b-97c0-6a8a5d113ab8 · outbound

This paper cites Eberhardt, Phillip Atiba Goff, Valerie J.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Eberhardt, Phillip Atiba Goff, Valerie J

Reference 14

Resolution
verified exact
doi, observed 2026-08-07T12:13:20.812539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:13:14.750088Z digest=sha256:617edf0760bce46f91adb25a43b189e067c1a744c723c8b89239f0a0086c1abb

Observation 58b48d5e-0cb3-401c-a28c-3804aefd0595 · outbound

This paper cites Causal Abstractions of Neural Networks.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Causal Abstractions of Neural Networks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:14.859364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:14.859364Z digest=sha256:786f55b1a6296b7b9b45485d1ffd73f8e47aaa1325f90aaf7a90f327dff925d2

Observation 78865775-a7b6-4d8b-a38d-51be5ae4aba3 · outbound

This paper cites Parameter-Efficient Fine-Tuning of LLaMA for the Clinical Domain.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Parameter-Efficient Fine-Tuning of LLaMA for the Clinical Domain

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:15.219580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:15.219580Z digest=sha256:c13da24aa16ba8dc5dc3c12eaf88cb5a01e596816a6d00bf9a560469c9491497

Observation 4586dff6-61a8-4446-bff3-2bb97f195644 · outbound

This paper cites Dissecting Recall of Factual Associations in Auto-Regressive Language Models.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Dissecting Recall of Factual Associations in Auto-Regressive Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:16.064697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:16.064697Z digest=sha256:f2fd441390fc7702320fa9b7e553c7a66b419ddecee9a585245ac46d8feefb93

Observation df06e2fa-0b1f-4ae8-969c-f9f447510e4a · outbound

This paper cites Patchscopes: A Unifying Framework for Inspecting Hidden Representations of Language Models.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Patchscopes: A Unifying Framework for Inspecting Hidden Representations of Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:16.597821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:16.597821Z digest=sha256:1a33ed092cec29c112a418356205241d28785ea5c54a7486ccd5b1eff50fe3e5

Observation ff1a6bf1-8d9a-4d56-a162-33a4b15fe7ab · outbound

This paper cites Greenwald and Mahzarin R.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Greenwald and Mahzarin R

Reference 19

Resolution
verified exact
doi, observed 2026-08-07T12:13:20.659024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:13:16.686678Z digest=sha256:ae2e55f00f67a14964fe05013dd5270ceda23628d7c9390e84c1fea28e584151

Observation 35e296b4-9900-4eca-ad87-17e749474d98 · outbound

This paper cites Greenwald, Debbie E.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Greenwald, Debbie E

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:16.789954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:16.789954Z digest=sha256:0b0bc564ea1d9fa2519c23dda2312aff6a4a7e948e7d79dbd1b22492a4f66c91

Observation 224ed99f-e562-4cb1-b849-0c55c27b0dff · outbound

This paper cites Model Editing Harms General Abilities of Large Language Models: Regularization to the Rescue.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Model Editing Harms General Abilities of Large Language Models: Regularization to the Rescue

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:16.945425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:16.945425Z digest=sha256:77e9cbc6f3924ef36cf0725c04edc43032737cbf42a12da06ed187bc11e4f729

Observation d49ceaf5-1daa-4fd0-a90d-dba6ef10d58f · outbound

This paper cites Language Models Represent Space and Time.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Language Models Represent Space and Time

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:17.067308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:17.067308Z digest=sha256:b9211fcc3a237821f71917b0c2e08144fff9c3469d621dde2867b4651439bd52

Observation b85fd806-9985-45e6-81e9-159e2ab2f9e0 · outbound

This paper cites How does GPT-2 compute greater-than?: Interpreting mathematical abilities in a pre-trained language model.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race How does GPT-2 compute greater-than?: Interpreting mathematical abilities in a pre-trained language model

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:17.160012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:17.160012Z digest=sha256:03cd77f33b8c06c788910302de1bbfae9781473ee27dff4dadeb4d676bc4e82e

Observation d2637206-1289-405a-96cc-507c0d022e51 · outbound

This paper cites How to use and interpret activation patching.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race How to use and interpret activation patching

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:17.249944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:17.249944Z digest=sha256:7838de02dd7c768f80bc5176395ab6d16cdb7f6fa4bd26b6eb09fc1974e14900

Observation b4570df3-89e4-456c-a36b-1130f1215455 · outbound

This paper cites an unresolved cited work.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Unresolved cited work

Reference 25

Resolution
verified exact
doi, observed 2026-08-07T12:13:20.549384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:13:17.354469Z digest=sha256:c38d2595e48c08ccfa6151419a8059eeb3f04bc3c11c30a0cbe2e21ede032c57

Observation 3acbc310-04b9-4d3a-8ac8-f5ef511f3555 · outbound

This paper cites Generative Models as a Complex Systems Science: How can we make sense of large language model behavior?.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Generative Models as a Complex Systems Science: How can we make sense of large language model behavior?

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:17.479369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:17.479369Z digest=sha256:7a2ca38eed711dfb55374ebff6ee048fd5ae767a294d6a4882e891ff89c3c65f

Observation 480c2b11-832e-45c7-b6fd-5f6b41f0b59d · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race LoRA: Low-Rank Adaptation of Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:17.522209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:17.522209Z digest=sha256:0c60247f72d9c826b4f7f9eea8fb65151f0995c93f71acfcb117115e23b8e500

Observation 243aac18-e9b7-4429-92a4-baab3b8899e9 · outbound

This paper cites Auxiliary task demands mask the capabilities of smaller language models.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Auxiliary task demands mask the capabilities of smaller language models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:17.581753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:17.581753Z digest=sha256:d9fd3d8c56988213bf4456a6862f91e62ff859da37033d40c4929b5f93d7f377

Observation 37a768e2-1405-43ac-9b6f-612938c3823b · outbound

This paper cites Interpretability Beyond Feature Attribution: Quantitative Testing with Concept Activation Vectors (TCAV).

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Interpretability Beyond Feature Attribution: Quantitative Testing with Concept Activation Vectors (TCAV)

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:17.616483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:17.616483Z digest=sha256:3f6e4b42e9a6eba811fd1105f150620f39c790bbf9e7d752754caa97cddd6bcd

Observation 334c5b38-8c20-4ff2-9902-768826231db7 · outbound

This paper cites Investigating Implicit Bias in Large Language Models: A Large-Scale Study of Over 50 LLMs.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Investigating Implicit Bias in Large Language Models: A Large-Scale Study of Over 50 LLMs

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:13:22.496299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:13:17.673015Z digest=sha256:56e8aae225bf65471ab1cd381881f9a962f33ed1f3dd5dea19449d00dbc9616c

Observation e8c1a24c-31d4-49f9-8c43-df6bbe522c10 · outbound

This paper cites A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:17.735447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:17.735447Z digest=sha256:4b98140a5ae5d98d5c442d284955a7f796105e2ab03981c941306b4e360ffa14

Observation c8fc1369-dbc1-4c58-9090-afafa544b429 · outbound

This paper cites Levinson, Huajian Cai, and Danielle Young.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Levinson, Huajian Cai, and Danielle Young

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:13:23.767619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:13:17.789820Z digest=sha256:f59908e780889c720529a45c5a6507ca5e42609ec5086ad286c519894c2daf1e

Observation 66717e8c-5316-418f-89fe-dd5ecaf976b1 · outbound

This paper cites Inference-Time Intervention: Eliciting Truthful Answers from a Language Model.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Inference-Time Intervention: Eliciting Truthful Answers from a Language Model

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:17.840290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:17.840290Z digest=sha256:17317cb5cd0fbd2f5add1ac7ecd03bb02c06dbbe99948227f4d9806cc0c41326

Observation a41f2c33-8fd2-409f-bcd0-ee0654ea817c · outbound

This paper cites Revisiting Jailbreaking for Large Language Models: A Representation Engineering Perspective.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Revisiting Jailbreaking for Large Language Models: A Representation Engineering Perspective

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:17.876984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:17.876984Z digest=sha256:f5b745c57a8f9183e36b38aaab89d9cf88374d62ff68f29662708a71abfd265c

Observation 2a9d59b5-2f6e-49be-a7f3-b9a841bab127 · outbound

This paper cites Label Supervised LLaMA Finetuning.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Label Supervised LLaMA Finetuning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:17.913243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:17.913243Z digest=sha256:4296cc00b5860d373a137ad5ec5c771ce9076b72e81f5ce0b3e1a5f595f336b6

Observation 862d51d8-9230-4423-a187-bc9d42e4369f · outbound

This paper cites Does Circuit Analysis Interpretability Scale? Evidence from Multiple Choice Capabilities in Chinchilla.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Does Circuit Analysis Interpretability Scale? Evidence from Multiple Choice Capabilities in Chinchilla

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:17.953463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:17.953463Z digest=sha256:41d73af21c11ed868ecdbf66c6d59169b88711ebaf42417a000453ab4a3e0a6a

Observation 2a1acba3-eaf7-4cb5-be78-f6d35f2f4183 · outbound

This paper cites Assessing Dialect Fairness and Robustness of Large Language Models in Reasoning Tasks.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Assessing Dialect Fairness and Robustness of Large Language Models in Reasoning Tasks

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:17.999979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:17.999979Z digest=sha256:78302e9e75f576dda860aca5d39811623caca28b5c01852d9ddd0bf5499cd75f

Observation 1e5db672-887f-4e75-befc-cd4d69882a92 · outbound

This paper cites Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:18.050267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:18.050267Z digest=sha256:da6d5a2c2ea1457585fd00054890dc14b8709bfecbff5b4363a6e49e1c7c8c67

Observation e298aa2c-a53b-4891-9675-da7ad6e89c15 · outbound

This paper cites The Llama 3 Herd of Models.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race The Llama 3 Herd of Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:18.104681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:18.104681Z digest=sha256:4fdac4633c0efbf743969571f4bace6b8f1806fe01d651a572c1bb7b8694c6b9

Observation 4a7b237b-bf63-47e0-ad39-b82784fff478 · outbound

This paper cites Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:18.154269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:18.154269Z digest=sha256:2293fd90fa6b8f03081896c145b378f38068a21fc014d3c35d8a161dd6dae44a

Observation 147bb845-edff-422e-92ce-19d5829b3ec8 · outbound

This paper cites Locating and Editing Factual Associations in GPT.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Locating and Editing Factual Associations in GPT

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:18.187628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:18.187628Z digest=sha256:53a9b1601bb668f806ff9bb4103eeadcfaf7581e6d76ccfce7f4e29635a072e0

Observation aebf94c6-0dcc-4975-9a8d-b3ee07eb6fc7 · outbound

This paper cites Language Models Learn Rare Phenomena from Less Rare Phenomena: The Case of the Missing AANNs.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Language Models Learn Rare Phenomena from Less Rare Phenomena: The Case of the Missing AANNs

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:18.204599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:18.204599Z digest=sha256:8c6ad96be95ddc08c34563bb24e9f6175db578c25289736e3f99820e895e3b80

Observation cc38fca4-fe45-431c-9c57-0e03e2cae629 · outbound

This paper cites Progress measures for grokking via mechanistic interpretability.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Progress measures for grokking via mechanistic interpretability

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:18.242414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:18.242414Z digest=sha256:c695eac83d1eb97c187b5c2a0ec13f45918e32f1a245ddb054b3e139bdd0bd08

Observation c2dfeb36-b961-4fc3-8601-ecc849ada23a · outbound

This paper cites an unresolved cited work.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Unresolved cited work

Reference 44

Resolution
verified exact
doi, observed 2026-08-07T12:13:20.396510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:13:18.308267Z digest=sha256:889b405c18cc3670df1a18342f4d8547186e0ae5e83651e07a0214bcc33d79c2

Observation a4e02824-91e2-44eb-8315-2d05e6bbe815 · outbound

This paper cites Norton, Samuel R.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Norton, Samuel R

Reference 45

Resolution
verified exact
raw_fallback, observed 2026-08-07T12:13:22.193037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:13:18.353714Z digest=sha256:cd7036d6e24de7afbc1d6b1b25c80dc62976e7fb9e0ce0d03f07e62d7de37c22

Observation d5d56511-f8a6-4701-b478-93e5dfc9d302 · outbound

This paper cites Nosek, Anthony G.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Nosek, Anthony G

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:13:23.606008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:13:18.428035Z digest=sha256:440f37acd0cf22392fc9b4b88de65b9b3cc4c26b7144b665bd873b1722b9b210

Observation 663161dd-2dd9-4ce3-b2c0-a9d32f946e9e · outbound

This paper cites an unresolved cited work.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:13:23.461680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:13:18.453442Z digest=sha256:ebafc796606c1e07c297465c0f5e3a133ee54323b177e1a000a56c7a1b371377

Observation d7e20d09-c0fc-438d-8b05-7fd6ebaa2492 · outbound

This paper cites an unresolved cited work.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:13:23.346962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:13:18.494148Z digest=sha256:05efcbef00cbf4c5d20f116b676bac05a1bfbd9b77e0fcde45e4a1fa52e4e8b5

Observation f53dca34-7fcd-425f-add1-d764bb3a832f · outbound

This paper cites Steering Llama 2 via Contrastive Activation Addition.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Steering Llama 2 via Contrastive Activation Addition

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:18.522808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:18.522808Z digest=sha256:9f147443dd706dd71fe9cd75ef0af0d43216d85f671372a689e42b3b7bd5615d

Observation fa03f3aa-6535-4e1f-814a-d6d57c3fa12a · outbound

This paper cites BBQ: A Hand-Built Bias Benchmark for Question Answering.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race BBQ: A Hand-Built Bias Benchmark for Question Answering

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:18.554467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:18.554467Z digest=sha256:1813f6bce8939f44cc6627751435d385734cc0c13fa2785fee7bdd0c10467112

Observation 7f9206dc-82bb-4d97-bbb5-fddde5ac9efc · outbound

This paper cites an unresolved cited work.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Unresolved cited work

Reference 51

Resolution
verified exact
doi, observed 2026-08-07T12:13:20.248908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:13:18.586340Z digest=sha256:ba5bfc268a8025e553ed3b3201fd6223219ddcc5a42e9dfc1afdda67eb8bed42

Observation d18a9a1e-ae54-49ad-95ad-42e3b8f0955e · outbound

This paper cites Interpreting Bias in Large Language Models: A Feature-Based Approach.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Interpreting Bias in Large Language Models: A Feature-Based Approach

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:13:21.933327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:13:18.626848Z digest=sha256:865c9538e92facbb8054eab8a98a35f19cb910a5078827f62e3076873674fdf3

Observation bd3f5835-517f-4287-8a21-971b9f3e7c35 · outbound

This paper cites an unresolved cited work.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:13:23.195270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:13:18.673799Z digest=sha256:9698593dc47afc08b04dd3e8ea8ac288068be0dbbc8f7434370b0f76ef60d9c8

Observation 2d48be1c-cd16-4ea7-a872-ed5fd0e79589 · outbound

This paper cites IssueBench: Millions of Realistic Prompts for Measuring Issue Bias in LLM Writing Assistance.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race IssueBench: Millions of Realistic Prompts for Measuring Issue Bias in LLM Writing Assistance

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:18.714338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:18.714338Z digest=sha256:b50b0fba0b73ed0cefa998f8dc35c84d3e695d8e83b940f205ad7636608be038

Observation 9d141c4d-440b-4d20-8b30-a8982e54cb9d · outbound

This paper cites Efficient RLHF: Reducing the Memory Usage of PPO.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Efficient RLHF: Reducing the Memory Usage of PPO

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:18.743392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:18.743392Z digest=sha256:c5514759396bb6487c978b7cc3e7158db77fbee410414b157487b04c625d908c

Observation a58d3885-7902-467a-b295-3112b6003fad · outbound

This paper cites Parameter Efficient Reinforcement Learning from Human Feedback.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Parameter Efficient Reinforcement Learning from Human Feedback

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:18.772325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:18.772325Z digest=sha256:13a555edb9917daf3b58362075d387174e286e77da6d7fdc0cc57bef1a76f51e

Observation e5212170-b945-4119-b102-0a9a94a2f5bc · outbound

This paper cites Stevens, Victoria C.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Stevens, Victoria C

Reference 57

Resolution
verified exact
doi, observed 2026-08-07T12:13:20.147034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:13:18.804298Z digest=sha256:9531510c16c364782ee6e0c4dd1b17f8a65f177228d112391bb6e798043d355a

Observation 9657af22-7b98-451f-a8a8-63c6060afb66 · outbound

This paper cites Improving Instruction-Following in Language Models through Activation Steering.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Improving Instruction-Following in Language Models through Activation Steering

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:18.854878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:18.854878Z digest=sha256:90245494dc0d21f0415dca16b2f86c6c94523b5bb42d8c52d113028ae2271532

Observation d67a5ee7-ab89-4cbd-8a08-ebfbdb2eecfb · outbound

This paper cites Exploring the impact of low-rank adaptation on the performance, efficiency, and regularization of RLHF.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Exploring the impact of low-rank adaptation on the performance, efficiency, and regularization of RLHF

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:18.882947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:18.882947Z digest=sha256:06852610bbb9ba0557804e345e71f25eac224688e1829ebe0a0ac41282fa5b5b

Observation 38452e32-1802-4a5e-b743-99108b95aee0 · outbound

This paper cites SuryaKiran at MEDIQA-Sum 2023: Leveraging LoRA for Clinical Dialogue Summarization.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race SuryaKiran at MEDIQA-Sum 2023: Leveraging LoRA for Clinical Dialogue Summarization

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:13:21.647590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:13:18.924741Z digest=sha256:78c1874819f529d0a9b4b8d93ade6a72bd75f3a922a021fc5d2b981ff232b59e

Observation 68ee41b9-5b47-418d-b855-3d6afbe5dc61 · outbound

This paper cites Evaluating and Mitigating Discrimination in Language Model Decisions.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Evaluating and Mitigating Discrimination in Language Model Decisions

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:18.979922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:18.979922Z digest=sha256:3b3ab7cd56a48a3521311db0f73eff4b73b4c22d061b9f70b7d49778208350b6

Observation 2e7df569-9d9b-41c6-9fa8-1a2fe05f9e2e · outbound

This paper cites Clinical Camel: An Open Expert-Level Medical Language Model with Dialogue-Based Knowledge Encoding.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Clinical Camel: An Open Expert-Level Medical Language Model with Dialogue-Based Knowledge Encoding

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:19.036340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:19.036340Z digest=sha256:1cc250fef0b9a3f5de13ffdb9ce30039fa6029c5ced234f7902c326b980af5ac

Observation 33d965b8-0bd6-42eb-8d4c-014e25b9a186 · outbound

This paper cites Steering Language Models With Activation Engineering.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Steering Language Models With Activation Engineering

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:19.106538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:19.106538Z digest=sha256:4297e2e458fbe3768fde7cf116c6c2ed0e3c100c99feca671ce07dc510e0c4b2

Observation 2274232a-1bde-4b51-a43f-f0d4ec7d6e0f · outbound

This paper cites an unresolved cited work.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:13:23.082467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:13:19.178953Z digest=sha256:d28ccc477bbf1ae1b42368edb2566a699d2b049c2324aa688f4e4ac975485e3d

Observation 5724ed5a-dc59-49c0-9876-5279146e77aa · outbound

This paper cites DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:19.236371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:19.236371Z digest=sha256:058c49a3bc8d4c9cfc4862bb24ff89a1c1ba0d6a9e020862bb74d03075994214

Observation 25085deb-1076-4066-a0b5-1f3681d0afe1 · outbound

This paper cites Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:19.352932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:19.352932Z digest=sha256:e6fbb560f215bbd48e03261915dba8ebed8b2c7f79243b1b48a54c7249288342

Observation a9b299b5-3400-43f7-846d-22f590859600 · outbound

This paper cites Jailbroken: How Does LLM Safety Training Fail?.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Jailbroken: How Does LLM Safety Training Fail?

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:19.432193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:19.432193Z digest=sha256:49417365e54c3fbabb098319c9c3fbc2472f720dfa6b6b2684fa56e4d2c3f11e

Observation 5e43b040-e749-4864-b1e7-84ea0d259d51 · outbound

This paper cites Fundamental Limitations of Alignment in Large Language Models.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Fundamental Limitations of Alignment in Large Language Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:19.496083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:19.496083Z digest=sha256:2eed1a84394b4a987a8ec84d25ffe7eae9f5c8b6fbb99f664395d2261e242c6f

Observation 23e4f887-4e7a-45ad-ad63-9c1688b592c9 · outbound

This paper cites AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:19.564287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:19.564287Z digest=sha256:f3db98ec4f6b39abd49a1d2a30b407f16cb9c829f4c148279b973ca940e59f82

Observation 5864e4db-8df3-4dce-830e-4957a71edba1 · outbound

This paper cites Uncovering Safety Risks of Large Language Models through Concept Activation Vector.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Uncovering Safety Risks of Large Language Models through Concept Activation Vector

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:19.638500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:19.638500Z digest=sha256:2be050c923450bc3b28b7b82fbb3601eb73c670d618cfdfccfc16dbf1e67891b

Observation 2a13aa06-7e8b-4911-9eec-797247b2bfa4 · outbound

This paper cites AutoRE: Document-Level Relation Extraction with Large Language Models.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race AutoRE: Document-Level Relation Extraction with Large Language Models

Reference 71

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:13:21.360912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:13:19.714890Z digest=sha256:c242294016f60214e88dd536a6f82bc891000fa6023612e980db072ffc3beddf

Observation 94cf689b-5052-4457-8597-023c50383804 · outbound

This paper cites Understanding and Mitigating Gender Bias in LLMs via Interpretable Neuron Editing.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Understanding and Mitigating Gender Bias in LLMs via Interpretable Neuron Editing

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:19.780988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:19.780988Z digest=sha256:4dcaa6d290e4f0d56795d858b1a218c53110f96c9c9c1704fb5a90749150ef18

Observation 60034b32-2269-4448-a321-ae31c92f7089 · outbound

This paper cites Towards Best Practices of Activation Patching in Language Models: Metrics and Methods.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Towards Best Practices of Activation Patching in Language Models: Metrics and Methods

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:19.843762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:19.843762Z digest=sha256:2ef18082afeaba57f3993e13c6286ecf2b332544c3bed05b3eb89af9eff42205

Observation 6a303adf-c8b0-47a8-90e1-7d81d48f201d · outbound

This paper cites DialogueLLM: Context and Emotion Knowledge-Tuned Large Language Models for Emotion Recognition in Conversations.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race DialogueLLM: Context and Emotion Knowledge-Tuned Large Language Models for Emotion Recognition in Conversations

Reference 74

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:13:21.217774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:13:19.901587Z digest=sha256:451f412d7c6610d4c69d1e0f1f958e343c677d9ef40537e9a064822d7d987bb3

Observation 24a1d6c1-3256-4b83-bf4b-bf345a072fad · outbound

This paper cites The Clock and the Pizza: Two Stories in Mechanistic Explanation of Neural Networks.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race The Clock and the Pizza: Two Stories in Mechanistic Explanation of Neural Networks

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:19.962960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:19.962960Z digest=sha256:db79d4bb4fc818c3ab6258f052f067db567e4ab134a51794ec72d498e7bbbe25

Observation c4d24a5c-c423-4b35-bcbe-afa331c8565e · outbound

This paper cites online" 'onlinestring :=.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race online" 'onlinestring :=

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:20.045313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:20.045313Z digest=sha256:9f851d9c708c10c84f90d7faf8d913ac143c3485d4b3ae1a615719b25200dfde

Observation 76c487d5-46d7-409a-9876-765a1bab2955 · outbound

This paper cites write newline.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race write newline

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:20.054591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:20.054591Z digest=sha256:e551761c8a175bf56c12fd9a2ab7dfac3187af6b326e897a7f45cdae4185c8ca

Pith citing papers

No inbound Pith citation observations are available.