Pith. sign in

Paper Citation Record · LEDGER

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race

As of 7 August 2026, this Paper Citation Record lists 77 of 77 outbound references and 0 inbound Pith citation observations for arXiv:2506.00253.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.00253 v3

Coverage vector

measured 77 of 77 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:13:20.054591Z

measured 77 of 77 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

77 of 77 outbound references displayed

  • verified exact14
  • verified fuzzy2
  • unresolved61
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 584af120-bfe6-40c1-b645-8947fdc8da08 · outbound

This paper cites Physics of Language Models: Part 3.1, Knowledge Storage and Extraction.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Physics of Language Models: Part 3.1, Knowledge Storage and Extraction

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:13.741659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:13.741659Z digest=sha256:c044d075b1e20156057705e8da31f41a2faa552527ec89cf72fd28088111ae6b

Observation 6693e333-a4bc-4cd5-aab8-f4c3287e2506 · outbound

This paper cites Apfelbaum, Michael I.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Apfelbaum, Michael I

Reference 2

Resolution
verified exact
doi, observed 2026-08-07T12:13:21.009857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:13:13.794504Z digest=sha256:002f7ea602740dd10ca80181aa6ee30e75441a8464ec514f6ef838345650bcf1

Observation 6959ccea-181f-4160-a58e-67650f41031f · outbound

This paper cites Griffiths.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Griffiths

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:13.852499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:13.852499Z digest=sha256:98ac402c985e5d3c38e2f3db5212d277c0409bb8dfe2cbc2e78045cca3c58ff9

Observation 79169ae5-5c3e-4610-9b62-d5849c130c93 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Constitutional AI: Harmlessness from AI Feedback

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:13.935489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:13.935489Z digest=sha256:3ad21d46b7869b6543c1d5e00e9c604c7205be3a6bfe4197c9c05eb8837be181

Observation dee7fa67-5010-44a5-a867-7fc4e4cf00ed · outbound

This paper cites Eliciting Latent Predictions from Transformers with the Tuned Lens.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Eliciting Latent Predictions from Transformers with the Tuned Lens

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:14.003675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:14.003675Z digest=sha256:7113511bf996d92aa01c69f3a77068dcafa703ab8a68a57bdf8878188f24b4a8

Observation 25500e5a-d219-4d35-a161-615a1926a79d · outbound

This paper cites Mechanistic Interpretability for AI Safety -- A Review.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Mechanistic Interpretability for AI Safety -- A Review

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:14.094600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:14.094600Z digest=sha256:379fbcbed5677da6bf54481021da93da28ef2e5aa165f830c8b3eb3f70436903

Observation b9021306-b3f1-4f09-a292-b7889ada3133 · outbound

This paper cites an unresolved cited work.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:13:24.122596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:13:14.161233Z digest=sha256:5c789af8ae780e6dd099296636ec90adcdc39264ec1bd33d68e1c68a55b526ff

Observation ec3a1bb6-f211-4382-a1a5-f523cad9fb58 · outbound

This paper cites an unresolved cited work.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:13:23.955589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:13:14.260481Z digest=sha256:03f6038298b82bf63e8b6f74865bbcaa0707f668c3703519ab19a747bd34a96a

Observation b587f0b4-5934-44b7-83d5-8e94695885a6 · outbound

This paper cites Bryson, and Arvind Narayanan.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Bryson, and Arvind Narayanan

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:14.350152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:14.350152Z digest=sha256:b4546bae439b3cdf0cef781d9da7d692851b453d73a2deeb1b2783acb3ba0056

Observation 00cfff6c-03a5-4672-a484-d602f0ba15bd · outbound

This paper cites SelfIE: Self-Interpretation of Large Language Model Embeddings.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race SelfIE: Self-Interpretation of Large Language Model Embeddings

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:14.424151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:14.424151Z digest=sha256:23bcb6dcbef9f1dd19d0da32448272370bf86bd9ca28f4d39f842c128c1d99c9

Observation 38c54000-6257-4f66-ab11-430cc0e4cb18 · outbound

This paper cites an unresolved cited work.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Unresolved cited work

Reference 11

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-07T12:13:22.867513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:13:14.501115Z digest=sha256:7aa22c6599ad6322253c88f3659545d8d20bd8749f2a2554e027e56c04e3e7ff

Observation 08e949fc-017d-46d7-a534-d69180fd89ee · outbound

This paper cites Mitigating Social Biases in Language Models through Unlearning.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Mitigating Social Biases in Language Models through Unlearning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:14.600807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:14.600807Z digest=sha256:8982edd48b13275e2c1018c3e75cd7538c23514d7f73b01312ccc8f2fad6191e

Observation b2146111-85f9-44e2-b22e-f2e110c592eb · outbound

This paper cites Towards A Rigorous Science of Interpretable Machine Learning.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Towards A Rigorous Science of Interpretable Machine Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:14.667487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:14.667487Z digest=sha256:9827482b0865b039c6c9fe6c3c49d893c8544304f98b9e01e6769ee8c73050b3

Observation afe9ad33-06bc-464b-97c0-6a8a5d113ab8 · outbound

This paper cites Eberhardt, Phillip Atiba Goff, Valerie J.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Eberhardt, Phillip Atiba Goff, Valerie J

Reference 14

Resolution
verified exact
doi, observed 2026-08-07T12:13:20.812539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:13:14.750088Z digest=sha256:bffb00477e27e8f3328f51f438ed9094583011db508cc64165b26caaae6998fa

Observation 58b48d5e-0cb3-401c-a28c-3804aefd0595 · outbound

This paper cites Causal Abstractions of Neural Networks.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Causal Abstractions of Neural Networks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:14.859364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:14.859364Z digest=sha256:a5b0b24a6abebf90c1459f2345b5545f4385b8c81e6dd4600b1dc1b6d6193183

Observation 78865775-a7b6-4d8b-a38d-51be5ae4aba3 · outbound

This paper cites Parameter-Efficient Fine-Tuning of LLaMA for the Clinical Domain.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Parameter-Efficient Fine-Tuning of LLaMA for the Clinical Domain

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:15.219580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:15.219580Z digest=sha256:b96ec182a0ca424deaa1af9f2dd5fb032fbcbccbe072219422a3e44848d3ea91

Observation 4586dff6-61a8-4446-bff3-2bb97f195644 · outbound

This paper cites Dissecting Recall of Factual Associations in Auto-Regressive Language Models.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Dissecting Recall of Factual Associations in Auto-Regressive Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:16.064697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:16.064697Z digest=sha256:02ecc8513524705607be46373ced5da89b577defe100eadac5e3459c89e4b0f2

Observation df06e2fa-0b1f-4ae8-969c-f9f447510e4a · outbound

This paper cites Patchscopes: A Unifying Framework for Inspecting Hidden Representations of Language Models.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Patchscopes: A Unifying Framework for Inspecting Hidden Representations of Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:16.597821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:16.597821Z digest=sha256:f7e6c47ef4e05052b1004a173931b4ef9be3253e0b882bac83529c7b6728b8a5

Observation ff1a6bf1-8d9a-4d56-a162-33a4b15fe7ab · outbound

This paper cites Greenwald and Mahzarin R.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Greenwald and Mahzarin R

Reference 19

Resolution
verified exact
doi, observed 2026-08-07T12:13:20.659024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:13:16.686678Z digest=sha256:9ee4975e149e90c87f28cade00062f1347114f59c520d558f690360302aa299f

Observation 35e296b4-9900-4eca-ad87-17e749474d98 · outbound

This paper cites Greenwald, Debbie E.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Greenwald, Debbie E

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:16.789954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:16.789954Z digest=sha256:4c8c28b846bbed7381db059ba6d95e83347988f75ebd46b51fc8fbe5135ab68b

Observation 224ed99f-e562-4cb1-b849-0c55c27b0dff · outbound

This paper cites Model Editing Harms General Abilities of Large Language Models: Regularization to the Rescue.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Model Editing Harms General Abilities of Large Language Models: Regularization to the Rescue

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:16.945425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:16.945425Z digest=sha256:f973560c151a3301d4aa216750b1992a9ec785333525648a4b667cac997e1af4

Observation d49ceaf5-1daa-4fd0-a90d-dba6ef10d58f · outbound

This paper cites Language Models Represent Space and Time.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Language Models Represent Space and Time

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:17.067308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:17.067308Z digest=sha256:e30d5655fdb950966a9c7e9a52e3b040765d9d73e0ce743a62ac623ca5a13313

Observation b85fd806-9985-45e6-81e9-159e2ab2f9e0 · outbound

This paper cites How does GPT-2 compute greater-than?: Interpreting mathematical abilities in a pre-trained language model.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race How does GPT-2 compute greater-than?: Interpreting mathematical abilities in a pre-trained language model

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:17.160012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:17.160012Z digest=sha256:6703c4c4eaaba078f914c361e39df65a15341f017f936f23e470962f76614e29

Observation d2637206-1289-405a-96cc-507c0d022e51 · outbound

This paper cites How to use and interpret activation patching.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race How to use and interpret activation patching

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:17.249944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:17.249944Z digest=sha256:87bf7a32d66a1bf0e8bf36c8534669206bacc8aeeb6c2fac07e72ba4a59ac0b5

Observation b4570df3-89e4-456c-a36b-1130f1215455 · outbound

This paper cites an unresolved cited work.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Unresolved cited work

Reference 25

Resolution
verified exact
doi, observed 2026-08-07T12:13:20.549384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:13:17.354469Z digest=sha256:9072a5de80b6f1a62bcd6a7f3a3069a4bd04320a3b510555bd24fed431619c77

Observation 3acbc310-04b9-4d3a-8ac8-f5ef511f3555 · outbound

This paper cites Generative Models as a Complex Systems Science: How can we make sense of large language model behavior?.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Generative Models as a Complex Systems Science: How can we make sense of large language model behavior?

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:17.479369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:17.479369Z digest=sha256:ad3d3a306b2ee4eccdc4dcd206c7c51d8b1a93141fc194798a97627ca9a6322b

Observation 480c2b11-832e-45c7-b6fd-5f6b41f0b59d · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race LoRA: Low-Rank Adaptation of Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:17.522209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:17.522209Z digest=sha256:341055b1214de24d09d007cf914cee07101aa81459793dcc545742c75048816c

Observation 243aac18-e9b7-4429-92a4-baab3b8899e9 · outbound

This paper cites Auxiliary task demands mask the capabilities of smaller language models.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Auxiliary task demands mask the capabilities of smaller language models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:17.581753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:17.581753Z digest=sha256:f9c341c55218c198ad7a5fc2b99272282e7d557f0a42da97835d0b0513ccc105

Observation 37a768e2-1405-43ac-9b6f-612938c3823b · outbound

This paper cites Interpretability Beyond Feature Attribution: Quantitative Testing with Concept Activation Vectors (TCAV).

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Interpretability Beyond Feature Attribution: Quantitative Testing with Concept Activation Vectors (TCAV)

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:17.616483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:17.616483Z digest=sha256:bae1a4f6e4184d05edafdb39430545476f928753fdc0ed43987c4ce01a4bbdd9

Observation 334c5b38-8c20-4ff2-9902-768826231db7 · outbound

This paper cites Investigating Implicit Bias in Large Language Models: A Large-Scale Study of Over 50 LLMs.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Investigating Implicit Bias in Large Language Models: A Large-Scale Study of Over 50 LLMs

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:13:22.496299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:13:17.673015Z digest=sha256:767232bfdf209d3430e23d8dda9e310573bce02dc6ddfce87de3f31c03de2e3e

Observation e8c1a24c-31d4-49f9-8c43-df6bbe522c10 · outbound

This paper cites A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:17.735447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:17.735447Z digest=sha256:95762c08b86355b7ec8dde119bd269403060974218fef9d0b5db56c025d68f8b

Observation c8fc1369-dbc1-4c58-9090-afafa544b429 · outbound

This paper cites Levinson, Huajian Cai, and Danielle Young.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Levinson, Huajian Cai, and Danielle Young

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:13:23.767619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:13:17.789820Z digest=sha256:f7890698798ca8002a038ab5ddb4fc41084dee3078f529c7b30d5acfbdd139e5

Observation 66717e8c-5316-418f-89fe-dd5ecaf976b1 · outbound

This paper cites Inference-Time Intervention: Eliciting Truthful Answers from a Language Model.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Inference-Time Intervention: Eliciting Truthful Answers from a Language Model

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:17.840290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:17.840290Z digest=sha256:cfae886a56ff0a0710eaa76ade22f3c1f539e06f6bcbfc8c2aa37d36fca6a346

Observation a41f2c33-8fd2-409f-bcd0-ee0654ea817c · outbound

This paper cites Revisiting Jailbreaking for Large Language Models: A Representation Engineering Perspective.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Revisiting Jailbreaking for Large Language Models: A Representation Engineering Perspective

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:17.876984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:17.876984Z digest=sha256:0214c98f17c37982be106eec9c50dd4a92bcb2a399e2bdd0664f988337cc8b2d

Observation 2a9d59b5-2f6e-49be-a7f3-b9a841bab127 · outbound

This paper cites Label Supervised LLaMA Finetuning.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Label Supervised LLaMA Finetuning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:17.913243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:17.913243Z digest=sha256:7a47dfed65eda9bda6b9cecdc307bd4f9a515c04e429eb60167a9c703ec4bae1

Observation 862d51d8-9230-4423-a187-bc9d42e4369f · outbound

This paper cites Does Circuit Analysis Interpretability Scale? Evidence from Multiple Choice Capabilities in Chinchilla.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Does Circuit Analysis Interpretability Scale? Evidence from Multiple Choice Capabilities in Chinchilla

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:17.953463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:17.953463Z digest=sha256:98d74926772ef4024f6c1992d83376bc67e95cea5c9cd7eab878ffda165cd14f

Observation 2a1acba3-eaf7-4cb5-be78-f6d35f2f4183 · outbound

This paper cites Assessing Dialect Fairness and Robustness of Large Language Models in Reasoning Tasks.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Assessing Dialect Fairness and Robustness of Large Language Models in Reasoning Tasks

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:17.999979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:17.999979Z digest=sha256:76a60130cc4f5508fcfc7152a828e1d4c941a68d3d95a871a4dc2da7156e634b

Observation 1e5db672-887f-4e75-befc-cd4d69882a92 · outbound

This paper cites Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:18.050267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:18.050267Z digest=sha256:fe3a80afd17ccf84bd219045a1d3d3aa7acef99f53f0b28a9765dd1ad851fd54

Observation e298aa2c-a53b-4891-9675-da7ad6e89c15 · outbound

This paper cites The Llama 3 Herd of Models.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race The Llama 3 Herd of Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:18.104681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:18.104681Z digest=sha256:4d16138030c06585e1a1be57333f764111fbe999807234603c08d80a5350df99

Observation 4a7b237b-bf63-47e0-ad39-b82784fff478 · outbound

This paper cites Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:18.154269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:18.154269Z digest=sha256:2e383108be62bfdc92bd1bd8dc0c931786fde8f316a4b61c90e0f2b1338ad9c1

Observation 147bb845-edff-422e-92ce-19d5829b3ec8 · outbound

This paper cites Locating and Editing Factual Associations in GPT.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Locating and Editing Factual Associations in GPT

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:18.187628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:18.187628Z digest=sha256:da0aa2da9e5818cdd74afbcb44e6de61dc85215a041a7c97424ac85654a22d7c

Observation aebf94c6-0dcc-4975-9a8d-b3ee07eb6fc7 · outbound

This paper cites Language Models Learn Rare Phenomena from Less Rare Phenomena: The Case of the Missing AANNs.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Language Models Learn Rare Phenomena from Less Rare Phenomena: The Case of the Missing AANNs

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:18.204599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:18.204599Z digest=sha256:1a8fed158874f397c69f61cf9262b5c863f1dafba71edda8326e1770c268bb2e

Observation cc38fca4-fe45-431c-9c57-0e03e2cae629 · outbound

This paper cites Progress measures for grokking via mechanistic interpretability.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Progress measures for grokking via mechanistic interpretability

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:18.242414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:18.242414Z digest=sha256:18a06874867b9c81b491dbc7f481400b92da3fa2cdd53d9213f323847213e016

Observation c2dfeb36-b961-4fc3-8601-ecc849ada23a · outbound

This paper cites an unresolved cited work.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Unresolved cited work

Reference 44

Resolution
verified exact
doi, observed 2026-08-07T12:13:20.396510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:13:18.308267Z digest=sha256:c44cd30c4cbc75c237418314ce306628e466bdf86ac45d7db9f2fc54e1f647a2

Observation a4e02824-91e2-44eb-8315-2d05e6bbe815 · outbound

This paper cites Norton, Samuel R.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Norton, Samuel R

Reference 45

Resolution
verified exact
raw_fallback, observed 2026-08-07T12:13:22.193037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:13:18.353714Z digest=sha256:a2e479b833fc976e85d0385b9b17837778a03fcd2b9702782a8128a1aae1df8f

Observation d5d56511-f8a6-4701-b478-93e5dfc9d302 · outbound

This paper cites Nosek, Anthony G.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Nosek, Anthony G

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:13:23.606008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:13:18.428035Z digest=sha256:925bca7ba47327f10ceff59069f25ec4e0b1c75b508d539929847a2d6b35575b

Observation 663161dd-2dd9-4ce3-b2c0-a9d32f946e9e · outbound

This paper cites an unresolved cited work.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:13:23.461680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:13:18.453442Z digest=sha256:c7243e143bfa67a50eae692c3f29d301911b991343e14ebc01ebb58c043fc449

Observation d7e20d09-c0fc-438d-8b05-7fd6ebaa2492 · outbound

This paper cites an unresolved cited work.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:13:23.346962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:13:18.494148Z digest=sha256:d0a84e1adbe23a11102bac8ee5715bbe420a8994d1508985ca3af6dbb1c51716

Observation f53dca34-7fcd-425f-add1-d764bb3a832f · outbound

This paper cites Steering Llama 2 via Contrastive Activation Addition.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Steering Llama 2 via Contrastive Activation Addition

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:18.522808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:18.522808Z digest=sha256:1532da3dcf4e3e20f20b730e4e07d61db467da96d3b2e9bc58a555bf0f3803be

Observation fa03f3aa-6535-4e1f-814a-d6d57c3fa12a · outbound

This paper cites BBQ: A Hand-Built Bias Benchmark for Question Answering.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race BBQ: A Hand-Built Bias Benchmark for Question Answering

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:18.554467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:18.554467Z digest=sha256:7de15fbdde27b476c4b9c077f3e61915f90c2d7e201c4a0fe526a9f068809b36

Observation 7f9206dc-82bb-4d97-bbb5-fddde5ac9efc · outbound

This paper cites an unresolved cited work.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Unresolved cited work

Reference 51

Resolution
verified exact
doi, observed 2026-08-07T12:13:20.248908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:13:18.586340Z digest=sha256:56b05ae56c303887a4708363c7fbb14aa537566626355e2a891510af61d59952

Observation d18a9a1e-ae54-49ad-95ad-42e3b8f0955e · outbound

This paper cites Interpreting Bias in Large Language Models: A Feature-Based Approach.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Interpreting Bias in Large Language Models: A Feature-Based Approach

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:13:21.933327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:13:18.626848Z digest=sha256:e1f7ed5f78d45ba31e1963ef86b1b7d5d0e14c1ee087775e97432d491fe6da76

Observation bd3f5835-517f-4287-8a21-971b9f3e7c35 · outbound

This paper cites an unresolved cited work.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:13:23.195270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:13:18.673799Z digest=sha256:a3c0d561f6b5af3e1b657d8b6b4bfae728a79b61f597410f69023c092e0b5ee6

Observation 2d48be1c-cd16-4ea7-a872-ed5fd0e79589 · outbound

This paper cites IssueBench: Millions of Realistic Prompts for Measuring Issue Bias in LLM Writing Assistance.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race IssueBench: Millions of Realistic Prompts for Measuring Issue Bias in LLM Writing Assistance

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:18.714338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:18.714338Z digest=sha256:bf0b2d427d68a08e060536c504b1ed5e9b9d25946a25b0a4f5ab9189b758f594

Observation 9d141c4d-440b-4d20-8b30-a8982e54cb9d · outbound

This paper cites Efficient RLHF: Reducing the Memory Usage of PPO.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Efficient RLHF: Reducing the Memory Usage of PPO

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:18.743392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:18.743392Z digest=sha256:a1096b3433cbb50577cb2526d9c9de05c599ca6354907828d35d40812d492bb9

Observation a58d3885-7902-467a-b295-3112b6003fad · outbound

This paper cites Parameter Efficient Reinforcement Learning from Human Feedback.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Parameter Efficient Reinforcement Learning from Human Feedback

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:18.772325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:18.772325Z digest=sha256:567accfb3cf35c2416a8a5cad4cf1469e76fa6c00fa23fe2c2bfca3653245c2d

Observation e5212170-b945-4119-b102-0a9a94a2f5bc · outbound

This paper cites Stevens, Victoria C.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Stevens, Victoria C

Reference 57

Resolution
verified exact
doi, observed 2026-08-07T12:13:20.147034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:13:18.804298Z digest=sha256:d26908649b9687acfeabf06a1166f9579ce50733861ee6652ef9d7a7007ad62e

Observation 9657af22-7b98-451f-a8a8-63c6060afb66 · outbound

This paper cites Improving Instruction-Following in Language Models through Activation Steering.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Improving Instruction-Following in Language Models through Activation Steering

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:18.854878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:18.854878Z digest=sha256:c6c854bce1903f8877cbb16f24e8c0b098e914d33aa8f41a1ef2b8a216fbc243

Observation d67a5ee7-ab89-4cbd-8a08-ebfbdb2eecfb · outbound

This paper cites Exploring the impact of low-rank adaptation on the performance, efficiency, and regularization of RLHF.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Exploring the impact of low-rank adaptation on the performance, efficiency, and regularization of RLHF

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:18.882947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:18.882947Z digest=sha256:a840656f4b0c74243d1db5417466919d00a8f47afa6b8e705636f3360dbaa24e

Observation 38452e32-1802-4a5e-b743-99108b95aee0 · outbound

This paper cites SuryaKiran at MEDIQA-Sum 2023: Leveraging LoRA for Clinical Dialogue Summarization.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race SuryaKiran at MEDIQA-Sum 2023: Leveraging LoRA for Clinical Dialogue Summarization

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:13:21.647590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:13:18.924741Z digest=sha256:15cef761da05f6acc495b4e15c4696cfa24bebf14d80f261b04b9cd21e3d4f47

Observation 68ee41b9-5b47-418d-b855-3d6afbe5dc61 · outbound

This paper cites Evaluating and Mitigating Discrimination in Language Model Decisions.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Evaluating and Mitigating Discrimination in Language Model Decisions

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:18.979922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:18.979922Z digest=sha256:97fd1e3f06931c238b0a37d1e2e942448a54da094b29e49ce7f2e1107857d753

Observation 2e7df569-9d9b-41c6-9fa8-1a2fe05f9e2e · outbound

This paper cites Clinical Camel: An Open Expert-Level Medical Language Model with Dialogue-Based Knowledge Encoding.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Clinical Camel: An Open Expert-Level Medical Language Model with Dialogue-Based Knowledge Encoding

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:19.036340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:19.036340Z digest=sha256:2b66ffecd254dcde9c05793787f8019f0de26a7d7eb24ad31547bfddfe55c33d

Observation 33d965b8-0bd6-42eb-8d4c-014e25b9a186 · outbound

This paper cites Steering Language Models With Activation Engineering.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Steering Language Models With Activation Engineering

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:19.106538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:19.106538Z digest=sha256:8b7d35f43b9df7b6ddb71cb5bfb976036a8c15d4433a2213ae2c2e81b5291b81

Observation 2274232a-1bde-4b51-a43f-f0d4ec7d6e0f · outbound

This paper cites an unresolved cited work.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:13:23.082467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:13:19.178953Z digest=sha256:16278d07b7ffe4b58aa2b4b53347931616a1deada1d86b80e2862bd771665eb3

Observation 5724ed5a-dc59-49c0-9876-5279146e77aa · outbound

This paper cites DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:19.236371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:19.236371Z digest=sha256:3b7871b34b7fe79883742ad40eea6545f9bbc23a6037682fb0c57143eeb8d2f4

Observation 25085deb-1076-4066-a0b5-1f3681d0afe1 · outbound

This paper cites Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:19.352932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:19.352932Z digest=sha256:4a5782ececfd2ae0344df5e1471abdc51543282e516bbf15ee1e6eabc62b2188

Observation a9b299b5-3400-43f7-846d-22f590859600 · outbound

This paper cites Jailbroken: How Does LLM Safety Training Fail?.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Jailbroken: How Does LLM Safety Training Fail?

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:19.432193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:19.432193Z digest=sha256:1ed7baeda234353a53bab828a8edf2ea3d6b15a23486b54a65ed3363c1ada990

Observation 5e43b040-e749-4864-b1e7-84ea0d259d51 · outbound

This paper cites Fundamental Limitations of Alignment in Large Language Models.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Fundamental Limitations of Alignment in Large Language Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:19.496083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:19.496083Z digest=sha256:9a57b39e1d5d2f6d8a667f86bcde274048793aeb1eacef4d8dac00ec13024f54

Observation 23e4f887-4e7a-45ad-ad63-9c1688b592c9 · outbound

This paper cites AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:19.564287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:19.564287Z digest=sha256:7ef65c9f3dd3b255f8f41998d7a138224472239938c4d020880fd5cd927e765b

Observation 5864e4db-8df3-4dce-830e-4957a71edba1 · outbound

This paper cites Uncovering Safety Risks of Large Language Models through Concept Activation Vector.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Uncovering Safety Risks of Large Language Models through Concept Activation Vector

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:19.638500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:19.638500Z digest=sha256:03a10c5799fd4552594cf27715110e04e8b251106860a37c7166fd38969ec31c

Observation 2a13aa06-7e8b-4911-9eec-797247b2bfa4 · outbound

This paper cites AutoRE: Document-Level Relation Extraction with Large Language Models.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race AutoRE: Document-Level Relation Extraction with Large Language Models

Reference 71

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:13:21.360912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:13:19.714890Z digest=sha256:fbdb127d3e052c219ee669312f784714b4aa096562d3a7f96a076166f1b1ac36

Observation 94cf689b-5052-4457-8597-023c50383804 · outbound

This paper cites Understanding and Mitigating Gender Bias in LLMs via Interpretable Neuron Editing.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Understanding and Mitigating Gender Bias in LLMs via Interpretable Neuron Editing

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:19.780988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:19.780988Z digest=sha256:1f5709fb0418ce99bd1604c90f042ef54d0225346bf8d422bd603f7d55a28ac1

Observation 60034b32-2269-4448-a321-ae31c92f7089 · outbound

This paper cites Towards Best Practices of Activation Patching in Language Models: Metrics and Methods.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Towards Best Practices of Activation Patching in Language Models: Metrics and Methods

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:19.843762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:19.843762Z digest=sha256:47c1dc2a42b4a7c6a42a6e281eba45ec8b929eed3dabdabfb81a00eb32be9925

Observation 6a303adf-c8b0-47a8-90e1-7d81d48f201d · outbound

This paper cites DialogueLLM: Context and Emotion Knowledge-Tuned Large Language Models for Emotion Recognition in Conversations.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race DialogueLLM: Context and Emotion Knowledge-Tuned Large Language Models for Emotion Recognition in Conversations

Reference 74

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:13:21.217774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:13:19.901587Z digest=sha256:77f78bea9884f13dd471b07e96f0ac6dd9a7a895bc263fd77ee23d4bf5ecc0b1

Observation 24a1d6c1-3256-4b83-bf4b-bf345a072fad · outbound

This paper cites The Clock and the Pizza: Two Stories in Mechanistic Explanation of Neural Networks.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race The Clock and the Pizza: Two Stories in Mechanistic Explanation of Neural Networks

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:19.962960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:19.962960Z digest=sha256:86312ac285f999adf5762a7d4779890962a3a7efd45bf9a8b4e931e8fe1eafe7

Observation c4d24a5c-c423-4b35-bcbe-afa331c8565e · outbound

This paper cites online" 'onlinestring :=.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race online" 'onlinestring :=

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:20.045313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:20.045313Z digest=sha256:8fce3a85c5728868ecbc2d3c4799309fd93fb13d43ca9227e21c647fa1d2351a

Observation 76c487d5-46d7-409a-9876-765a1bab2955 · outbound

This paper cites write newline.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race write newline

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:20.054591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:20.054591Z digest=sha256:0f7b44033be9a3930ce6953953c4e97cb699291abb065cb4e680c447f35096c3

Pith citing papers

No inbound Pith citation observations are available.