Pith. sign in

Paper Citation Record · LEDGER

Analysing Moral Bias in Finetuned LLMs through Mechanistic Interpretability

As of 22 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 2 inbound Pith citation observations for arXiv:2510.12229.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2510.12229 v3

Coverage vector

measured 15 of 15 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T10:04:08.124917Z

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T22:20:53.464810Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-09T05:46:01.579330Z

Reference resolution

15 of 15 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 19963716-f3d0-4db2-a663-672329b7f802 · outbound

This paper cites Addressing Moral Uncertainty using Large Language Models for Ethical Decision-Making.

Analysing Moral Bias in Finetuned LLMs through Mechanistic Interpretability Addressing Moral Uncertainty using Large Language Models for Ethical Decision-Making

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T10:04:07.130183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:04:07.130183Z digest=sha256:8cc6502459be618e1749c7b7f25fa52b6636c67770dc2f64ffc45b9ca0077de0

Observation a334bba8-d244-413f-9e79-27b7ee69afae · outbound

This paper cites Mistral 7B.

Analysing Moral Bias in Finetuned LLMs through Mechanistic Interpretability Mistral 7B

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T10:04:07.482408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:04:07.482408Z digest=sha256:5be1c30e356f8d156858ac77b5e23c15562247f2df956b74b02818fab51100d2

Observation 61b60131-a68b-46c5-9592-7dc3bb7f12f7 · outbound

This paper cites URL https://distill.pub/ 2020/circuits/zoom-in/.

Analysing Moral Bias in Finetuned LLMs through Mechanistic Interpretability URL https://distill.pub/ 2020/circuits/zoom-in/

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T10:04:07.655520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:04:07.655520Z digest=sha256:175d1cacd4d0885417d584d444cb211c1bd67bdd790e0d666ce684b6e86c51aa

Observation e1b90ac2-af4b-4981-91d6-2c43468eb6e7 · outbound

This paper cites Can LLMs Simulate Human Behavioral Variability? A Case Study in the Phonemic Fluency Task.

Analysing Moral Bias in Finetuned LLMs through Mechanistic Interpretability Can LLMs Simulate Human Behavioral Variability? A Case Study in the Phonemic Fluency Task

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T10:04:07.846936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:04:07.846936Z digest=sha256:72de39f3972a7f940eeb0dc2394b392853bec81766b2b4ac41132117c250f962

Observation d4d23b16-2f90-42ed-a028-165c546272cf · outbound

This paper cites Sarfati, Y ., Hardy-Bayl´e, M.

Analysing Moral Bias in Finetuned LLMs through Mechanistic Interpretability Sarfati, Y ., Hardy-Bayl´e, M

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T10:04:07.898295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:04:07.898295Z digest=sha256:62e3b39a70b44a013039eba6749627a99e07fe4df69ec3c2914213be89132847

Observation 9563ab4b-bcc7-4fbd-b8be-149b0fb75df1 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Analysing Moral Bias in Finetuned LLMs through Mechanistic Interpretability Gemma 2: Improving Open Language Models at a Practical Size

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T10:04:07.979364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:04:07.979364Z digest=sha256:6f0f4f071638c02529e26c83e1b8f51a5d86e9d6f53e16b355e1cafb144077cd

Observation 425f9994-2666-456d-9267-2c710b28a798 · outbound

This paper cites Taxonomy of risks posed by lan- guage models.

Analysing Moral Bias in Finetuned LLMs through Mechanistic Interpretability Taxonomy of risks posed by lan- guage models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T10:04:08.039228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:04:08.039228Z digest=sha256:56dcaff2d1dcdbdba10dd861378383a2cd75ee23dc94fae2e06da3ba2b8a8741

Observation 457e262e-5331-48a7-ba88-877c2af58e10 · outbound

This paper cites acl-long.893/.

Analysing Moral Bias in Finetuned LLMs through Mechanistic Interpretability acl-long.893/

Reference 893

Resolution
unresolved
no resolver link, observed 2026-08-04T10:04:06.984040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:04:06.984040Z digest=sha256:40d8a5e2d921525aa2a0d98e6d74e23af65e0d10f0c2d212fd422e9c23cd412d

Observation f1e57758-7181-4be4-8d6a-e94627284b89 · outbound

This paper cites Could a Large Language Model be Conscious?.

Analysing Moral Bias in Finetuned LLMs through Mechanistic Interpretability Could a Large Language Model be Conscious?

Reference 2005

Resolution
unresolved
no resolver link, observed 2026-08-04T10:04:06.629384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:04:06.629384Z digest=sha256:940fa21c9e14a5d1282a946b3705bc4d16ba9ed50d8ad42b6333b5acd6ba685c

Observation 1741f6ed-a444-4df9-9ef3-1a013553e5a7 · outbound

This paper cites Towards Best Practices of Activation Patching in Language Models: Metrics and Methods.

Analysing Moral Bias in Finetuned LLMs through Mechanistic Interpretability Towards Best Practices of Activation Patching in Language Models: Metrics and Methods

Reference 2010

Resolution
unresolved
no resolver link, observed 2026-08-04T10:04:08.124917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:04:08.124917Z digest=sha256:4be2731d3722cdf337339758a4838fd6b84584d27275f464c3c6477293333427

Observation ec1a5368-4d93-46ef-9348-bc03d92d5fcf · outbound

This paper cites The Llama 3 Herd of Models.

Analysing Moral Bias in Finetuned LLMs through Mechanistic Interpretability The Llama 3 Herd of Models

Reference 2013

Resolution
unresolved
no resolver link, observed 2026-08-04T10:04:07.321726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:04:07.321726Z digest=sha256:c0f86519e7bf4252df97c39e8a3c88721bba66fc9a9c4610bef7cda88ef07817

Observation 0a68d1b7-cf90-46d8-9ab4-3e89cbdff978 · outbound

This paper cites M., Gebru, T., McMillan-Major, A., and Shmitchell, S.

Analysing Moral Bias in Finetuned LLMs through Mechanistic Interpretability M., Gebru, T., McMillan-Major, A., and Shmitchell, S

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-04T10:04:06.560346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:04:06.560346Z digest=sha256:9a7d78934ca58bd53ac3840e7576d49e002ece70e771e50805896c62d4eb2864

Observation 59110f0d-e0b7-48c3-95af-bfa9a5a84f07 · outbound

This paper cites doi: 10.18653/v1/2023.acl-long.

Analysing Moral Bias in Finetuned LLMs through Mechanistic Interpretability doi: 10.18653/v1/2023.acl-long

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T10:04:06.846244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:04:06.846244Z digest=sha256:59f875b5080d172cf5a844f60b5837f802ad2fc29a1d3859c2e03eaa00dfba34

Observation f5192ad6-599f-412c-8d58-0d48776bb6ed · outbound

This paper cites Interpreting Bias in Large Language Models: A Feature-Based Approach.

Analysing Moral Bias in Finetuned LLMs through Mechanistic Interpretability Interpreting Bias in Large Language Models: A Feature-Based Approach

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T10:04:07.765256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:04:07.765256Z digest=sha256:e1d88a9a8c70a7785c2b63bfb759e3cf0796543e9478674bd73cefdcb64e4a54

Observation 6a2ef335-0406-48f7-87ce-ccb2a49563dd · outbound

This paper cites Dissecting Bias in LLMs: A Mechanistic Interpretability Perspective.

Analysing Moral Bias in Finetuned LLMs through Mechanistic Interpretability Dissecting Bias in LLMs: A Mechanistic Interpretability Perspective

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T10:04:06.753097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:04:06.753097Z digest=sha256:b6210beff4375f5b8251996dedaa7dd656d12c949c2c45911c9c67881d67cef6

Pith citing papers

Observation f702573e-df5f-4169-b928-02847faba96c · inbound

Mechanistic Interpretability of Cognitive Complexity in LLMs via Linear Probing using Bloom's Taxonomy cites this paper.

Mechanistic Interpretability of Cognitive Complexity in LLMs via Linear Probing using Bloom's Taxonomy Analysing Moral Bias in Finetuned LLMs through Mechanistic Interpretability

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T22:20:53.464810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T22:20:53.464810Z digest=sha256:da045bdc58e8f616fa1a68bde5be119c752dca84e8f92be48947de6cea105298

Observation 55550e9d-3148-46e8-95a7-d0387ade3479 · inbound

User identity conditions moral wrongness ratings in non-reasoning large language models cites this paper.

User identity conditions moral wrongness ratings in non-reasoning large language models Analysing Moral Bias in Finetuned LLMs through Mechanistic Interpretability

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-20T02:18:20.628178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T05:37:58.875536Z digest=sha256:877142138c05f0af214322e03551e0baf8688f2b4f16125bc85d53a0f82932ce