Pith. sign in

Paper Citation Record · LEDGER

TrustGPT: A Benchmark for Trustworthy and Responsible Large Language Models

As of 12 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 19 inbound Pith citation observations for arXiv:2306.11507.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2306.11507 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T18:15:02.605649Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T19:50:11.115863Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 59f1e04a-4e79-4a20-a729-eae710f09163 · inbound

Moral Persuasion in Large Language Models: Evaluating Susceptibility and Ethical Alignment cites this paper.

Moral Persuasion in Large Language Models: Evaluating Susceptibility and Ethical Alignment TrustGPT: A Benchmark for Trustworthy and Responsible Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T18:15:02.605649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:15:02.605649Z digest=sha256:eea00aa36c6539691cef9972f8a66895087874e6ef837f87b116d0c51b7c1bb2

Observation 250faf4d-a769-4771-a602-a0044531f234 · inbound

Observing Micromotives and Macrobehavior of Large Language Models cites this paper.

Observing Micromotives and Macrobehavior of Large Language Models TrustGPT: A Benchmark for Trustworthy and Responsible Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T18:24:05.924949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T18:24:05.924949Z digest=sha256:7b15d5c34f9f8c9ad4093681c8f9aea6ac3382e86a6cb5ec13a9416260e424da

Observation d2119278-9c00-4249-a600-99afeee24c86 · inbound

LangFair: A Python Package for Assessing Bias and Fairness in Large Language Model Use Cases cites this paper.

LangFair: A Python Package for Assessing Bias and Fairness in Large Language Model Use Cases TrustGPT: A Benchmark for Trustworthy and Responsible Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T21:56:29.752606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:56:29.752606Z digest=sha256:a80550c991e478a0ee4133d9e06049a04338399abc98160079d764a634f6b548

Observation a93fecbc-57ab-4676-a326-5f88daccd6ea · inbound

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values cites this paper.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values TrustGPT: A Benchmark for Trustworthy and Responsible Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:59.701166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:53:59.701166Z digest=sha256:51a7b7e2b1c34e786a39e3c9299a3d13127e6d7a8539655c1ec37a8b0f9c2547

Observation a5beb4ef-85cb-4626-be2f-77ddb7b527a6 · inbound

Fairness through Difference Awareness: Measuring Desired Group Discrimination in LLMs cites this paper.

Fairness through Difference Awareness: Measuring Desired Group Discrimination in LLMs TrustGPT: A Benchmark for Trustworthy and Responsible Large Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-09T14:03:44.434667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:03:44.434667Z digest=sha256:eae9fb368f54493f6fe83036b21ad95a0be599ff422bc458a4d5353a85373839

Observation cd7d5d5c-ea78-42e4-b482-e3af89ecce4c · inbound

Artificial Intelligence in Spectroscopy: Advancing Chemistry from Prediction to Generation and Beyond cites this paper.

Artificial Intelligence in Spectroscopy: Advancing Chemistry from Prediction to Generation and Beyond TrustGPT: A Benchmark for Trustworthy and Responsible Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T20:11:11.994909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:11:11.994909Z digest=sha256:be7b2d6bb6ea750f9adeb7fc130a114e62512dc2c6c9a6e253d6256704476081

Observation 162d771a-418c-4b4e-a320-87e3c17d1a75 · inbound

When Ethics and Payoffs Diverge: LLM Agents in Morally Charged Social Dilemmas cites this paper.

When Ethics and Payoffs Diverge: LLM Agents in Morally Charged Social Dilemmas TrustGPT: A Benchmark for Trustworthy and Responsible Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:55.580256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:55.580256Z digest=sha256:9ddf2929d9060f25bc1a6340681ecbadea8dba06dde0cb4ef18b710472070e8e

Observation 69224819-50eb-4cda-833f-5ad01ab6c371 · inbound

OpenFActScore: Open-Source Atomic Evaluation of Factuality in Text Generation cites this paper.

OpenFActScore: Open-Source Atomic Evaluation of Factuality in Text Generation TrustGPT: A Benchmark for Trustworthy and Responsible Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T19:20:16.536212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:20:16.536212Z digest=sha256:4fced331b7b6a5c0a574926bae9440747139f18f49a122b2af7c377510720be3

Observation 8b639068-14e2-43aa-bcb9-7271a1414879 · inbound

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM cites this paper.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM TrustGPT: A Benchmark for Trustworthy and Responsible Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:00.005776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:00.005776Z digest=sha256:6399c22f58c83a423782eba2656aec0a81b1a15e4573ebab5ec9c627743bd13b

Observation bb7343bc-d8fc-43bc-8867-3e4323e6d0e2 · inbound

EPT Benchmark: Evaluation of Persian Trustworthiness in Large Language Models cites this paper.

EPT Benchmark: Evaluation of Persian Trustworthiness in Large Language Models TrustGPT: A Benchmark for Trustworthy and Responsible Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T23:03:16.421385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:03:16.421385Z digest=sha256:34556c974a22743a8b814f364ca020c4c70b146333ef36760b3603e151fe1338

Observation 29799214-f1d3-4d1d-9649-d2b5ecd02cb2 · inbound

Safety Is Not Universal: The Selective Safety Trap in LLM Alignment cites this paper.

Safety Is Not Universal: The Selective Safety Trap in LLM Alignment TrustGPT: A Benchmark for Trustworthy and Responsible Large Language Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-16T16:08:04.813096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T16:03:21.391883Z digest=sha256:09a91140bdef64fb2b8c66a81f43a8fdefa32a07b15ca12c3724e6a270ef3403

Observation f9420c9b-c3fc-4a4f-a9f3-13835ab1c1fa · inbound

Safety Is Not Universal: The Selective Safety Trap in LLM Alignment cites this paper.

Safety Is Not Universal: The Selective Safety Trap in LLM Alignment TrustGPT: A Benchmark for Trustworthy and Responsible Large Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T12:08:44.317070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:08:44.317070Z digest=sha256:8fd26692c77f3a04f0e61c7c9740337c9fcf66d3039492b0e3782a5780ae1790

Observation 98745a4a-0145-48e9-b9ee-89b9cb6c67e5 · inbound

From RAG to Agentic RAG for Faithful Islamic Question Answering cites this paper.

From RAG to Agentic RAG for Faithful Islamic Question Answering TrustGPT: A Benchmark for Trustworthy and Responsible Large Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T11:08:23.419771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:08:23.419771Z digest=sha256:082f2360109dbbdc2bae98b0dc2750136706597f530b89a06c33cb7aca0cb325

Observation 379c3a04-0776-476f-aa29-4ebb71d48d29 · inbound

Insight: Enhancing Mobile Accessibility for Blind and Visually Impaired Users with LLMs cites this paper.

Insight: Enhancing Mobile Accessibility for Blind and Visually Impaired Users with LLMs TrustGPT: A Benchmark for Trustworthy and Responsible Large Language Models

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-12T02:11:15.302288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-12T02:11:09.787984Z digest=sha256:f53bcf7a08dbc44a81a84fad9e2a7d60c38e9a89c242c97d83ba1922d8fa983b

Observation 051922d9-5140-4680-97b0-23317a76243a · inbound

StereoTales: A Multilingual Framework for Open-Ended Stereotype Discovery in LLMs cites this paper.

StereoTales: A Multilingual Framework for Open-Ended Stereotype Discovery in LLMs TrustGPT: A Benchmark for Trustworthy and Responsible Large Language Models

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:51:27.418348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-12T04:50:17.399580Z digest=sha256:a2e8d7135f2d50efb0ac3bd7fc41b958399c8ac26db0f23d016a56450b2c583f

Observation 79ca632e-1935-4bf4-88a9-a6ef4e3eb072 · inbound

StereoTales: A Multilingual Framework for Open-Ended Stereotype Discovery in LLMs cites this paper.

StereoTales: A Multilingual Framework for Open-Ended Stereotype Discovery in LLMs TrustGPT: A Benchmark for Trustworthy and Responsible Large Language Models

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:22:29.071324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-13T07:20:32.494840Z digest=sha256:9cd3ecbceeafd8c50c6c54a230f944a721b106a44fa6856b924d53778b15ec1a

Observation 1be6c546-25c7-4a0b-8668-7cfc56a920eb · inbound

A Red Teaming Framework for Large Language Models: A Case Study on Faithfulness Evaluation cites this paper.

A Red Teaming Framework for Large Language Models: A Case Study on Faithfulness Evaluation TrustGPT: A Benchmark for Trustworthy and Responsible Large Language Models

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T19:50:11.118037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-25T20:58:53.119386Z digest=sha256:9cc1d817e73f64ecb69ee0ed9df06c5c34cd5ee8fbd135bd63bd1c1f9b22c7e9

Observation efcd11f5-2eff-45e5-985a-5ec9819e94c7 · inbound

A Red Teaming Framework for Large Language Models: A Case Study on Faithfulness Evaluation cites this paper.

A Red Teaming Framework for Large Language Models: A Case Study on Faithfulness Evaluation TrustGPT: A Benchmark for Trustworthy and Responsible Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T10:16:36.314418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:16:36.314418Z digest=sha256:9d6f7a208837b85bdc3188a31e35e49358c9b5aa7ee1157a43500499056784e1

Observation 726877a8-a677-4263-843b-061db4ead206 · inbound

Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges cites this paper.

Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges TrustGPT: A Benchmark for Trustworthy and Responsible Large Language Models

Reference 123

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:28.889774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T00:55:28.889774Z digest=sha256:1e8e7dd82422e274301316800e255291d3451a574f7af845d9fcb58cc5801975