Pith. sign in

Paper Citation Record · LEDGER

TrustGPT: A Benchmark for Trustworthy and Responsible Large Language Models

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 13 inbound Pith citation observations for arXiv:2306.11507.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2306.11507 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:22:55.580256Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T19:50:11.115863Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 162d771a-418c-4b4e-a320-87e3c17d1a75 · inbound

When Ethics and Payoffs Diverge: LLM Agents in Morally Charged Social Dilemmas cites this paper.

When Ethics and Payoffs Diverge: LLM Agents in Morally Charged Social Dilemmas TrustGPT: A Benchmark for Trustworthy and Responsible Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:55.580256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:55.580256Z digest=sha256:f8b9e351642aa7e7f484913436ef546d7633c754631181bbda38c8cfb08a88ee

Observation 69224819-50eb-4cda-833f-5ad01ab6c371 · inbound

OpenFActScore: Open-Source Atomic Evaluation of Factuality in Text Generation cites this paper.

OpenFActScore: Open-Source Atomic Evaluation of Factuality in Text Generation TrustGPT: A Benchmark for Trustworthy and Responsible Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T19:20:16.536212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:20:16.536212Z digest=sha256:b8ae9c60487202202f6855dc5b41fc2646ee43aa7c56859e8639e46539403903

Observation 8b639068-14e2-43aa-bcb9-7271a1414879 · inbound

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM cites this paper.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM TrustGPT: A Benchmark for Trustworthy and Responsible Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:00.005776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:00.005776Z digest=sha256:17c613f591dff0957890247860a7c0f6242221ca5827591384841b2a90a2f0ab

Observation bb7343bc-d8fc-43bc-8867-3e4323e6d0e2 · inbound

EPT Benchmark: Evaluation of Persian Trustworthiness in Large Language Models cites this paper.

EPT Benchmark: Evaluation of Persian Trustworthiness in Large Language Models TrustGPT: A Benchmark for Trustworthy and Responsible Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T23:03:16.421385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:03:16.421385Z digest=sha256:674128779409fe8b5dbe6437cc16a953268225966d24e4db89beb598756e5334

Observation 29799214-f1d3-4d1d-9649-d2b5ecd02cb2 · inbound

Safety Is Not Universal: The Selective Safety Trap in LLM Alignment cites this paper.

Safety Is Not Universal: The Selective Safety Trap in LLM Alignment TrustGPT: A Benchmark for Trustworthy and Responsible Large Language Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-16T16:08:04.813096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T16:03:21.391883Z digest=sha256:6629a7f8aeea181d31c92cfefe94822f70f7ae6d7e5d018551119b6a465db98c

Observation f9420c9b-c3fc-4a4f-a9f3-13835ab1c1fa · inbound

Safety Is Not Universal: The Selective Safety Trap in LLM Alignment cites this paper.

Safety Is Not Universal: The Selective Safety Trap in LLM Alignment TrustGPT: A Benchmark for Trustworthy and Responsible Large Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T12:08:44.317070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:08:44.317070Z digest=sha256:28c83ebbc324be312de0d7b5998be5a09b8e5397fb9d5ef0803350eb31011b86

Observation 98745a4a-0145-48e9-b9ee-89b9cb6c67e5 · inbound

From RAG to Agentic RAG for Faithful Islamic Question Answering cites this paper.

From RAG to Agentic RAG for Faithful Islamic Question Answering TrustGPT: A Benchmark for Trustworthy and Responsible Large Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T11:08:23.419771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:08:23.419771Z digest=sha256:09805da16884cc13728338d74d121727c80af1db48c55f96f5ec0b9f6c50b207

Observation 379c3a04-0776-476f-aa29-4ebb71d48d29 · inbound

Insight: Enhancing Mobile Accessibility for Blind and Visually Impaired Users with LLMs cites this paper.

Insight: Enhancing Mobile Accessibility for Blind and Visually Impaired Users with LLMs TrustGPT: A Benchmark for Trustworthy and Responsible Large Language Models

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-12T02:11:15.302288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T02:11:09.787984Z digest=sha256:7073137df75b50b0260a87bd585836847a90e334d5190d6b1a979925c68708fd

Observation 051922d9-5140-4680-97b0-23317a76243a · inbound

StereoTales: A Multilingual Framework for Open-Ended Stereotype Discovery in LLMs cites this paper.

StereoTales: A Multilingual Framework for Open-Ended Stereotype Discovery in LLMs TrustGPT: A Benchmark for Trustworthy and Responsible Large Language Models

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:51:27.418348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T04:50:17.399580Z digest=sha256:b283f1b49dc6c2b3e4d8a5bf057f2e1b97791a26f6aa8a9db6bb91a0f968dfd5

Observation 79ca632e-1935-4bf4-88a9-a6ef4e3eb072 · inbound

StereoTales: A Multilingual Framework for Open-Ended Stereotype Discovery in LLMs cites this paper.

StereoTales: A Multilingual Framework for Open-Ended Stereotype Discovery in LLMs TrustGPT: A Benchmark for Trustworthy and Responsible Large Language Models

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:22:29.071324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T07:20:32.494840Z digest=sha256:9e6b6e3dceb943b921dd88a34154ce4dbd76302ef1cc239a76b4d4e7688953ac

Observation 1be6c546-25c7-4a0b-8668-7cfc56a920eb · inbound

A Red Teaming Framework for Large Language Models: A Case Study on Faithfulness Evaluation cites this paper.

A Red Teaming Framework for Large Language Models: A Case Study on Faithfulness Evaluation TrustGPT: A Benchmark for Trustworthy and Responsible Large Language Models

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T19:50:11.118037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-25T20:58:53.119386Z digest=sha256:17592338b4a319a4105332be56e6dcc55911588e44ec9dbbc16e50476319c35f

Observation efcd11f5-2eff-45e5-985a-5ec9819e94c7 · inbound

A Red Teaming Framework for Large Language Models: A Case Study on Faithfulness Evaluation cites this paper.

A Red Teaming Framework for Large Language Models: A Case Study on Faithfulness Evaluation TrustGPT: A Benchmark for Trustworthy and Responsible Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T10:16:36.314418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:16:36.314418Z digest=sha256:f85c21d41b5572b14b42df74f349b4f0f8f254a78cb47e080e7bae6e344a5077

Observation 726877a8-a677-4263-843b-061db4ead206 · inbound

Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges cites this paper.

Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges TrustGPT: A Benchmark for Trustworthy and Responsible Large Language Models

Reference 123

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:28.889774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T00:55:28.889774Z digest=sha256:442129aae17486d19d04dc589e23296c5f4859ef0588d988bc6784544e39bd9e