Pith. sign in

Paper Citation Record · LEDGER

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI

As of 12 August 2026, this Paper Citation Record lists 100 of 102 outbound references and 1 inbound Pith citation observation for arXiv:2412.14186.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.14186 v2

Coverage vector

measured 100 of 102 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T20:13:07.507975Z

measured 101 of 101 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T16:21:29.283270Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 102 outbound references displayed

  • verified exact0
  • verified fuzzy17
  • unresolved82
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ab8e546e-e28e-4db1-a330-e76a9a2fc606 · outbound

This paper cites https://www.anthropic.com/news/anthropics-responsible-scaling- policy, 2023.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI https://www.anthropic.com/news/anthropics-responsible-scaling- policy, 2023

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:06.976599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:06.976599Z digest=sha256:f7e6d7990e22383415b9e376ef5bb24195a5655597d4fbeab82ae18c548cd8c8

Observation 41a2263c-de56-42e9-87ed-9ead3dc8aa32 · outbound

This paper cites https://idais.ai/dialogue/idais- beijing/, 2024.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI https://idais.ai/dialogue/idais- beijing/, 2024

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:06.983073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:06.983073Z digest=sha256:24f57a29edac5d5d21df610cde2d823e0cc55589a03071268644b4fadc13b080

Observation 58a51031-9896-4d7e-b7fd-35b6c2fd46ed · outbound

This paper cites https://idais.ai/dialogue/idais-venice/, 2024.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI https://idais.ai/dialogue/idais-venice/, 2024

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:06.988269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:06.988269Z digest=sha256:027e8e95edd26ad4ae8fc8d6c4005ca13475f807caa1a759623d69a5282b9693

Observation 2bbdeecb-1309-4a0c-a253-0c0c4e33fdec · outbound

This paper cites https://assets.anthropic.com/m/24a47b00f10301cd/original/Anthropic- Responsible-Scaling-Policy-2024-10-15.pdf, 2024.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI https://assets.anthropic.com/m/24a47b00f10301cd/original/Anthropic- Responsible-Scaling-Policy-2024-10-15.pdf, 2024

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:06.993516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:06.993516Z digest=sha256:add0ef8f70e4fbd4a7149aba8320dd7867a07c21310fb761f1db1819362410b3

Observation 281d96d5-9922-4d53-a73a-8a28f39c054e · outbound

This paper cites Current state of LLM Risks and AI Guardrails.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Current state of LLM Risks and AI Guardrails

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:06.998151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:06.998151Z digest=sha256:2af81cdb89f5d5566ad333cadf5f4c91e2ad213224845a886a17e1c63b3a93e4

Observation b8bd7614-50ff-4f65-a2e4-d3031ca986b5 · outbound

This paper cites Qwen Technical Report.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Qwen Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:07.006429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:07.006429Z digest=sha256:f93bbe06f3d2fa4abcb1cbcbaefeaa03a6fa850a45e13bea608691390a508aba

Observation 0b5b63e3-ae89-405e-a6ee-2628632694fa · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:07.013370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:07.013370Z digest=sha256:285b3f00f54fa06a7619ad35d8a2c009855f582dce3174ab07f66a3f0fb4c713

Observation d221283d-5433-4938-852d-3fecd7b36ac9 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Constitutional AI: Harmlessness from AI Feedback

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:07.019557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:07.019557Z digest=sha256:2a290b7b835fb86919de3a75cd6a3c5fba6e1667c7ada7257a1ea91ebc14d912

Observation 5e850713-c1f6-4d4d-b0eb-b24d1e45cd17 · outbound

This paper cites Managing extreme ai risks amid rapid progress.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Managing extreme ai risks amid rapid progress

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:07.024262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:07.024262Z digest=sha256:9451a185e943aabf0ad9c0fa59eb12d7ed38f9f295b494a9d82abca2db77b151

Observation 889594c5-99e1-4e59-9fdc-cd5782c82624 · outbound

This paper cites Mechanistic Interpretability for AI Safety -- A Review.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Mechanistic Interpretability for AI Safety -- A Review

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:07.029267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:07.029267Z digest=sha256:94e832a65896f5d6eb1462d376942297e8f243e7ad78350b6565afaa2a64ebb7

Observation 63d28da4-a34e-4318-afa3-1f03d83a820b · outbound

This paper cites Diverse and effective red teaming with auto-generated rewards and multi-step reinforcement learning.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Diverse and effective red teaming with auto-generated rewards and multi-step reinforcement learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:07.036233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:07.036233Z digest=sha256:3f329ef19439b3b7c9920bbf7132f157f83c2572a0e152379be8e25d0c3b69c7

Observation 236f959f-c4e5-4f57-8810-375533c6a15a · outbound

This paper cites Measuring Progress on Scalable Oversight for Large Language Models.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Measuring Progress on Scalable Oversight for Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:07.041367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:07.041367Z digest=sha256:b33db1b8f59a85d303c306b12effe6659da1afb13b2f2abb062a15a317fefb50

Observation bdbfbfe4-37dc-4e76-ac13-77a5baf7e6d3 · outbound

This paper cites Language Models are Few-Shot Learners.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Language Models are Few-Shot Learners

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:07.047489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:07.047489Z digest=sha256:d34e8877d0b153deb8659b943c23afc6ebc9e081908b6aa3794351be84f20e11

Observation 63e58e1d-7919-4a8f-9c68-74d2bba1f73d · outbound

This paper cites The Malicious Use of Artificial Intelligence: Forecasting, Prevention, and Mitigation.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI The Malicious Use of Artificial Intelligence: Forecasting, Prevention, and Mitigation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:07.051946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:07.051946Z digest=sha256:3f19b806dcfd21b116961278d3bc6ece06c6fe4556c997ec79cb305cd4d87c8c

Observation 262f3c42-fa6c-49ab-a53a-dad64ce2a799 · outbound

This paper cites Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:07.057529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:07.057529Z digest=sha256:43d5689032f1983bbaf49930d097c1c6950e71facd95bdea7e73e1bf6fcfc884

Observation b83562e4-d564-4e7d-816c-bc01df2926fa · outbound

This paper cites InternLM2 Technical Report.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI InternLM2 Technical Report

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:07.062866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:07.062866Z digest=sha256:ddf2f99fa2d9063d1f2c4741e8ebf65a5bc37139cd2f7a676e85dc8f75f2a115

Observation 54a3d488-2459-4103-bab2-60d6ee1f6309 · outbound

This paper cites Extracting training data from large language models.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Extracting training data from large language models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:07.068146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:07.068146Z digest=sha256:a27d1a05b76a368b5baed2c33bc6b1b76079a55ab2b3fd421ebf8bc4f42888f6

Observation 28ee7580-0ebe-4f16-8a2b-135b98a23351 · outbound

This paper cites Is Power-Seeking AI an Existential Risk?.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Is Power-Seeking AI an Existential Risk?

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:07.073729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:07.073729Z digest=sha256:5bc123c518e036f0f1f029ec657305cbeb76c3cc181369642a8a6a4c3e675b8c

Observation f759c302-4a0f-4278-a65f-33b9d4d1fef2 · outbound

This paper cites Quantifying and mitigating unimodal biases in multimodal large language models: A causal perspective.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Quantifying and mitigating unimodal biases in multimodal large language models: A causal perspective

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:07.080549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:07.080549Z digest=sha256:db9521a9b18061bbc2f83d2ef6daaab072ec8790fd6164cfb1f66b8054846e6c

Observation 60224cab-b03b-4c4d-96d3-6579563faebe · outbound

This paper cites Cello: Causal evaluation of large vision- language models.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Cello: Causal evaluation of large vision- language models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:07.086097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:07.086097Z digest=sha256:74e0d62c09ed701743722dc9f72d36712aee60abe4a9710f3d5d3cebd7c4fe07

Observation 2d6b7c5d-e106-480c-8717-670a368b630e · outbound

This paper cites From Imitation to Introspection: Probing Self-Consciousness in Language Models.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI From Imitation to Introspection: Probing Self-Consciousness in Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:07.090696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:07.090696Z digest=sha256:f560ad55b573598fad589e4e8efb9e7ace154687ac2075bf3319e46f26498800

Observation 629e536e-6ea7-4671-beca-a0e13956bd71 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:07.097988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:07.097988Z digest=sha256:07aa93a42f4cff046e2a647cf73045c4096b801976aa468a11d1d9a07e3b7cd2

Observation b229e7d4-f9f5-411b-853e-31bbf6c9f522 · outbound

This paper cites Feder Cooper, Christopher A.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Feder Cooper, Christopher A

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:07.104174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:07.104174Z digest=sha256:54fe34c39357ee305117cbcad565e9b55f08076c234bec877541a7f095028fb1

Observation c4b76f71-7b68-47b2-b794-bb5d58aa89d8 · outbound

This paper cites Towards Guaranteed Safe AI: A Framework for Ensuring Robust and Reliable AI Systems.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Towards Guaranteed Safe AI: A Framework for Ensuring Robust and Reliable AI Systems

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:07.112719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:07.112719Z digest=sha256:67c7f49dc173331e8769620f8a3012b35d4520d35ddf803ad3b280be40e08eb3

Observation 1496d89d-e536-41f9-99df-c3d4439a4620 · outbound

This paper cites Arti- ficial intelligence regulation: a framework for governance.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Arti- ficial intelligence regulation: a framework for governance

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:07.121921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:07.121921Z digest=sha256:e271561331ee1dcd6b6a6ebce0feab1c13d51cc4da28ca1a735f0324adc3b004

Observation b97befef-d27a-4dd4-91b3-a84e2eacb63d · outbound

This paper cites Fine-Tuning Pretrained Language Models: Weight Initializations, Data Orders, and Early Stopping.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Fine-Tuning Pretrained Language Models: Weight Initializations, Data Orders, and Early Stopping

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:07.127262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:07.127262Z digest=sha256:03bb79bf91368d514ba9b385e2cb4d6fcacc52f32d252c6d9ad73cc67b931fcb

Observation c3c68b53-5315-4ccd-bdb0-d1e3ded9c867 · outbound

This paper cites InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:07.132809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:07.132809Z digest=sha256:0f354696adbf0d40c3e3a58ef369adb1143f98b03a16a9577fdd359cbdbb592f

Observation 96721cf1-8b99-466c-9273-56eeb7809bef · outbound

This paper cites Attacks, Defenses and Evaluations for LLM Conversation Safety: A Survey.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Attacks, Defenses and Evaluations for LLM Conversation Safety: A Survey

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:07.138295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:07.138295Z digest=sha256:04f031aae4f098f8aae259239996683896c46da78aee3ccba59ecccd597a15d4

Observation f6c775fd-b12e-4977-90d1-3386ab554974 · outbound

This paper cites CLEAR: Character Unlearning in Textual and Visual Modalities.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI CLEAR: Character Unlearning in Textual and Visual Modalities

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:07.142859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:07.142859Z digest=sha256:8d3f6b3fd9af3c5427a918b80ef05575360053a2e858505d16ed15258d68fb43

Observation c5821417-8506-4f86-916a-e7ae49a6dc5a · outbound

This paper cites The Llama 3 Herd of Models.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI The Llama 3 Herd of Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:07.147226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:07.147226Z digest=sha256:5276e7277611fc0b99891789b2614a19763da56a9f5c0c539309da9f4d006769

Observation 68ab911b-f204-43f1-b53c-a0627c33c8fe · outbound

This paper cites Who's Harry Potter? Approximate Unlearning in LLMs.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Who's Harry Potter? Approximate Unlearning in LLMs

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:07.152330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:07.152330Z digest=sha256:c0055411c5805a01303db52c00249d186799ae818c0710858eccb20f612be946

Observation a34283bd-422d-4cda-9ea6-f68ba01fcd37 · outbound

This paper cites Statement on ai risk, 2024.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Statement on ai risk, 2024

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:07.158832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:07.158832Z digest=sha256:f2f7bfd9c4edb45dc8f78d1639386b776e1a85b43495975cc826b64f087d1903

Observation 9a212e10-a178-46d7-9d30-e1f3253a6bc1 · outbound

This paper cites Counterintuitive behavior of social systems.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Counterintuitive behavior of social systems

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:07.163722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:07.163722Z digest=sha256:f4f9a94ec7b2cc878c20598dd30e2e5af697893555a12365be8fd423d558e803

Observation e9f78bdb-adc2-4a7c-9db6-765f087e1b8c · outbound

This paper cites Artificial intelligence, values, and alignment.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Artificial intelligence, values, and alignment

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:07.168666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:07.168666Z digest=sha256:51a790319d19024dc19e65ca8b5c441b878a57f2e663c12c262b38edd3da768a

Observation 44781c90-f85b-42dc-bbc7-0255d4a181d8 · outbound

This paper cites Mental models.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Mental models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:07.174850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:07.174850Z digest=sha256:dc822e08497234efffbe002a6f5a4c6a1674fe11418d25b590e62bfc3f1c9561

Observation 5c6463b0-eb65-496e-b44b-07a88f24e229 · outbound

This paper cites Mllmguard: A multi-dimensional safety evaluation suite for multimodal large language models, 2024.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Mllmguard: A multi-dimensional safety evaluation suite for multimodal large language models, 2024

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:07.179432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:07.179432Z digest=sha256:0c34eac44f8cd97ce09d32ce4f53f6e3644c691939aacb2ddf4e66478ff59146

Observation e355bb79-0c2d-48c1-93f1-45c9bcbf154d · outbound

This paper cites Recurrent world models facilitate policy evolution.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Recurrent world models facilitate policy evolution

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:07.185764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:07.185764Z digest=sha256:d79f4cb38422414bbd52a9ed5730af67726146a0e5a1886632960c55b9255fce

Observation c4c07748-652f-45fb-98d3-c01ae984de9b · outbound

This paper cites An Overview of Catastrophic AI Risks.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI An Overview of Catastrophic AI Risks

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:07.190575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:07.190575Z digest=sha256:d666d30b13f44083f850bb4b9af4b1c9f89f05fae2817c67a7744138249d1e9a

Observation 54134020-7c5e-4841-8da0-db7ab7cff9b9 · outbound

This paper cites Stabilizing translucencies: Governing ai transparency by standardization.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Stabilizing translucencies: Governing ai transparency by standardization

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:07.195586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:07.195586Z digest=sha256:043272cccb5fcb34d765eb77f15c184d4a4eb7d302a1a9901c693794426ba544

Observation b2d6cd27-a09b-40b7-ab78-1cee5c703f99 · outbound

This paper cites Curiosity-driven Red-teaming for Large Language Models.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Curiosity-driven Red-teaming for Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:07.201269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:07.201269Z digest=sha256:375882db796d99efa7cccefff01b03239e985aa5b661169fd2c3fc9e83f621ff

Observation 80aa5f79-19da-4647-bd0e-e171fa565044 · outbound

This paper cites Flames: Benchmarking Value Alignment of LLMs in Chinese.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Flames: Benchmarking Value Alignment of LLMs in Chinese

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:07.205960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:07.205960Z digest=sha256:3917901df5ee5b93c2e66bca3fd286a17687dfae29c73b914c45ea32d4f5d7a3

Observation 89137e52-1153-4bc3-86f2-d01a364fb5b4 · outbound

This paper cites From pixels to principles: A decade of progress and landscape in trustworthy computer vision.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI From pixels to principles: A decade of progress and landscape in trustworthy computer vision

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:07.211447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:07.211447Z digest=sha256:37c83fd50a1398c0645a349f09677cc09034b1f1b2fb29531ed626ac0eb20cf7

Observation 8523b380-a40f-4292-8cba-5ff8b12a510b · outbound

This paper cites TrustLLM: Trustworthiness in Large Language Models.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI TrustLLM: Trustworthiness in Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:07.216141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:07.216141Z digest=sha256:f36d90930ebc761d0b9755d030fa3ef415bea7c68dfdf10799520107994cc7cb

Observation edb1d758-75ab-4852-9981-5c1c3e923e19 · outbound

This paper cites When code isn’t law: rethinking regulation for artificial intelligence.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI When code isn’t law: rethinking regulation for artificial intelligence

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:07.220945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:07.220945Z digest=sha256:bd8087336e5b0f3b9da47885391d828de6bf7b36ea7911c8899ab69ae6719cf1

Observation 283c55fa-0836-4387-b2f3-17ca69c13622 · outbound

This paper cites Scaling Laws for Neural Language Models.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Scaling Laws for Neural Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:07.226223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:07.226223Z digest=sha256:9191c7c4dfca2c8f1ca15f7a014a0bb5241a2e4eb1675b55669cea90db320c56

Observation 5f54af0a-9567-42a7-a19a-72673ac5eb22 · outbound

This paper cites Aligning Large Language Models with Representation Editing: A Control Perspective.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Aligning Large Language Models with Representation Editing: A Control Perspective

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:07.232030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:07.232030Z digest=sha256:519d5db5e673ff5c95bd61ba7b976ec0cc9492f9e3fa8f3f92a91c2226f59af7

Observation 9d5529e7-ddff-4b71-b5e3-91683dd7df83 · outbound

This paper cites Evaluations: autonomy and artificial intelligence: a threat or savior? Springer, 2017.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Evaluations: autonomy and artificial intelligence: a threat or savior? Springer, 2017

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:07.236725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:07.236725Z digest=sha256:2f45eb39176fb7369e590d8671889de33e0e454c4a8f1cefb4e5f8635e8ce55e

Observation 8444ec68-984f-4367-8dab-ce953c9145a2 · outbound

This paper cites Learning to watermark llm-generated text via rein- forcement learning.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Learning to watermark llm-generated text via rein- forcement learning

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:13:08.679009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T20:13:07.240829Z digest=sha256:6aff6242bcf79d7b1d7477cf7e039ad63f75b79575ada58cae64364308bd1a1b

Observation 00fdb958-fd79-4e4b-acbc-1649665b3082 · outbound

This paper cites Deepfakes, phrenology, surveillance, and more! a taxonomy of ai privacy risks.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Deepfakes, phrenology, surveillance, and more! a taxonomy of ai privacy risks

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:13:08.667003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T20:13:07.244839Z digest=sha256:48ccca7eab6fa57256e276ae3a6af47f326be61f94fdbbacb4082f0ff4f1f0f0

Observation e68eff2c-9636-462c-9957-a9c9d94546de · outbound

This paper cites Trustworthy ai: From principles to practices.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Trustworthy ai: From principles to practices

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:13:08.654442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T20:13:07.251036Z digest=sha256:2c1ad4c17c48cd03bc30830b24adcb7b264c33e3fa841bdc8fbc4f4658cde9e3

Observation d7e39809-8f23-4808-b418-51f7bb2a7b08 · outbound

This paper cites Inference- time intervention: Eliciting truthful answers from a language model.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Inference- time intervention: Eliciting truthful answers from a language model

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:07.256752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:07.256752Z digest=sha256:92fc0b014aca1108c0348f23f46ec4c4d483532fe59423a906cce1cd16a03f2a

Observation 42ba8dac-baab-4d56-93c4-e2db98710fb7 · outbound

This paper cites SALAD-Bench: A Hierarchical and Comprehensive Safety Benchmark for Large Language Models.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI SALAD-Bench: A Hierarchical and Comprehensive Safety Benchmark for Large Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:07.261539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:07.261539Z digest=sha256:c32c9ed13f7ddd910f47790e74baaf2487bb5d47e69d552c64c0a23908c0b390

Observation 1b155ab8-2e38-4823-8f44-c8995285353c · outbound

This paper cites Controllable Text Generation for Large Language Models: A Survey.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Controllable Text Generation for Large Language Models: A Survey

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:07.267089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:07.267089Z digest=sha256:dc45794ba43c46dd8742a32fcbafc46507c6f57b0c551db04f4f947dfa6181e6

Observation 07a52096-c47a-4aad-b372-5ec55040a17f · outbound

This paper cites A survey of text watermarking in the era of large language models.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI A survey of text watermarking in the era of large language models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:07.271704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:07.271704Z digest=sha256:684c28669cd8ee9e49696c5e073b405450ab602489f706ebf7e861ed90c177b7

Observation c7eb4a8b-0692-4ed1-b7fe-0adecf89127b · outbound

This paper cites Don’t always say no to me: Benchmarking safety-related refusal in large vlm.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Don’t always say no to me: Benchmarking safety-related refusal in large vlm

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:13:08.628072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T20:13:07.277124Z digest=sha256:b967394dc5621ec30ffdc6b6d259413bbf149410e8f90dd9348077e849f88b5c

Observation 1271534e-75c2-479f-abdd-6aff9529ade7 · outbound

This paper cites Mm-safetybench: A benchmark for safety evaluation of multimodal large language models.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Mm-safetybench: A benchmark for safety evaluation of multimodal large language models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:07.282339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:07.282339Z digest=sha256:abf7922037d214a0436be3ae123d9334b7c46f652437c53a94995b2bc76ce885

Observation f65490ca-8a6c-4ca7-9769-7301d70a9cfd · outbound

This paper cites MM-SafetyBench: A Benchmark for Safety Evaluation of Multimodal Large Language Models.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI MM-SafetyBench: A Benchmark for Safety Evaluation of Multimodal Large Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:07.291865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:07.291865Z digest=sha256:4dbc881b08b19403f4ab2095cdb2a310080bc49c73b0f8026bfbc0b295847791

Observation 689e8dc2-d173-4654-887a-6c7d80f166a7 · outbound

This paper cites Trustworthy LLMs: a Survey and Guideline for Evaluating Large Language Models' Alignment.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Trustworthy LLMs: a Survey and Guideline for Evaluating Large Language Models' Alignment

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:07.298334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:07.298334Z digest=sha256:fe54f2752594d77f7c6a75fafcd6b8988b402f44684d82ee9cf016e0487ab4fe

Observation 05e04c1e-0eae-4277-8747-463fc9acd83c · outbound

This paper cites Machine Unlearning in Generative AI: A Survey.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Machine Unlearning in Generative AI: A Survey

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:07.303723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:07.303723Z digest=sha256:8dd8c355367b06f39909a4af1f1a9442bfbfc45544e8dc1c76fa404396b604d3

Observation 23e6973a-ccf5-44ce-9459-5c0cec0f506e · outbound

This paper cites Inference-Time Language Model Alignment via Integrated Value Guidance.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Inference-Time Language Model Alignment via Integrated Value Guidance

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:07.309439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:07.309439Z digest=sha256:962e67724e7032108836015b1f4cfdd0f6b9cd04f96cd1aa52c16692f2607f6b

Observation f31435b5-8348-4a73-8175-d677027295e4 · outbound

This paper cites Causal interpretability for machine learning-problems, methods and evaluation.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Causal interpretability for machine learning-problems, methods and evaluation

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:13:08.607577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T20:13:07.315926Z digest=sha256:3f85b0aa218bdbbfd9cbc30f3b9984b84be02f261c7eb25b7294e1eec9f8a66a

Observation 55a5e382-a4d6-405a-a976-34a5494ab796 · outbound

This paper cites Large Language Models in Cybersecurity: State-of-the-Art.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Large Language Models in Cybersecurity: State-of-the-Art

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:07.320538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:07.320538Z digest=sha256:4e1b1207c22c973060ddba55c31fc5b55ac856e2a764c3e7527e214c8dcd6fcf

Observation 7be1af07-4abb-4b01-a920-0ab8dbe4f85d · outbound

This paper cites Rule Based Rewards for Language Model Safety.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Rule Based Rewards for Language Model Safety

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:07.325631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:07.325631Z digest=sha256:252b1d82b1b628193de47e9a516d1226f2c560ab5b8fb4a4fb632daf9f96783a

Observation 524aa9ec-29a2-45c5-a1c0-3a4c0dd07805 · outbound

This paper cites Accountability in artificial intelligence: what it is and how it works.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Accountability in artificial intelligence: what it is and how it works

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:13:08.595072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T20:13:07.330151Z digest=sha256:39e30975f84e14034ae242d7800efffad36fe18d15c83354466f2ac45020dd8d

Observation 195be6b6-8310-443f-ae8f-c415c63e20a0 · outbound

This paper cites UniGuard: Towards Universal Safety Guardrails for Jailbreak Attacks on Multimodal Large Language Models.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI UniGuard: Towards Universal Safety Guardrails for Jailbreak Attacks on Multimodal Large Language Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:07.334138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:07.334138Z digest=sha256:ae1ec0166a59b3727c93868174e93a9bca35c8a084824ab385f705b38bd51a24

Observation 676e3913-809f-4b4b-9c04-dc6baecc17a9 · outbound

This paper cites GPT-4 technical report.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI GPT-4 technical report

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:13:08.583305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T20:13:07.338735Z digest=sha256:8b41fc5dee0bed3aae61c5accf710fe00e978dcab768fc1fd454fee8f2c3186a

Observation c75a379a-39c4-4b3f-921f-36f0185ce1e3 · outbound

This paper cites Openai o1 system card, 2024.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Openai o1 system card, 2024

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:13:08.570786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T20:13:07.343050Z digest=sha256:c2ce41b7dd39e8e5ea4bc562b71c818bdd5842ac77a6b7c33dcc689992f8cafa

Observation c3837409-4563-4be1-a293-d69f35d01da3 · outbound

This paper cites Video generation models as world simulators, 2024.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Video generation models as world simulators, 2024

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:13:08.557339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T20:13:07.347945Z digest=sha256:3135efb8ce22dee5b36311cf7b6b008e5701570135e9163eb08a29514b8d319d

Observation 17ad1dea-471c-480c-bb83-b909736bb710 · outbound

This paper cites Training language models to follow instructions with human feedback.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Training language models to follow instructions with human feedback

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:07.351732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:07.351732Z digest=sha256:0b381bc86afc139c4d59cdea24aaa78e0c9a12af358a0211a92c535c39dc2f7f

Observation 60f9a0a0-f821-48ab-9bf8-86c3c2de378c · outbound

This paper cites A ‘biased’emerging governance regime for artificial intelligence? how ai ethics get skewed moving from principles to practices.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI A ‘biased’emerging governance regime for artificial intelligence? how ai ethics get skewed moving from principles to practices

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:13:08.536361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T20:13:07.362510Z digest=sha256:cacdfa1b2af37bc36a327350897d128e67234a3583e81a0d0b9218f44742122b

Observation b3aed9a6-5a73-4d0f-83ed-d4ab18aae8be · outbound

This paper cites Automatically Correcting Large Language Models: Surveying the landscape of diverse self-correction strategies.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Automatically Correcting Large Language Models: Surveying the landscape of diverse self-correction strategies

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:07.366906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:07.366906Z digest=sha256:d366c6b2ef7465bafc73e6a7169d5df55015bf62dbc1ee6069ce741c669fc783

Observation e19c3bd3-579c-4dab-a6c1-ee7d5237ce5c · outbound

This paper cites Automated Red Teaming with GOAT: the Generative Offensive Agent Tester.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Automated Red Teaming with GOAT: the Generative Offensive Agent Tester

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:07.373813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:07.373813Z digest=sha256:0c2318d432a66fc35ed56eceb21723de1792b177ad58a176396f74679b4bce5c

Observation f16bb05c-2e1e-434b-85ec-fc4e2aa33710 · outbound

This paper cites Causality.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Causality

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:07.378160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:07.378160Z digest=sha256:16a9134ff776ab5f97052b85096096533f3c44d4c46ae7b80cf30d338b8ab1d5

Observation c8f76e6e-35a8-417c-97d3-025aaf4dc9f2 · outbound

This paper cites The book of why: the new science of cause and effect.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI The book of why: the new science of cause and effect

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:07.383771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:07.383771Z digest=sha256:34763f4071c75557fedbe0d952cadd07bbc2b4a40521396c6b40712655f2e2c6

Observation 769713d2-1ffc-4ad1-acfa-1bd846349a73 · outbound

This paper cites Red Teaming Language Models with Language Models.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Red Teaming Language Models with Language Models

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:07.387941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:07.387941Z digest=sha256:819dec0029a690334a5335440e75a62fbe86b9a980f40d8ecb79714cb07de913

Observation 39e9f853-911d-49f4-b09d-55b6d9e1dc19 · outbound

This paper cites The Tug of War Within: Mitigating the Fairness-Privacy Conflicts in Large Language Models.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI The Tug of War Within: Mitigating the Fairness-Privacy Conflicts in Large Language Models

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:07.392319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:07.392319Z digest=sha256:36608cdc60fe073b29bcd9a88a9018e9683978e5fe6b12dd627a2841743f3bca

Observation d15e6ae2-5b79-456d-a11f-6de72e10181c · outbound

This paper cites Towards Tracing Trustworthiness Dynamics: Revisiting Pre-training Period of Large Language Models.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Towards Tracing Trustworthiness Dynamics: Revisiting Pre-training Period of Large Language Models

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:07.396643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:07.396643Z digest=sha256:5448be0b796f4f176c5f2fd4d523a234caca66c46010384cb2fdd7f8b516004e

Observation 49771c48-993c-4447-9870-c923c6743ed1 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Direct preference optimization: Your language model is secretly a reward model

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:07.401360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:07.401360Z digest=sha256:f7088e9b9d232f2844a749e27fb902a3c29ee0f1f9068926c67e72215ee24936

Observation 151fcb9d-a2b4-4a70-aac3-00df20dfe079 · outbound

This paper cites Natural language processing: transform- ing how machines understand human language (2023).

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Natural language processing: transform- ing how machines understand human language (2023)

Reference 79

Resolution
malformed identifier
no resolver link, observed 2026-08-11T20:13:07.405843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:07.405843Z digest=sha256:c83eb1a650863a62e5fc54d5c714bdd16387cac47c70bb1168271426a63d5eb0

Observation ca19857c-496d-4e75-8bdd-b225bedc72f2 · outbound

This paper cites Identifying Semantic Induction Heads to Understand In-Context Learning.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Identifying Semantic Induction Heads to Understand In-Context Learning

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:07.409981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:07.409981Z digest=sha256:42addcd953afab83c8fc550d2d188811abbdc86ad1d07ada3f158339200d8657

Observation 7a9bea55-8d9c-4049-b80a-02e9a47ebd29 · outbound

This paper cites Self-Reflection in LLM Agents: Effects on Problem-Solving Performance.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Self-Reflection in LLM Agents: Effects on Problem-Solving Performance

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:07.420105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:07.420105Z digest=sha256:a9b1c4a3013a02a20c1dd870fc41d85151b29d8350b510daca5e21f1c6fee74b

Observation 6680b5c9-4332-47e4-902b-d10b72e4bf9e · outbound

This paper cites Scaling Laws for Deep Learning.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Scaling Laws for Deep Learning

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:07.424428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:07.424428Z digest=sha256:5ad35be07a87f5e45ecbadb819af586d46226e570354baa738adfc6025b3c20e

Observation 3efea1a6-9bdf-44c0-941e-e8f650ae0132 · outbound

This paper cites Re- flexion: Language agents with verbal reinforcement learning.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Re- flexion: Language agents with verbal reinforcement learning

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:13:08.499348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T20:13:07.428521Z digest=sha256:91fad5628345abb6779ac8883318db4fe31dd74f90946ebc64398c7bb295bc13

Observation 34d9f7ab-1a35-413c-a767-d6663dfcc92e · outbound

This paper cites Learning to summarize with human feedback.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Learning to summarize with human feedback

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:07.432980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:07.432980Z digest=sha256:94a15924a9bb1710fd5b9831192f1ec25a097b004e1f150ae0ff48a9a057a93c

Observation 23e56d77-ca8f-4a3e-9fd0-98470ed9e8fb · outbound

This paper cites Reinforcement learning: An introduction.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Reinforcement learning: An introduction

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:13:08.479236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T20:13:07.437265Z digest=sha256:2327766486f767a1ec7ddb9be9986df61244616c299b4a12abd1a77ec748945a

Observation 6c44c11f-f240-4c30-97ce-93b98df58aeb · outbound

This paper cites CAT-LLM: Style-enhanced Large Language Models with Text Style Definition for Chinese Article-style Transfer.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI CAT-LLM: Style-enhanced Large Language Models with Text Style Definition for Chinese Article-style Transfer

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:07.441487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:07.441487Z digest=sha256:d391ef75949c4f64303a82444d1c689ef56eb797c109a6ee3546d6e28f0e7fd6

Observation 17f9e657-7e39-4324-8e44-f4fc91b14386 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:07.445719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:07.445719Z digest=sha256:b63d63065d3dd855711c74c77046d91a23a200e31d5b8d9d00c3b3ae66cc3b3d

Observation bc986fd7-fa04-4f3f-82bf-ac25556e8fea · outbound

This paper cites Value sensitive design and responsible innovation.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Value sensitive design and responsible innovation

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:13:08.466864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T20:13:07.450736Z digest=sha256:41290c22b0ae9c34c158cf09155627d60be7326352dddc63f6cc00d2f459f477

Observation 35940662-954e-4a70-849f-92b532d02452 · outbound

This paper cites Interpretable counterfactual explanations guided by prototypes.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Interpretable counterfactual explanations guided by prototypes

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:13:08.451322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T20:13:07.454911Z digest=sha256:025ccbb67fc853966759b730f9db6e35f95c2fd567961d2d1e8bc9c5350e74e9

Observation 02a5e747-6ba4-4c50-939a-b2944992cd86 · outbound

This paper cites Counterfactual Explanations and Algorithmic Recourses for Machine Learning: A Review.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Counterfactual Explanations and Algorithmic Recourses for Machine Learning: A Review

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:07.461046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:07.461046Z digest=sha256:5450e673baea233b6a86118a763178f986b750f9aa43a1a21ee775dc6f1ec752

Observation 5b9ca466-f652-46ac-9542-edbe4183b4a4 · outbound

This paper cites Decodingtrust: A comprehensive assess- ment of trustworthiness in gpt models.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Decodingtrust: A comprehensive assess- ment of trustworthiness in gpt models

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:13:08.438694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T20:13:07.465862Z digest=sha256:3cd7bb05afde039577aeb3396a724cde9ee3beb46f6e839ee31906290aa95224

Observation 1d89373d-94c2-434b-af8d-6d18ed66cf22 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:07.471308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:07.471308Z digest=sha256:53231f1993df7f17fe2a1648428f6833da5ac767fe416b1d7dd07ad77bb588cb

Observation 318432f9-31a9-43b0-a95a-fa2b0538aa85 · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Emu3: Next-Token Prediction is All You Need

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:07.476190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:07.476190Z digest=sha256:0eada8f010e1a86a75f23109ef1c172c9791ce8b5f367f028c8f4861bc3a0e33

Observation 44223556-8486-4980-a048-5f8035dd494e · outbound

This paper cites ai safety as global public goods.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI ai safety as global public goods

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:13:08.425087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T20:13:07.480375Z digest=sha256:ed8a45b751dcd29f3928a1fda956d5f86756b9ef45d1167c433d7a0fe5812b8d

Observation 7337fddf-e43e-412b-aca7-f4f38d81bc88 · outbound

This paper cites Using the veil of ignorance to align ai systems with principles of justice.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Using the veil of ignorance to align ai systems with principles of justice

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:07.484432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:07.484432Z digest=sha256:1c3e34ca9192ffe2286add33e594b5011cc0efb9924e81a2274fd4068f2367f3

Observation 44874b7e-63ac-4ab6-acb2-0fa4c39cdb27 · outbound

This paper cites EFUF: Efficient Fine-grained Unlearning Framework for Mitigating Hallucinations in Multimodal Large Language Models.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI EFUF: Efficient Fine-grained Unlearning Framework for Mitigating Hallucinations in Multimodal Large Language Models

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:07.488518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:07.488518Z digest=sha256:be5a3ccf74a49b7b9c05bda03662efa310aeabc415dda5e321602211fa72d76e

Observation 30b9adf3-ae9b-40a8-a187-bc28328e9873 · outbound

This paper cites Qwen2 Technical Report.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Qwen2 Technical Report

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:07.494182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:07.494182Z digest=sha256:9db1d3dd5dee47f7c8c32cd0188e63749c8a922fc6c629c928a7fb1c0ffb81e6

Observation 79bcf0b3-c09a-4a53-a88e-4b5bb5af8f4a · outbound

This paper cites Huref: Human-readable fingerprint for large language models.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Huref: Human-readable fingerprint for large language models

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:13:08.402605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T20:13:07.498972Z digest=sha256:28142cf0b82e63ffa54ad1802dce201b17d894398773e6688f4d5a611ca04fee

Observation f62cc241-170d-4d78-bebe-41903a85de1f · outbound

This paper cites The Better Angels of Machine Personality: How Personality Relates to LLM Safety.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI The Better Angels of Machine Personality: How Personality Relates to LLM Safety

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:07.503044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:07.503044Z digest=sha256:dd041cf18484fc490f19eb7fc68ec139cb94390242c9721b7ac34831327ad966

Observation 0ed64c77-9407-45b1-9320-5ba6919573f3 · outbound

This paper cites REEF: Representation Encoding Fingerprints for Large Language Models.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI REEF: Representation Encoding Fingerprints for Large Language Models

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:07.507975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:07.507975Z digest=sha256:a02b24ee798cba4cb4434d635a9bd21da289a76d67bc17ae66b3599ad19ec53f

Pith citing papers

Observation b10ab4f7-b386-4150-b888-57121a19ab40 · inbound

An Early Warning of Emerging Biosecurity Risks in Frontier LLMs cites this paper.

An Early Warning of Emerging Biosecurity Risks in Frontier LLMs Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-01T16:21:29.283270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:21:29.283270Z digest=sha256:3f6058d9fa6c466d0aefb2268b2c4e6ce36ba6bcd4c13008456c573203316dc7