Pith. sign in

Paper Citation Record · LEDGER

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values

As of 24 August 2026, this Paper Citation Record lists 88 of 88 outbound references and 3 inbound Pith citation observations for arXiv:2501.07071.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.07071 v3

Coverage vector

measured 88 of 88 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T20:53:59.936442Z

measured 91 of 91 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:46:04.231544Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T19:21:09.039911Z

Reference resolution

88 of 88 outbound references displayed

  • verified exact4
  • verified fuzzy0
  • unresolved84
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation dd8d6cff-19a7-4799-9a23-a2146dfe54df · outbound

This paper cites Moral Foundations of Large Language Models.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values Moral Foundations of Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:59.429593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:53:59.429593Z digest=sha256:8b25276ede4c06674b44d01a8d6ea09a87af207b42275bc3534588d4094a7460

Observation 90125a4d-96ab-4dae-8f09-e54829f43ebb · outbound

This paper cites an unresolved cited work.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:59.437040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:53:59.437040Z digest=sha256:46fb6c101f2bad34e409cafadf3b14d162cdab6683c2b6cdb7d52e61eb200092

Observation a35c3028-2628-4057-a4f6-96a9642f965e · outbound

This paper cites an unresolved cited work.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:59.441226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:53:59.441226Z digest=sha256:7cb4d6a3bb33890f274367ef674a9009136aa9d28dc4e8c36f92867b759beb20

Observation e4a518c0-b5fd-489e-b26d-4b5e375eac42 · outbound

This paper cites Measuring Implicit Bias in Explicitly Unbiased Large Language Models.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values Measuring Implicit Bias in Explicitly Unbiased Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:59.445391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:53:59.445391Z digest=sha256:bf4e8146ac4a9b6609ab1ec3eaac5268c26c473bc88604f187fa5f81ab6b0582

Observation 2152b734-53f1-4b02-8d5b-040a0e4847bb · outbound

This paper cites an unresolved cited work.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:59.449984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:53:59.449984Z digest=sha256:3739a2f20a3aca5d61e433a4af6ff07e3df32740d5c80156b8938c0176c72e11

Observation de2b5f2b-a18e-49d4-af8a-6abc410f2ac2 · outbound

This paper cites an unresolved cited work.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:59.454773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:53:59.454773Z digest=sha256:f124b746e4456b35f4363b9414bd046da66c829f829b3042f7d167c4158fd1f7

Observation 02afae19-2142-4cc8-87d5-c03d7ce3bbd1 · outbound

This paper cites Beyond Human Norms: Unveiling Unique Values of Large Language Models through Interdisciplinary Approaches.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values Beyond Human Norms: Unveiling Unique Values of Large Language Models through Interdisciplinary Approaches

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:59.458896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:53:59.458896Z digest=sha256:e4d8d74beac7c726cf9a395a6f354a373aeccd93a75e4d6985ab42d6f69daaa3

Observation e5998d49-aa8f-434b-8cbb-310803b99f2a · outbound

This paper cites an unresolved cited work.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:59.463351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:53:59.463351Z digest=sha256:2b4d97aedc712717d9f04a4f9506db5c20181a34e17af158c9be14183974a008

Observation 1e507837-2256-487c-b89a-bb35617d590f · outbound

This paper cites an unresolved cited work.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:59.467465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:53:59.467465Z digest=sha256:b9f4befcb04f279fad3c0f3809532c0dc524f1df6c925c7fd416d9f18d0fc160

Observation 387bb31f-05d9-458f-aef4-8e260b7cec4a · outbound

This paper cites an unresolved cited work.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:59.625463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:53:59.625463Z digest=sha256:0350651f9e36ac470d2fc624db53d70bb43566d9ae0edb6a1407793d9b6d0d95

Observation 5e49c73a-3dd4-44e7-9a57-3a0757b5e38f · outbound

This paper cites Sparks of Artificial General Intelligence: Early experiments with GPT-4.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values Sparks of Artificial General Intelligence: Early experiments with GPT-4

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:59.629480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:53:59.629480Z digest=sha256:5423987416dcdc0663ec672fcf1e68b8823007338ef8c546f3f0c025a408d160

Observation 3965dec3-5f17-4f77-9e4b-9f7cf92dd639 · outbound

This paper cites an unresolved cited work.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:59.633460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:53:59.633460Z digest=sha256:8f5dca614ae1807a879283ea1e9afd76cd9fb4f57d5f046e4870752c1f3f65ae

Observation 1bd5cd1d-8a6a-4068-9989-60c15703c4c4 · outbound

This paper cites an unresolved cited work.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:59.637094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:53:59.637094Z digest=sha256:01c9137e8b0db1ca80e23ca82b67c1c47830106c16824c34b05101635246a9cf

Observation 1cc5fb84-910a-472e-a815-4e50599cfbf7 · outbound

This paper cites an unresolved cited work.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:59.640835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:53:59.640835Z digest=sha256:383cac738aace8246e06e14c9dc63a5807be14920b47a27f74438fff13191787

Observation 9fc528a2-7809-47d8-b0ff-764a294e0ce6 · outbound

This paper cites an unresolved cited work.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:54:00.894458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-10T20:53:59.646451Z digest=sha256:857230ba087fc20ecc5a0cf243d23ffce971c2a32242809715795dc0c4d10354

Observation f1874cb0-53ba-4c1c-86a8-5dc646a6405c · outbound

This paper cites an unresolved cited work.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:54:00.883350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-10T20:53:59.650780Z digest=sha256:916c8a49756e89854d6e23299f4aa28792a30be0ac0b23518d813403dcbe80cc

Observation df82eac1-aaa2-4eee-b422-ebb9336a8db7 · outbound

This paper cites an unresolved cited work.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:59.654576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:53:59.654576Z digest=sha256:0f7c83e893048dc0239ab95d05c34315c13457c180d6f8c76e33f7a570aaeeb6

Observation a301184e-0c90-44c5-b7b6-75ec68c7f740 · outbound

This paper cites Generalization or Memorization: Data Contamination and Trustworthy Evaluation for Large Language Models.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values Generalization or Memorization: Data Contamination and Trustworthy Evaluation for Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:59.658656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:53:59.658656Z digest=sha256:644c59cbe63357c5dfac6cf82e938a253cb87f2ae6fbd90f673c98e38ce020a8

Observation ab5e3816-e522-4bb7-aa0e-f6f6b9f3b044 · outbound

This paper cites Denevil: Towards Deciphering and Navigating the Ethical Values of Large Language Models via Instruction Learning.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values Denevil: Towards Deciphering and Navigating the Ethical Values of Large Language Models via Instruction Learning

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-10T20:54:00.524349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-10T20:53:59.661902Z digest=sha256:6d3d26746bffa168f13a01f423b34f8b938d2a89295011f1d0612d581cb4b7f3

Observation f8497d2e-90e8-48cc-96fd-8af105de3a85 · outbound

This paper cites The Llama 3 Herd of Models.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values The Llama 3 Herd of Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:59.665341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:53:59.665341Z digest=sha256:b813c4c9f0927d8c997e6e1469d95058c8cfb69b007196529a4ebd339b44fd05

Observation 920b5441-97a4-4cdb-a422-383ac4007d7e · outbound

This paper cites NPHardEval: Dynamic Benchmark on Reasoning Ability of Large Language Models via Complexity Classes.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values NPHardEval: Dynamic Benchmark on Reasoning Ability of Large Language Models via Complexity Classes

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:59.668297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:53:59.668297Z digest=sha256:eb19e434e9e7131161fc4bc59d84a9e6eef127f851edb9488c85024e6c9c2869

Observation 6bb883c4-043e-4b3e-ab2e-54c88fb9f76f · outbound

This paper cites Does Moral Code Have a Moral Code? Probing Delphi's Moral Philosophy.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values Does Moral Code Have a Moral Code? Probing Delphi's Moral Philosophy

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:59.671503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:53:59.671503Z digest=sha256:64db7122005e834cc8ab378990857eff79355842594917e19ef09b33a6c82e31

Observation e6343fe4-6cfe-4513-aaab-e46cd783c690 · outbound

This paper cites RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:59.674695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:53:59.674695Z digest=sha256:e60eb99b3bcb9248105afd182dffec31db8712a44c3de08b572c33d0a3840c0c

Observation 2362bf30-31d0-4643-9423-d4b696de4a08 · outbound

This paper cites an unresolved cited work.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:59.678537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:53:59.678537Z digest=sha256:376c2a9d2995dc7a5e59f7689b88ab3e484bec81915582303d78d6f86f036e16

Observation 557eb183-0902-4450-aa0e-35727c0a5394 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:59.682240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:53:59.682240Z digest=sha256:2285e9ddb5e90b82fbf0e3bc0d56ba62876f1fb6ecba26a76c01b91d1444f00c

Observation e562fc96-9710-42ce-a379-d6470b3f0c13 · outbound

This paper cites Aligning AI With Shared Human Values.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values Aligning AI With Shared Human Values

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:59.686181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:53:59.686181Z digest=sha256:089a4ece61db307de4adf06237c6f1abfbf6ab6e39a44e76ff96f138123194e9

Observation b60e7b9d-07b3-49a1-9589-a6aec22a6586 · outbound

This paper cites an unresolved cited work.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:54:00.858628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-10T20:53:59.689923Z digest=sha256:476cc500f9f63e72ace4354a64fca16367deed5a0dbe8f53ce7a472167f2ddc1

Observation 7c8218de-e527-4482-acc0-339adebc83cf · outbound

This paper cites an unresolved cited work.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values Unresolved cited work

Reference 28

Resolution
verified exact
doi, observed 2026-08-10T20:53:59.981021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-10T20:53:59.693591Z digest=sha256:776711721c14dc17f71472224801cd93c8a4f0e253918ea120126db196f41ef3

Observation 0726bc81-d9ad-4529-8851-d12004ebf73b · outbound

This paper cites an unresolved cited work.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:54:00.846106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-10T20:53:59.697363Z digest=sha256:432d0f10d66770206ec1aac6edf9590b3c2f0f7154c970dcbf29c9395b640607

Observation a93fecbc-57ab-4676-a326-5f88daccd6ea · outbound

This paper cites TrustGPT: A Benchmark for Trustworthy and Responsible Large Language Models.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values TrustGPT: A Benchmark for Trustworthy and Responsible Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:59.701166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:53:59.701166Z digest=sha256:fec8c82c1d1f08035954504313d65a6bddf7e089b4b71144a0785a1952dd83b4

Observation 0aa2ac91-789b-4524-bb3d-1cf9ca964083 · outbound

This paper cites GPT-4o System Card.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values GPT-4o System Card

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:59.704944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:53:59.704944Z digest=sha256:2584105450db43e2625a09847abf322249dd2b0b6bbe250b01aca9f05329bf8e

Observation d325663a-f4da-4045-980f-87e2da218a26 · outbound

This paper cites MoralBench: Moral Evaluation of LLMs.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values MoralBench: Moral Evaluation of LLMs

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:59.708695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:53:59.708695Z digest=sha256:488501d72b0793b2cf1dd3bf4547f98a3049252704ea26d50fdc3bf7926bad03

Observation 35a4ac06-1d4d-4a8b-a532-5cb3c5444303 · outbound

This paper cites Raising the Bar: Investigating the Values of Large Language Models via Generative Evolving Testing.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values Raising the Bar: Investigating the Values of Large Language Models via Generative Evolving Testing

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:59.712610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:53:59.712610Z digest=sha256:22970642092de1585aded787f754a2a517af4a2f63209a5e3afbe0bb56337dcc

Observation 5eeabbc7-5fc9-4684-99cf-052fc09fb29c · outbound

This paper cites Can Machines Learn Morality? The Delphi Experiment.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values Can Machines Learn Morality? The Delphi Experiment

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:59.716667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:53:59.716667Z digest=sha256:9b701983314768e98876496901a725dde5af8db62996792549d6b13542bc4bdd

Observation 81e4d102-f62d-4f69-ad8b-25bda705a52d · outbound

This paper cites Scaling Laws for Neural Language Models.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values Scaling Laws for Neural Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:59.720353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:53:59.720353Z digest=sha256:123f0f354de1bf6f3e94b08835ed2ea4c28deb9726833facef4a8ccfb39a381e

Observation 7e0382f0-aff1-4646-9045-9abe5db402a5 · outbound

This paper cites an unresolved cited work.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:54:00.835245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-10T20:53:59.724133Z digest=sha256:5752c91b59e55f831218591239890d683c37bbc96373e596556a828bdf13c542

Observation d2c5eb2a-fff2-45d8-9a89-edbe51a009fc · outbound

This paper cites an unresolved cited work.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:54:00.824234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-10T20:53:59.728298Z digest=sha256:17ca8117ae5fc0a8acde5f04f97f007bc35012a53c299dbe821485fc8e9edbde

Observation 1c6c52ea-5495-427d-b620-6428913a1a76 · outbound

This paper cites SALAD-Bench: A Hierarchical and Comprehensive Safety Benchmark for Large Language Models.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values SALAD-Bench: A Hierarchical and Comprehensive Safety Benchmark for Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:59.732280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:53:59.732280Z digest=sha256:75762c0ca58faf87c1721fa35b28e1a3526aeb0edec6f4655c8d8c3a1a4640f7

Observation 9d6439d2-e77b-4d61-a9ac-b786c99d9dec · outbound

This paper cites Holistic Evaluation of Language Models.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values Holistic Evaluation of Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:59.736259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:53:59.736259Z digest=sha256:0e0efce323cbad4403dfab20ae5e7946be641a19cad698df2741075ed5f69692

Observation 45977ecc-aacb-4fae-9985-8f9d6709ea8f · outbound

This paper cites an unresolved cited work.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:54:00.813196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-10T20:53:59.740028Z digest=sha256:275ee59f92d1ec1b94313f16e0b00ad790908b7e0629b0878290866a42625edf

Observation 8bb3c084-bbd9-44f6-9d3a-b3bdc459f127 · outbound

This paper cites an unresolved cited work.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:54:00.801054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-10T20:53:59.743970Z digest=sha256:4ab106ee2f3106b51ce183937e958ebaa8c5b4387688ce1e6bc51908847fe12a

Observation 08f63be4-8dcd-45f5-af5a-96b6b12a737a · outbound

This paper cites an unresolved cited work.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:59.747894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:53:59.747894Z digest=sha256:a05f537e6c4df772726722328cd6b32cf45b4385c26a5f4176a0ae0e4ea9b431

Observation 87665183-cee0-4915-8e44-a3e63cd4ddd5 · outbound

This paper cites Cultural Alignment in Large Language Models: An Explanatory Analysis Based on Hofstede's Cultural Dimensions.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values Cultural Alignment in Large Language Models: An Explanatory Analysis Based on Hofstede's Cultural Dimensions

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:59.751693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:53:59.751693Z digest=sha256:2d916343bc5148fa3d38df29132bae8317d0a4aa4300d9328383690f15b3dc54

Observation c47c6702-2324-45b8-97e7-3cc6ee6de9a1 · outbound

This paper cites Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:59.756986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:53:59.756986Z digest=sha256:f6c9e8179225d0d9a1bef1a24d9ca2323ae0b0884f7df9f974502784804d3a89

Observation d8105e7d-a15d-4fe0-8362-bd2309c3a094 · outbound

This paper cites LocalValueBench: A Collaboratively Built and Extensible Benchmark for Evaluating Localized Value Alignment and Ethical Safety in Large Language Models.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values LocalValueBench: A Collaboratively Built and Extensible Benchmark for Evaluating Localized Value Alignment and Ethical Safety in Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:59.762857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:53:59.762857Z digest=sha256:6851be33e1c5972407680bf2f415a9f66d0b9e506948b9cbe7cf59156a075e7e

Observation 3462bc9c-7cd5-4206-8253-aac61f24ca3f · outbound

This paper cites SG-Bench: Evaluating LLM Safety Generalization Across Diverse Tasks and Prompt Types.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values SG-Bench: Evaluating LLM Safety Generalization Across Diverse Tasks and Prompt Types

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:59.766722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:53:59.766722Z digest=sha256:b1ca6f36eafa9bdcdcb05704c6f2341c1aa9590d1b0cebd1ea4cfafee1f4bc6d

Observation 002cfe78-1e13-4c77-9045-4f4c15bbb648 · outbound

This paper cites CrowS-Pairs: A Challenge Dataset for Measuring Social Biases in Masked Language Models.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values CrowS-Pairs: A Challenge Dataset for Measuring Social Biases in Masked Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:59.770377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:53:59.770377Z digest=sha256:a0c42b2471ab414212e0a0b8a62fa779cb2ad43c80cfe68aaa6f68b36525876c

Observation 7158a678-927b-4410-8307-058b714631cf · outbound

This paper cites an unresolved cited work.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values Unresolved cited work

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:59.773824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:53:59.773824Z digest=sha256:5d2afa2e5e7826a2b0879533a0e39b64557f3ae7be43f3717a6e5f11df1d1e21

Observation 7ee81e38-c4a8-4e42-9cbc-621fdd6ee537 · outbound

This paper cites BBQ: A Hand-Built Bias Benchmark for Question Answering.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values BBQ: A Hand-Built Bias Benchmark for Question Answering

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:59.777521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:53:59.777521Z digest=sha256:57b76ba9a8d06d9574a3a946068f89280bc8057399f9877eddda612ba69d5dcb

Observation 2d0985df-4f72-4989-9231-393c3ac33011 · outbound

This paper cites an unresolved cited work.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:54:00.773044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-10T20:53:59.781392Z digest=sha256:f15bcb37fb02d0373f08c91122dfc0383006ae23f12279ef5992515f9a07b0eb

Observation 4dad50ce-a0c8-4b81-a557-b05893cbfabf · outbound

This paper cites an unresolved cited work.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:54:00.759174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-10T20:53:59.785097Z digest=sha256:8c089815d66489f6776a0aaba0d65f7752f3b6f570250bc5171b0d718b976d00

Observation 90b27071-bc62-43d2-8f29-ed12296a46d3 · outbound

This paper cites ValueBench: Towards Comprehensively Evaluating Value Orientations and Understanding of Large Language Models.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values ValueBench: Towards Comprehensively Evaluating Value Orientations and Understanding of Large Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:59.788544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:53:59.788544Z digest=sha256:91c22bd50ab5cd566795b9f620c6b8556b45a7d9d186bd13a3ae8dfa70d01f9d

Observation 09d7e566-cda0-4dda-82ab-35cd44f4bac0 · outbound

This paper cites an unresolved cited work.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:54:00.743503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-10T20:53:59.791934Z digest=sha256:304ee72a67344d52fd78c559b10f161370169e91f83bcc9d92577a1936549d22

Observation 228822bd-e71e-4187-bc37-bf449c904394 · outbound

This paper cites an unresolved cited work.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values Unresolved cited work

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:59.795899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:53:59.795899Z digest=sha256:679a3198e045762a7e8c93d700e34dd6df85d4b4c91c569eab52e88a84cba526

Observation 9890f054-8132-4903-93d6-ad6fc2168552 · outbound

This paper cites an unresolved cited work.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values Unresolved cited work

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:59.799510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:53:59.799510Z digest=sha256:026a267c4a57215abf9f9e7139b211b680eabb3ce293ed3461e124d9f5e396d2

Observation 67d50a26-7c78-4332-8a94-6dc5cba44f15 · outbound

This paper cites an unresolved cited work.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values Unresolved cited work

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:59.803413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:53:59.803413Z digest=sha256:85b84b48f7ec5d9a8f532bca7210e7bff515658bf8c08b27ba4adff991556b9f

Observation 0dfca55d-17f4-428e-aaf0-dfa116b7c506 · outbound

This paper cites an unresolved cited work.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:54:00.709427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-10T20:53:59.807191Z digest=sha256:0defaeb22265492f7d3a9e7a9b1dc72f7f24bc4833b87dbceb761b76cbe37ce3

Observation 02f62884-0fa9-4d41-9ca0-bfeeab3bea14 · outbound

This paper cites Moral Mimicry: Large Language Models Produce Moral Rationalizations Tailored to Political Identity.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values Moral Mimicry: Large Language Models Produce Moral Rationalizations Tailored to Political Identity

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:59.811260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:53:59.811260Z digest=sha256:eca56817a70995bd52bf73a991aa9b0d635da9766affa74741d363c748a04228

Observation c4028256-53a1-4374-9c62-2183500790b7 · outbound

This paper cites an unresolved cited work.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:54:00.697999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-10T20:53:59.814793Z digest=sha256:7aeafe7fa38e2aafe015e744926272e021a902bfcc129599b0e3df99bb41caee

Observation 5bcaeb03-1931-486f-959c-ad3f510f22e3 · outbound

This paper cites an unresolved cited work.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:54:00.685493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-10T20:53:59.819166Z digest=sha256:48958e334e0c5f83b31145b31168f40a686dbeac37bc37b142bc8f7021947cc0

Observation 47b6bbea-6854-45dc-a796-53692e89be89 · outbound

This paper cites an unresolved cited work.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:54:00.674668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-10T20:53:59.825603Z digest=sha256:9a8db23dd1b5036a94a493ad57be52246cbcb4de4bfb677748c68c3720372dff

Observation 9eb9ba33-b22c-4f49-806b-c3010c53a40a · outbound

This paper cites Safety Assessment of Chinese Large Language Models.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values Safety Assessment of Chinese Large Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:59.829561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:53:59.829561Z digest=sha256:3264814e9574c301faa472d8b64f8dde389933432e3160b21c775c040a069530

Observation 1ebd31de-47d1-48c1-b9c8-0ed0de7d8330 · outbound

This paper cites TrustLLM: Trustworthiness in Large Language Models.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values TrustLLM: Trustworthiness in Large Language Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:59.833852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:53:59.833852Z digest=sha256:365e0c99d0158a7bae14fbd4bda1e8bea93bcb22581d4d684667d8f1f8e53f35

Observation 516397f5-0974-43c6-a3e9-06628db84389 · outbound

This paper cites an unresolved cited work.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:54:00.662700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-10T20:53:59.839050Z digest=sha256:4cb5c4dbae89e244615a8080bdaad80057e2233c26e6145e4ae9e35698b04957

Observation a7f07b41-bfd9-4aff-a4f7-4e11e6cf1e9e · outbound

This paper cites an unresolved cited work.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values Unresolved cited work

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:59.843196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:53:59.843196Z digest=sha256:e9f55cec0d5f33f9c051bcda7b26803ac79e3cba3de9113bddefe01a2d9a0c39

Observation d0f45bd8-bffb-4ab8-97b2-f7f0d736b85f · outbound

This paper cites The information bottleneck method.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values The information bottleneck method

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:59.847196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:53:59.847196Z digest=sha256:ae19af8ece314aeb09f6f525b2ffd908ebe11bce1792959e75b34cb08859a6bc

Observation cd03390e-9659-4c26-8995-3b5cd38df0b3 · outbound

This paper cites DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:59.851548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:53:59.851548Z digest=sha256:14f656f7cbf1fe0860952ed656e3df7cf1113deddedc908ab8c90d114e23646a

Observation 6092e811-ef8e-41cf-bfaf-922c1024dde9 · outbound

This paper cites Self-Instruct: Aligning Language Models with Self-Generated Instructions.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values Self-Instruct: Aligning Language Models with Self-Generated Instructions

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:59.855283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:53:59.855283Z digest=sha256:427b11b8e17630f9edffec8549ef533f102f0dfcd7f52f0823b05252feb322a2

Observation c92bd5c0-e3ed-4b80-abeb-485b19cb41af · outbound

This paper cites Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:59.858945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:53:59.858945Z digest=sha256:df5a0789145ac32c55850243e82e60839f6f5842fdbd6a04e73e8f79d8ab718e

Observation f649b0b2-3fa8-4401-8d6e-0fb492700756 · outbound

This paper cites Emergent Abilities of Large Language Models.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values Emergent Abilities of Large Language Models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:59.862849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:53:59.862849Z digest=sha256:7b2cf49d389c9248177acfd2da907c37577167e056b7c1cce389dcef1087e537

Observation 2f765dec-5099-4bf4-a8b7-c62e802f8b3b · outbound

This paper cites Ethical and social risks of harm from Language Models.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values Ethical and social risks of harm from Language Models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:59.866678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:53:59.866678Z digest=sha256:e6b6f0745aec7ff55e40c41d555b4a446a833fd1cc57369dbb32cc3bc4d4a53a

Observation 97a88101-800e-49dd-ba55-7f351f035a2f · outbound

This paper cites an unresolved cited work.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values Unresolved cited work

Reference 72

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:54:00.643551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-10T20:53:59.871527Z digest=sha256:914dd537c2ab55e21154ba49fe8ccf1b386211febd842a43567468e1dad86317

Observation c076d058-3e67-42b3-8dbc-b7ca71eba2c2 · outbound

This paper cites Evaluating Evaluation Metrics: A Framework for Analyzing NLG Evaluation Metrics using Measurement Theory.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values Evaluating Evaluation Metrics: A Framework for Analyzing NLG Evaluation Metrics using Measurement Theory

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:59.875120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:53:59.875120Z digest=sha256:1f6b3c91f6f25657967751f250d3f5e73275cba526933c92f18e9d5a8ad03777

Observation 47c34cd6-e8b0-4aea-9644-d83618796c28 · outbound

This paper cites CValues: Measuring the Values of Chinese Large Language Models from Safety to Responsibility.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values CValues: Measuring the Values of Chinese Large Language Models from Safety to Responsibility

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:59.879566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:53:59.879566Z digest=sha256:e08759b33ec4ec6c7052e2c670d813068ab978b6476125f469317437f904287c

Observation 87e1d68f-20c8-4e7a-933b-82eebca4ad30 · outbound

This paper cites SC-Safety: A Multi-round Open-ended Question Adversarial Safety Benchmark for Large Language Models in Chinese.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values SC-Safety: A Multi-round Open-ended Question Adversarial Safety Benchmark for Large Language Models in Chinese

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:59.883464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:53:59.883464Z digest=sha256:5ed75c4f1632a488f792f4ce57930e65f75f506af0545339fdfbe19501865850

Observation a0ef0138-546d-4999-a425-0aca1a3cf8bb · outbound

This paper cites Value FULCRA: Mapping Large Language Models to the Multidimensional Spectrum of Basic Human Values.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values Value FULCRA: Mapping Large Language Models to the Multidimensional Spectrum of Basic Human Values

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:59.887140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:53:59.887140Z digest=sha256:57161e4fb2a2e95550bec21222bff7f91f50e66ffc158f6631786b69850a704b

Observation b2bc02b3-2ba4-424d-864b-b090b13fc88d · outbound

This paper cites CLAVE: An Adaptive Framework for Evaluating Values of LLM Generated Responses.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values CLAVE: An Adaptive Framework for Evaluating Values of LLM Generated Responses

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:59.891154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:53:59.891154Z digest=sha256:d57ee782b5359519885bfb6a2cad12e7cb17b006fa30c3e0f8635aecf790cc47

Observation 02337954-07a9-47a4-befb-ad8e200b2773 · outbound

This paper cites R-Judge: Benchmarking Safety Risk Awareness for LLM Agents.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values R-Judge: Benchmarking Safety Risk Awareness for LLM Agents

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:59.894939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:53:59.894939Z digest=sha256:459da0400e944150686bd94d39a6c53e1ad52ccddde87a60334ea2b9a09c45b2

Observation b794624a-4476-4186-9a07-f96a5c272ab1 · outbound

This paper cites S-Eval: Towards Automated and Comprehensive Safety Evaluation for Large Language Models.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values S-Eval: Towards Automated and Comprehensive Safety Evaluation for Large Language Models

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:59.898841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:53:59.898841Z digest=sha256:2e6a562d41399ccdc1a82b4dfda246323aa48bb39a0f54e1a3d82edc80398d3e

Observation 997b0e67-c338-4f4a-83ac-50fd34814006 · outbound

This paper cites Quantifying Risk Propensities of Large Language Models: Ethical Focus and Bias Detection through Role-Play.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values Quantifying Risk Propensities of Large Language Models: Ethical Focus and Bias Detection through Role-Play

Reference 80

Resolution
verified exact
local_arxiv, observed 2026-08-10T20:54:00.079928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-10T20:53:59.902745Z digest=sha256:0909475d362e845df22b977d280cd458d5531a39e66cf98ed875bf466d471657

Observation b4df0278-4455-45e0-9d3b-66eac63bc6c5 · outbound

This paper cites SafetyBench: Evaluating the Safety of Large Language Models.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values SafetyBench: Evaluating the Safety of Large Language Models

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:59.906558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:53:59.906558Z digest=sha256:0da4b7c4447173baf0bb2b572694b128c2f774de183cd1ddc24605caab68d2d3

Observation 84d81f90-ef08-4a43-916e-e914e35e0b2d · outbound

This paper cites an unresolved cited work.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values Unresolved cited work

Reference 82

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:54:00.633860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-10T20:53:59.910983Z digest=sha256:da370e3ddda678e6e014356fbdcebd53fdb3c4e274a0cf6f6480a93cea759542

Observation 70411700-5270-4d20-95a0-527d5562ab3b · outbound

This paper cites WorldValuesBench: A Large-Scale Benchmark Dataset for Multi-Cultural Value Awareness of Language Models.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values WorldValuesBench: A Large-Scale Benchmark Dataset for Multi-Cultural Value Awareness of Language Models

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:59.915037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:53:59.915037Z digest=sha256:adcca6937de185bfeddfaf6f4b9ab84590e01597b0d70eb16390dfa7d73fc811

Observation 744f3462-4ac5-4d01-8d7d-5ee7317e11e1 · outbound

This paper cites an unresolved cited work.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values Unresolved cited work

Reference 84

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:54:00.622560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-10T20:53:59.919801Z digest=sha256:5f28940de304d81db2943c27306e82dc908410b3c6c9a912ac6d89c4129795f7

Observation 8e1aec2d-a416-4962-9679-d82b97301936 · outbound

This paper cites The Moral Integrity Corpus: A Benchmark for Ethical Dialogue Systems.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values The Moral Integrity Corpus: A Benchmark for Ethical Dialogue Systems

Reference 85

Resolution
verified exact
local_arxiv, observed 2026-08-10T20:54:00.019800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-10T20:53:59.923515Z digest=sha256:d200ba8c78191dec4475392d711a6e742e4ed972d50f6657b4c66bd3888380fa

Observation bf421a58-3622-49d4-a643-151e5d1ba0a8 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:59.927521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:53:59.927521Z digest=sha256:8889391bf138b68fe1b7614ba5abe11ac6d098a61627b245ed3512e389ede9a9

Observation 7afd67a7-80e7-4cfc-bd6f-f7020f1bd535 · outbound

This paper cites online" 'onlinestring :=.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values online" 'onlinestring :=

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:59.931772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:53:59.931772Z digest=sha256:0362b4d54c6acca15d7ed3aed1f70514257e172db743d93f503b4eedaaa93d98

Observation ad811448-7a59-442c-9e58-bc63cdf0a137 · outbound

This paper cites write newline.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values write newline

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:59.936442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:53:59.936442Z digest=sha256:58a7748acf19aa585622a870f095d0f56ea1c40f4bed26933c40a1a50b09c1f5

Pith citing papers

Observation ea1647ac-5d7f-4717-a1ea-ba916413f40d · inbound

Multi-level Value Alignment in Agentic AI Systems: Survey and Perspectives cites this paper.

Multi-level Value Alignment in Agentic AI Systems: Survey and Perspectives Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T04:46:04.231544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:46:04.231544Z digest=sha256:ed18a3cc853f3bc6078a97c25bd5b864e2add208d345a0a3e3a5347235eb55bf

Observation b16fd9f5-104b-49ca-a883-69d8469619d2 · inbound

Domain Specific Benchmarks for Evaluating Multimodal Large Language Models cites this paper.

Domain Specific Benchmarks for Evaluating Multimodal Large Language Models Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values

Reference 134

Resolution
unresolved
no resolver link, observed 2026-08-07T00:39:42.140955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:39:42.140955Z digest=sha256:4729aac357995481b5fa6b1a3e7364c21dc0856ac23b890a0b8b6f74d5fabccd

Observation ba55b8d0-fe57-48a5-ab9c-94f77a0b1f71 · inbound

When AI Speaks, Whose Values Does It Express? A Cross-Cultural Audit of Individualism-Collectivism Bias in Large Language Models cites this paper.

When AI Speaks, Whose Values Does It Express? A Cross-Cultural Audit of Individualism-Collectivism Bias in Large Language Models Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:21:09.045430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-08T12:08:47.116158Z digest=sha256:c5c995ea32e9ace0ac80e47781eaf39cfe84026d149feba2fcd674a1759585be