Pith. sign in

Paper Citation Record · LEDGER

Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2304.03279.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2304.03279 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T12:46:26.589118Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

28
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d6e829ee-a8f4-4e19-ac1d-65563706e160 · inbound

A Roadmap to Pluralistic Alignment cites this paper.

A Roadmap to Pluralistic Alignment Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark

Reference 190

Resolution
verified exact
arxiv_id, observed 2026-05-16T14:37:53.622855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-16T14:37:53.279275Z digest=sha256:db3c3346a70a5f576256b3890cc5d7d2caadc92abfe37ba267ec1861f2d0ae31

Observation 88056ed4-f8c0-44ed-bb6c-0d29398f998f · inbound

Deception in LLMs: Self-Preservation and Autonomous Goals in Large Language Models cites this paper.

Deception in LLMs: Self-Preservation and Autonomous Goals in Large Language Models Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T12:46:26.589118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T12:46:26.589118Z digest=sha256:7cc7372f816dd9cb5367c5d4fc16d14cc82acdc582bc51f2b1d86fa86896f012

Observation 99247684-0c9e-447f-aafd-d8371dd15337 · inbound

The Odyssey of the Fittest: Can Agents Survive and Still Be Good? cites this paper.

The Odyssey of the Fittest: Can Agents Survive and Still Be Good? Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T19:23:42.468613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T19:23:42.468613Z digest=sha256:64ffa86110eceaf58073047d9144a8738dcfe4412d6387309586e1ad069eeff5

Observation b307d4f7-9636-4544-bb80-90daabd2a075 · inbound

Compromising Honesty and Harmlessness in Language Models via Deception Attacks cites this paper.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T05:42:43.435854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:42:43.435854Z digest=sha256:c84a5ab3a5c68e4f4f31da1f08335f6ce01a5e6b0562fb55a33d4151fa4a4d7e

Observation 1192a071-1ed1-46a6-a603-41d70e10d646 · inbound

Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs cites this paper.

Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-08T00:04:57.367167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:04:57.367167Z digest=sha256:28105417b4831196c1fae4e1c1cca146591ab4c78c835e81c43f5b4911f5587f

Observation 7f55633a-4be4-43d8-adfc-2160525c1239 · inbound

Moral Susceptibility and Robustness under Persona Role-Play in Large Language Models cites this paper.

Moral Susceptibility and Robustness under Persona Role-Play in Large Language Models Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-17T23:20:27.811163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-17T23:16:56.957905Z digest=sha256:8319b57edbac820f1c3bba75174fa4c5f61b7885e974acf43c4ac1ee8de13042

Observation 3249d592-9433-4f85-99ea-2d2ebb4485fe · inbound

Beyond Benchmark Islands: Toward Representative Trustworthiness Evaluation for Agentic AI cites this paper.

Beyond Benchmark Islands: Toward Representative Trustworthiness Evaluation for Agentic AI Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-22T10:21:23.244031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-22T10:19:56.003219Z digest=sha256:6118ab2c205a73e237b9ce18c7c7976cee758bb14f9ff080eaa8730ef7605dce

Observation 3b8da89d-87e0-4d07-b336-25fb556d38f0 · inbound

Restoration, Exploration and Transformation: How Youth Engage Character.AI Chatbots for Feels, Fun and Finding themselves cites this paper.

Restoration, Exploration and Transformation: How Youth Engage Character.AI Chatbots for Feels, Fun and Finding themselves Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark

Reference 64

Resolution
malformed identifier
arxiv_id, observed 2026-05-15T12:45:37.283306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T12:44:53.677818Z digest=sha256:80ddcf700e190561b95015067403f3b33681af2dec2a2d2b2450ef774ea42648

Observation ff9e1600-92e5-4847-8f14-6e8ccc1694b7 · inbound

Positive Alignment: Artificial Intelligence for Human Flourishing cites this paper.

Positive Alignment: Artificial Intelligence for Human Flourishing Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark

Reference 148

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:59:47.953274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-15T05:56:56.902705Z digest=sha256:e72202adcaf788999db85c5a93717162148002b3be3b0fed13e6ffa3b5ef87d8

Observation ce6bc9a0-d7d8-4f36-a07d-38740ddff390 · inbound

Correcting Influence: Unboxing LLM Outputs with Orthogonal Latent Spaces cites this paper.

Correcting Influence: Unboxing LLM Outputs with Orthogonal Latent Spaces Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark

Reference 228

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T20:17:55.545244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-14T20:17:01.224864Z digest=sha256:bcca4749ab6c83381736e9cae0818555c5f685d503017422e55b73315ca655e3

Observation e8a3e7e2-a852-48ae-8c54-3e7a98dff719 · inbound

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents cites this paper.

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-21T01:43:56.709741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-21T01:42:55.693115Z digest=sha256:7cf9144dce5d7be2913076a221b4a4f015f3487b309f4c24d36a48d3e41cec84

Observation d0da8262-22b1-4148-8e7c-96665a13f2c1 · inbound

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety cites this paper.

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:51:08.014677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-22T05:50:28.114140Z digest=sha256:8f874acc87775007f9b2a140079b314d96c5363594b66a8e5de34436a2f1ceaf

Observation 8d56c348-3b2b-49cd-a2bb-bd8e7a2fa2cb · inbound

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety cites this paper.

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-25T06:06:43.199764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-25T06:05:27.736494Z digest=sha256:ef0df645021f6764bab51602432be6c729656daca00358446bc790cb3c4dec0c

Observation 3420d241-80b2-4375-a7a1-d63611d127b2 · inbound

Beyond Goodhart's Law: A Dynamic Benchmark for Evaluating Compliance in Multi-Agent Systems cites this paper.

Beyond Goodhart's Law: A Dynamic Benchmark for Evaluating Compliance in Multi-Agent Systems Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:47:17.747048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-27T21:53:37.616447Z digest=sha256:8f1f0ad656e373435e173d9f040fa9f89576cb8ad4a63fb4d74d3911531de630

Observation 14430cf5-520c-4fa0-b301-bf36fb5c54a2 · inbound

Decoding Hidden Deception in Reasoning LLMs: Activation Explainers for Deception Auditing cites this paper.

Decoding Hidden Deception in Reasoning LLMs: Activation Explainers for Deception Auditing Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:18:56.662641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-27T01:28:38.810889Z digest=sha256:41b9d3e798d549133b26998531f06470e7866b94141d56e4a353403bcd20ab9d

Observation e3602deb-3912-42eb-ab3c-2899864aa92c · inbound

Engineering Trustworthy Agentic AI for Critical Systems cites this paper.

Engineering Trustworthy Agentic AI for Critical Systems Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-01T15:07:30.716801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:07:30.716801Z digest=sha256:fb4d2b355292213b9adbe725eb25d8732e1e6322b0ba6ea11ed4db58b2580676