Pith. sign in

Paper Citation Record · LEDGER

Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2304.03279.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2304.03279 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T15:07:30.716801Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

28
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d6e829ee-a8f4-4e19-ac1d-65563706e160 · inbound

A Roadmap to Pluralistic Alignment cites this paper.

A Roadmap to Pluralistic Alignment Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark

Reference 190

Resolution
verified exact
arxiv_id, observed 2026-05-16T14:37:53.622855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-16T14:37:53.279275Z digest=sha256:9a92e518c13ceb8a7be7433488c91d7ae67a4119e678e4a0133516778c4cb648

Observation 7f55633a-4be4-43d8-adfc-2160525c1239 · inbound

Moral Susceptibility and Robustness under Persona Role-Play in Large Language Models cites this paper.

Moral Susceptibility and Robustness under Persona Role-Play in Large Language Models Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-17T23:20:27.811163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T23:16:56.957905Z digest=sha256:b39731110840036be0e3f9598db59ed67ea5f911bd452feaeae38fe942e94bfe

Observation 3249d592-9433-4f85-99ea-2d2ebb4485fe · inbound

Beyond Benchmark Islands: Toward Representative Trustworthiness Evaluation for Agentic AI cites this paper.

Beyond Benchmark Islands: Toward Representative Trustworthiness Evaluation for Agentic AI Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-22T10:21:23.244031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T10:19:56.003219Z digest=sha256:d284dce989dcee177add896b8d945fba7eb5e4160b503bf0689d700e81dd2781

Observation 3b8da89d-87e0-4d07-b336-25fb556d38f0 · inbound

Restoration, Exploration and Transformation: How Youth Engage Character.AI Chatbots for Feels, Fun and Finding themselves cites this paper.

Restoration, Exploration and Transformation: How Youth Engage Character.AI Chatbots for Feels, Fun and Finding themselves Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark

Reference 64

Resolution
malformed identifier
arxiv_id, observed 2026-05-15T12:45:37.283306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:44:53.677818Z digest=sha256:a1cf7e655f084af9ed15a0da0c9a729e203c36c68f7fbb8ec48f5809f13fd68d

Observation ff9e1600-92e5-4847-8f14-6e8ccc1694b7 · inbound

Positive Alignment: Artificial Intelligence for Human Flourishing cites this paper.

Positive Alignment: Artificial Intelligence for Human Flourishing Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark

Reference 148

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:59:47.953274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T05:56:56.902705Z digest=sha256:c8c4d1cf29b6c4b11e4e0128d5b16b75ed0eee7d636e0d28c46641b0d4e46050

Observation ce6bc9a0-d7d8-4f36-a07d-38740ddff390 · inbound

Correcting Influence: Unboxing LLM Outputs with Orthogonal Latent Spaces cites this paper.

Correcting Influence: Unboxing LLM Outputs with Orthogonal Latent Spaces Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark

Reference 228

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T20:17:55.545244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-14T20:17:01.224864Z digest=sha256:37d13b2ca3d1bb9028b7a3c2d4db82ee65634682dd40341e915cec94ea0aa5ee

Observation e8a3e7e2-a852-48ae-8c54-3e7a98dff719 · inbound

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents cites this paper.

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-21T01:43:56.709741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-21T01:42:55.693115Z digest=sha256:fb1544247aae82198b099d9e5c64b1817c9a888173d93ce613fe221921bdcd9d

Observation d0da8262-22b1-4148-8e7c-96665a13f2c1 · inbound

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety cites this paper.

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:51:08.014677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T05:50:28.114140Z digest=sha256:53704166a290f66b7ca26d318efc3eb6a3b21f522c3bc56c6e4005638b0b24df

Observation 8d56c348-3b2b-49cd-a2bb-bd8e7a2fa2cb · inbound

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety cites this paper.

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-25T06:06:43.199764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-25T06:05:27.736494Z digest=sha256:8ec2b33bc739ce69475ee278fa347ad74bb9f95b611779091c9a0c851740697d

Observation 3420d241-80b2-4375-a7a1-d63611d127b2 · inbound

Beyond Goodhart's Law: A Dynamic Benchmark for Evaluating Compliance in Multi-Agent Systems cites this paper.

Beyond Goodhart's Law: A Dynamic Benchmark for Evaluating Compliance in Multi-Agent Systems Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:47:17.747048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T21:53:37.616447Z digest=sha256:e938bebef1a45dbb7fdbf49e8e6040f7a1e584ede1a28e43e5c5fa214d74114e

Observation 14430cf5-520c-4fa0-b301-bf36fb5c54a2 · inbound

Decoding Hidden Deception in Reasoning LLMs: Activation Explainers for Deception Auditing cites this paper.

Decoding Hidden Deception in Reasoning LLMs: Activation Explainers for Deception Auditing Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:18:56.662641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T01:28:38.810889Z digest=sha256:1381379e4f517af014434ff2e84c7cf706f6121ff0764aee2099e5988ad5612a

Observation e3602deb-3912-42eb-ab3c-2899864aa92c · inbound

Engineering Trustworthy Agentic AI for Critical Systems cites this paper.

Engineering Trustworthy Agentic AI for Critical Systems Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-01T15:07:30.716801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:07:30.716801Z digest=sha256:0d9536c73fd0ba6f6f9fa1df9f7b7a8e9f4b0cb1816a365dd240d95ff07335d4