Pith. sign in

Paper Citation Record · LEDGER

Large Language Models Hack Rewards, and Society

As of 8 August 2026, this Paper Citation Record lists 75 of 75 outbound references and 0 inbound Pith citation observations for arXiv:2606.04075.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.04075 v2

Coverage vector

measured 75 of 75 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-28T11:30:35.285902Z

measured 75 of 75 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

75 of 75 outbound references displayed

  • verified exact4
  • verified fuzzy0
  • unresolved53
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch18

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9efdbd50-f087-4816-a5da-91dea8936e6c · outbound

This paper cites Concrete Problems in AI Safety.

Large Language Models Hack Rewards, and Society Concrete Problems in AI Safety

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-02T01:46:26.951300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T11:30:35.285902Z digest=sha256:149927351e1d3c8230cd9f4721d96383f13b6c9e4e3280e66b6e436049287f0c

Observation 54b98638-9d67-420f-afb9-6aeb6d0faf41 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Large Language Models Hack Rewards, and Society Advances in Neural Information Processing Systems , volume=

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-28T11:30:35.285902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T11:30:35.285902Z digest=sha256:8399c91669444e9138edd825a5eb0e471f86d2ef3d7cef94191d1480d6b28e0b

Observation 7c5a8d67-acc0-40ef-a2ec-93e5c1e1ac91 · outbound

This paper cites International Conference on Learning Representations , year=.

Large Language Models Hack Rewards, and Society International Conference on Learning Representations , year=

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-28T11:30:35.285902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T11:30:35.285902Z digest=sha256:1aa19c768579aecbf2941e989163592413a62656f1bd4f4e66d2d7ef7837feeb

Observation 584b8cfb-ef6e-4f39-a031-7560d8c3f61d · outbound

This paper cites Specification gaming: the flip side of.

Large Language Models Hack Rewards, and Society Specification gaming: the flip side of

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-28T11:30:35.285902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T11:30:35.285902Z digest=sha256:e1198be6153b8b36bdf90cd11d13e4f0e6f077e280b96c31b899b8b3d31c0748

Observation f8a451c3-6c4c-4b86-8401-3be7115aae5e · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Large Language Models Hack Rewards, and Society Advances in Neural Information Processing Systems , volume=

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-28T11:30:35.285902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T11:30:35.285902Z digest=sha256:afdeda748bbd58b6daff5c4b478ee946b631129182c24720e54f81a08d3fce80

Observation ffb2b9c2-dc12-4877-85fd-c94abdde8eb9 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Large Language Models Hack Rewards, and Society Advances in Neural Information Processing Systems , volume=

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-28T11:30:35.285902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T11:30:35.285902Z digest=sha256:95fe3243b213c6f0b1506fcf4371ffb45aeafc8ef21712f313cea7260116786e

Observation a22189c9-f4c2-4e04-abb8-56a805d470f9 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Large Language Models Hack Rewards, and Society Advances in Neural Information Processing Systems , volume=

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-28T11:30:35.285902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T11:30:35.285902Z digest=sha256:4fb002e0a5f8c7532b246a0234887849446e4b37e7421033cb63693200391f5d

Observation c85e5c66-30a2-43f7-abe0-e5797eb5a412 · outbound

This paper cites Constitutional.

Large Language Models Hack Rewards, and Society Constitutional

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-28T11:30:35.285902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T11:30:35.285902Z digest=sha256:d5f1ad67fd2287207a0cb9ecf96ed5c5642b8aafad9b4ca9678e7f9648eb50da

Observation 5afb8ea0-e3b8-48da-8dd8-a819b67635f0 · outbound

This paper cites an unresolved cited work.

Large Language Models Hack Rewards, and Society Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-28T11:30:35.285902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T11:30:35.285902Z digest=sha256:aff6312b4b8458f16f7e718ee15b6091bc02e137c01f23a3936823b62b2681d0

Observation 8eea5c98-e98e-447e-8471-a3d9a23c1c4c · outbound

This paper cites Conference on Empirical Methods in Natural Language Processing , year=.

Large Language Models Hack Rewards, and Society Conference on Empirical Methods in Natural Language Processing , year=

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-28T11:30:35.285902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T11:30:35.285902Z digest=sha256:1f70cb58915a3df70cf9a84f630cb80b8bf0202fdd03cb6ff447f0edfb5754ea

Observation 6a023afe-25e1-47b6-989f-e0237c8e63be · outbound

This paper cites Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned.

Large Language Models Hack Rewards, and Society Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T01:46:26.954082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T11:30:35.285902Z digest=sha256:3fd792d89cab69e13b7490891f1676d35278e5bc754bff54624bc3ed6f856ba7

Observation 6546ecd0-79b7-457b-8cb9-68acc82e76e3 · outbound

This paper cites an unresolved cited work.

Large Language Models Hack Rewards, and Society Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-28T11:30:35.285902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T11:30:35.285902Z digest=sha256:aa0cd7fa00fb1ac65dd314628a0c69ffa73eab4bda161d8a339065983a60edff

Observation 1e62de06-0edc-4f60-866c-1d25240d7d1f · outbound

This paper cites Efficient memory management for large language model serving with.

Large Language Models Hack Rewards, and Society Efficient memory management for large language model serving with

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-28T11:30:35.285902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T11:30:35.285902Z digest=sha256:5eee8d61dd996dd0818b54a4ac8854031567e001230f56b87e9e7d302de30d9c

Observation 8cc6291f-ad3a-456a-9a5d-0e34b35ed174 · outbound

This paper cites Artificial Intelligence and Law , year=.

Large Language Models Hack Rewards, and Society Artificial Intelligence and Law , year=

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-28T11:30:35.285902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T11:30:35.285902Z digest=sha256:10a192b6e78a1d6b1e1332ac228f44a81281077eaaf0342b9376de7cccbfffcf

Observation 41dcb44f-03ae-4f4f-a95f-0e0edbacdee8 · outbound

This paper cites an unresolved cited work.

Large Language Models Hack Rewards, and Society Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-28T11:30:35.285902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T11:30:35.285902Z digest=sha256:dc9991245f1264dcb79d31f82896429e6637822b496e0d5ff4b62d1d69d48c3d

Observation 5a45b5d8-65b1-432d-9b0d-25da7a571b89 · outbound

This paper cites Artificial Life , volume=.

Large Language Models Hack Rewards, and Society Artificial Life , volume=

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-28T11:30:35.285902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T11:30:35.285902Z digest=sha256:5347d3ce19885b6cf5eaf352be0df12fe7611ee8659265da1af4494c23f41099

Observation 8dbdf55d-109e-4ade-a36c-a1a34b1c19ee · outbound

This paper cites Categorizing variants of.

Large Language Models Hack Rewards, and Society Categorizing variants of

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-28T11:30:35.285902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T11:30:35.285902Z digest=sha256:cfe9d29ddc9ad0a452a2ab1fad81c1d7d119e7004e6a8ff251c18efa3c9c7b91

Observation 16b4e23d-d159-438f-a465-a7cd95822a0c · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Large Language Models Hack Rewards, and Society Fine-Tuning Language Models from Human Preferences

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T01:46:27.020310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T11:30:35.285902Z digest=sha256:7e935452e73b9fa64316ee08d97f20710d59b2762ef9c7b3ba2320a97e693b2e

Observation 6e94129f-99c1-43dc-b5eb-4819b7e238ba · outbound

This paper cites Transactions on Machine Learning Research , year=.

Large Language Models Hack Rewards, and Society Transactions on Machine Learning Research , year=

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-28T11:30:35.285902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T11:30:35.285902Z digest=sha256:1ff1313fe780a064ce11022e9398abcb8aec1a73b0318f903ebec8ebf653d1ca

Observation 8357985a-befe-4cf5-acfb-d58ffb1a555b · outbound

This paper cites International Conference on Machine Learning , year=.

Large Language Models Hack Rewards, and Society International Conference on Machine Learning , year=

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-28T11:30:35.285902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T11:30:35.285902Z digest=sha256:aeb49e0c6f0cbccd257ce92d94685c380ef0411fb622839cde0ca75112dfd484

Observation 066e16c3-29bb-4dd0-8ef1-cdb84634eff6 · outbound

This paper cites Joint European conference on machine learning and knowledge discovery in databases , pages=.

Large Language Models Hack Rewards, and Society Joint European conference on machine learning and knowledge discovery in databases , pages=

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-28T11:30:35.285902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T11:30:35.285902Z digest=sha256:ce580a9c4be34bede12665d6fd879eb02bbbc82a7bb30333027ad5f420281489

Observation cb0636b4-d0cc-44f8-a2db-478e03d02a7e · outbound

This paper cites 2008 , publisher=.

Large Language Models Hack Rewards, and Society 2008 , publisher=

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-28T11:30:35.285902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T11:30:35.285902Z digest=sha256:929597bab9a744c39852d6d6a9d9bbd1efbf3972ce1921afd397d2440451a13d

Observation 22b9cc5d-0635-44cf-be67-8fbffd195530 · outbound

This paper cites Journal of Banking & Finance , volume=.

Large Language Models Hack Rewards, and Society Journal of Banking & Finance , volume=

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-28T11:30:35.285902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T11:30:35.285902Z digest=sha256:9da8a91310ce3f2aa1b7a62717bc495c0ba32348e0a6df2f7426870cc742880d

Observation eb146931-4c8a-43ee-b05e-707bc7ad60bf · outbound

This paper cites 2017 ieee symposium on security and privacy (sp) , pages=.

Large Language Models Hack Rewards, and Society 2017 ieee symposium on security and privacy (sp) , pages=

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-28T11:30:35.285902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T11:30:35.285902Z digest=sha256:365c1ea6f589e5a74e429582f38fd77640b3178bd43cb667ef62e841c841d517

Observation 0d3cb87f-dd97-4014-be20-5b22bb76bcb9 · outbound

This paper cites Proceedings of the national academy of sciences , volume=.

Large Language Models Hack Rewards, and Society Proceedings of the national academy of sciences , volume=

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-28T11:30:35.285902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T11:30:35.285902Z digest=sha256:b168b7a26132a342d1f842bb526aac3aa56c5b6bd9fb066fb64c4658928ec997

Observation c92f32e6-971b-4d1f-aae4-fcea7f15585d · outbound

This paper cites The Quarterly Journal of Economics , volume=.

Large Language Models Hack Rewards, and Society The Quarterly Journal of Economics , volume=

Reference 26

Resolution
unresolved
no resolver link, observed 2026-06-28T11:30:35.285902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T11:30:35.285902Z digest=sha256:7cf20119b928925bc32f4c6cbd7eb11c767565f953ae63a00ce15f83cea8a00d

Observation 75ae0e75-9a6e-4d43-b03d-2a1de4c82b93 · outbound

This paper cites 2011 , publisher=.

Large Language Models Hack Rewards, and Society 2011 , publisher=

Reference 27

Resolution
unresolved
no resolver link, observed 2026-06-28T11:30:35.285902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T11:30:35.285902Z digest=sha256:5a2e2e76af600358bb935416081ab8f622d9757a6a473938050d4b822c44b355

Observation 854eab67-52f8-4ce1-91d1-788c5cf33e5f · outbound

This paper cites IEEE Transactions on Software Engineering , volume=.

Large Language Models Hack Rewards, and Society IEEE Transactions on Software Engineering , volume=

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-28T11:30:35.285902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T11:30:35.285902Z digest=sha256:8186b7522f982fd05b313307568ea197ac4a38e519ab8b767a755755ee306afa

Observation b8a62c8c-423f-431f-a4dd-870060063d99 · outbound

This paper cites Advances in neural information processing systems , volume=.

Large Language Models Hack Rewards, and Society Advances in neural information processing systems , volume=

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-28T11:30:35.285902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T11:30:35.285902Z digest=sha256:5c10a3ffcdc21dc4b3281375742e18f222acf6b8463eb07350e4fd204bfa13c2

Observation 19c1af63-037c-46cf-9f48-b89d7de7acef · outbound

This paper cites International conference on foundations of software technology and theoretical computer science , pages=.

Large Language Models Hack Rewards, and Society International conference on foundations of software technology and theoretical computer science , pages=

Reference 30

Resolution
unresolved
no resolver link, observed 2026-06-28T11:30:35.285902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T11:30:35.285902Z digest=sha256:5a4bef0c7a7952489513c198dcb9fc5ae5332a5e1b886b82231a3f629a1271d0

Observation 5bcc912b-3254-4736-9c87-7b21ff450b35 · outbound

This paper cites ACM Computing Surveys , volume=.

Large Language Models Hack Rewards, and Society ACM Computing Surveys , volume=

Reference 31

Resolution
unresolved
no resolver link, observed 2026-06-28T11:30:35.285902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T11:30:35.285902Z digest=sha256:4c964da8c5104e159702f16e4a46c8d4af8221933539008683933d0aa71c05ca

Observation 8a1ae977-4cf2-433a-9b2a-32ce414c877e · outbound

This paper cites ACM Computing Surveys (CSUR) , volume=.

Large Language Models Hack Rewards, and Society ACM Computing Surveys (CSUR) , volume=

Reference 32

Resolution
unresolved
no resolver link, observed 2026-06-28T11:30:35.285902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T11:30:35.285902Z digest=sha256:2c8e77ced37eb03a2ee3c2fed48e7c585a0d72589afa328635add6d2f89af188

Observation c70e1219-0983-4038-8c46-e39919b6a405 · outbound

This paper cites Advances in neural information processing systems , volume=.

Large Language Models Hack Rewards, and Society Advances in neural information processing systems , volume=

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-28T11:30:35.285902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T11:30:35.285902Z digest=sha256:5f10297136787d6b57304ca3f6bff697ef61e029471189cfbbc1b7c6774ab7b6

Observation 67f54172-0119-4fef-8d09-c76d2c08939a · outbound

This paper cites TradingAgents: Multi-Agents LLM Financial Trading Framework.

Large Language Models Hack Rewards, and Society TradingAgents: Multi-Agents LLM Financial Trading Framework

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T01:46:27.003033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T11:30:35.285902Z digest=sha256:1ee842ea1d29ea9f001797f405958b1420436a5a110517c7f5266ce39627c363

Observation ba3d1739-5bfe-48e8-a4e8-57442e83be39 · outbound

This paper cites Findings of the Association for Computational Linguistics: ACL 2024 , pages=.

Large Language Models Hack Rewards, and Society Findings of the Association for Computational Linguistics: ACL 2024 , pages=

Reference 35

Resolution
unresolved
no resolver link, observed 2026-06-28T11:30:35.285902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T11:30:35.285902Z digest=sha256:110694517051271a944a1b5624d55480860e6b062d613342f174e51930080017

Observation c0b76931-b613-41f8-a053-5128de6b9ac5 · outbound

This paper cites Political Analysis , volume=.

Large Language Models Hack Rewards, and Society Political Analysis , volume=

Reference 36

Resolution
unresolved
no resolver link, observed 2026-06-28T11:30:35.285902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T11:30:35.285902Z digest=sha256:e16931c448f4685113172e38261022ac21f3c80401f0e9627ec6f22aeadb9898

Observation 1ad8bde6-bf97-4eda-b8fc-2c3f85f9c586 · outbound

This paper cites AgentSociety: Large-Scale Simulation of LLM-Driven Generative Agents Advances Understanding of Human Behaviors and Society.

Large Language Models Hack Rewards, and Society AgentSociety: Large-Scale Simulation of LLM-Driven Generative Agents Advances Understanding of Human Behaviors and Society

Reference 37

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T01:46:27.006925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T11:30:35.285902Z digest=sha256:77a92154c6ac684a1a74e7fa62d820edc398e48113d103d16a4c58422e878e40

Observation cb9b8cd8-ff78-489a-b0f7-9da875eb5642 · outbound

This paper cites Findings of the Association for Computational Linguistics: EMNLP 2024 , pages=.

Large Language Models Hack Rewards, and Society Findings of the Association for Computational Linguistics: EMNLP 2024 , pages=

Reference 38

Resolution
unresolved
no resolver link, observed 2026-06-28T11:30:35.285902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T11:30:35.285902Z digest=sha256:b69140c2379f5510d4ec7bf566a9d70e0810da61708ad508e7abf282c90a2e49

Observation 5d17cf83-4421-420e-86b0-585a22b214fc · outbound

This paper cites Science Advances , volume=.

Large Language Models Hack Rewards, and Society Science Advances , volume=

Reference 39

Resolution
unresolved
no resolver link, observed 2026-06-28T11:30:35.285902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T11:30:35.285902Z digest=sha256:325659e77c5bb248a4737fffd3b6de7630232f21ed471f1fbda14928f75d585c

Observation 255e9924-10d8-406f-8d6c-6455072278fa · outbound

This paper cites Generative Language Models and Automated Influence Operations: Emerging Threats and Potential Mitigations.

Large Language Models Hack Rewards, and Society Generative Language Models and Automated Influence Operations: Emerging Threats and Potential Mitigations

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T01:46:27.015636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T11:30:35.285902Z digest=sha256:69850adf7684b1e45a41fac0c639d123ec2551a912dabb11c41e67f2e83a544f

Observation 2e361583-db4b-4110-8795-6bf517c122d8 · outbound

This paper cites Navigating the Risks: A Survey of Security, Privacy, and Ethics Threats in LLM-Based Agents.

Large Language Models Hack Rewards, and Society Navigating the Risks: A Survey of Security, Privacy, and Ethics Threats in LLM-Based Agents

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T01:46:27.010929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T11:30:35.285902Z digest=sha256:55d47d05bbe5029b98dda480ad98332c595a95a0f1c583241180d014ad1fff4a

Observation 9e33b3c3-f666-429f-b3a8-0f14e9ecc732 · outbound

This paper cites an unresolved cited work.

Large Language Models Hack Rewards, and Society Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-06-28T11:30:35.285902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T11:30:35.285902Z digest=sha256:6cd3527be1f8e21475f0eee898bc09ffa524f09c56b3134d173ae677c8f438da

Observation 9ee9eaa9-831d-41c4-8411-ffeff97c5c48 · outbound

This paper cites 2024 , eprint=.

Large Language Models Hack Rewards, and Society 2024 , eprint=

Reference 43

Resolution
unresolved
no resolver link, observed 2026-06-28T11:30:35.285902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T11:30:35.285902Z digest=sha256:f99d099ad001cd5efc9fe03384d17e68c916a6ed679502b73f4e807abaf8a1da

Observation 01738f1b-68fc-4a16-b4e5-74b2db40a2ce · outbound

This paper cites Public Administration Review , volume=.

Large Language Models Hack Rewards, and Society Public Administration Review , volume=

Reference 44

Resolution
unresolved
no resolver link, observed 2026-06-28T11:30:35.285902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T11:30:35.285902Z digest=sha256:00902a6b2fb54061fcc121a82f5b8a1141bb5e687f376404776e16639b98850b

Observation 1b310ca2-d417-4290-a21d-3ddaeb7c9b1e · outbound

This paper cites American sociological review , volume=.

Large Language Models Hack Rewards, and Society American sociological review , volume=

Reference 45

Resolution
unresolved
no resolver link, observed 2026-06-28T11:30:35.285902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T11:30:35.285902Z digest=sha256:55276e325d439535b7b7f23f2805c5875349231e46c17b145cf34584ce1273e5

Observation f87a040c-ddc9-495d-9ef1-53b93b73626d · outbound

This paper cites New York: Russell Sage Foundation , year=.

Large Language Models Hack Rewards, and Society New York: Russell Sage Foundation , year=

Reference 46

Resolution
unresolved
no resolver link, observed 2026-06-28T11:30:35.285902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T11:30:35.285902Z digest=sha256:8a795dc6e3e969c1bdc33e8fb39a9f5a2483d934ca2077ffe2c8f4b7236c9d5b

Observation 02022335-1413-419c-9494-6bb26874e09a · outbound

This paper cites short-termism.

Large Language Models Hack Rewards, and Society short-termism

Reference 47

Resolution
unresolved
no resolver link, observed 2026-06-28T11:30:35.285902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T11:30:35.285902Z digest=sha256:d650edc88e965c31d4c0f75a9c8dd349119f2351ef197abbb4497820077353f1

Observation 4f82fd7d-0e14-4daf-b20b-0b5ac41f2e5d · outbound

This paper cites Monetary theory and practice: The UK experience , pages=.

Large Language Models Hack Rewards, and Society Monetary theory and practice: The UK experience , pages=

Reference 48

Resolution
unresolved
no resolver link, observed 2026-06-28T11:30:35.285902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T11:30:35.285902Z digest=sha256:2c05ca75977e0e00603668dac9ee6ab45dfac5d5a06ac31ae3687ffd62e03561

Observation 5e52f5f1-3b63-4760-8fe1-e88529dc5060 · outbound

This paper cites RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback.

Large Language Models Hack Rewards, and Society RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback

Reference 49

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T01:46:27.024786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T11:30:35.285902Z digest=sha256:c7b05dddeb2b158f4d1d642ef33de3dd3113fce7853ab0c830f470007aa496ce

Observation 73bf1c18-c609-4837-87ed-c7dfba1e7e44 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Large Language Models Hack Rewards, and Society DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 50

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T01:46:27.033328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T11:30:35.285902Z digest=sha256:7a398704b8593b2edcefda097f4785f4b6b788e58d1a1d6bb944c2ec423fd6fb

Observation eb18a35c-3c6f-4fad-8258-7d320d3709c0 · outbound

This paper cites Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges.

Large Language Models Hack Rewards, and Society Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges

Reference 51

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T01:46:26.988139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T11:30:35.285902Z digest=sha256:2c8dba02927392375331ab2c7116e4536e42cfc7bca0640a7b0f13c162f9d9b4

Observation 8624b820-36b4-4fcf-8606-1021e9b205ae · outbound

This paper cites A Long Way to Go: Investigating Length Correlations in RLHF.

Large Language Models Hack Rewards, and Society A Long Way to Go: Investigating Length Correlations in RLHF

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T01:46:26.991887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T11:30:35.285902Z digest=sha256:54426a1dc94740240b508795cad34a14c0bf7be4c4c47df37a99c3dfba8d10f6

Observation fc16fbca-74f8-41dc-90dd-e3b490d9c523 · outbound

This paper cites Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models.

Large Language Models Hack Rewards, and Society Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

Reference 53

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T01:46:26.998730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T11:30:35.285902Z digest=sha256:00a261b9826f5e908c72794f6af7aaf18c6511ab00b8283fd30cbc536f256d8d

Observation a821e81f-971b-4b59-8eea-821bd0920e4d · outbound

This paper cites Natural Emergent Misalignment from Reward Hacking in Production RL.

Large Language Models Hack Rewards, and Society Natural Emergent Misalignment from Reward Hacking in Production RL

Reference 54

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T01:46:26.977064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T11:30:35.285902Z digest=sha256:00fb8ea9f16a6706e30a417ada9cd1ed9c178a6d7e21927e8e192b140d768f4b

Observation 8e531e21-b025-4d4f-b37c-bb78efe65315 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Large Language Models Hack Rewards, and Society Advances in Neural Information Processing Systems , volume=

Reference 55

Resolution
unresolved
no resolver link, observed 2026-06-28T11:30:35.285902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T11:30:35.285902Z digest=sha256:8883278d875f12590e973f0f9e0c452407a7775e5b6fe566e0b653b269c7a39d

Observation 275e842f-49d1-46a3-9022-42411fff8769 · outbound

This paper cites an unresolved cited work.

Large Language Models Hack Rewards, and Society Unresolved cited work

Reference 56

Resolution
unresolved
no resolver link, observed 2026-06-28T11:30:35.285902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T11:30:35.285902Z digest=sha256:c34436313c7078cf1a450308bb216157061457c85dec0a7cedbacae6a2a94dc5

Observation e47899e1-f552-4569-9203-bb1ba7c8a990 · outbound

This paper cites Learning to Discover at Test Time.

Large Language Models Hack Rewards, and Society Learning to Discover at Test Time

Reference 57

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T01:46:26.983990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T11:30:35.285902Z digest=sha256:53f62275c694f8e1f4ee78a010b83b53686333115dd97515e5077753afb9c33e

Observation f84771c3-3d0f-43a8-be0f-38c68c4f85f2 · outbound

This paper cites Proceedings of the AAAI Conference on Artificial Intelligence , volume=.

Large Language Models Hack Rewards, and Society Proceedings of the AAAI Conference on Artificial Intelligence , volume=

Reference 58

Resolution
unresolved
no resolver link, observed 2026-06-28T11:30:35.285902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T11:30:35.285902Z digest=sha256:79968579cf012655a872a470b8c9d8b8030ac0f1eaa0a9dfbed22020b0a7b7be

Observation dcce7423-9742-4dff-ab3d-f34fc17c9a39 · outbound

This paper cites Algorithmic Collusion by Large Language Models.

Large Language Models Hack Rewards, and Society Algorithmic Collusion by Large Language Models

Reference 59

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T01:46:26.995173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T11:30:35.285902Z digest=sha256:0661f4b7ad286869c4706dbd6c338d63265787b1a54002aa833221575035e0f3

Observation 09ea372e-6966-4b61-88c7-46ddd4e38b8c · outbound

This paper cites Can AI expose tax loopholes? Towards a new generation of legal policy assistants.

Large Language Models Hack Rewards, and Society Can AI expose tax loopholes? Towards a new generation of legal policy assistants

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-07-02T01:46:26.980528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T11:30:35.285902Z digest=sha256:22dbb2b72ab9344a7cf89b414457ad359de03e368c51af3632d4596d3de92a69

Observation a92c3ba6-6776-40c3-bb1e-8971b9910b20 · outbound

This paper cites arXiv preprint arXiv:2603.20281 , year=.

Large Language Models Hack Rewards, and Society arXiv preprint arXiv:2603.20281 , year=

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-07-02T01:46:27.029572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T11:30:35.285902Z digest=sha256:297b83c451dc2141631542334d29bdaaf3c5a833ccf5744ca6c1e3eb496c08fb

Observation 3236340e-b708-4bf3-a947-2e2b2deb9d45 · outbound

This paper cites The Twelfth International Conference on Learning Representations , year=.

Large Language Models Hack Rewards, and Society The Twelfth International Conference on Learning Representations , year=

Reference 62

Resolution
unresolved
no resolver link, observed 2026-06-28T11:30:35.285902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T11:30:35.285902Z digest=sha256:aeb0df3e71294a432f18b37412cefaa2a48fa2436a785bf55d7822ddf764fc02

Observation b36843a0-d119-4598-84e5-d47493a012c9 · outbound

This paper cites arXiv preprint arXiv:2507.08068 , year=.

Large Language Models Hack Rewards, and Society arXiv preprint arXiv:2507.08068 , year=

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-07-02T01:46:26.966168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T11:30:35.285902Z digest=sha256:bd345c4d6c3dc0f974d51f399c116b45a99b74045d70611d26b41efc5e4035de

Observation 8e7d8c9d-acb6-4c4f-8d97-58ddd8ac2d83 · outbound

This paper cites Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages=.

Large Language Models Hack Rewards, and Society Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages=

Reference 64

Resolution
unresolved
no resolver link, observed 2026-06-28T11:30:35.285902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T11:30:35.285902Z digest=sha256:77c137ef8c0bb04942ea965df2bab72338b5a748ff535b4555ce06704c89f709

Observation 5b67b3ce-bc6b-4faf-977d-e3fb44c47ec3 · outbound

This paper cites Qwen3 Technical Report.

Large Language Models Hack Rewards, and Society Qwen3 Technical Report

Reference 65

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T01:46:26.970140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T11:30:35.285902Z digest=sha256:1363cfbc54d7baf8e6ca6b15e6ea1377dff4bef79c6453e91a1d2cecabf9851b

Observation fca1f5f6-26be-4f95-bc46-a5bbed9aab11 · outbound

This paper cites 2025 , url =.

Large Language Models Hack Rewards, and Society 2025 , url =

Reference 66

Resolution
unresolved
no resolver link, observed 2026-06-28T11:30:35.285902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T11:30:35.285902Z digest=sha256:d29e693441a2d0d969fb6c835c0cae2c9f665bb6077c2ac1ab6d40f7aa442da3

Observation 0b01f78f-a34b-41a9-9a90-6ab7e0a4249e · outbound

This paper cites , title =.

Large Language Models Hack Rewards, and Society , title =

Reference 67

Resolution
unresolved
no resolver link, observed 2026-06-28T11:30:35.285902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T11:30:35.285902Z digest=sha256:2c4bdaa2d4017ae0e7e6cac281d9dc2043e8dbad696933eb9038393eb7cc15b1

Observation 0e2c5db9-676d-4e0e-b5b7-1995683c2eb2 · outbound

This paper cites , title =.

Large Language Models Hack Rewards, and Society , title =

Reference 68

Resolution
unresolved
no resolver link, observed 2026-06-28T11:30:35.285902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T11:30:35.285902Z digest=sha256:e85a5ec3b86839bcbb107d988c18356fae7f4d625096b81c77495e8c5c81c1d3

Observation 3d768157-7af4-4210-83a6-6e629150c6ed · outbound

This paper cites HealthBench: Evaluating Large Language Models Towards Improved Human Health.

Large Language Models Hack Rewards, and Society HealthBench: Evaluating Large Language Models Towards Improved Human Health

Reference 69

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T01:46:26.973660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T11:30:35.285902Z digest=sha256:75d7b818f9130d86f3e0d78341c4932cb0a74167c15cd740002ac80130cec11f

Observation ef2fee43-d1f8-4bcb-9dcd-b053df353693 · outbound

This paper cites Richard and Koch, Gary G.

Large Language Models Hack Rewards, and Society Richard and Koch, Gary G

Reference 70

Resolution
unresolved
no resolver link, observed 2026-06-28T11:30:35.285902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T11:30:35.285902Z digest=sha256:2c94a85b88c649e6d18cebe29dac6323eb58c559190236b7a11fcb5a0af2281b

Observation 4a0a8b9f-ffdd-4448-9cb7-15f7a2fce038 · outbound

This paper cites Emergent Misalignment : Narrow finetuning can produce broadly misaligned LLMs , May 2025.

Large Language Models Hack Rewards, and Society Emergent Misalignment : Narrow finetuning can produce broadly misaligned LLMs , May 2025

Reference 71

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T01:46:26.958324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T11:30:35.285902Z digest=sha256:796f06a3c9edcbd268b2a5ae53b7edebec455b82931972d60c52a47876b57627

Observation 9a77f0fc-8f8c-492d-bf1f-8786a37c02b3 · outbound

This paper cites OpenAI GPT-5 System Card.

Large Language Models Hack Rewards, and Society OpenAI GPT-5 System Card

Reference 72

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T01:46:26.962571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T11:30:35.285902Z digest=sha256:8476a9e3ab3bf3b88e009de2970cbeef1020cd578fde00a11c2a2b1275ef5f5b

Observation 0eee1109-9cca-4592-a37b-6bb32e06bce9 · outbound

This paper cites an unresolved cited work.

Large Language Models Hack Rewards, and Society Unresolved cited work

Reference 73

Resolution
unresolved
no resolver link, observed 2026-06-28T11:30:35.285902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T11:30:35.285902Z digest=sha256:c12f5ad5d8535169b06a461fa5cc78f9eb39d11bb5c23b54772e2779bd01d66f

Observation d1d78bf0-6b90-4310-9ad1-40dff0108329 · outbound

This paper cites The Fourteenth International Conference on Learning Representations , year=.

Large Language Models Hack Rewards, and Society The Fourteenth International Conference on Learning Representations , year=

Reference 74

Resolution
unresolved
no resolver link, observed 2026-06-28T11:30:35.285902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T11:30:35.285902Z digest=sha256:a368944e8c5755c1bb95875b24dd2957df13f3713e2fd95efeee2906c41ae0f0

Observation 3e9f0df3-28c3-40e4-8bf1-0ef793de5762 · outbound

This paper cites The International Conference on Learning Representations (ICLR) Blog Post Track , year=.

Large Language Models Hack Rewards, and Society The International Conference on Learning Representations (ICLR) Blog Post Track , year=

Reference 75

Resolution
unresolved
no resolver link, observed 2026-06-28T11:30:35.285902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T11:30:35.285902Z digest=sha256:45353d9b0233ca4da6ebef79c68f6f9456312964a05fe8cddcff090d9510f938

Pith citing papers

No inbound Pith citation observations are available.