Pith. sign in

Paper Citation Record · LEDGER

Agent Identity Evals: Measuring Agentic Identity

As of 10 August 2026, this Paper Citation Record lists 99 of 99 outbound references and 1 inbound Pith citation observation for arXiv:2507.17257.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.17257 v1

Coverage vector

measured 99 of 99 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T14:57:07.014317Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-29T00:26:54.019256Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

99 of 99 outbound references displayed

  • verified exact5
  • verified fuzzy27
  • unresolved65
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 90bd6375-8c91-4597-a574-77070359b937 · outbound

This paper cites AI Agents That Matter.

Agent Identity Evals: Measuring Agentic Identity AI Agents That Matter

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T14:57:06.463141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:57:06.463141Z digest=sha256:3c3113eedaac2800d03fd0fac3c470f602d70389092c00b24afcab9214494f7e

Observation 945f1a7b-6fd7-4e19-ac75-042c8c279e6e · outbound

This paper cites Introducing Devin, the first AI software engineer, March 2024.

Agent Identity Evals: Measuring Agentic Identity Introducing Devin, the first AI software engineer, March 2024

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T14:57:06.475113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:57:06.475113Z digest=sha256:1393c89f9228c6b719219f1b230b704e6924a937586fb87d2dba0ec073ba3f16

Observation 69472970-6998-4831-9543-b426b146d6de · outbound

This paper cites GAIA: a benchmark for General AI Assistants, November 2023.

Agent Identity Evals: Measuring Agentic Identity GAIA: a benchmark for General AI Assistants, November 2023

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T14:57:06.480272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:57:06.480272Z digest=sha256:40a962a80296e3197b461135cc89b1d7be38f6a290316b4941ca3e8aeb634c52

Observation 8985ac9c-af6b-44df-94c6-5f026bf80eae · outbound

This paper cites ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities.

Agent Identity Evals: Measuring Agentic Identity ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T14:57:06.486170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:57:06.486170Z digest=sha256:98c6b13838621bea9b9e9e8799edeac48298616fb132c347569c3ae53226d673

Observation a2e33d96-9c84-4793-a4d4-08cedb088a61 · outbound

This paper cites Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models.

Agent Identity Evals: Measuring Agentic Identity Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T14:57:06.491612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:57:06.491612Z digest=sha256:5cd3538a2fe1e79a3fffd251f516284c2097026a61c157c68954f3963c549589

Observation 32ab33e2-503c-4e5c-8737-589af7213ef5 · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

Agent Identity Evals: Measuring Agentic Identity SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T14:57:06.497559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:57:06.497559Z digest=sha256:2568a92c55bce330532e1bd5f264c0ecf134f7cfe854e796b014cc8b13806b42

Observation c5116462-c38d-4f22-9bee-eaecab64161e · outbound

This paper cites Jimenez, John Yang, Kevin Liu, and Aleksander Madry.

Agent Identity Evals: Measuring Agentic Identity Jimenez, John Yang, Kevin Liu, and Aleksander Madry

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T14:57:06.502779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:57:06.502779Z digest=sha256:3201c0772550f60884b555e8176f754673509041f783fdbaa206bcbcd6af2786

Observation e10fb22f-ea94-4a24-88c8-0d62453d73a5 · outbound

This paper cites A Real-World WebAgent with Planning, Long Context Understanding, and Program Synthesis.

Agent Identity Evals: Measuring Agentic Identity A Real-World WebAgent with Planning, Long Context Understanding, and Program Synthesis

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T14:57:06.508586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:57:06.508586Z digest=sha256:ec2b7a2276317563558e60124486d4c6d59b937cca48216cc7203e2410a48fef

Observation 7434111f-2dea-4750-9cbc-7bc751762b1c · outbound

This paper cites MultiOn AI, 2024.

Agent Identity Evals: Measuring Agentic Identity MultiOn AI, 2024

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T14:57:06.513987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:57:06.513987Z digest=sha256:da3f7bb45c0cc2accdd1e845ebf2243cca408f461b440974bfd7442d8c82f047

Observation 8b2999e5-6fd0-465a-8655-893edee84222 · outbound

This paper cites Agents that reduce work and information overload.Communications of the ACM, 37(7):30–40, July 1994.

Agent Identity Evals: Measuring Agentic Identity Agents that reduce work and information overload.Communications of the ACM, 37(7):30–40, July 1994

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T14:57:06.519681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:57:06.519681Z digest=sha256:2e99f92f304ad1127e3e2e1cd93942395d00e92ca4b120d587499a4ec6bd6daf

Observation a7759664-acff-462f-b635-aac716f94f0a · outbound

This paper cites Artificial life meets entertainment: lifelike autonomous agents.Communications of the ACM, 38(11):108–114, November 1995.

Agent Identity Evals: Measuring Agentic Identity Artificial life meets entertainment: lifelike autonomous agents.Communications of the ACM, 38(11):108–114, November 1995

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T14:57:06.525291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:57:06.525291Z digest=sha256:7ad1dd118638d3820c1ddd5bac504ad6b9be4a03425d6870f84790feb100ed23

Observation db8ed391-0c32-4d5a-82b0-45fed6ecde76 · outbound

This paper cites Autonomous interface agents.

Agent Identity Evals: Measuring Agentic Identity Autonomous interface agents

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T14:57:06.530778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:57:06.530778Z digest=sha256:a3744dfb278ed313b42a5ba3522e216ae895e15c42e9727067399ea5ee3a9f29

Observation a4f0d0d0-972e-4622-9d3f-d5e27980c9d1 · outbound

This paper cites A roadmap of agent research and development.Autonomous agents and multi-agent systems, 1:7–38, 1998.

Agent Identity Evals: Measuring Agentic Identity A roadmap of agent research and development.Autonomous agents and multi-agent systems, 1:7–38, 1998

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T14:57:06.535664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:57:06.535664Z digest=sha256:3ab96d163870a92a174390296e54ab0ad984368107085378bca0189c73872bee

Observation d0e3df6e-dbf9-44d7-98e4-270a481e4533 · outbound

This paper cites an unresolved cited work.

Agent Identity Evals: Measuring Agentic Identity Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T14:57:06.541028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:57:06.541028Z digest=sha256:431cb07ecbe465b32a18ce6c48396acd39da984cb7bdc5bd1d29c26bbc6d5b20

Observation 83473c7d-0f6a-4595-a04b-0cbccf4774c7 · outbound

This paper cites Sutton and Andrew G.

Agent Identity Evals: Measuring Agentic Identity Sutton and Andrew G

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T14:57:06.546230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:57:06.546230Z digest=sha256:6ea4633b7c07af07fcaccad9839b7163e83fe6672be4e39de385d3b3f2437424

Observation 672f404c-3772-4a1a-abcd-3d86ff36d32b · outbound

This paper cites Russell and Peter Norvig.Artificial Intelligence: A Modern Approach.

Agent Identity Evals: Measuring Agentic Identity Russell and Peter Norvig.Artificial Intelligence: A Modern Approach

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T14:57:06.551410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:57:06.551410Z digest=sha256:462e192fe6be04c326e48e9da11d1fd73d36be5967773d34a74df442cf14ac3d

Observation de2bfb9b-e387-4eff-8d0c-2dc7856635f9 · outbound

This paper cites Harms from Increasingly Agentic Algorithmic Systems.

Agent Identity Evals: Measuring Agentic Identity Harms from Increasingly Agentic Algorithmic Systems

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T14:57:06.556692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:57:06.556692Z digest=sha256:261fb261089ef7c752f129ffccd390a6a86997fcb059f4dd8d544591dac3f094

Observation 746aebd6-f3a8-416f-a363-6e72ba19984a · outbound

This paper cites AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation.

Agent Identity Evals: Measuring Agentic Identity AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T14:57:06.562823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:57:06.562823Z digest=sha256:86c6334eb8c2a8755ee5884d8cc2b22a8929a234c10b4a2f609d1c350b954e7b

Observation 2c1ea291-854c-4179-95d8-01de49566168 · outbound

This paper cites OpenAI Charter, 2018.

Agent Identity Evals: Measuring Agentic Identity OpenAI Charter, 2018

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T14:57:06.569557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:57:06.569557Z digest=sha256:98b5ab2a57fd08f23c2b979d392e928dbe1241ea40107b97d18ca6787898f5f0

Observation 1326a40c-96ed-4430-8f13-415b0a409744 · outbound

This paper cites The Ethics of Advanced AI Assistants.

Agent Identity Evals: Measuring Agentic Identity The Ethics of Advanced AI Assistants

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T14:57:06.575199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:57:06.575199Z digest=sha256:2e120d0dafbd39c56735de6ce475f4a8b74f61fdb58a872f1a9246d5236f18f7

Observation d75cdf1d-b80b-47e8-b82b-df8b4258541e · outbound

This paper cites Governing AI Agents, April 2024.

Agent Identity Evals: Measuring Agentic Identity Governing AI Agents, April 2024

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T14:57:06.581377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:57:06.581377Z digest=sha256:34318b353fa4b216ee3a5dcc7b2ca23cd842e15acae645b9e99e4b59bd383c06

Observation 0a6431b4-4476-4047-b056-7635743f9d17 · outbound

This paper cites Potter and Kevin J.

Agent Identity Evals: Measuring Agentic Identity Potter and Kevin J

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T14:57:06.588723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:57:06.588723Z digest=sha256:8ecae328d8208a1a76e894cf6cc82e86087ccea38523dc70f7f59c9b70e453c1

Observation 17a9dce9-82a0-4fdc-ae1a-ae463859cd48 · outbound

This paper cites Position: Stop acting like language model agents are normal agents, 2025.

Agent Identity Evals: Measuring Agentic Identity Position: Stop acting like language model agents are normal agents, 2025

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T14:57:06.593854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:57:06.593854Z digest=sha256:97cfaf06031a78d3fc538b01b20a7324a3c177bbbaedc0e726595ff3935ad9d4

Observation d3dc9ae4-b0f9-40aa-8832-27e2b4d228dc · outbound

This paper cites Emergent causality and the foundation of consciousness.

Agent Identity Evals: Measuring Agentic Identity Emergent causality and the foundation of consciousness

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T14:57:06.599169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:57:06.599169Z digest=sha256:c358fc7fa0b6fffe7086b9c56e14d26149275b695cac9a964506f24af141c65e

Observation ec33ae29-47cb-4f7f-9595-e578ed8ad0d4 · outbound

This paper cites Compression, the fermi paradox and artificial super-intelligence.

Agent Identity Evals: Measuring Agentic Identity Compression, the fermi paradox and artificial super-intelligence

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T14:57:06.604343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:57:06.604343Z digest=sha256:4e06b2bd1d6812aaf6ede230060fc8dadd986d4ada83252b982c4b74ccc6c9a0

Observation 7889ef22-788c-4144-8b6f-8664c3391ed8 · outbound

This paper cites World Scientific, 2013.

Agent Identity Evals: Measuring Agentic Identity World Scientific, 2013

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T14:57:06.609869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:57:06.609869Z digest=sha256:9e22bc145831f81c0b2b9afd8b89dcf6edc34390b854540c8765b39e54f18750

Observation cf7a54d8-162a-401d-bd87-f4d51fce04f4 · outbound

This paper cites Thorisson.A New Constructivist AI: From Manual Methods to Self-Constructive Systems, pages 145–171.

Agent Identity Evals: Measuring Agentic Identity Thorisson.A New Constructivist AI: From Manual Methods to Self-Constructive Systems, pages 145–171

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:57:08.627558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:57:06.614775Z digest=sha256:5d351432e1e1468e656ac06f00ccbf105a42ea2d0fd974dd85e75d0be7feac28

Observation bc84d72b-86d4-4056-87a0-310a48c5281c · outbound

This paper cites Artificial general intelligence: Concept, state of the art.Journal of Artificial General Intelligence, 5(1):1–48, 2014.

Agent Identity Evals: Measuring Agentic Identity Artificial general intelligence: Concept, state of the art.Journal of Artificial General Intelligence, 5(1):1–48, 2014

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:57:08.607831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:57:06.620260Z digest=sha256:d2e82f0ac42bfa826bb79dbb53228ffb7d6ef6b0c8b62cb30e0077b7d9e26052

Observation e2bcb39f-eb72-46e0-be5f-d87dec9d79ce · outbound

This paper cites AgentBench: Evaluating LLMs as Agents.

Agent Identity Evals: Measuring Agentic Identity AgentBench: Evaluating LLMs as Agents

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T14:57:06.625605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:57:06.625605Z digest=sha256:a6470a49f3f5aa447076e9bb0d6fcbeba64c66c177cfc6909c9612422354cb0d

Observation 52b961ae-a22f-47fa-96da-851c8e5427a0 · outbound

This paper cites GAIA: a benchmark for General AI Assistants.

Agent Identity Evals: Measuring Agentic Identity GAIA: a benchmark for General AI Assistants

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T14:57:06.633930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:57:06.633930Z digest=sha256:b1faf6d5d3dfa455a55874133aa349a3900e1c21ac0527b5475f9a0a78ecf14b

Observation 43152c97-83f1-4231-afce-e6da05582938 · outbound

This paper cites MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation.

Agent Identity Evals: Measuring Agentic Identity MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T14:57:06.640490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:57:06.640490Z digest=sha256:cf18c1379622ebb3eeb48b6a354bf38f1e7743a272ad17efedbc1876ba00261f

Observation 33891f3f-acec-4c08-bd5c-dcd9349a5315 · outbound

This paper cites AgentSims: An Open-Source Sandbox for Large Language Model Evaluation.

Agent Identity Evals: Measuring Agentic Identity AgentSims: An Open-Source Sandbox for Large Language Model Evaluation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T14:57:06.646251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:57:06.646251Z digest=sha256:bee33796065c4b247af1eccb2b95dc6ab4fc050cd445140a45bb03ebb8b661c1

Observation c549e7f8-995f-4bfa-ac27-180421bcdbd8 · outbound

This paper cites CharacterEval: A Chinese Benchmark for Role-Playing Conversational Agent Evaluation.

Agent Identity Evals: Measuring Agentic Identity CharacterEval: A Chinese Benchmark for Role-Playing Conversational Agent Evaluation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T14:57:06.652456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:57:06.652456Z digest=sha256:eb15ca650437f72288d9b5a15e56cc3352a5e686b2acde8faab5ba5cb53f37ec

Observation 3ef074a3-eee1-4112-841a-28577879b915 · outbound

This paper cites CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities.

Agent Identity Evals: Measuring Agentic Identity CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T14:57:06.658280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:57:06.658280Z digest=sha256:3cb9dbce450297fc89055116e06395f534b8e0cdde156143bb092744402be3fd

Observation d02d8bf0-e065-4fe5-8a9a-ad4782b7ba78 · outbound

This paper cites MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents.

Agent Identity Evals: Measuring Agentic Identity MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T14:57:06.663765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:57:06.663765Z digest=sha256:5ecb5500583516963270d37f53021bbcbb35516a30c1ce000690676f5b1c9ca5

Observation 804ceb17-0626-4071-be1b-d4827f4bb80f · outbound

This paper cites ELT-Bench: An End-to-End Benchmark for Evaluating AI Agents on ELT Pipelines.

Agent Identity Evals: Measuring Agentic Identity ELT-Bench: An End-to-End Benchmark for Evaluating AI Agents on ELT Pipelines

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T14:57:06.669214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:57:06.669214Z digest=sha256:d72c5bf6d54c5d37b432551b4969c7be0843dcfc504b684d2d40aafdb7c51985

Observation a9ee71b9-a33c-4ed2-bda6-c1bd3e44cf3c · outbound

This paper cites Benchmarking Agentic Workflow Generation.

Agent Identity Evals: Measuring Agentic Identity Benchmarking Agentic Workflow Generation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T14:57:06.674390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:57:06.674390Z digest=sha256:ad11d0e3b7b1eb33930a389722751b8c63adf157a3baf9768c7104713a58cd74

Observation f9a3f08d-56b6-4d43-b786-a86940cb4b05 · outbound

This paper cites PARTNR: A Benchmark for Planning and Reasoning in Embodied Multi-agent Tasks.

Agent Identity Evals: Measuring Agentic Identity PARTNR: A Benchmark for Planning and Reasoning in Embodied Multi-agent Tasks

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T14:57:06.679450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:57:06.679450Z digest=sha256:e6ef18447bb9b405fe49027f793639bbb8ce38b5e48bac2bfdd046699e5ccf97

Observation 2b1e2002-fd6c-43b5-aad9-15252b844b2d · outbound

This paper cites John wiley & sons, 2009.

Agent Identity Evals: Measuring Agentic Identity John wiley & sons, 2009

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T14:57:06.684556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:57:06.684556Z digest=sha256:95d7729a78abaa7563f83d3e3f0df0b1b4bcfff5c2db8184308ae622a37aeb59

Observation cd8da897-fddb-42aa-acf7-3c7fd8515160 · outbound

This paper cites Is it an agent, or just a program? a taxonomy for autonomous agents.

Agent Identity Evals: Measuring Agentic Identity Is it an agent, or just a program? a taxonomy for autonomous agents

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:57:08.573766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:57:06.689804Z digest=sha256:f85bcef1bb445022f5a17953ab4e0e46f4e31c71ec79b7003ee0737fd80c89f6

Observation 04543301-01cb-4902-a72a-ad8da4b74a81 · outbound

This paper cites Jennings.

Agent Identity Evals: Measuring Agentic Identity Jennings

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T14:57:06.694956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:57:06.694956Z digest=sha256:1a81a548995676ea6a54c460e523c72b593182bb0ef8e739fc835956f83c0c1c

Observation a607cefc-00b4-419f-940b-2a2a753c0a26 · outbound

This paper cites Adversarial nli: A new benchmark for natural language understanding.

Agent Identity Evals: Measuring Agentic Identity Adversarial nli: A new benchmark for natural language understanding

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:57:08.539061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:57:06.700379Z digest=sha256:5606a464052fa2222eaf24753b9ea0341393ef05e1b8cee74b5fffa4c4b04fb5

Observation 95826660-379d-4eb6-9756-e95ab877cf27 · outbound

This paper cites Metagpt: Meta programming for multi-agent collaborative framework.

Agent Identity Evals: Measuring Agentic Identity Metagpt: Meta programming for multi-agent collaborative framework

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:57:08.519669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:57:06.705353Z digest=sha256:e22c3e1feb12ec1c8f06cc93c730b46f8138a2dc729234d8cb58f0c9695bd60e

Observation 0d7e05da-9a35-4c94-9b5d-ee9effd20baf · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Agent Identity Evals: Measuring Agentic Identity Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T14:57:06.710356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:57:06.710356Z digest=sha256:60e1129b0243040b44e2881eb4d2d25569d2e877cbc1bfd18d09f8da37e18cf5

Observation f23d8a28-6c94-450f-98d5-75788a28f913 · outbound

This paper cites Langchain.https://github.com/langchain-ai/langchain, 2022.

Agent Identity Evals: Measuring Agentic Identity Langchain.https://github.com/langchain-ai/langchain, 2022

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T14:57:06.716016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:57:06.716016Z digest=sha256:3b33cdf66910ce94b12f996e55ad34e457c443b622b09f018d9735b0b696283c

Observation a4a7f49f-a8d2-47bd-bf03-dd47b7d09ada · outbound

This paper cites Thórisson.

Agent Identity Evals: Measuring Agentic Identity Thórisson

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:57:08.480959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:57:06.721374Z digest=sha256:5de412c5f2fe3f8dfc25a3657078ce99c2e938afe0e959708e532b8882c55a2e

Observation 74d8fa55-dc95-48bc-bf98-dda8ff0f92f4 · outbound

This paper cites Autocatalytic endogenous reflective architecture.

Agent Identity Evals: Measuring Agentic Identity Autocatalytic endogenous reflective architecture

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:57:08.461528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:57:06.726076Z digest=sha256:5f33e6b1e4cfb77e7dc6b1fa902078c349970407c928b26fd62e9bdc9f59c3d1

Observation a2c2b775-dd2f-49a8-afee-f14d11658249 · outbound

This paper cites Wang.Rigid Flexibility: The Logic of Intelligence.

Agent Identity Evals: Measuring Agentic Identity Wang.Rigid Flexibility: The Logic of Intelligence

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:57:08.441732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:57:06.730962Z digest=sha256:e1368c06f7dbfa80aa66fa8b029e017a0154bc31596572d5015b32ab8556e544

Observation 2bcc24a6-5df1-4d5c-a8be-4f9cece8686c · outbound

This paper cites ‘opennars for applications’: Architecture and control.

Agent Identity Evals: Measuring Agentic Identity ‘opennars for applications’: Architecture and control

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:57:08.422928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:57:06.736311Z digest=sha256:221519a5f360f3609ad7dcf39906d2dcf8be987d097c0434dbc7c7040eee55b0

Observation cdf351cb-faeb-43dc-81e8-c24744186362 · outbound

This paper cites The general theory of general intelligence: A pragmatic patternist perspective.

Agent Identity Evals: Measuring Agentic Identity The general theory of general intelligence: A pragmatic patternist perspective

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:57:08.404633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:57:06.742133Z digest=sha256:794027bb4e7bd02b8ccb1cd3f25c178cc308715f3aad60a46e2377899614f821

Observation da8d3d15-f4dc-45e4-b0a8-05e76f77a4d3 · outbound

This paper cites Opencog hyperon: A framework for agi at the human level and beyond.

Agent Identity Evals: Measuring Agentic Identity Opencog hyperon: A framework for agi at the human level and beyond

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:57:08.384987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:57:06.747813Z digest=sha256:7a45d8930d9bf55859f6e17dda13600a4a1beb5788298088a318ed5defad9a39

Observation 915ff8e5-45f6-4f21-bc6c-1d10ed3a60b4 · outbound

This paper cites Cognitive Architectures for Language Agents.

Agent Identity Evals: Measuring Agentic Identity Cognitive Architectures for Language Agents

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T14:57:06.753218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:57:06.753218Z digest=sha256:891fc840190522f34c5ac66176a2fef0ac9144b7cdb1ea6e1d424cf4a0f540d1

Observation 92bad81b-a4ac-45d9-ac1b-510ddd997125 · outbound

This paper cites Language agents in the digital world: Opportunities and risks.princeton-nlp.github.io, Jul 2023.

Agent Identity Evals: Measuring Agentic Identity Language agents in the digital world: Opportunities and risks.princeton-nlp.github.io, Jul 2023

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:57:08.363823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:57:06.758612Z digest=sha256:c0a852b0f99aea4a5006759206d9825e8d234c6c120fc1a1be40ee43d54eb7ba

Observation 5f24bec1-a7ad-4d23-b8df-a93f33229745 · outbound

This paper cites Practices for governing agentic ai systems.

Agent Identity Evals: Measuring Agentic Identity Practices for governing agentic ai systems

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:57:08.345775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:57:06.763988Z digest=sha256:ea7d43fc8b6c8ffe96cbc301542b7e5b1d5cc10b187eb56fbcaf9260a6cab5c1

Observation fbd1cec1-c232-4682-a0d0-5ce705925f94 · outbound

This paper cites Langlois, Pedro A.

Agent Identity Evals: Measuring Agentic Identity Langlois, Pedro A

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:57:08.324834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:57:06.769739Z digest=sha256:65f80dfbaf67a8d0c61c1f8df490f9ff98ef1463f2f328c7be123f03368f6bc9

Observation b1c14acb-291a-4f21-870f-6d2b5435512e · outbound

This paper cites Visibility into AI Agents.

Agent Identity Evals: Measuring Agentic Identity Visibility into AI Agents

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T14:57:06.774550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:57:06.774550Z digest=sha256:3793e95c8e094aa703b5427573d01836e5f8c31d33f1a49e7b9598f91a2c6167

Observation 6946e8d5-3a53-4839-abbb-c76404867d5e · outbound

This paper cites Attention is all you need.

Agent Identity Evals: Measuring Agentic Identity Attention is all you need

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:57:08.306755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:57:06.780280Z digest=sha256:d8210d919852953c69ad536983e5cdcf8527e2843ad990c7af8105ba397d4542

Observation e7c8b012-c9bf-4d9a-88c3-6474a32cc3c2 · outbound

This paper cites "Do you follow me?": A Survey of Recent Approaches in Dialogue State Tracking.

Agent Identity Evals: Measuring Agentic Identity "Do you follow me?": A Survey of Recent Approaches in Dialogue State Tracking

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T14:57:06.785581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:57:06.785581Z digest=sha256:d806d998a0b46cbffd62d4529c2b31bca1c3b3532adde9010f456166a469164d

Observation f3d31cfa-957b-495e-8293-0e4e54dccf72 · outbound

This paper cites Multi-domain dialogue state tracking with disentangled domain-slot attention.

Agent Identity Evals: Measuring Agentic Identity Multi-domain dialogue state tracking with disentangled domain-slot attention

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:57:08.287517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:57:06.791317Z digest=sha256:d9d0a26dbf2db2f531ab57c856314d24388ded4719da22730ab63ee9594e04b8

Observation ec5e05ec-665d-4be0-a55b-700914390034 · outbound

This paper cites Towards LLM-driven Dialogue State Tracking.

Agent Identity Evals: Measuring Agentic Identity Towards LLM-driven Dialogue State Tracking

Reference 61

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T14:57:07.477381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:57:06.796551Z digest=sha256:c76dae846bf43b8558d965d97e2083dce766f814c7af7fa3607822c178688717

Observation 8a17dc82-efc5-4e51-8b2c-2887ac3113ff · outbound

This paper cites Object-Centric Learning with Slot Attention.

Agent Identity Evals: Measuring Agentic Identity Object-Centric Learning with Slot Attention

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T14:57:06.801744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:57:06.801744Z digest=sha256:10eb796d051c0b296f301512c160e2d1b4fac674555fc011cecdf4d1e339cd8c

Observation db263df3-14bd-4d11-91be-81faad2405bc · outbound

This paper cites 4D Panoptic Scene Graph Generation.

Agent Identity Evals: Measuring Agentic Identity 4D Panoptic Scene Graph Generation

Reference 63

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T14:57:07.434467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:57:06.806851Z digest=sha256:53646d4ef7ac61c39dbb37da907c8df764581cd523cc58c7f238a2d530d0973b

Observation 43676c60-28f5-4613-8984-92cb44f28a37 · outbound

This paper cites Scene Graph Generation: A Comprehensive Survey.

Agent Identity Evals: Measuring Agentic Identity Scene Graph Generation: A Comprehensive Survey

Reference 64

Resolution
verified exact
local_arxiv, observed 2026-08-06T14:57:07.410678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:57:06.811983Z digest=sha256:955d3a347045bd73d8c6e05939d817d7b24b899134e2db6b9f63a6725314ca45

Observation 13c86a5d-bfd6-4355-ba66-b2356f171d86 · outbound

This paper cites From Pixels to Graphs: Open-Vocabulary Scene Graph Generation with Vision-Language Models.

Agent Identity Evals: Measuring Agentic Identity From Pixels to Graphs: Open-Vocabulary Scene Graph Generation with Vision-Language Models

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-08-06T14:57:07.383206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:57:06.817174Z digest=sha256:1daa68e7c406f13cad0fe079bf8b7c66f7f9be62d65b7a27cd833204fe56642a

Observation cf69d3a4-dcf2-485f-a310-0cd7f694a459 · outbound

This paper cites Dichroic cavity mode splitting and lifetimes from interactions with a ferromagnetic metal.

Agent Identity Evals: Measuring Agentic Identity Dichroic cavity mode splitting and lifetimes from interactions with a ferromagnetic metal

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-08-06T14:57:07.356457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:57:06.822275Z digest=sha256:642c2a238072123f643258d5889eabb4fb6ea07f62d11156aca6e7129b0d76bc

Observation da7ce7d5-1f88-400f-b71a-6ad4a2bc369a · outbound

This paper cites Position: Stop Acting Like Language Model Agents Are Normal Agents.

Agent Identity Evals: Measuring Agentic Identity Position: Stop Acting Like Language Model Agents Are Normal Agents

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T14:57:06.827850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:57:06.827850Z digest=sha256:f1ee363b4a5170c04a86715a0dd397ac62cc10351514cd19218139dbb5a6dbf1

Observation 88506cc3-0bc2-4c28-ad0e-e502e65a3106 · outbound

This paper cites The Illusion of State in State-Space Models.

Agent Identity Evals: Measuring Agentic Identity The Illusion of State in State-Space Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T14:57:06.833193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:57:06.833193Z digest=sha256:b5ca6f74ae46086b92b95f715a601f1bf871bcce756d5a1196480b2fba64e3c7

Observation 342b2bb9-3a86-4035-ba76-ff087d74e944 · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

Agent Identity Evals: Measuring Agentic Identity Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T14:57:06.838393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:57:06.838393Z digest=sha256:6d07d9a7056b8fe039601a286f248b03c9e5265178dff54d64746d7587a41559

Observation d8f299ec-a635-494c-b359-b66aedde4a9d · outbound

This paper cites Tree of thoughts: Deliberate problem solving with large language models.Ad- vances in neural information processing systems, 36:11809–11822, 2023.

Agent Identity Evals: Measuring Agentic Identity Tree of thoughts: Deliberate problem solving with large language models.Ad- vances in neural information processing systems, 36:11809–11822, 2023

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T14:57:06.843816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:57:06.843816Z digest=sha256:97fc31cf8950d281ba05694de44b6790b3df2f911a898eb917b92c48acaf4f5d

Observation 92925f36-7b58-42f9-805f-1c5aec1f6bbc · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Agent Identity Evals: Measuring Agentic Identity DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T14:57:06.849365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:57:06.849365Z digest=sha256:6f65cc7ae8b925b38bc78aa3253e626cd464c8c8c5ff4fed5c76b1b035da0294

Observation 00196c89-0ca0-4602-9251-e23e8c0f2a11 · outbound

This paper cites Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell.

Agent Identity Evals: Measuring Agentic Identity Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-06T14:57:06.854681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:57:06.854681Z digest=sha256:d595fba62df44f4688add162258fa20ea36868412656b556763c191360d2de1d

Observation b5ac2ec1-4433-42c9-b7f8-85b93680d51f · outbound

This paper cites QLoRA: Efficient Finetuning of Quantized LLMs.

Agent Identity Evals: Measuring Agentic Identity QLoRA: Efficient Finetuning of Quantized LLMs

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T14:57:06.860316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:57:06.860316Z digest=sha256:02730ed97bce9a24b0f8c3dbaae98505001aa6b277a3a1e7ce4485a2735edda9

Observation e2f44cc2-f6ae-4848-8367-9cb6094efe55 · outbound

This paper cites Do language models have bayesian brains? distinguishing stochastic and deterministic decision patterns within large language models.

Agent Identity Evals: Measuring Agentic Identity Do language models have bayesian brains? distinguishing stochastic and deterministic decision patterns within large language models

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:57:08.240504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:57:06.865678Z digest=sha256:1a3ccff00def4b48036dc36550b3cff21c2f8ce0082d2283f52a114ac6096153

Observation 0092c7d4-4ae0-402a-a1ba-e18115443018 · outbound

This paper cites Do I Know This Entity? Knowledge Awareness and Hallucinations in Language Models.

Agent Identity Evals: Measuring Agentic Identity Do I Know This Entity? Knowledge Awareness and Hallucinations in Language Models

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-06T14:57:06.871290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:57:06.871290Z digest=sha256:cc0a5990d70ec22e3c7753fbb32c93b6a6a583447ae72b14f9813d1c6519ac8d

Observation a81a66a5-757a-4faa-8a14-ebda8925b122 · outbound

This paper cites Exploring Autonomous Agents through the Lens of Large Language Models: A Review.

Agent Identity Evals: Measuring Agentic Identity Exploring Autonomous Agents through the Lens of Large Language Models: A Review

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-06T14:57:06.876711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:57:06.876711Z digest=sha256:f55b87c3778877d2d3c5a94cf8662db73a91dc1e7665c6edb34bdb43124ac919

Observation 15fb529e-3caf-4a08-bb0b-0efca3b5e99f · outbound

This paper cites PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts.

Agent Identity Evals: Measuring Agentic Identity PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T14:57:06.882060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:57:06.882060Z digest=sha256:ec77e9465ef42bfa413e3cc7886c7c1c545162dfaddef6cc8ab55545d36cefc7

Observation b5aa7e15-cdf3-4244-9f65-81628866973c · outbound

This paper cites Robustifying Language Models with Test-Time Adaptation.

Agent Identity Evals: Measuring Agentic Identity Robustifying Language Models with Test-Time Adaptation

Reference 78

Resolution
verified exact
local_arxiv, observed 2026-08-06T14:57:07.180393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:57:06.887492Z digest=sha256:f9dfb92ec23854747e22b6eed62eadde7a94b6f272e2cee2a2c9370434dd498b

Observation 5ac6a081-2131-489e-aab5-3c81dd804c36 · outbound

This paper cites Evaluating the Robustness of Neural Language Models to Input Perturbations.

Agent Identity Evals: Measuring Agentic Identity Evaluating the Robustness of Neural Language Models to Input Perturbations

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-06T14:57:06.893117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:57:06.893117Z digest=sha256:0d8c9881d68806d52ab849b833b886924ab2a96871e5e889748278991acae5c3

Observation 3f07845c-48b5-43dd-8869-ee4845427d18 · outbound

This paper cites KGPA: Robustness Evaluation for Large Language Models via Cross-Domain Knowledge Graphs.

Agent Identity Evals: Measuring Agentic Identity KGPA: Robustness Evaluation for Large Language Models via Cross-Domain Knowledge Graphs

Reference 80

Resolution
verified exact
local_arxiv, observed 2026-08-06T14:57:07.137196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:57:06.898735Z digest=sha256:8b28f50d6ce63af02d7d2d17e0abbd5758a37d2b294536ee4c30e080410e83c3

Observation 069695b2-e419-4f83-b243-d8af4455dc52 · outbound

This paper cites Measuring the Inconsistency of Large Language Models in Preferential Ranking.

Agent Identity Evals: Measuring Agentic Identity Measuring the Inconsistency of Large Language Models in Preferential Ranking

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-06T14:57:06.905560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:57:06.905560Z digest=sha256:84049cfbae49fa539d170229878b5d7c0ab5808581cbe9fc16ebd1b27969def4

Observation 33e30a3e-e233-4d4b-9455-e98db6691ece · outbound

This paper cites ASSERT: Automated Safety Scenario Red Teaming for Evaluating the Robustness of Large Language Models.

Agent Identity Evals: Measuring Agentic Identity ASSERT: Automated Safety Scenario Red Teaming for Evaluating the Robustness of Large Language Models

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-06T14:57:06.911179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:57:06.911179Z digest=sha256:b2513b77f424f3174155e1b2b31e50ff1fc8fd798394d45d5247e0bad527ce19

Observation 53160d10-d113-4d5c-80ca-3e1483657778 · outbound

This paper cites Large language models can be easily distracted by irrelevant context.

Agent Identity Evals: Measuring Agentic Identity Large language models can be easily distracted by irrelevant context

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-06T14:57:06.916840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:57:06.916840Z digest=sha256:1ad56d6f316864c00be00fca3f58a7503373df48c4b1b9079e280bbfdc711721

Observation 8d24e04b-5c05-4f81-8347-4c434e4ca36f · outbound

This paper cites Long Context RAG Performance of Large Language Models.

Agent Identity Evals: Measuring Agentic Identity Long Context RAG Performance of Large Language Models

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-06T14:57:06.923301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:57:06.923301Z digest=sha256:aca14158c6ed6a8a39f5b39718ca49a6cf3fbb79efadf9b0eab659f2afabb3e5

Observation c0c529d7-a28a-4b70-b831-5a3e2e20eb1f · outbound

This paper cites Lost in the middle: How language models use long contexts.Transactions of the Association for Computational Linguistics, 12:157–173, 2024.

Agent Identity Evals: Measuring Agentic Identity Lost in the middle: How language models use long contexts.Transactions of the Association for Computational Linguistics, 12:157–173, 2024

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-06T14:57:06.928664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:57:06.928664Z digest=sha256:b6644740695d0b008a703ad214613dd53c4525f98e8fe4d495431e5834e54257

Observation ff543c29-3253-40c5-b5bd-e37b1f23c550 · outbound

This paper cites Computational dualism and objective superintelligence.

Agent Identity Evals: Measuring Agentic Identity Computational dualism and objective superintelligence

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:57:08.195060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:57:06.933739Z digest=sha256:19639b35f6ec09854309d7520258dc750d0def3010df0ee48eec208abd8da62f

Observation aa3e6cd3-4a6d-47e7-ad1c-2760f91419a5 · outbound

This paper cites Are biological systems more intelligent than artificial intelligence? 2025.

Agent Identity Evals: Measuring Agentic Identity Are biological systems more intelligent than artificial intelligence? 2025

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:57:08.176790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:57:06.939038Z digest=sha256:697e9b2b509f673eaa0b01af27a6c97b476f3637f460a1069c0305d20fc98517

Observation 7a2d9398-0423-400d-86b8-19aefa192ab0 · outbound

This paper cites Philosophical specification of empathetic ethical artificial intelligence.IEEE Transactions on Cognitive and Developmental Systems, 14(2):292–300, 2022.

Agent Identity Evals: Measuring Agentic Identity Philosophical specification of empathetic ethical artificial intelligence.IEEE Transactions on Cognitive and Developmental Systems, 14(2):292–300, 2022

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:57:08.158661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:57:06.943812Z digest=sha256:7f14bcbd7b086f936aadb76b3320834a6824e7a30d8cc2ee1aa863a06cc60e90

Observation 507122a8-1b37-403a-b3f8-50df7bced55a · outbound

This paper cites an unresolved cited work.

Agent Identity Evals: Measuring Agentic Identity Unresolved cited work

Reference 89

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:57:08.139110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:57:06.949190Z digest=sha256:8c7745fea5083db6169a2ce9519daba7d71fa0b6d63278cbdd53faaf535751bc

Observation e17834ce-3c35-4e96-bf23-288aedb93d66 · outbound

This paper cites an unresolved cited work.

Agent Identity Evals: Measuring Agentic Identity Unresolved cited work

Reference 90

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:57:08.120137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:57:06.954486Z digest=sha256:86845fbd2ff03b36c3c6da565daa3eed92ddcf04a7072205f6a12aa4c68349d8

Observation e7b478e5-fa0e-4c70-9b2f-7e170ea32c92 · outbound

This paper cites De- velop a 3-stage marketing strategy for a new eco-friendly coffee brand,.

Agent Identity Evals: Measuring Agentic Identity De- velop a 3-stage marketing strategy for a new eco-friendly coffee brand,

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:57:08.101068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:57:06.960486Z digest=sha256:e3882975c8dafbf74c9c4d836a855a6f0642b4a24bf710ae6ca4a3e4bce4031a

Observation 2a7d42a1-3a74-4474-8bc2-b42ea5b98c08 · outbound

This paper cites –Perform the standardised planning task.

Agent Identity Evals: Measuring Agentic Identity –Perform the standardised planning task

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:57:08.082506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:57:06.966448Z digest=sha256:6e084c491700a0d4f4fdf33a2eb268a59e9e812b67e59bf215c2b714b6dc2b22

Observation a08a87dd-dda9-4de4-964f-62a6da37fcd8 · outbound

This paper cites Planning Performance Scores for Configi).

Agent Identity Evals: Measuring Agentic Identity Planning Performance Scores for Configi)

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:57:08.064659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:57:06.972275Z digest=sha256:e69628688175098987eee5e47866ff6f16e5a6a329b57c0604d463f8d758c80c

Observation 953d708c-93c4-4e7d-a751-82584a7b432c · outbound

This paper cites an unresolved cited work.

Agent Identity Evals: Measuring Agentic Identity Unresolved cited work

Reference 94

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:57:08.046352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:57:06.978251Z digest=sha256:7fd69b98641f7e7f8536a5defb25e0de26706a849df2545733b56c165d1d6727

Observation f1d46ca2-af28-4f1f-adf7-84060a0c1d79 · outbound

This paper cites an unresolved cited work.

Agent Identity Evals: Measuring Agentic Identity Unresolved cited work

Reference 95

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:57:08.020161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:57:06.984262Z digest=sha256:ba25045bfaad2955ffe4c1ef55055c95eae4f53eb18ed7f4950164e8d149c6c8

Observation ebe7b66b-9a3b-452b-bc7d-ca80c71b002f · outbound

This paper cites every k ticks.

Agent Identity Evals: Measuring Agentic Identity every k ticks

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:57:08.000276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:57:06.989789Z digest=sha256:6fb94b3f247074cd78e8313676273543a7cdd4813ebee943e7b5a6041f3acf03

Observation d4870296-af78-41d6-b209-2a9057d762d0 · outbound

This paper cites LLMs do not retain information across separate inference instances [ 68, 58].

Agent Identity Evals: Measuring Agentic Identity LLMs do not retain information across separate inference instances [ 68, 58]

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:57:07.983099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:57:06.996541Z digest=sha256:21e1976524ff4f7558676ed37286fed233c48fd6e0b67dc143d60a663b806102

Observation 53e6b534-a56c-4092-b748-3d901feaa208 · outbound

This paper cites LLM outputs are typically probabilistic [72, 73, 74], meaning the same query can yield varying or even incorrect results on different runs [ 75].

Agent Identity Evals: Measuring Agentic Identity LLM outputs are typically probabilistic [72, 73, 74], meaning the same query can yield varying or even incorrect results on different runs [ 75]

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:57:07.963009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:57:07.002510Z digest=sha256:a41aaf7962843eb4781e61afd73bde7b7937e26b8658eca9deeea47e07533d74

Observation 1483fa5c-bb67-46e7-84ba-be185531ebd6 · outbound

This paper cites an unresolved cited work.

Agent Identity Evals: Measuring Agentic Identity Unresolved cited work

Reference 99

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:57:07.944459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:57:07.008612Z digest=sha256:15909218aad6d5b999a7371c0a7828a1125f2a0096f5c61dc329a678c7217bc8

Observation 509b5f59-6a6a-4195-9454-48d853ad1cdb · outbound

This paper cites All interaction with an LLM is text-based: agent definitions, environmental factors, and actions are translated into tokens, which the LLM interprets to produce responses in kind.

Agent Identity Evals: Measuring Agentic Identity All interaction with an LLM is text-based: agent definitions, environmental factors, and actions are translated into tokens, which the LLM interprets to produce responses in kind

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:57:07.926581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:57:07.014317Z digest=sha256:42fcb47a720ca40797d2a266b8f3141c65e26b3eb62dc3c6c65d30a1ceb0c803

Pith citing papers

Observation e75ce735-be26-4773-b040-641474872e2b · inbound

Dissociative Identity: Language Model Agents Lack Grounding for Reputation Mechanisms cites this paper.

Dissociative Identity: Language Model Agents Lack Grounding for Reputation Mechanisms Agent Identity Evals: Measuring Agentic Identity

Reference 101

Resolution
verified exact
arxiv_id, observed 2026-06-29T00:32:53.095668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T00:26:54.019256Z digest=sha256:3fe4cd8a52d295a238d9f3cda1c84a70eca79d92bf87ae693bc48023c284ba6d