Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-30T16:13:43.527362Z
Paper Citation Record · LEDGER
As of 20 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 2 inbound Pith citation observations for arXiv:2605.23657.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-30T16:13:43.527362Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T00:35:53.030478Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-03T04:57:38.275398Z
47 of 47 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation b75c7611-4df1-4cb6-8a1f-87ba7f386b1a · outbound
OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents GPT-5.4 thinking system card
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 1a67702c-f8b8-41ef-ac56-2b263c546dec · outbound
OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents System card: Claude Opus 4.6
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 159a3b8f-3929-4fe5-a190-9498cf2e2d6b · outbound
OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents Claude code by anthropic | ai coding agent, terminal, ide
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 4f7bf070-115a-4f84-a142-23b5e1acff31 · outbound
OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents Codex by openai | ai coding agent.https://openai.com/codex/
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 50003525-bc7b-41db-ad5f-bb9d5d109e01 · outbound
OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents Equipping agents for the real world with agent skills
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 4977c240-d9a0-4f4b-9400-05254e0e78f9 · outbound
OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents Harbor: A framework for evaluating and optimizing agents and models in container environments, January 2026
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation d43f0d5f-c501-45f7-9c2e-438fc31a1b12 · outbound
OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents Pptagent: Generating and evaluating presentations beyond text-to-slides
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation c4809828-f9c2-4c2c-ba1d-ecf176040395 · outbound
OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents Frabench and ufeval: Unified fine-grained evaluation with task and aspect generalization
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation c9fcc1d4-ab15-49b5-a11b-89a5cc64649b · outbound
OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents Frabench and ufeval: Unified fine-grained evaluation with task and aspect generalization
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation ba44fca2-d872-4ffe-a7f1-78867099894e · outbound
OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents Webarena: A realistic web environment for build- ing autonomous agents
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 87a7d8b4-8cea-4975-bf2a-e0908dabff8f · outbound
OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents GPT-5.3-Codex system card
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 886612b5-d755-44c9-8f2b-14727bc8017f · outbound
OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents Gemini CLI
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation b7e87c06-b283-4d8d-8515-1055fc8f7dd8 · outbound
OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents Gemini 3.1 pro model card
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation d04d6cfc-d3b4-454f-bd4c-18bd4f4c9135 · outbound
OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents Kimi code CLI.https://github.com/MoonshotAI/kimi-cli
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation cff4c5a4-cd66-43b5-8ab1-2e3bafaca1d1 · outbound
OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents Kimi K2: Open Agentic Intelligence
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation b7649d9c-82a9-4b50-80d0-8ffcd34b92e1 · outbound
OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents MiniMax M2.7: Early echoes of self-evolution
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 73968c50-681f-4204-90d1-f84edc767fb8 · outbound
OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents Deepseek-v4: Towards highly efficient million-token context intelligence
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 5891fc9b-7c23-4bbd-8cec-3ae562cdf2b3 · outbound
OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents GLM-5: from Vibe Coding to Agentic Engineering
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation db150676-9cfe-4aa1-87fc-99e0b4311957 · outbound
OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents Intuitive or dependent? investigating LLMs’ behavior style to conflicting prompts
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 4be3eeb2-de98-4f1b-9fe9-815ece0142e8 · outbound
OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents Why claude code skills don’t activate and how to fix it
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation d23f6910-880d-4ba6-80b2-413723754f72 · outbound
OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents OckBench: Measuring the Efficiency of LLM Reasoning
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation e1d08491-2902-4bcf-bc0d-be9edbc51656 · outbound
OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents Measuring Style Similarity in Diffusion Models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 5585cee7-6fd2-42c5-b9dc-cd83449c9d7d · outbound
OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents Agent Skills for Large Language Models: Architecture, Acquisition, Security, and the Path Forward
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 9978f370-df0a-403b-b6ab-3255dd4319f2 · outbound
OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents EvoSkill: Automated Skill Discovery for Multi-Agent Systems
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation f4b2c33e-c5fc-4a9b-a7de-1a0dfb7a9f29 · outbound
OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents Autoskill: Experience-driven lifelong learning via skill self-evolution
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 01a69210-82ef-43f5-895f-def70f343e13 · outbound
OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents SkillWeaver: Web Agents can Self-Improve by Discovering and Honing Skills
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation c7f09166-40c4-4c25-be4d-e1700d3e3007 · outbound
OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 466544ab-b010-4fa9-a672-1b2ff509cbaa · outbound
OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents PinchBench: Real-world benchmarks for AI coding agents
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 0be73169-c536-46d3-aa2b-e056835a5d1c · outbound
OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents Wildclawbench
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 9b50b325-656f-4223-8701-724cb4692bda · outbound
OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents Swe-bench: Can language models resolve real-world github issues? 2023
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation e3b918cc-6bf3-4356-8681-2dc7caca94ec · outbound
OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents Agentbench: Evaluating llms as agents
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 9b742e26-6c99-428d-a1ca-37d918b187c9 · outbound
OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents Toward Generalizable Evaluation in the LLM Era: A Survey Beyond Benchmarks
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation cd20161a-c3ed-40fb-96c8-7169f93cb83a · outbound
OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents Automating dataset updates towards reliable and timely evaluation of large language models
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 48aaae0c-4e2e-43ce-adf8-b10ce95860bc · outbound
OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents URL https://proceedings.neurips.cc/paper_ files/paper/2024/file/1e89c12621c0315373f20f0aeabe5dbe-Paper-Datasets_ and_Benchmarks_Track.pdf
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 35bfca28-bdc0-41dc-82fb-7b610e01bf1a · outbound
OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents EvoWiki: Evaluating LLMs on evolving knowledge
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation d158a6ca-d315-4274-a761-fd2453d3fdf8 · outbound
OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents Livebench: A challenging, contamination-free LLM benchmark
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation b3f14967-defb-48a3-b0c5-eb550d7bdd8a · outbound
OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 96af274b-df03-4a5d-9a72-25d26464b3cd · outbound
OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents application
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 9d90ce09-2278-41f9-9864-87ab83bf4559 · outbound
OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents application
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 7025c0db-c37d-4754-8adf-40e01baea2ba · outbound
OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents application
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 8c65efdb-65ca-44b0-b727-de6659235c25 · outbound
OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents application
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 96e4e560-34d2-4efd-af91-ccbd414de3ef · outbound
OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents application
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation ecc33d8e-8316-48dc-b538-fd916959119a · outbound
OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents expressed
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation a041606d-dc8a-43ca-9078-53150be84cae · outbound
OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents score": <1-5>
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation ea1bedc8-5afd-443d-987c-75d5723b61ba · outbound
OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents score": <1-5>
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 3a54be36-86d9-41ba-9789-739b8c6d4f50 · outbound
OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents score": <1-5>
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation c161e18f-d64d-4f5a-91e5-5d9e44dcc1e5 · outbound
OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents score": <1-5>
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 455fcb44-311a-421e-9bc2-2a45139a5146 · inbound
Skill Coverage: A Test Adequacy Metric for Agent Skills OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation fca0b68c-e694-4d6d-a48d-84bda6132c8f · inbound
Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.