Pith. sign in

Paper Citation Record · LEDGER

OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents

As of 20 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 2 inbound Pith citation observations for arXiv:2605.23657.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.23657 v2

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-30T16:13:43.527362Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:35:53.030478Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-03T04:57:38.275398Z

Reference resolution

47 of 47 outbound references displayed

  • verified exact13
  • verified fuzzy32
  • unresolved0
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b75c7611-4df1-4cb6-8a1f-87ba7f386b1a · outbound

This paper cites GPT-5.4 thinking system card.

OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents GPT-5.4 thinking system card

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T14:14:59.495126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T16:13:43.527362Z digest=sha256:a8d946074bc9e92d2b030826fca2b9b2931c4964fa085f1a7a301b6f94c35bbb

Observation 1a67702c-f8b8-41ef-ac56-2b263c546dec · outbound

This paper cites System card: Claude Opus 4.6.

OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents System card: Claude Opus 4.6

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T14:14:59.541895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T16:13:43.527362Z digest=sha256:a74aa98771fc752fb0ee2cedfeaa4487b9076a19f10c264635669d4a0c790672

Observation 159a3b8f-3929-4fe5-a190-9498cf2e2d6b · outbound

This paper cites Claude code by anthropic | ai coding agent, terminal, ide.

OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents Claude code by anthropic | ai coding agent, terminal, ide

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T14:14:59.548908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T16:13:43.527362Z digest=sha256:7fe05a1d85afc7deb235226148b8b850a7318ab4db85e4e9f97698e2077ef867

Observation 4f7bf070-115a-4f84-a142-23b5e1acff31 · outbound

This paper cites Codex by openai | ai coding agent.https://openai.com/codex/.

OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents Codex by openai | ai coding agent.https://openai.com/codex/

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T14:14:59.496971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T16:13:43.527362Z digest=sha256:2447acaa8ef63db722b56cdfcdb2981fc58e289797e5fb2ffa3c84df3e5ca93e

Observation 50003525-bc7b-41db-ad5f-bb9d5d109e01 · outbound

This paper cites Equipping agents for the real world with agent skills.

OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents Equipping agents for the real world with agent skills

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T14:14:59.532642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T16:13:43.527362Z digest=sha256:5a8d9d124f902cadce4d6dc60448835ca25e84cb76b1fce4c1eaec2f66e64c66

Observation 4977c240-d9a0-4f4b-9400-05254e0e78f9 · outbound

This paper cites Harbor: A framework for evaluating and optimizing agents and models in container environments, January 2026.

OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents Harbor: A framework for evaluating and optimizing agents and models in container environments, January 2026

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T14:14:59.534514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T16:13:43.527362Z digest=sha256:2ef62c4fcb7f514018bb8f8fa44b40f6a76e8403fc1610c117d606482b9fdea2

Observation d43f0d5f-c501-45f7-9c2e-438fc31a1b12 · outbound

This paper cites Pptagent: Generating and evaluating presentations beyond text-to-slides.

OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents Pptagent: Generating and evaluating presentations beyond text-to-slides

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T14:14:59.543607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T16:13:43.527362Z digest=sha256:fa2be199b3777b5d5423c622ec167962f0bccbfbbcac487e34fb8cc6473b712e

Observation c4809828-f9c2-4c2c-ba1d-ecf176040395 · outbound

This paper cites Frabench and ufeval: Unified fine-grained evaluation with task and aspect generalization.

OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents Frabench and ufeval: Unified fine-grained evaluation with task and aspect generalization

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T14:14:59.485827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T16:13:43.527362Z digest=sha256:a6aec1409c0276f833c50ac253f24a90dafa5988b7473ba75f061edac2a165c9

Observation c9fcc1d4-ab15-49b5-a11b-89a5cc64649b · outbound

This paper cites Frabench and ufeval: Unified fine-grained evaluation with task and aspect generalization.

OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents Frabench and ufeval: Unified fine-grained evaluation with task and aspect generalization

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-06-30T16:14:53.084734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T16:13:43.527362Z digest=sha256:bd4687d64f6934f0ed7ee118bfbb3957944c9df167779cce58185b1132e8233e

Observation ba44fca2-d872-4ffe-a7f1-78867099894e · outbound

This paper cites Webarena: A realistic web environment for build- ing autonomous agents.

OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents Webarena: A realistic web environment for build- ing autonomous agents

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T14:14:59.536400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T16:13:43.527362Z digest=sha256:2cafdc7db6774662bbf33a84cdd470680c0e37d3184f519b9f6d96c707b7e5e3

Observation 87a7d8b4-8cea-4975-bf2a-e0908dabff8f · outbound

This paper cites GPT-5.3-Codex system card.

OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents GPT-5.3-Codex system card

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T14:14:59.489631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T16:13:43.527362Z digest=sha256:8e701c0118cf6e9e7d2319ed8b822204daedf71a007fc371626a40d1d68f7a90

Observation 886612b5-d755-44c9-8f2b-14727bc8017f · outbound

This paper cites Gemini CLI.

OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents Gemini CLI

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T14:14:59.513425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T16:13:43.527362Z digest=sha256:dd7f232c9435c2a39103bd5393d1e40df1fcaba2ca58b85c9fb140f83cd9e798

Observation b7e87c06-b283-4d8d-8515-1055fc8f7dd8 · outbound

This paper cites Gemini 3.1 pro model card.

OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents Gemini 3.1 pro model card

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T14:14:59.522823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T16:13:43.527362Z digest=sha256:082f030ecccf3ca47e5682e95348b7d27694aad54a33bacf75b9c098d53abb02

Observation d04d6cfc-d3b4-454f-bd4c-18bd4f4c9135 · outbound

This paper cites Kimi code CLI.https://github.com/MoonshotAI/kimi-cli.

OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents Kimi code CLI.https://github.com/MoonshotAI/kimi-cli

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T14:14:59.483864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T16:13:43.527362Z digest=sha256:4fc1d9164732ea2c4478509fd31c41d073da141a89ee723a786c4cb297640b86

Observation cff4c5a4-cd66-43b5-8ab1-2e3bafaca1d1 · outbound

This paper cites Kimi K2: Open Agentic Intelligence.

OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents Kimi K2: Open Agentic Intelligence

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-06-30T16:14:53.081856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T16:13:43.527362Z digest=sha256:4c45ce4581903eddf15db49c8d9914f08f8d779a0161f1e65e563b50e1a29456

Observation b7649d9c-82a9-4b50-80d0-8ffcd34b92e1 · outbound

This paper cites MiniMax M2.7: Early echoes of self-evolution.

OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents MiniMax M2.7: Early echoes of self-evolution

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T14:14:59.491424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T16:13:43.527362Z digest=sha256:2cdb823090db3eb7b0507bd345a5a3c3fa36c0529033ca6cb19bdd7d34111d72

Observation 73968c50-681f-4204-90d1-f84edc767fb8 · outbound

This paper cites Deepseek-v4: Towards highly efficient million-token context intelligence.

OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents Deepseek-v4: Towards highly efficient million-token context intelligence

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T14:14:59.551140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T16:13:43.527362Z digest=sha256:516db59a7609e9789c29dad3e9b1c59286050c218a87cf0fdbb03cf6df3d43ac

Observation 5891fc9b-7c23-4bbd-8cec-3ae562cdf2b3 · outbound

This paper cites GLM-5: from Vibe Coding to Agentic Engineering.

OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents GLM-5: from Vibe Coding to Agentic Engineering

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-06-30T16:14:53.078272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T16:13:43.527362Z digest=sha256:ad1193b70b038a4777bebf982b94c165627e74412860bae1f39f1dd98a7d4907

Observation db150676-9cfe-4aa1-87fc-99e0b4311957 · outbound

This paper cites Intuitive or dependent? investigating LLMs’ behavior style to conflicting prompts.

OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents Intuitive or dependent? investigating LLMs’ behavior style to conflicting prompts

Reference 19

Resolution
malformed identifier
raw_fallback, observed 2026-07-08T14:14:59.502646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T16:13:43.527362Z digest=sha256:2713f1d6e530f448cae553dfc1c324e3c98bc8484fc858d60d527b32f28a0296

Observation 4be3eeb2-de98-4f1b-9fe9-815ece0142e8 · outbound

This paper cites Why claude code skills don’t activate and how to fix it.

OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents Why claude code skills don’t activate and how to fix it

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T14:14:59.525180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T16:13:43.527362Z digest=sha256:8c1faf4335cb9763fb05f2ebdb07600faa02fd2499b4af13cc5064abdc25d2c7

Observation d23f6910-880d-4ba6-80b2-413723754f72 · outbound

This paper cites OckBench: Measuring the Efficiency of LLM Reasoning.

OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents OckBench: Measuring the Efficiency of LLM Reasoning

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-06-30T16:14:53.076015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T16:13:43.527362Z digest=sha256:a7862adf5b8228d0de48bf9d58fef068a4f459aad83dbbd2603dbb483635eaa6

Observation e1d08491-2902-4bcf-bc0d-be9edbc51656 · outbound

This paper cites Measuring Style Similarity in Diffusion Models.

OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents Measuring Style Similarity in Diffusion Models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-06-30T16:14:53.070417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T16:13:43.527362Z digest=sha256:8a43839a9c87e8c31166ae58d90fabbbb34e980114e22e9c22b3a1219b07307f

Observation 5585cee7-6fd2-42c5-b9dc-cd83449c9d7d · outbound

This paper cites Agent Skills for Large Language Models: Architecture, Acquisition, Security, and the Path Forward.

OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents Agent Skills for Large Language Models: Architecture, Acquisition, Security, and the Path Forward

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-06-30T16:14:53.068148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T16:13:43.527362Z digest=sha256:11abd85c3fe817ba5d4d3a6431a61df1273ae0a4e5d51e01d5506d7f16b3743a

Observation 9978f370-df0a-403b-b6ab-3255dd4319f2 · outbound

This paper cites EvoSkill: Automated Skill Discovery for Multi-Agent Systems.

OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents EvoSkill: Automated Skill Discovery for Multi-Agent Systems

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-06-30T16:14:53.070717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T16:13:43.527362Z digest=sha256:a7327bd1fb55857e90d3cb982d8a6f3601d125388b2cda521c663ee32a7391c8

Observation f4b2c33e-c5fc-4a9b-a7de-1a0dfb7a9f29 · outbound

This paper cites Autoskill: Experience-driven lifelong learning via skill self-evolution.

OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents Autoskill: Experience-driven lifelong learning via skill self-evolution

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-06-30T16:14:53.089748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T16:13:43.527362Z digest=sha256:33ea76500e091347b70bf770f1dc8b074f8dece1d32d1ded117287496a2e68e6

Observation 01a69210-82ef-43f5-895f-def70f343e13 · outbound

This paper cites SkillWeaver: Web Agents can Self-Improve by Discovering and Honing Skills.

OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents SkillWeaver: Web Agents can Self-Improve by Discovering and Honing Skills

Reference 26

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T16:14:53.092208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T16:13:43.527362Z digest=sha256:4bbaf5f908c450b0f4fc8c6f0783da7e1d55dce7a3b33264e40bd45321d65f84

Observation c7f09166-40c4-4c25-be4d-e1700d3e3007 · outbound

This paper cites SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks.

OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-06-30T16:14:53.084252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T16:13:43.527362Z digest=sha256:887533f0c078ba77a71fca96dc957cf27158d3a7e88c37800b56ce97d5df08bd

Observation 466544ab-b010-4fa9-a672-1b2ff509cbaa · outbound

This paper cites PinchBench: Real-world benchmarks for AI coding agents.

OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents PinchBench: Real-world benchmarks for AI coding agents

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T14:14:59.498851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T16:13:43.527362Z digest=sha256:904b615d9baecb0b611b13e4fb188a5d7f81cd5f1b7d9bb46168801b8101fc61

Observation 0be73169-c536-46d3-aa2b-e056835a5d1c · outbound

This paper cites Wildclawbench.

OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents Wildclawbench

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T14:14:59.530826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T16:13:43.527362Z digest=sha256:1ba44ff5dd8280c96ef7c055e8242c6fe808649421a2b0000a0a617fc203be11

Observation 9b50b325-656f-4223-8701-724cb4692bda · outbound

This paper cites Swe-bench: Can language models resolve real-world github issues? 2023.

OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents Swe-bench: Can language models resolve real-world github issues? 2023

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T14:14:59.540226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T16:13:43.527362Z digest=sha256:37a3678e6fcf484f1cb628d6f854c7717d59bac5813268615ea676d05022a58f

Observation e3b918cc-6bf3-4356-8681-2dc7caca94ec · outbound

This paper cites Agentbench: Evaluating llms as agents.

OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents Agentbench: Evaluating llms as agents

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T14:14:59.493270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T16:13:43.527362Z digest=sha256:934c215c580a4437735291231ebb45ed24c66aa9e559f5ab01b0972b1bcefca9

Observation 9b742e26-6c99-428d-a1ca-37d918b187c9 · outbound

This paper cites Toward Generalizable Evaluation in the LLM Era: A Survey Beyond Benchmarks.

OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents Toward Generalizable Evaluation in the LLM Era: A Survey Beyond Benchmarks

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-06-30T16:14:53.086828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T16:13:43.527362Z digest=sha256:18a86443814aefc22d9f0ff4b49494c92c576bf85788578bb03d1fee0b8d700a

Observation cd20161a-c3ed-40fb-96c8-7169f93cb83a · outbound

This paper cites Automating dataset updates towards reliable and timely evaluation of large language models.

OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents Automating dataset updates towards reliable and timely evaluation of large language models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T14:14:59.500837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T16:13:43.527362Z digest=sha256:843e359b5dd0b853083c482313e189ec19e19caefd4b8c82752ab06db514f5e6

Observation 48aaae0c-4e2e-43ce-adf8-b10ce95860bc · outbound

This paper cites URL https://proceedings.neurips.cc/paper_ files/paper/2024/file/1e89c12621c0315373f20f0aeabe5dbe-Paper-Datasets_ and_Benchmarks_Track.pdf.

OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents URL https://proceedings.neurips.cc/paper_ files/paper/2024/file/1e89c12621c0315373f20f0aeabe5dbe-Paper-Datasets_ and_Benchmarks_Track.pdf

Reference 34

Resolution
verified exact
doi, observed 2026-06-30T16:14:52.625105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T16:13:43.527362Z digest=sha256:cae5efd782b59b15c262937bf8f5e9a88dcd65cb9808a61ebe3724f25e6d25e7

Observation 35bfca28-bdc0-41dc-82fb-7b610e01bf1a · outbound

This paper cites EvoWiki: Evaluating LLMs on evolving knowledge.

OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents EvoWiki: Evaluating LLMs on evolving knowledge

Reference 35

Resolution
verified exact
doi, observed 2026-06-30T16:14:52.622408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T16:13:43.527362Z digest=sha256:8a568629d82c6cdb44ac36de798b9020c0d12fced07da47fab4a2943f4fb201c

Observation d158a6ca-d315-4274-a761-fd2453d3fdf8 · outbound

This paper cites Livebench: A challenging, contamination-free LLM benchmark.

OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents Livebench: A challenging, contamination-free LLM benchmark

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T14:14:59.527131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T16:13:43.527362Z digest=sha256:986f26701d3b996803fb1813da1154a38972187117fdebe01ce8e7192707e906

Observation b3f14967-defb-48a3-b0c5-eb550d7bdd8a · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-06-30T16:14:53.094645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T16:13:43.527362Z digest=sha256:1c57d36218d13862a62d631728c4ff8ffa5fd4d3bb93f062f7fbd44d50ed38e3

Observation 96af274b-df03-4a5d-9a72-25d26464b3cd · outbound

This paper cites application.

OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents application

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T14:14:59.487693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T16:13:43.527362Z digest=sha256:23a0268476669ac64498424cbd1f90c1aa76cf0ae8bf4c0e56387930180526f3

Observation 9d90ce09-2278-41f9-9864-87ab83bf4559 · outbound

This paper cites application.

OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents application

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T14:14:59.511213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T16:13:43.527362Z digest=sha256:5dc44aa926d6c14762b00ff278c634679eab8391c29671bb5d23c72a6ee6ed6e

Observation 7025c0db-c37d-4754-8adf-40e01baea2ba · outbound

This paper cites application.

OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents application

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T14:14:59.538226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T16:13:43.527362Z digest=sha256:69ad2583b7a47aaf66f84a1037c1225b9607fe5ea1d159076e57fabedc22b646

Observation 8c65efdb-65ca-44b0-b727-de6659235c25 · outbound

This paper cites application.

OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents application

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T14:14:59.515389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T16:13:43.527362Z digest=sha256:63d5bcebedb7ce0f51efa6cc79a5f23f8b32142f4b00f09fc49b2447deb73420

Observation 96e4e560-34d2-4efd-af91-ccbd414de3ef · outbound

This paper cites application.

OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents application

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T14:14:59.546965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T16:13:43.527362Z digest=sha256:816a3e5b9ed7beed48c53497279d6126a406fd0def8605b30ee5b3df16b0b208

Observation ecc33d8e-8316-48dc-b538-fd916959119a · outbound

This paper cites expressed.

OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents expressed

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T14:14:59.506288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T16:13:43.527362Z digest=sha256:4a2b7026f9bc5d8cb04763f0b38fcfbe54ab5d673f47ceba28e1f5c034fb3d4d

Observation a041606d-dc8a-43ca-9078-53150be84cae · outbound

This paper cites score": <1-5>.

OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents score": <1-5>

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T14:14:59.529023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T16:13:43.527362Z digest=sha256:c72edf6e3675702347001341942dc9420a991dd929fd1693fc4132f222f63f12

Observation ea1bedc8-5afd-443d-987c-75d5723b61ba · outbound

This paper cites score": <1-5>.

OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents score": <1-5>

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T14:14:59.482009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T16:13:43.527362Z digest=sha256:daac1c219567909b3793cdabaefd6eb77a5426166074afcc73d751e5f024669f

Observation 3a54be36-86d9-41ba-9789-739b8c6d4f50 · outbound

This paper cites score": <1-5>.

OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents score": <1-5>

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T14:14:59.504441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T16:13:43.527362Z digest=sha256:6303b7c372a3678d9bd9384946eb809a8a66dc63605075f818d246545c4e05fc

Observation c161e18f-d64d-4f5a-91e5-5d9e44dcc1e5 · outbound

This paper cites score": <1-5>.

OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents score": <1-5>

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T14:14:59.545213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T16:13:43.527362Z digest=sha256:279dfa65ca1ab11b63f0f100d96b5d3e0d6336e2af0da2b7cada69d89c8a3cc8

Pith citing papers

Observation 455fcb44-311a-421e-9bc2-2a45139a5146 · inbound

Skill Coverage: A Test Adequacy Metric for Agent Skills cites this paper.

Skill Coverage: A Test Adequacy Metric for Agent Skills OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-07-03T04:57:38.276563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T13:32:51.474878Z digest=sha256:5c60cf56dd5040b79dc8be2f01fd8817d348331e5bfda1ad4d592de6b18d44bc

Observation fca0b68c-e694-4d6d-a48d-84bda6132c8f · inbound

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents cites this paper.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T00:35:53.030478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:35:53.030478Z digest=sha256:d083328bfdab5a671aa08d73d6d2bea5b40136cef4e096da57a8610b41090385