Pith. sign in

Paper Citation Record · LEDGER

SkillJuror: Measuring How Agent Skill Organization Changes Runtime Behavior

As of 19 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 3 inbound Pith citation observations for arXiv:2606.11543.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.11543 v1

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-27T10:12:03.985549Z

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T10:20:42.015720Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T04:26:30.136710Z

Reference resolution

18 of 18 outbound references displayed

  • verified exact14
  • verified fuzzy0
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a962568c-13cb-4197-9c6f-8e31cc492908 · outbound

This paper cites Cost-of-Pass: An economic framework for evaluating language models.

SkillJuror: Measuring How Agent Skill Organization Changes Runtime Behavior Cost-of-Pass: An economic framework for evaluating language models

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T10:17:57.329983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T10:12:03.985549Z digest=sha256:372d99090aed68b6c589e19560a27e6ab6d2c7e247079820fafa2d8db4d3c4e0

Observation 721d9cef-ef29-416b-890e-f6d072e78851 · outbound

This paper cites AgentProcessBench: Diagnosing Step-Level Process Quality in Tool-Using Agents.

SkillJuror: Measuring How Agent Skill Organization Changes Runtime Behavior AgentProcessBench: Diagnosing Step-Level Process Quality in Tool-Using Agents

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-03T10:17:57.324765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T10:12:03.985549Z digest=sha256:6083989fccd633d2b46dfab20e79a6e1b25adbc3e63628b960c81ed59b194ed7

Observation a1e84a75-625a-4f58-99b3-7a3a9f4e1c02 · outbound

This paper cites Swe-skills-bench: Do agent skills actually help in real-world software engineering?.

SkillJuror: Measuring How Agent Skill Organization Changes Runtime Behavior Swe-skills-bench: Do agent skills actually help in real-world software engineering?

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:17:57.339606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T10:12:03.985549Z digest=sha256:e0342cd7fe7e8c67d174444fe6a54e8f66f8eef848d1688c651e3f95346640ee

Observation e5770f74-beaf-4028-9be9-a8a91e3ff81a · outbound

This paper cites arXiv preprint arXiv:2510.04550 , year=.

SkillJuror: Measuring How Agent Skill Organization Changes Runtime Behavior arXiv preprint arXiv:2510.04550 , year=

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:17:57.344341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T10:12:03.985549Z digest=sha256:b5cf772551c20d7b61c8943d244231d8b25b3b5ae6b835b823891a790b67b98f

Observation b07270af-cfda-4327-b90a-3601c691f388 · outbound

This paper cites Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented Agents.

SkillJuror: Measuring How Agent Skill Organization Changes Runtime Behavior Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented Agents

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-03T10:17:57.336615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T10:12:03.985549Z digest=sha256:09410fdc7432f228168982893787c076b6f104f47de7ab15a4f18d14ec6530bb

Observation 4ba12618-cb34-40bb-93d2-a66017a01337 · outbound

This paper cites SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks.

SkillJuror: Measuring How Agent Skill Organization Changes Runtime Behavior SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-03T10:17:57.334435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T10:12:03.985549Z digest=sha256:59f105efa2e712e15ebaf12823bb24f6f26f16a47f44d2e8965c5ba98c4b9c72

Observation 97ef49c5-1581-4ff3-ae95-88bc3fa855fd · outbound

This paper cites Large Language Model Agent: A Survey on Methodology, Applications and Challenges.

SkillJuror: Measuring How Agent Skill Organization Changes Runtime Behavior Large Language Model Agent: A Survey on Methodology, Applications and Challenges

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-07-03T10:17:57.317585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T10:12:03.985549Z digest=sha256:e2ac15eeb5d158b1ec043915bbb0595df950172dba003c88285f699379b0fceb

Observation 38197f54-18ce-4289-80f2-786cfad5cfab · outbound

This paper cites SkillGen: Verified Inference-Time Agent Skill Synthesis.

SkillJuror: Measuring How Agent Skill Organization Changes Runtime Behavior SkillGen: Verified Inference-Time Agent Skill Synthesis

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-03T10:17:57.341738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T10:12:03.985549Z digest=sha256:cb8a764afbfc935f912dd3b8372960e9d6d82f639d80a9684110bd6bc6e3e507

Observation 4b598877-4146-4006-a466-6af2f89edeb5 · outbound

This paper cites Skill-Pro: Learning Reusable Skills from Experience via Non-Parametric PPO for LLM Agents.

SkillJuror: Measuring How Agent Skill Organization Changes Runtime Behavior Skill-Pro: Learning Reusable Skills from Experience via Non-Parametric PPO for LLM Agents

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-03T10:17:57.313157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T10:12:03.985549Z digest=sha256:6076ffd071c393cccdeae62ac6cbb4f145f1ef72d513e2764b1d76e79a692fbe

Observation 91bb7898-74ce-40c8-a992-a35f91275ffc · outbound

This paper cites Quantifying language models’ sensitivity tospuriousfeaturesinpromptdesignor: Howilearnedtostartworryingaboutpromptformatting.

SkillJuror: Measuring How Agent Skill Organization Changes Runtime Behavior Quantifying language models’ sensitivity tospuriousfeaturesinpromptdesignor: Howilearnedtostartworryingaboutpromptformatting

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-27T10:12:03.985549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:12:03.985549Z digest=sha256:69fe78e1579710c2736ff2c6b1bc4dac18d141b3b6e8712b5d6b3f8181f61d8e

Observation 4e5bd309-c953-4979-84d7-7e5260cf1fd1 · outbound

This paper cites More Skills, Worse Agents? Skill Shadowing Degrades Performance When Expanding Skill Libraries.

SkillJuror: Measuring How Agent Skill Organization Changes Runtime Behavior More Skills, Worse Agents? Skill Shadowing Degrades Performance When Expanding Skill Libraries

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-03T10:17:57.315348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T10:12:03.985549Z digest=sha256:5b7461e1eef967cfa30bed7445cefad62e8ccf4607b2dedf1470b710bea6763c

Observation a27af5af-d43c-47a3-908f-74bde58c1899 · outbound

This paper cites RestGPT: Connecting Large Language Models with Real-World RESTful APIs.

SkillJuror: Measuring How Agent Skill Organization Changes Runtime Behavior RestGPT: Connecting Large Language Models with Real-World RESTful APIs

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:17:57.320186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T10:12:03.985549Z digest=sha256:6060faa63125f1672407d0ca65ac7670270958c3e1b5d1e613dbc24811aa517b

Observation 0efd1c0a-6b14-4b83-ac0a-8f137913f47e · outbound

This paper cites Voyager: An Open-Ended Embodied Agent with Large Language Models.

SkillJuror: Measuring How Agent Skill Organization Changes Runtime Behavior Voyager: An Open-Ended Embodied Agent with Large Language Models

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-03T10:17:57.322340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T10:12:03.985549Z digest=sha256:2639780d325d0d19c0194e0b0683c518731b73592c8ab3e58a88668df93a72e1

Observation b81fb300-c2e5-42e7-b7e5-f75951e3964f · outbound

This paper cites Large language models as optimizers.

SkillJuror: Measuring How Agent Skill Organization Changes Runtime Behavior Large language models as optimizers

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-27T10:12:03.985549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:12:03.985549Z digest=sha256:32a686d0516cd5eb8a3bbc22403368c48478e799a2afc7843b9921ba85739a35

Observation fa88216c-5ed6-4203-8edb-303120e3c9e2 · outbound

This paper cites SkillOpt: Executive Strategy for Self-Evolving Agent Skills.

SkillJuror: Measuring How Agent Skill Organization Changes Runtime Behavior SkillOpt: Executive Strategy for Self-Evolving Agent Skills

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-03T10:17:57.310907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T10:12:03.985549Z digest=sha256:7d7b9f779fd6283b8423a5ca1cea8365c8ff06c0cceeb87106c472267790cb2b

Observation 04ba4c68-b8c6-496a-86d9-48e973877bce · outbound

This paper cites SkillLearnBench: Benchmarking Continual Learning Methods for Agent Skill Generation on Real-World Tasks.

SkillJuror: Measuring How Agent Skill Organization Changes Runtime Behavior SkillLearnBench: Benchmarking Continual Learning Methods for Agent Skill Generation on Real-World Tasks

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-03T10:17:57.332246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T10:12:03.985549Z digest=sha256:703e4971997e4a093962d104b9a1e49258b516145d2e0a050f0442231ebc8c56

Observation 224359bf-e6d4-44f7-a1f4-3da4386da2f4 · outbound

This paper cites Externalization in LLM Agents: A Unified Review of Memory, Skills, Protocols and Harness Engineering.

SkillJuror: Measuring How Agent Skill Organization Changes Runtime Behavior Externalization in LLM Agents: A Unified Review of Memory, Skills, Protocols and Harness Engineering

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-07-03T10:17:57.327382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T10:12:03.985549Z digest=sha256:4caf08ac9b913343c93a1d12b86db64b336aa61be644e24f131e760fc6e79e8e

Observation cf9e92dc-5ef9-48c4-82b9-cc707219105e · outbound

This paper cites The ERU- rate denominator excludes unknown events.

SkillJuror: Measuring How Agent Skill Organization Changes Runtime Behavior The ERU- rate denominator excludes unknown events

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-27T10:12:03.985549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:12:03.985549Z digest=sha256:607ee78f599924c5413a24f10a3901628144ad98b4d2361187482056a64c5ca1

Pith citing papers

Observation a535cdee-e078-49be-b9bc-64cd775d6504 · inbound

Skill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving Skills cites this paper.

Skill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving Skills SkillJuror: Measuring How Agent Skill Organization Changes Runtime Behavior

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T04:28:45.144472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:28:45.144472Z digest=sha256:9affe43e552cdb376362146e15f73014d5fb4a9d3b330b191230774e687d6d64

Observation 18af5fd2-88ac-4fc7-ac73-bc12688d9576 · inbound

Rethinking Self-Evolving Agent Skills: Feedback Dynamics over Multiple Rounds cites this paper.

Rethinking Self-Evolving Agent Skills: Feedback Dynamics over Multiple Rounds SkillJuror: Measuring How Agent Skill Organization Changes Runtime Behavior

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-05T04:26:30.139623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-05T04:26:29.268566Z digest=sha256:898e9694645126a6fe5619bbc22c0f481e4f9127d2f02648321405316eccb4d7

Observation 99ce2704-f7f9-4a00-b166-14c7106acc03 · inbound

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution cites this paper.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution SkillJuror: Measuring How Agent Skill Organization Changes Runtime Behavior

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-08T10:20:42.015720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T10:20:42.015720Z digest=sha256:28f9efc28d6b6719c8300f1d2046d7d7379db24c0641b73a7e4d58216f140512