Pith. sign in

Paper Citation Record · LEDGER

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning

As of 15 August 2026, this Paper Citation Record lists 87 of 87 outbound references and 0 inbound Pith citation observations for arXiv:2608.05139.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.05139 v1

Coverage vector

measured 87 of 87 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T04:30:47.578235Z

measured 87 of 87 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

87 of 87 outbound references displayed

  • verified exact13
  • verified fuzzy13
  • unresolved60
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 22f954c5-f7fa-4896-b4de-52cadc159226 · outbound

This paper cites AgentCoMa: A Compositional Benchmark Mixing Commonsense and Mathematical Reasoning in Real-World Scenarios.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning AgentCoMa: A Compositional Benchmark Mixing Commonsense and Mathematical Reasoning in Real-World Scenarios

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:38.138268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:38.138268Z digest=sha256:436e2ff78b2da1ea0979eef6745e277c879f54d0611f4cf86383cf7ebdbf881f

Observation 85344c15-b040-40d8-8ff8-eda6b9ff7f74 · outbound

This paper cites STaD: Scaffolded Task Design for Identifying Compositional Skill Gaps in LLMs.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning STaD: Scaffolded Task Design for Identifying Compositional Skill Gaps in LLMs

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-06T04:30:49.765028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T04:30:38.264262Z digest=sha256:508a641283ff5b0dabd11c93b1c2b07a03d59e4ceb1c19ee0f63f02b6dff3faa

Observation 2fc5af15-6fa6-4c42-acba-1ed19670ae6c · outbound

This paper cites HeroBench: A Benchmark for Long-Horizon Planning and Structured Reasoning in Virtual Worlds.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning HeroBench: A Benchmark for Long-Horizon Planning and Structured Reasoning in Virtual Worlds

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:38.368705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:38.368705Z digest=sha256:63a5edd4aabc26ed7b432070f600d2c388ec0738382256c500fc000db2749f97

Observation d332b2ae-1cc2-4932-99a7-15062ac05017 · outbound

This paper cites Concepts or Skills? Rethinking Instruction Selection for Multi-modal Models.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Concepts or Skills? Rethinking Instruction Selection for Multi-modal Models

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-06T04:30:49.742374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T04:30:38.448384Z digest=sha256:cd624e935597d64cba0bb8a21845921dfd3767475c0b5d728dacceee630ac6fe

Observation 314b5d5d-6905-4b0c-94ec-bb31b4f714f4 · outbound

This paper cites Skill-Based Mixture-of-Experts: Adaptive Routing for Heterogeneous Reasoning via Inferred Skills.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Skill-Based Mixture-of-Experts: Adaptive Routing for Heterogeneous Reasoning via Inferred Skills

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:38.535214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:38.535214Z digest=sha256:912de9d33209642af793d0bde881308942db840a8dc5891e2671cb6c6044eefc

Observation 7726b0ef-4a0a-42e4-8438-80c868ed471b · outbound

This paper cites SkillCraft: Can LLM agents learn to use tools skillfully?arXiv preprint arXiv:2603.00718, 2026.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning SkillCraft: Can LLM agents learn to use tools skillfully?arXiv preprint arXiv:2603.00718, 2026

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:38.636140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:38.636140Z digest=sha256:c066c5face87b6d0bb5523237f9a9978776304462d7c723254f29d45de3904da

Observation 79047070-6705-4482-9ea9-ef5261f62b8a · outbound

This paper cites Self-evolving curriculum for LLM reasoning.arXiv preprint arXiv:2505.14970, 2025.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Self-evolving curriculum for LLM reasoning.arXiv preprint arXiv:2505.14970, 2025

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:38.777948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:38.777948Z digest=sha256:c3347cdb0b7ac27b09c63090011bcdaec098a2f8252139d6e17ad7784fc0980e

Observation af86f2c9-eed4-4a18-90d8-f13ccc6943a8 · outbound

This paper cites WebSRC: A Dataset for Web-Based Structural Reading Comprehension.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning WebSRC: A Dataset for Web-Based Structural Reading Comprehension

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:38.889955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:38.889955Z digest=sha256:9fe82919f261e2295b3a7031c06821c60ba7d6cdfe187a920934e8d85eb8a7c7

Observation e00d57f2-0c17-4e4c-b441-2e87735c546b · outbound

This paper cites Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:38.938928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:38.938928Z digest=sha256:e7a2a73d16f91851c11868968d1737fc85fa7fbf8154cf055e24c7063c22d10e

Observation 89a2579b-c321-4849-8b09-1032a7c6b001 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Training Verifiers to Solve Math Word Problems

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:39.031956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:39.031956Z digest=sha256:2c34fb0bf42763dd381f951608db7b2ced31ba8b66977ecb4c2b22163d6f0373

Observation ed4106fb-7792-4233-8261-434ad4b6bfd6 · outbound

This paper cites Metacognitive capabilities of LLMs: An exploration in mathematical problem solving.Advances in Neural Information Processing Systems, 2024.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Metacognitive capabilities of LLMs: An exploration in mathematical problem solving.Advances in Neural Information Processing Systems, 2024

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:39.141611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:39.141611Z digest=sha256:37dbd2218dd5729766f461d645f536eb3557691233d01ac23c91fcfd75bb589f

Observation 4e249b7d-bd6a-40e0-871e-0056f67c682c · outbound

This paper cites Thinkless: LLM Learns When to Think.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Thinkless: LLM Learns When to Think

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:39.238053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:39.238053Z digest=sha256:4236c911761b21ac406df9b1afad66d610ba0d1b80d9e81aa215bc5ed647e7b2

Observation 74652f66-0821-4cb1-a765-b61f5f08c2e9 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:39.338585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:39.338585Z digest=sha256:99b6eb737e9117c8bd6db0308b7c5646669eb8b2b8327a7582491de9c589f2cf

Observation e38f4f09-afe9-4b91-9cad-06f220e0180c · outbound

This paper cites R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:39.422268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:39.422268Z digest=sha256:1d9da8d70badfff039c15fbbe6aa664bd3d4344213246957b1cc062484062a8d

Observation 43c2c010-9b67-4309-a4ea-a3cba7e046f5 · outbound

This paper cites AdaptMI: Adaptive Skill-based In-context Math Instruction for Small Language Models.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning AdaptMI: Adaptive Skill-based In-context Math Instruction for Small Language Models

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-06T04:30:49.550875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T04:30:39.563825Z digest=sha256:c099a3c208ac38db23e1c4f9677f6f7b19c96eb6227002828e4c5799e468a4b4

Observation a21e7b44-59e0-4b5a-b43d-5d0ed2e6e5a6 · outbound

This paper cites STAT: Skill-targeted adaptive training.arXiv preprint arXiv:2510.10023, 2025.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning STAT: Skill-targeted adaptive training.arXiv preprint arXiv:2510.10023, 2025

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:39.641296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:39.641296Z digest=sha256:eb3c5713ad2b1f4574e869481a38c705396d339c68dc6922da7d1a04aff91629

Observation fdfaa97d-e9d8-461e-bef5-8363bdb39ffc · outbound

This paper cites Self-Distillation Zero: Self-Revision Turns Binary Rewards into Dense Supervision.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Self-Distillation Zero: Self-Revision Turns Binary Rewards into Dense Supervision

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:39.705631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:39.705631Z digest=sha256:7e2aaf2c06fc208c8aa1281385ee966e10b5289375be71ce7fb7cc7f01cd1c7b

Observation d40a405c-c813-416e-a2f7-7e2774c6a670 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Measuring Mathematical Problem Solving With the MATH Dataset

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:39.862494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:39.862494Z digest=sha256:339a6ad5d591e440ed47d5d71da07aeecbdde22d23a58793638667c3afce8ce8

Observation 772fc02a-582d-4e54-b85c-a10830a34981 · outbound

This paper cites Evaluating Memory in LLM Agents via Incremental Multi-Turn Interactions.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Evaluating Memory in LLM Agents via Incremental Multi-Turn Interactions

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:39.961952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:39.961952Z digest=sha256:2586452b5d3732055e5015ac4a835dd99b35a96b2c9c64f78cfb9fe6963a713d

Observation 15ec8f37-481f-4885-b375-d5cad83cc672 · outbound

This paper cites Open-R1: A fully open reproduction of DeepSeek-R1.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Open-R1: A fully open reproduction of DeepSeek-R1

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:40.042883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:40.042883Z digest=sha256:15b2317afb5a2a4c855f1d93277e3da02c7266b1e5eb1e474e7d57b58f6ba9a4

Observation ea8571ec-8a18-4a8d-84f0-c290b8239f7e · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:40.119348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:40.119348Z digest=sha256:9ac023b16b692f3fd934e2871685f56710865070ff1151c0bcd1f36930e8b331

Observation cef7bf9b-2075-44b1-95b1-70fdfc3554a4 · outbound

This paper cites DRP: Distilled Reasoning Pruning with Skill-aware Step Decomposition for Efficient Large Reasoning Models.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning DRP: Distilled Reasoning Pruning with Skill-aware Step Decomposition for Efficient Large Reasoning Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:40.212488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:40.212488Z digest=sha256:ccf544e3102add0044b783dbccb29316def45195f927955318a33037709a5801

Observation eaf7ed6d-2b24-4875-a2cc-0e669720ea6d · outbound

This paper cites Instruct-SkillMix: A Powerful Pipeline for LLM Instruction Tuning.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Instruct-SkillMix: A Powerful Pipeline for LLM Instruction Tuning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:40.386775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:40.386775Z digest=sha256:065968c121681513540e2632803a48fb63e7149a86246fc45919298690bef412

Observation 79592456-402e-4e60-bf62-d166c363c77b · outbound

This paper cites BIG-Bench Extra Hard.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning BIG-Bench Extra Hard

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:40.572955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:40.572955Z digest=sha256:a977dcdf143450d77fcb2dc1ce2b423ba1c717d3d973b1198fac941b5f90bdab

Observation 634ed69b-c440-426e-8986-7e327cee4b8f · outbound

This paper cites Benchmark profiling: Mechanistic diagnosis of LLM benchmarks.arXiv preprint arXiv:2510.01232, 2025.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Benchmark profiling: Mechanistic diagnosis of LLM benchmarks.arXiv preprint arXiv:2510.01232, 2025

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:40.730539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:40.730539Z digest=sha256:e166cf0fd2453b7c07618e0a501a66215eb2debece402e5db74318e6daa3e484

Observation c7f1954f-b3c2-46f5-b7cf-9da893fc87d2 · outbound

This paper cites MSCoRe: A benchmark for multi-stage collaborative reasoning in LLM agents.arXiv preprint arXiv:2509.17628, 2025.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning MSCoRe: A benchmark for multi-stage collaborative reasoning in LLM agents.arXiv preprint arXiv:2509.17628, 2025

Reference 26

Resolution
verified exact
raw_fallback, observed 2026-08-06T04:30:49.323037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T04:30:40.868199Z digest=sha256:55b82731e0bc84e82738c5bada0a11e8539158e240f01a3023a374c59bd5e250

Observation 8c72a2b4-2d84-42ef-aa85-1ee0fb024670 · outbound

This paper cites START: Self-taught Reasoner with Tools.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning START: Self-taught Reasoner with Tools

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:41.015987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:41.015987Z digest=sha256:d65406d568c318bd799638f923b7a77cf3c0e8e2a80fd2c5f0d8ba080ee0d298

Observation bc749d60-dded-40be-b893-bf3265216b73 · outbound

This paper cites SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:41.170677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:41.170677Z digest=sha256:fabd1bd7de3729095fa64284725b38d9e7da029991ad0fb5d8d89820927d04d0

Observation b6b5672e-cab7-459a-bf83-17daea23ab7b · outbound

This paper cites Benchmark test-time scaling of general LLM agents.arXiv preprint arXiv:2602.18998, 2026.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Benchmark test-time scaling of general LLM agents.arXiv preprint arXiv:2602.18998, 2026

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:41.324235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:41.324235Z digest=sha256:4796c300567e3c37b1174253ebf9aabca21b01533e6d16aa1a58a663ccc4c77d

Observation 7f1b4aa7-6cba-4dd4-972b-d20686689bf4 · outbound

This paper cites MTR-Bench: A Comprehensive Benchmark for Multi-Turn Reasoning Evaluation.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning MTR-Bench: A Comprehensive Benchmark for Multi-Turn Reasoning Evaluation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:41.455573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:41.455573Z digest=sha256:91700ac0cb0e8ed4c02f6ec6e219a5619d354c41d88cc52a75ce1160b931140a

Observation d1a728ed-df21-49e1-950f-bbadfa80112b · outbound

This paper cites Let's Verify Step by Step.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Let's Verify Step by Step

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:41.616138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:41.616138Z digest=sha256:df24666c222effe1792265ed75dfabc805e14021212de88135fb2c0920cd328f

Observation cad6590c-cf1c-48fa-a5a2-3e51e0b2dc1d · outbound

This paper cites ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:41.776241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:41.776241Z digest=sha256:2ed48592dcfa61cb367e2592287d3a40a38e27aed33a84d0ad8b39041742f4f1

Observation a80a4082-64be-48d1-b0ba-890af7af79d9 · outbound

This paper cites AgentBench: Evaluating LLMs as Agents.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning AgentBench: Evaluating LLMs as Agents

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:41.907092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:41.907092Z digest=sha256:2f79455015dbd3adcabdde60b5c93c2acdd19a85da455965f0080bda3ebc2d73

Observation 5a04cf71-b807-41e6-aed6-8baef6e5fe51 · outbound

This paper cites GAIA: a benchmark for General AI Assistants.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning GAIA: a benchmark for General AI Assistants

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:42.065260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:42.065260Z digest=sha256:3befa71c93cb37a41691cca6a8f5b396069dcc849521dcf9f5416958c3eee4d9

Observation a0bb827e-e8e3-4fb8-9109-87a6f223097e · outbound

This paper cites Benchmarking and Understanding Compositional Relational Reasoning of LLMs.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Benchmarking and Understanding Compositional Relational Reasoning of LLMs

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-08-06T04:30:49.095104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T04:30:42.215749Z digest=sha256:5d1c229c805bc863c9ebf52a5e6b60dd31b1b4472be68836b3449f248c25c6ae

Observation c135f8fc-0df8-42c5-849d-509b8887d953 · outbound

This paper cites Reasoning curriculum: Bootstrapping broad LLM reasoning from math.arXiv preprint arXiv:2510.26143, 2025.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Reasoning curriculum: Bootstrapping broad LLM reasoning from math.arXiv preprint arXiv:2510.26143, 2025

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:42.370880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:42.370880Z digest=sha256:8920091983df86303303febb0b5f6ed8de935660f51d15928080652b0bc72147

Observation 263f343b-08a9-4a51-abf6-b814cf1ab837 · outbound

This paper cites Compositional Semantic Parsing on Semi-Structured Tables.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Compositional Semantic Parsing on Semi-Structured Tables

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:42.504186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:42.504186Z digest=sha256:1815a40bf5a0de41de07ffb5f86aa972469ad5f4966791c0bf457f0a4043eb8a

Observation d070f3e4-b2fb-4427-96fe-6e06ad3690c8 · outbound

This paper cites Learning to reason across parallel samples for LLM reasoning.arXiv preprint arXiv:2506.09014, 2025.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Learning to reason across parallel samples for LLM reasoning.arXiv preprint arXiv:2506.09014, 2025

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:42.598191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:42.598191Z digest=sha256:248011a35b38da59c83ee64afe87d51c62a4efa471db548c1f959f1fc33eabe2

Observation d5fb6985-d76d-40f3-a2a0-a51a09c8b4e2 · outbound

This paper cites EmoAgent: Assessing and safeguarding human-AI interaction for mental health safety.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning EmoAgent: Assessing and safeguarding human-AI interaction for mental health safety

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:42.722899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:42.722899Z digest=sha256:a72d31380f47e1396baa2600e82964266aeb8ac79c520ee04dc014cbc5c25159

Observation 47ee61b7-1e6d-428b-8a50-87643f3dfebc · outbound

This paper cites LogicSkills: A structured benchmark for formal reasoning in large language models.arXiv preprint arXiv:2602.06533, 2026.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning LogicSkills: A structured benchmark for formal reasoning in large language models.arXiv preprint arXiv:2602.06533, 2026

Reference 40

Resolution
verified exact
raw_fallback, observed 2026-08-06T04:30:48.916500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T04:30:42.911167Z digest=sha256:f1f0847db0664b5dd2eb90e8f7606342329c2f0af2c6daccfb52bb26ddc5d7fb

Observation 5bb94826-154a-4760-abbf-48449447200f · outbound

This paper cites Reasoning models are test exploiters: Rethinking multiple-choice.arXiv preprint arXiv:2507.15337, 2025.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Reasoning models are test exploiters: Rethinking multiple-choice.arXiv preprint arXiv:2507.15337, 2025

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:43.060338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:43.060338Z digest=sha256:6f4330580f170597098f38c209ffed35a04ef9ae8bc27b439cf5eca01fd0eb2b

Observation 10a0be47-6ecc-48d0-9f7c-af1da1abf9c6 · outbound

This paper cites Revisiting Compositional Generalization Capability of Large Language Models Considering Instruction Following Ability.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Revisiting Compositional Generalization Capability of Large Language Models Considering Instruction Following Ability

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-08-06T04:30:48.758172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T04:30:43.189580Z digest=sha256:91c8ea116883a9d8682e9179e581f17141a3ef9ee5ed390ca2722cf6b6115abe

Observation 3769a8a2-921a-42ea-a2ef-7aab904d3024 · outbound

This paper cites AI-Assisted Generation of Difficult Math Questions.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning AI-Assisted Generation of Difficult Math Questions

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:43.301574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:43.301574Z digest=sha256:28e74b2a234b3fe3fad5ec073b2800bd26aff9a4c26c5fa89f2babe97da39375

Observation 47094969-dc1f-4766-9d03-82be32a2124c · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:43.408023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:43.408023Z digest=sha256:8ec801baf80c4a5e329560813b9ca9780b81c2f6f085fbe0acc795b37e387e52

Observation fd41d836-26bf-4ccb-8a8d-7e2b61b4918a · outbound

This paper cites DARE- bench: Evaluating modeling and instruction fidelity of LLMs in data science.arXiv preprint arXiv:2602.24288, 2026.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning DARE- bench: Evaluating modeling and instruction fidelity of LLMs in data science.arXiv preprint arXiv:2602.24288, 2026

Reference 45

Resolution
verified exact
raw_fallback, observed 2026-08-06T04:30:48.717785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T04:30:43.543973Z digest=sha256:87382f146faac301d8d63de3bcee30f35cb35dfb32ca5d75313809763c2981f4

Observation 6c3ae39e-671d-4674-a715-775352b9cce4 · outbound

This paper cites PaperBench: Evaluating AI's Ability to Replicate AI Research.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:43.695729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:43.695729Z digest=sha256:ad529f194fc1c972e317b38daf80964412e35d9aa0c0cffcde316345ad33961e

Observation 705f7630-3321-4561-bf0d-3c7c08bd223a · outbound

This paper cites Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:43.796711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:43.796711Z digest=sha256:df1547aac3568be989d48c6ff6842506517fba913c40afd26b8347787788de3a

Observation 7834b505-9f05-4d7a-8e38-ead4733a71c1 · outbound

This paper cites Reinforcement Learning for Self-Improving Agent with Skill Library.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Reinforcement Learning for Self-Improving Agent with Skill Library

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:43.949208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:43.949208Z digest=sha256:24e540500a0196da526f3a462880a2c83df29b2836f9d75d201932178399ead2

Observation a28e7c8e-8e10-49e3-937b-da961f2cf94c · outbound

This paper cites Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:44.073590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:44.073590Z digest=sha256:d6e34843286816d0050491f6499402076c921592e97ec96a48e040207eda9b0d

Observation 9eac2e05-e348-4e38-86bb-fc182ea509c7 · outbound

This paper cites OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:44.200213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:44.200213Z digest=sha256:2588e06d2bf58e17313f486bd2d0f65b80546a4700e9c70fb6412a70ce07e623

Observation 05187c29-552d-42d3-83c5-fd0d9b42a8ec · outbound

This paper cites MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:44.345592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:44.345592Z digest=sha256:e0982bcecac5f397000a2c38ee8b1fc69a80b9e1f87e5f02e6db509e115d33b9

Observation aa1f2ede-0d00-40fd-92b6-f54e5b855943 · outbound

This paper cites Reinforcingmulti-turn reasoning in LLM agents via turn-level reward design.arXiv preprint arXiv:2505.11821, 2025.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Reinforcingmulti-turn reasoning in LLM agents via turn-level reward design.arXiv preprint arXiv:2505.11821, 2025

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:44.416172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:44.416172Z digest=sha256:4f85a31c9b4f724f55d6061571f0fe51312e4c039bb7a7a87dd112bbdc9d89ed

Observation fb388906-b5c2-4b7d-8ce8-4ffefc7cc252 · outbound

This paper cites Towardscompositionalgeneralization of LLMs via skill taxonomy guided data synthesis.arXiv preprint arXiv:2601.03676, 2026.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Towardscompositionalgeneralization of LLMs via skill taxonomy guided data synthesis.arXiv preprint arXiv:2601.03676, 2026

Reference 53

Resolution
verified exact
raw_fallback, observed 2026-08-06T04:30:48.497373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T04:30:44.481066Z digest=sha256:2a522aadc09877a66738b5d6aa426619fd646309708391e6bacb1c22b67ad898

Observation b0e4634c-9404-4cd6-a681-ab2c31072565 · outbound

This paper cites Hi-ToM: A benchmark for evaluating higher-order theory of mind reasoning in large language models.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Hi-ToM: A benchmark for evaluating higher-order theory of mind reasoning in large language models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:44.546621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:44.546621Z digest=sha256:1c5223c3fa11df61cba973062fb4df7e882fae726f91b93fd976474f22f3504a

Observation 81e1505c-7efa-49cc-9bca-9615802b1c9b · outbound

This paper cites CritICL: Inference-time weak-to-strong generalization from small language model failure modes.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning CritICL: Inference-time weak-to-strong generalization from small language model failure modes

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:44.623109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:44.623109Z digest=sha256:a1254b60b7034254d902c97f92cd39132e14221586a8f86127aecdb784e6706e

Observation 564db2cd-3f30-4d90-aa42-6f26843d43a8 · outbound

This paper cites SkillRL: Evolving agents via recursive skill-augmented reinforcement learning, 2026.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning SkillRL: Evolving agents via recursive skill-augmented reinforcement learning, 2026

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:44.689567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:44.689567Z digest=sha256:c427e91769a1aadbeef6e52a9afeaab59e398ff95a36cb575adaff9a6b41e988

Observation 4e35a80f-b213-4d98-93a5-f462b54b3f1b · outbound

This paper cites LaRS: Latent Reasoning Skills for Chain-of-Thought Reasoning.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning LaRS: Latent Reasoning Skills for Chain-of-Thought Reasoning

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-08-06T04:30:48.427051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T04:30:44.748285Z digest=sha256:6380c28133f9ddc9073485a2d0bf8dff735abba60df411fc2065d05ddab37781

Observation 4be92d90-751e-4fa9-88f4-907765b67579 · outbound

This paper cites DeepCritic: Deliberate Critique with Large Language Models.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning DeepCritic: Deliberate Critique with Large Language Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:44.844166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:44.844166Z digest=sha256:c502ccb1253852d1267db9a3a1a97b59b17afc05abecffc7a9811cbd8bd754b6

Observation 5fdab5d0-8554-4199-93e0-da0123f8527f · outbound

This paper cites Transferability for General Reasoning: An Automated Curriculum for Multi-Domain RLVR.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Transferability for General Reasoning: An Automated Curriculum for Multi-Domain RLVR

Reference 59

Resolution
verified exact
local_arxiv, observed 2026-08-06T04:30:48.399971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T04:30:44.909089Z digest=sha256:e8a6fbc1bf6cc351ab42b6419f307c03d38b6a0fe2b3b5fd125bc8083f50f57f

Observation 4f045811-840f-41c5-a0ce-eebebd0315f5 · outbound

This paper cites LongProc: Benchmarking long-context language models on long procedural generation.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning LongProc: Benchmarking long-context language models on long procedural generation

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:44.984844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:44.984844Z digest=sha256:6e234b3aae407a4cd8bb79f2d205fa33446a4dafa00794c332a8fc09fa27ca07

Observation 8eb82e20-7176-43d1-a1af-6dfee17fb294 · outbound

This paper cites Skill-Mix: a Flexible and Expandable Family of Evaluations for AI models.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Skill-Mix: a Flexible and Expandable Family of Evaluations for AI models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:45.049452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:45.049452Z digest=sha256:fe587ac2cf12e6b2e68724893fc3c362cd9707a625ad1a53226b296161b34be9

Observation ba4c8a55-fae2-4392-a94e-b9b9559cf248 · outbound

This paper cites From𝑓(𝑥) and 𝑔(𝑥) to 𝑓(𝑔(𝑥)) : LLMs learn new skills in RL by composing old ones.arXiv preprint arXiv:2509.25123, 2025.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning From𝑓(𝑥) and 𝑔(𝑥) to 𝑓(𝑔(𝑥)) : LLMs learn new skills in RL by composing old ones.arXiv preprint arXiv:2509.25123, 2025

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:45.118337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:45.118337Z digest=sha256:bf7fc4d21dc5354c692637532407435f84ee30cf62c48802b74f290dba1b6c0a

Observation 4f441f9d-74ea-40fd-baaf-85ae5922310c · outbound

This paper cites Skill-aware data selection and fine-tuning for data-efficient reasoning distillation.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Skill-aware data selection and fine-tuning for data-efficient reasoning distillation

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:45.193661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:45.193661Z digest=sha256:4be1db11e0e3ad481350900c27f82eae72e7fc3be5e7e1ad939ebed50fcb8e56

Observation 7ecd2ff6-3be4-4195-b060-beed5ce26159 · outbound

This paper cites Skill-awaredataselectionandfine-tuning for data-efficient reasoning distillation, 2026.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Skill-awaredataselectionandfine-tuning for data-efficient reasoning distillation, 2026

Reference 64

Resolution
verified exact
raw_fallback, observed 2026-08-06T04:30:48.157980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T04:30:45.286495Z digest=sha256:6b8ab97616cb66399d44c2b704498c0a7a7e22a212e038390f250cb60dde6f94

Observation 9b43e4de-0cd7-482c-9948-122df95a7a07 · outbound

This paper cites Lee, Chenlei Leng, and Fanghui Liu.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Lee, Chenlei Leng, and Fanghui Liu

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:45.384534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:45.384534Z digest=sha256:e095a7e1a0cac1abe623de3634df79b784c04cc27f0a2f2f3e5898e1314d7854

Observation 0506bd30-63bc-448a-84f7-bf8d05965137 · outbound

This paper cites RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:45.506154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:45.506154Z digest=sha256:8865e9b8673b2e88094b21150c56410cf1f9e4594016e441b542ee1962bd4079

Observation 985fb27d-d840-456c-bdf2-e1b7307a67ec · outbound

This paper cites Can Models Learn Skill Composition from Examples?.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Can Models Learn Skill Composition from Examples?

Reference 67

Resolution
verified exact
local_arxiv, observed 2026-08-06T04:30:47.894990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T04:30:45.607770Z digest=sha256:69ccfd84826ab199f453c2f3fefe0540ad4fde7bd24fbbd76870ee5d2ee07258

Observation 51c4fc5c-8ba0-4182-8d2b-6574abb25afd · outbound

This paper cites A Survey of Process Reward Models: From Outcome Signals to Process Supervisions for Large Language Models.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning A Survey of Process Reward Models: From Outcome Signals to Process Supervisions for Large Language Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:45.669787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:45.669787Z digest=sha256:cdf16f88ba1d54d2efd4d1bd62d431d9347284cde3a62ede51793ee177e77553

Observation c0bce8bd-4215-492e-b0f8-185d3f0d0725 · outbound

This paper cites NATURAL PLAN: Benchmarking LLMs on Natural Language Planning.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning NATURAL PLAN: Benchmarking LLMs on Natural Language Planning

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:45.725769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:45.725769Z digest=sha256:69e6d515de1ae1f41ee93e6765cbd10545f7adb412a553164ffccd483d604437

Observation 3be68408-0eae-4f51-8cee-15b8fb01407e · outbound

This paper cites SkillRouter: Skill Routing for LLM Agents at Scale.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning SkillRouter: Skill Routing for LLM Agents at Scale

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:45.785107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:45.785107Z digest=sha256:036dca135ae88968bc881c55bbefffef4304c89bb9f77603f22946156cfa206b

Observation f26bf2fa-4768-4abe-a813-b75e88bafcd9 · outbound

This paper cites SkillLearnBench: Benchmarking Continual Learning Methods for Agent Skill Generation on Real-World Tasks.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning SkillLearnBench: Benchmarking Continual Learning Methods for Agent Skill Generation on Real-World Tasks

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:45.842422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:45.842422Z digest=sha256:42f1606cf87392028ac04a18ddc5be670366f2938453d2e54ead95adbb871d8c

Observation 63be7e85-8388-469b-a518-0c586be89495 · outbound

This paper cites It clearly outlines the key factors and their interrelationships.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning It clearly outlines the key factors and their interrelationships

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:30:49.927126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T04:30:45.896367Z digest=sha256:654799bd5777717523b8759f7d8574ba12eeab8f30317c7090ca72f46645693d

Observation d34e5cd0-02c8-4f10-990b-62371eb35b0e · outbound

This paper cites Specific examples from the data are used to support the narrative.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Specific examples from the data are used to support the narrative

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:30:49.918035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T04:30:45.957925Z digest=sha256:e4e35aeb0d2f5004964cdab108590488e77c36abb3271cd8c16f2b1b89a85a02

Observation 72849c37-fbd7-4519-87c9-f82842a368fb · outbound

This paper cites It captures the reader’s interest and effectively conveys the potential crisis.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning It captures the reader’s interest and effectively conveys the potential crisis

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:30:49.909058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T04:30:46.040304Z digest=sha256:80bdfd8873d56417c356adc5d7db22ab46888d183daad0aef9650c96803769e7

Observation 46ee3a08-804c-4b2e-a87f-28fe61b52bd8 · outbound

This paper cites It uses the data and insights from the previous steps to construct a plausible and coherent narrative.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning It uses the data and insights from the previous steps to construct a plausible and coherent narrative

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:30:49.900061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T04:30:46.110513Z digest=sha256:ce9e46b73688681871d710729a699b6c42a915863cd7430a65a339a0545de705

Observation 7d98e6aa-0390-4d6f-b55d-63eab9cd9a22 · outbound

This paper cites It should reflect the brand’s commitment to sustainability.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning It should reflect the brand’s commitment to sustainability

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:30:49.890925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T04:30:46.212338Z digest=sha256:7a7566885d17d0df0a1426c9c8f915e4267ec6c6f147fb53a2901d97c67a6d9d

Observation 862d3e81-c652-41de-a436-10c783881f85 · outbound

This paper cites It should not exceed 10 words.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning It should not exceed 10 words

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:30:49.881322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T04:30:46.277290Z digest=sha256:341d4ce6aa1d3a6e9b6b0b70a53d7f2a164b4a391ea6ac443636429a8210b295

Observation 66e12a0b-a7d2-4dd0-bfe5-a948738e9651 · outbound

This paper cites an unresolved cited work.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Unresolved cited work

Reference 78

Resolution
unresolved
raw_fallback, observed 2026-08-06T04:30:49.872486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T04:30:46.369607Z digest=sha256:ed9d1a8ab38125e1764c7d8c5aef97927b45569a182dd39df5422e1b514cbffe

Observation 8fa281e4-6a35-4557-93b4-bf1458f388eb · outbound

This paper cites no-interference.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning no-interference

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:30:49.863445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T04:30:46.432911Z digest=sha256:9e32ad384d93c5d3317fa42d4fa83914bb82e674719564db6c71a148222ff99a

Observation 505dddaa-bfa2-4486-a3a7-529a6a92c816 · outbound

This paper cites DON’T CHANGE THE ANSWER, CORE LOGIC, OR THE SKILL REQUIRED.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning DON’T CHANGE THE ANSWER, CORE LOGIC, OR THE SKILL REQUIRED

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:30:49.854613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T04:30:46.473121Z digest=sha256:a807405ff378797298128ae3738cd6fe194cc8fc4e6effb4007b3e21064cf68f

Observation aa927f39-65d7-40ad-a2d1-8e1547e1ae1b · outbound

This paper cites an unresolved cited work.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Unresolved cited work

Reference 81

Resolution
unresolved
raw_fallback, observed 2026-08-06T04:30:49.846320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T04:30:46.578880Z digest=sha256:ac46b319324ba005d057a3324b2b339cf933f092cba9e19f35de3102b734c025

Observation 97a2f555-a3a9-47c6-954f-b3375f1ef442 · outbound

This paper cites Conference trip on constraints.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Conference trip on constraints

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:30:49.837044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T04:30:46.749688Z digest=sha256:a9a625c0a3a5e22f11b63a7f573fc9fd12cd3788d6f5a418dcb31a12263f1775

Observation 6c420ad5-44a8-4954-8aef-9be629177b9d · outbound

This paper cites 29 Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning 2.Scenario consistent.All rewritten steps plausibly belong to the samescenario.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning 29 Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning 2.Scenario consistent.All rewritten steps plausibly belong to the samescenario

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:30:49.826925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T04:30:46.966544Z digest=sha256:09bb9926e571b45dfdb4b8c87d8f6c34186be97428829350028b8f69600887ea

Observation d60813c5-b66b-44ed-8dd6-4bcb8aca68b5 · outbound

This paper cites The core logic, numerical values, and (for multiple-choice) the option letters and contents must be unchanged.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning The core logic, numerical values, and (for multiple-choice) the option letters and contents must be unchanged

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:30:49.817051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T04:30:47.080943Z digest=sha256:d46a24368b69577644d5c9a67093cc247d69e76181f863816bd26ee544fdfcbe

Observation b72f18a9-bb5b-4b8c-88c0-516cb43f235d · outbound

This paper cites FAIL if any step references entities, settings, or framings that contradict the scenario or that read as an unrelated problem pasted in.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning FAIL if any step references entities, settings, or framings that contradict the scenario or that read as an unrelated problem pasted in

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:30:49.806753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T04:30:47.257228Z digest=sha256:173bfcba97c83e7c82dbf0057ee4f02f18e3eb2a24b14c51ab86f30414af63b4

Observation a66d5167-c914-4613-9b2f-da05bec9d10b · outbound

This paper cites using the value from the previous step.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning using the value from the previous step

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:30:49.796334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T04:30:47.390521Z digest=sha256:9e6f313efb88c9aa3d3711802433dbe9ad0a4a0616b89a705811229d4b3594f0

Observation 37ee22a3-c1f0-4ec2-9688-1ceb5fef7fab · outbound

This paper cites MAE” is the mean absolute error between the LLM and mean human score in[0, 1]. “Binary agree.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning MAE” is the mean absolute error between the LLM and mean human score in[0, 1]. “Binary agree

Reference 87

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T04:30:49.786033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T04:30:47.578235Z digest=sha256:4aa0de177360b863de32c6ed2463eb8c41d9c59d0ba670b57a971b31ec2e88e6

Pith citing papers

No inbound Pith citation observations are available.