Pith. sign in

Paper Citation Record · LEDGER

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks

As of 20 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 2 inbound Pith citation observations for arXiv:2603.22744.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2603.22744 v2

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-13T20:08:17.066898Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T05:13:41.056434Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-19T17:47:42.051078Z

Reference resolution

42 of 42 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved42
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d7fbd42c-d7d3-45c9-8d95-1b1252e38d56 · outbound

This paper cites $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:a6d5bf756ddf9bcb65cd94854c63602188db219f8e2c3354fb86ed11c2e6927e

Observation 244ca19c-b5c2-4ca4-be3e-cca0d8e05ff7 · outbound

This paper cites an unresolved cited work.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:b91441e7c0f6961b9919e89993ff18614e2d7a8f5c473a90ae4cfaa209225a8f

Observation 36a45d68-d64d-43c8-9037-f44fe6a53ecf · outbound

This paper cites an unresolved cited work.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:af185d10a1701acde2bccb43a415cc4a1f9a5bf4fb8c930ceb05fae4106dcbdc

Observation 0699bd2c-8021-4efb-8e8b-e014741fa273 · outbound

This paper cites an unresolved cited work.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:fd227e221376202cf45f2bf22b4b279ab9d8538a848996bc5c26add4ec63acb8

Observation 1ca76b21-4b78-47cd-b3b0-9599ad8d659b · outbound

This paper cites SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:b685fe5166fcca51f5844aa6cc807afa7e69476f6d566da728a94c617d59b00d

Observation dc06e516-a635-4405-88f6-b1309e324584 · outbound

This paper cites Laradji, Manuel Del Verme, Tom Marty, Léo Boisvert, Megh Thakkar, Quentin Cappart, David Vazquez, Nicolas Chapados, and Alexandre Lacoste.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks Laradji, Manuel Del Verme, Tom Marty, Léo Boisvert, Megh Thakkar, Quentin Cappart, David Vazquez, Nicolas Chapados, and Alexandre Lacoste

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:35dd40774e21fcb35101d4f5757d74d1c7c188e8e21e089eb737d51c8cb1b871

Observation d0e58077-934e-43e5-a8c3-24cef8486d2c · outbound

This paper cites InInternational Conference on Machine Learning (ICML).

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks InInternational Conference on Machine Learning (ICML)

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:5035fd482d965a5ab16096bc00c54ac9d50c4f34cb0515e41ac32c2e5838d050

Observation 56f252b3-a376-451d-9f64-a5ade0397673 · outbound

This paper cites A Survey on LLM-as-a-Judge.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks A Survey on LLM-as-a-Judge

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:3047274ac2d2d5b1c494ae2b2a67951a4a990e92f317e0cdb14fb4356a06ad85

Observation d01f52ba-03d7-43a8-844a-5c42cbaa0855 · outbound

This paper cites Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:02fddd6ec8e9dde898798bb4d2a45a66b13e7afaf3a318ea702a2a8ef61625cb

Observation 029ad0de-9241-4498-b7e1-d7836ed857a5 · outbound

This paper cites an unresolved cited work.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:a967a0103d9b792edb89923f8cbbba0237dea6a8d862f92b72492beee24f1588

Observation 5b4cd6d3-dfc4-4a6e-9ff5-87e141f12f5d · outbound

This paper cites an unresolved cited work.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:67e25a15c847118a7aa858f5d8e3a76d169326dd9b47b8b271c353d84f501cb3

Observation e3e38240-89fd-4922-a887-a005e3704b23 · outbound

This paper cites Richard Landis and Gary G.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks Richard Landis and Gary G

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:c98415e4a14b88f4e3582d603df0074db9794f96d12a5214895ddeb944e2e5f3

Observation b35ccd49-185e-4bc8-a492-0a209e7a8130 · outbound

This paper cites an unresolved cited work.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:f5d39023b5a6a88e48da47612e4c5858f6599530ff5f928b4452774808c5b694

Observation 400d0255-1a79-470d-aa5f-c8a680cdb8c2 · outbound

This paper cites SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:dcce07e4f2438072c1178bd8734766c4fb62e5a402141ef7df8dd7f2bebb6e35

Observation 8037ce73-bc4b-4e36-bf6a-42b001d72878 · outbound

This paper cites an unresolved cited work.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:57f82a7faed0cc118270de5688dc04ed9bcfa1a27dac1be32024ae2d7cd9bd15

Observation f6023b7d-4b66-403b-a774-1927ab8612f4 · outbound

This paper cites Fung, Chun Yuan, and Li Shen.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks Fung, Chun Yuan, and Li Shen

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:49a435f143b5708a0ab8f47981a2ce844bde2a9eda0587831e9f7632d229a31e

Observation 61888c53-ed6d-4790-a372-2d85eaea0dd1 · outbound

This paper cites an unresolved cited work.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:6d5aa6dbf12392c05f55860cd2b57b8852239f9a2296583609f126277c8edce5

Observation 8d368530-a062-478b-85e8-c2417287c8c1 · outbound

This paper cites an unresolved cited work.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:6905cf261489b0d9453c7f5abc55c1c3435a216e8fc30958ead1366971fb2787

Observation a4458278-742e-4962-9600-156361e1d2ef · outbound

This paper cites Hendryx, Brad Kenstler, and Bing Liu.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks Hendryx, Brad Kenstler, and Bing Liu

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:86e1318e8303f4062141ccb0272c2cabdc93806715c07f4a8d0824ff6da63836

Observation 344f05a3-a7a4-4c44-8491-3b539f241cbb · outbound

This paper cites an unresolved cited work.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:711228dab4fc56219bfb8f0f95ac4bf5e23fb74ec498cb35c08b31e9eab4ecc2

Observation 3b7189d9-5807-4d9e-aa2f-658142edfee5 · outbound

This paper cites Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:72fb0bf1a75ac33a4e431313fb73d90e3d536b6c40ba01bc354aa920a1eb52ee

Observation ed7b3964-4703-4a33-bb40-09cca2a2615a · outbound

This paper cites an unresolved cited work.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:371da73d7ed435dbce513e5e80b7304e7e2ffb2c502c593fe4d0801c609f578a

Observation f44f3ef9-fadc-4a67-a9a8-b92739647cc3 · outbound

This paper cites FullFront: Benchmarking MLLMs Across the Full Front-End Engineering Workflow.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks FullFront: Benchmarking MLLMs Across the Full Front-End Engineering Workflow

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:c352ccee9eed4fb436e8b6ae796c26aadc171a9df696418061bddd0a75f32509

Observation 155f65ea-abb3-4eec-bf17-4bce3d1a5f27 · outbound

This paper cites Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:2f78195efb2d804dfbeb43f485cb4ffd46bed736073507b416d73987ba2fb0fa

Observation 90cda441-fc70-486a-ae9e-1cfe9692fb85 · outbound

This paper cites an unresolved cited work.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:eec2600901402cd4716a20c7b359ad62ff1192d04958dbcafe3cec5340ad3290

Observation 08659bd8-a593-4980-8ef5-9d70334e8a5b · outbound

This paper cites an unresolved cited work.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:bda4ee6a9123d68958efcbb0266f0f1239970a4a45ff35c21b717d7681a2148b

Observation 042d3528-4b50-4f38-aa71-cf3d5f648a77 · outbound

This paper cites FronTalk: Benchmarking Front-End Development as Conversational Code Generation with Multi-Modal Feedback.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks FronTalk: Benchmarking Front-End Development as Conversational Code Generation with Multi-Modal Feedback

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:cb222b8908ffd3eb097f267bb53c263f77c94d81d111d3c4bf1551d66f1f7e92

Observation 611bd42f-a641-4568-baf4-bbd09117795f · outbound

This paper cites an unresolved cited work.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:603a9148478ddfaae030ed371c57b4b0c07c96f91d19bf37feb040d46ed89e46

Observation 7c22019b-63a9-4eeb-a07c-a2d650fa6d9c · outbound

This paper cites an unresolved cited work.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:d3763b5dd3c4954198d5801ed0eb9ebba64230ab3e1ee919fa49dc547e35fe82

Observation 59403e54-e626-4585-b7ff-e04ada0d2e2e · outbound

This paper cites $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:1ddffe3f8f3d9fd90b620d65396b5f1f983e3e162a80eaa4c5c575a8a58e7620

Observation bdfdcc78-df99-4818-ad2a-1b149c2c6aaa · outbound

This paper cites Xing, Hao Zhang, Joseph E.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks Xing, Hao Zhang, Joseph E

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:557a59a5ddc4bfdaead8d0f9407779f7f3b0da91c62710a252ab39cc97832d45

Observation f937b0d4-aada-488f-9716-a1885ed80966 · outbound

This paper cites Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:afe69c356d114eb2776b42547a7edb91586f32d6a9ec3a0459f0222dbe840bf8

Observation 71363743-5b87-42af-bc7c-e07ff116eb50 · outbound

This paper cites FrontendBench: A Benchmark for Evaluating LLMs on Front-End Development via Automatic Evaluation.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks FrontendBench: A Benchmark for Evaluating LLMs on Front-End Development via Automatic Evaluation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:81772cc9f64534e6f89402bcf45c151340767e4607f3d67b8c6df817d63633ef

Observation be060a8a-8ea2-4446-81de-c3a877926280 · outbound

This paper cites Score Description 1 Content is largely irrelevant or incoherent; fails to address the chapter instruction.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks Score Description 1 Content is largely irrelevant or incoherent; fails to address the chapter instruction

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:69664b2dc62966377295d15109e817af06d93dc4c59bbaec68194f2d59d87b5a

Observation 17c79079-b228-452b-ac59-6d3645994748 · outbound

This paper cites Score Description 1 Visuals are broken, missing, or unreadable; severe rendering artifacts.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks Score Description 1 Visuals are broken, missing, or unreadable; severe rendering artifacts

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:9408aa0e0ca800d27df9e910fd4c2e93c0c84ebfa3a149c8561d494b95017187

Observation 3e1b5513-247d-42dd-ad0a-71b48796cefb · outbound

This paper cites Score Description 1 No discernible teaching structure; concepts presented without context or progression.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks Score Description 1 No discernible teaching structure; concepts presented without context or progression

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:300be2d30a1c8b7728ed4a858dfd64980766d3c43ec1418d91da85497ae60a78

Observation b449af07-562b-4866-bf43-0da62021b8f8 · outbound

This paper cites Score Description 1 Severe desynchronization; narration and visuals are unrelated or offset by multiple seconds.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks Score Description 1 Severe desynchronization; narration and visuals are unrelated or offset by multiple seconds

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:bc309ee274c9415996fde34245185303bd42b10b2de6967398f44d2e1a8bbb02

Observation 2985bd7a-d448-4969-a971-f859e958f03e · outbound

This paper cites GRPO vs. PPO gradient flow.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks GRPO vs. PPO gradient flow

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:2fde981cf3e866425069538b87de238f59d4ce3bb2ca9b6cb6e8eeab5009201c

Observation e9118d6b-a471-48a6-bd43-c89eabda3870 · outbound

This paper cites an unresolved cited work.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:0f9945f785092fb30f3672dd58aeacc473ef7bacae11e6c8becfd681e800e31a

Observation c7986176-e56c-400d-b8db-0ea885c2e752 · outbound

This paper cites an unresolved cited work.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:8a8b820a1c786fa52c8be88834b206275bce86eba8a40134e7ddee0475564695

Observation 315e9659-0994-45b3-80ef-59c08094b135 · outbound

This paper cites Do NOT interpolate.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks Do NOT interpolate

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:16413d076950424c3a6d242212cb91223415241c17090b62cc4f2299e2b22595

Observation bde56565-cc51-4488-9f89-b8cfcc61b9dc · outbound

This paper cites rubric_scores.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks rubric_scores

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:586e1136860d91129f108df31ed7b1573c163b68136c37ae67620f7b1f8fdf34

Pith citing papers

Observation 9d53c220-e0b0-4420-9daa-00ecbf236258 · inbound

EnergyAgentBench: Benchmarking LLM Agents on Live Energy Infrastructure Data cites this paper.

EnergyAgentBench: Benchmarking LLM Agents on Live Energy Infrastructure Data LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-06-01T02:02:30.857220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T17:43:11.038874Z digest=sha256:9b6c356feb088a1ea72d0172b2947dbb8ce44bea0d7192b4ca06b6a42afe8305

Observation 3e3dde50-6ac1-476b-867f-2ca27d05bee4 · inbound

EnergyAgentBench: Benchmarking LLM Agents on Live Energy Infrastructure Data cites this paper.

EnergyAgentBench: Benchmarking LLM Agents on Live Energy Infrastructure Data LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T05:13:41.056434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:13:41.056434Z digest=sha256:a1352a2e05e1c8dd6dcfae663c201d6234f4c7fe7ad61b98651ffad83c0fd74c