Pith. sign in

Paper Citation Record · LEDGER

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks

As of 7 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 2 inbound Pith citation observations for arXiv:2603.22744.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2603.22744 v2

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-13T20:08:17.066898Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T05:13:41.056434Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-19T17:47:42.051078Z

Reference resolution

42 of 42 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved42
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d7fbd42c-d7d3-45c9-8d95-1b1252e38d56 · outbound

This paper cites $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:502deb141c26fb68d21b0ea35d38b0c881480e68a48ac15a0c0a2e15c3582926

Observation 244ca19c-b5c2-4ca4-be3e-cca0d8e05ff7 · outbound

This paper cites an unresolved cited work.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:71b6534db653a1cacec33907cfaa183711fab40f485f384b3d50707e0d229a4d

Observation 36a45d68-d64d-43c8-9037-f44fe6a53ecf · outbound

This paper cites an unresolved cited work.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:9efa9271e564f8af46520deefa09fd93e8ac9a634197cb16fd542fe311bce762

Observation 0699bd2c-8021-4efb-8e8b-e014741fa273 · outbound

This paper cites an unresolved cited work.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:acb5fc575f63e59107c82a333d28d3af6650aab4b3d71755fda511b45cd2ba97

Observation 1ca76b21-4b78-47cd-b3b0-9599ad8d659b · outbound

This paper cites SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:91c53a6f6e39ba40e9d66521ceb83c15d2d6228708af30de20f7f378b1ceaec9

Observation dc06e516-a635-4405-88f6-b1309e324584 · outbound

This paper cites Laradji, Manuel Del Verme, Tom Marty, Léo Boisvert, Megh Thakkar, Quentin Cappart, David Vazquez, Nicolas Chapados, and Alexandre Lacoste.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks Laradji, Manuel Del Verme, Tom Marty, Léo Boisvert, Megh Thakkar, Quentin Cappart, David Vazquez, Nicolas Chapados, and Alexandre Lacoste

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:0d87ef23b8478f0d21c0cff16061f33443128d470e423a87561a679f8feb8742

Observation d0e58077-934e-43e5-a8c3-24cef8486d2c · outbound

This paper cites InInternational Conference on Machine Learning (ICML).

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks InInternational Conference on Machine Learning (ICML)

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:fa1eec91ebfbfd52fc53e8588d6d73c4d48d0053a9a61dc3ab4f4a5bd78fac3a

Observation 56f252b3-a376-451d-9f64-a5ade0397673 · outbound

This paper cites A Survey on LLM-as-a-Judge.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks A Survey on LLM-as-a-Judge

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:06da57366457569336b0cd284ca2f0b3879746b3891ced873bbebac485cc6e2c

Observation d01f52ba-03d7-43a8-844a-5c42cbaa0855 · outbound

This paper cites Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:07be6193bfc394dfff9a7566667eefa3a2fea9eead6829cda4d665016f880e56

Observation 029ad0de-9241-4498-b7e1-d7836ed857a5 · outbound

This paper cites an unresolved cited work.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:dd8952d63c53e92948eaa809aab3c9fe26f7bd99ba9bbc55624fed4fe3b0cd22

Observation 5b4cd6d3-dfc4-4a6e-9ff5-87e141f12f5d · outbound

This paper cites an unresolved cited work.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:4a28c5dfa4bdfb175ecb961e3fa25e066a32bf9786aba3f0374f4b7f58ac0535

Observation e3e38240-89fd-4922-a887-a005e3704b23 · outbound

This paper cites Richard Landis and Gary G.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks Richard Landis and Gary G

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:16caa5882c28e02a28a3054abb9aa21b60a2969295c2fc1fc6e2f74d25ab06e4

Observation b35ccd49-185e-4bc8-a492-0a209e7a8130 · outbound

This paper cites an unresolved cited work.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:3c0656977ede127f3db8df8a223ed46fce663ddbf247a707e3b51643f7578707

Observation 400d0255-1a79-470d-aa5f-c8a680cdb8c2 · outbound

This paper cites SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:bb5f05fa772789241bf57758224d3a5d13bd68ac902fb381b858da3a90b6e8b9

Observation 8037ce73-bc4b-4e36-bf6a-42b001d72878 · outbound

This paper cites an unresolved cited work.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:4270f76d515a14533d546a91e39b1f5bb1712ac5b301da3f8458b4b26aced4d3

Observation f6023b7d-4b66-403b-a774-1927ab8612f4 · outbound

This paper cites Fung, Chun Yuan, and Li Shen.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks Fung, Chun Yuan, and Li Shen

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:5dcb9a410781b7298dfebfe8be148f34d918cae46c05020da8fee209f7d604fa

Observation 61888c53-ed6d-4790-a372-2d85eaea0dd1 · outbound

This paper cites an unresolved cited work.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:58a3730a6fe5c235f6b2ee5889b02c1100bbbe5ef63291038c1bcd70c83ae7e1

Observation 8d368530-a062-478b-85e8-c2417287c8c1 · outbound

This paper cites an unresolved cited work.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:711d93920b34997656b4222f42a288a06d5945f3847acc73f75575fef7ed66fc

Observation a4458278-742e-4962-9600-156361e1d2ef · outbound

This paper cites Hendryx, Brad Kenstler, and Bing Liu.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks Hendryx, Brad Kenstler, and Bing Liu

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:8e52d3ffc9f9e265233ab2959db517d63fc348375b5605cb74a7e10072c13566

Observation 344f05a3-a7a4-4c44-8491-3b539f241cbb · outbound

This paper cites an unresolved cited work.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:577b7365eae567e92dbd93a1a85fa41d1089b196802e2943b8c63a57ba4fbc97

Observation 3b7189d9-5807-4d9e-aa2f-658142edfee5 · outbound

This paper cites Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:f781a3c36a8a69fad8bf3f412f40fb2ed99e96c15cae851b709a705b1c80ebdf

Observation ed7b3964-4703-4a33-bb40-09cca2a2615a · outbound

This paper cites an unresolved cited work.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:d7a38925ac219ddbc1940a8b14aa83b09aedf039bf83d79402192754385fed91

Observation f44f3ef9-fadc-4a67-a9a8-b92739647cc3 · outbound

This paper cites FullFront: Benchmarking MLLMs Across the Full Front-End Engineering Workflow.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks FullFront: Benchmarking MLLMs Across the Full Front-End Engineering Workflow

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:20785233de86345504e8614166e667909d81f20dac9854baf70708340ef0108a

Observation 155f65ea-abb3-4eec-bf17-4bce3d1a5f27 · outbound

This paper cites Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:ee732db3f6309b1ce72c0406ae9849caa8db4532e598587f6d80455943c77ad8

Observation 90cda441-fc70-486a-ae9e-1cfe9692fb85 · outbound

This paper cites an unresolved cited work.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:3161aea0f129c4b8c1497eb66b4ca8337773d33b987aa7d05669071559832d6b

Observation 08659bd8-a593-4980-8ef5-9d70334e8a5b · outbound

This paper cites an unresolved cited work.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:5b2766fbd838a4364ff1a1e608d4e40399e22180fddd33181bc168c32984832f

Observation 042d3528-4b50-4f38-aa71-cf3d5f648a77 · outbound

This paper cites FronTalk: Benchmarking Front-End Development as Conversational Code Generation with Multi-Modal Feedback.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks FronTalk: Benchmarking Front-End Development as Conversational Code Generation with Multi-Modal Feedback

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:69e862dd9098ddf1803aee5102a637309b968c90401f2680c2179fb113f4dd6e

Observation 611bd42f-a641-4568-baf4-bbd09117795f · outbound

This paper cites an unresolved cited work.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:d8e32aa51e803b591b3e2ecde44db9dc045bdb8ccb4d443de166f3a81718dccb

Observation 7c22019b-63a9-4eeb-a07c-a2d650fa6d9c · outbound

This paper cites an unresolved cited work.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:774f7a965d98808df67d8f0118d1676fb9167eb3d45e8c4172519e04cc0dddd9

Observation 59403e54-e626-4585-b7ff-e04ada0d2e2e · outbound

This paper cites $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:feadf40a66c11478bc6480c3eaa031f5e2c36763bbe3669b648b8a20035b5793

Observation bdfdcc78-df99-4818-ad2a-1b149c2c6aaa · outbound

This paper cites Xing, Hao Zhang, Joseph E.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks Xing, Hao Zhang, Joseph E

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:dfdfb132e8ba5fcfbef2a331042dc0e8fb3a86ba9507c533b79a0c1d49964c0b

Observation f937b0d4-aada-488f-9716-a1885ed80966 · outbound

This paper cites Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:5f573401d2236321da3ffcae6903d094ca75ec11a2156ebe2a8e15bfba346da5

Observation 71363743-5b87-42af-bc7c-e07ff116eb50 · outbound

This paper cites FrontendBench: A Benchmark for Evaluating LLMs on Front-End Development via Automatic Evaluation.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks FrontendBench: A Benchmark for Evaluating LLMs on Front-End Development via Automatic Evaluation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:541687d57f875043a09848e6cffcc884e905e2f951c390048077fde2e2895bbf

Observation be060a8a-8ea2-4446-81de-c3a877926280 · outbound

This paper cites Score Description 1 Content is largely irrelevant or incoherent; fails to address the chapter instruction.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks Score Description 1 Content is largely irrelevant or incoherent; fails to address the chapter instruction

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:184b69d4e68dfe4a523eb39de3030bedec77358ccba31e9df6dc91f299a19b1e

Observation 17c79079-b228-452b-ac59-6d3645994748 · outbound

This paper cites Score Description 1 Visuals are broken, missing, or unreadable; severe rendering artifacts.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks Score Description 1 Visuals are broken, missing, or unreadable; severe rendering artifacts

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:c906e4df1adebf443fedd8939000af53c162a26c974664612bbeeded69ad30a5

Observation 3e1b5513-247d-42dd-ad0a-71b48796cefb · outbound

This paper cites Score Description 1 No discernible teaching structure; concepts presented without context or progression.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks Score Description 1 No discernible teaching structure; concepts presented without context or progression

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:172da3bd3099a2a471f400638e6303f1b64fb14b6dabf227fdb7e6bde1d509b1

Observation b449af07-562b-4866-bf43-0da62021b8f8 · outbound

This paper cites Score Description 1 Severe desynchronization; narration and visuals are unrelated or offset by multiple seconds.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks Score Description 1 Severe desynchronization; narration and visuals are unrelated or offset by multiple seconds

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:fbdef4e651092f69e52d4ee3023dde9f4a58b438821bccb45b275c8e0f0b707b

Observation 2985bd7a-d448-4969-a971-f859e958f03e · outbound

This paper cites GRPO vs. PPO gradient flow.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks GRPO vs. PPO gradient flow

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:a5a7f53c6779ff860b0d4e4eafbcecab172955c9d7bafd142fb6b28d460c22b9

Observation e9118d6b-a471-48a6-bd43-c89eabda3870 · outbound

This paper cites an unresolved cited work.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:e21911478ce7ef2af0ba774785db9ecf8b713cd83bf9b487c789397d976e145a

Observation c7986176-e56c-400d-b8db-0ea885c2e752 · outbound

This paper cites an unresolved cited work.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:0c92bfe94bac0d207c04018415e28febc75496ff2eecc24ae12961ccca73418c

Observation 315e9659-0994-45b3-80ef-59c08094b135 · outbound

This paper cites Do NOT interpolate.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks Do NOT interpolate

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:99a50cb155fe588b00aac0a546fe031d302a5abae4a2b9f4a2328b267cb2b785

Observation bde56565-cc51-4488-9f89-b8cfcc61b9dc · outbound

This paper cites rubric_scores.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks rubric_scores

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:534015ba35efd4364ba4829037a771a172348ad03960700781133f9393e79495

Pith citing papers

Observation 9d53c220-e0b0-4420-9daa-00ecbf236258 · inbound

EnergyAgentBench: Benchmarking LLM Agents on Live Energy Infrastructure Data cites this paper.

EnergyAgentBench: Benchmarking LLM Agents on Live Energy Infrastructure Data LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-06-01T02:02:30.857220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T17:43:11.038874Z digest=sha256:c9056a2ea7fc52d252f8ba46452d34c23d3b7cb8ecc94c7d89603519ee0fdcb4

Observation 3e3dde50-6ac1-476b-867f-2ca27d05bee4 · inbound

EnergyAgentBench: Benchmarking LLM Agents on Live Energy Infrastructure Data cites this paper.

EnergyAgentBench: Benchmarking LLM Agents on Live Energy Infrastructure Data LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T05:13:41.056434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:13:41.056434Z digest=sha256:59e8edc78fd0db0c6a321945700c0156ce43c306b049288aeaeb36206b920655