Pith. sign in

Paper Citation Record · LEDGER

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents

As of 20 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 0 inbound Pith citation observations for arXiv:2608.11727.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.11727 v1

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:35:53.062182Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

50 of 50 outbound references displayed

  • verified exact0
  • verified fuzzy29
  • unresolved20
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation dd50a0ba-0227-47cf-b045-7c1c0072b35d · outbound

This paper cites Claude code, 2024.https://claude.com/claude-code.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents Claude code, 2024.https://claude.com/claude-code

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:35:53.920647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:35:52.851448Z digest=sha256:596cb6dbd02c25fa21431ad9e16f43ea032ebd1a9fff3267f86d922bcc34cb36

Observation 7557ce4c-1c0b-48be-a577-805828f92a99 · outbound

This paper cites Models overview.https://platform.claude.com/docs/claude/docs/models-overview, 2026.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents Models overview.https://platform.claude.com/docs/claude/docs/models-overview, 2026

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:35:53.906832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:35:52.856428Z digest=sha256:53b197d36163f9407c238578ef9a135beb028975bec72baabce4202f7569925c

Observation 560ffd42-11fb-40ea-b414-76b7949bbc90 · outbound

This paper cites Building effective agents.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents Building effective agents

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:35:53.892417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:35:52.860808Z digest=sha256:d79f38111b60c9f0c7893ebb76a346c4f1bf2e80cea05687df78cd2a8121b130

Observation f2924955-bd89-4804-b29b-550ed3a30566 · outbound

This paper cites $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T00:35:52.865427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:35:52.865427Z digest=sha256:6aad0c5ef12239c3b0a90dfcd6e1e4d677a3a6380ecb452b3d711ef03b60eb4b

Observation 73fe18b2-1629-4890-a28d-c9f2fc1df10e · outbound

This paper cites MLE-bench: Evaluating machine learning agents on machine learning engineering.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents MLE-bench: Evaluating machine learning agents on machine learning engineering

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:35:53.878606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:35:52.870160Z digest=sha256:b7fa01b86e2861547d3cb4f01c830c500974244d506bcb7006d58208c496a35d

Observation e4edd941-19be-441f-87f2-4f4d0ff8f748 · outbound

This paper cites SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T00:35:52.874574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:35:52.874574Z digest=sha256:cd6e8c03ce49c4efcb4fb2d7a6f526c7e45df20d2c51834b0c96f6e48d93f11e

Observation e6047447-9ca6-4feb-8f1d-075c78ce9727 · outbound

This paper cites Datasheets for datasets.Communications of the ACM, 64(12):86–92, 2021.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents Datasheets for datasets.Communications of the ACM, 64(12):86–92, 2021

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T00:35:52.879593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:35:52.879593Z digest=sha256:2c9fde50da5b973c7158171f058ddf90d789b3766de45737861ff1736f2d1a06

Observation 3b2aadf6-456c-45d5-9404-a5c67300fee5 · outbound

This paper cites Gemini 3.1 Pro: Model card.https://deepmind.google/models/model-cards/ gemini-3-1-pro/, 2026.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents Gemini 3.1 Pro: Model card.https://deepmind.google/models/model-cards/ gemini-3-1-pro/, 2026

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:35:53.856111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:35:52.883711Z digest=sha256:af90d0232357c5154e73d97ed517e555a47e8819877a817b977cb46d62af1de4

Observation 94310d55-504c-4172-8c4c-91f3f2f8d770 · outbound

This paper cites Model Context Protocol (MCP) Tool Descriptions Are Smelly! Towards Improving AI Agent Efficiency with Augmented MCP Tool Descriptions.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents Model Context Protocol (MCP) Tool Descriptions Are Smelly! Towards Improving AI Agent Efficiency with Augmented MCP Tool Descriptions

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T00:35:52.887791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:35:52.887791Z digest=sha256:1626e270c98b3053d0bd06e7e6863107f4dbfba5c45d7ab56d992f3de6f11169

Observation d262a897-5a56-4027-8122-d9875218a7ec · outbound

This paper cites ProcCtrlBench: Evaluating Process-Level Defects and Control Preservation in LLM Coding Agents.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents ProcCtrlBench: Evaluating Process-Level Defects and Control Preservation in LLM Coding Agents

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T00:35:52.892302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:35:52.892302Z digest=sha256:80951fd56813ee966f96f0f55f875a68787a75f0bafe3d812c8654de21e07f10

Observation 1ec3ce90-7e45-4177-ba2f-09dbbbe802f9 · outbound

This paper cites Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T00:35:52.896736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:35:52.896736Z digest=sha256:abdb0791bb77c4188c85967e00efabf78c71bf81792449cfef882a0ff0a5b3aa

Observation 32704800-6895-4df4-93cc-8bb392eb205f · outbound

This paper cites MLAgentBench: Evaluating language agents on machine learning experimentation.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents MLAgentBench: Evaluating language agents on machine learning experimentation

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:35:53.843216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:35:52.901665Z digest=sha256:abd88c6442c284a7430e31fd94f8b80c1b658b45e898610cb334ba9af82cb578

Observation 8f2da929-ae61-4916-8c45-5c4111efbbbe · outbound

This paper cites FollowBench: A multi-level fine-grained constraints following benchmark for large language models.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents FollowBench: A multi-level fine-grained constraints following benchmark for large language models

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:35:53.829881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:35:52.905780Z digest=sha256:c3bebbf2159220812f5039bb9818d3140c72620aeb05f80b8aeafc24956fbe98

Observation 6b06da15-8e42-4bf1-8756-504e09057943 · outbound

This paper cites Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:35:53.816783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:35:52.909830Z digest=sha256:28bed1d92691a1f45cb44ca31821903bc5f0fb3bebd8f2990434045c5213b619

Observation 28198d65-8b6a-4848-b2d1-a9e0f1108062 · outbound

This paper cites AgentBench: Evaluating LLMs as agents.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents AgentBench: Evaluating LLMs as agents

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:35:53.803265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:35:52.913841Z digest=sha256:fe54df2d50fd6851de467666af102bec8acb134e9188f03fbfca6f387a736793

Observation df2458e0-0b6a-4ac4-9ef9-0fa7cf30241d · outbound

This paper cites Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T00:35:52.917646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:35:52.917646Z digest=sha256:0ce800292c8c8fb54eb4ab852c3e7d0a9383f70c1f94f5dc9080bba8ea0ce59f

Observation efb42d1a-3422-4ce0-9941-37f27846325d · outbound

This paper cites GAIA: A benchmark for general AI assistants.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents GAIA: A benchmark for general AI assistants

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:35:53.789693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:35:52.922037Z digest=sha256:d9929ff0e11b52024a02ca97558b1332bb7f4c8ec2b7d8988803579d558fb5fd

Observation d3118d91-10d5-4cf4-b83f-706f44eddd52 · outbound

This paper cites MiniMax M2.7: Model self-improvement.https://www.minimax.io/models/text/m27, 2026.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents MiniMax M2.7: Model self-improvement.https://www.minimax.io/models/text/m27, 2026

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:35:53.775883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:35:52.925970Z digest=sha256:e26cdf544f77d3faa211fca18893d00c8644aaa79730ecbec509f28b2368e076

Observation e054397b-6d58-42a5-8ac3-d116d6fe3abc · outbound

This paper cites Kimi K2.6 model card.https://huggingface.co/moonshotai/Kimi-K2.6, 2026.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents Kimi K2.6 model card.https://huggingface.co/moonshotai/Kimi-K2.6, 2026

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:35:53.762751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:35:52.930082Z digest=sha256:c43c1bcdf827614e72ce8c4999cd862d9426e629337e8e78c5c4e807e95d3693

Observation 5ae1cba8-2279-44c7-b66a-0cd602440146 · outbound

This paper cites GPT-5.5 model.https://developers.openai.com/api/docs/models/gpt-5.5/, 2026.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents GPT-5.5 model.https://developers.openai.com/api/docs/models/gpt-5.5/, 2026

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:35:53.749856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:35:52.934313Z digest=sha256:940b8b7f19772d3ec9fe9574316c8f54dff9605cc55e4a2e6575402831d441ac

Observation 7a44892d-486a-4291-9b8c-752d88af48bd · outbound

This paper cites Introducing SWE-bench verified.https://openai.com/index/ introducing-swe-bench-verified/, 2024.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents Introducing SWE-bench verified.https://openai.com/index/ introducing-swe-bench-verified/, 2024

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:35:53.736649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:35:52.938862Z digest=sha256:936dcb365f4849cff27ad8d3f07b3e9de3e5b9b2e26a5347e888a8fc2dd5853e

Observation ff6098a0-7b02-4743-a8fd-44d32f79c60b · outbound

This paper cites Patil, Tianjun Zhang, Xin Wang, and Joseph E.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents Patil, Tianjun Zhang, Xin Wang, and Joseph E

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:35:53.722887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:35:52.942970Z digest=sha256:43db914a81f6661cfb3d3f27d1efb2aaa93e5829b1bbd660e03d360bdacfa404

Observation 767cbce2-2a76-4483-a3df-ecf87742d934 · outbound

This paper cites Patil, Huanzhi Mao, Fanjia Yan, Charlie Ji, Vivek Suresh, Ion Stoica, and Joseph E.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents Patil, Huanzhi Mao, Fanjia Yan, Charlie Ji, Vivek Suresh, Ion Stoica, and Joseph E

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:35:53.709443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:35:52.946957Z digest=sha256:66f3663db710c3200c1d34eeafd04722b198d8aac968abdb8cce222038fad060

Observation e7ed8de2-7101-4010-ad4c-8a6e5f775006 · outbound

This paper cites Generalizing verifiable instruction following.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents Generalizing verifiable instruction following

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:35:53.696155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:35:52.951036Z digest=sha256:fb0ca6b8e0c0d0eb9a4388ef4b22b37ad46ec82746d43c15a5de9b0ec620ae62

Observation 45da52c9-23f5-44ad-961a-597793a83213 · outbound

This paper cites AgentIF: Bench- marking instruction following of large language models in agentic scenarios.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents AgentIF: Bench- marking instruction following of large language models in agentic scenarios

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:35:53.682836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:35:52.955062Z digest=sha256:893160bfc92c31f46f44889a9de5ebc99c504e1dc7e6952ce2f66de2b4a2ce68

Observation 87e0b0e3-8571-4ad4-9c60-ec7efb7e25f3 · outbound

This paper cites InFoBench: Evaluating Instruction Following Ability in Large Language Models.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents InFoBench: Evaluating Instruction Following Ability in Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T00:35:52.959112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:35:52.959112Z digest=sha256:fa025a97de5372bbcaaa922e488ccde3ec4ed43f55b26aa558f0ff2b4ce4e9ad

Observation d094cc47-8480-4ff5-8985-7aeb92da1199 · outbound

This paper cites Qwen3.6-Max-Preview released.https://qwen.ai/blog?id=qwen3.6-max-preview, 2026.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents Qwen3.6-Max-Preview released.https://qwen.ai/blog?id=qwen3.6-max-preview, 2026

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:35:53.669629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:35:52.963336Z digest=sha256:e59127e11c3f22a3c5b0a115e3f0297d7bdf6017ecb3cc2e44d77abb7d1da89d

Observation dbbf2186-5bb2-4760-95e3-1dfd29bb93e2 · outbound

This paper cites Toolformer: Language models can teach themselves to use tools.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents Toolformer: Language models can teach themselves to use tools

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:35:53.655948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:35:52.967780Z digest=sha256:8a6127eb2a3ab702e59cae681e9c0e0054c25d5319296ffee70900f395de6615

Observation 35d5b2e3-80c5-44ac-9dc9-57b4ad1685d0 · outbound

This paper cites Seed2.0 model card.https://yfz.ai/Seed2.0_Model_Card.pdf, 2026.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents Seed2.0 model card.https://yfz.ai/Seed2.0_Model_Card.pdf, 2026

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:35:53.642324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:35:52.971632Z digest=sha256:c868b52b04e17c3b152c741beae59ecea6f0642123fc57a026c819e65a1fe4bd

Observation 90050bbd-78d8-4c3a-828c-624629e18d88 · outbound

This paper cites OpenAI GPT-5 System Card.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents OpenAI GPT-5 System Card

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T00:35:52.975569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:35:52.975569Z digest=sha256:2c9e95513099f9e87be3676592227b59ec8235b6964fa7c65f3deb1c50230863

Observation 82bd7606-6097-4408-aa42-c0d2a839f38e · outbound

This paper cites PaperBench: Evaluating AI's Ability to Replicate AI Research.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T00:35:52.979863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:35:52.979863Z digest=sha256:e3bb5641435c5d13019ad857abb028d52838305915094510b3b4c445c08889f1

Observation 4a86962c-805c-4b22-bedf-f6a827ea6503 · outbound

This paper cites Step 3.5 Flash: Open frontier-level intelligence with 11b active parameters.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents Step 3.5 Flash: Open frontier-level intelligence with 11b active parameters

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T00:35:52.984189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:35:52.984189Z digest=sha256:1e8a1d8d8777c3ebb64b255369f2a13d32c3d48c8b6d942684254a4e59bf0142

Observation 64808c17-4970-4704-b91e-663890d60418 · outbound

This paper cites Tencent unveils Hy3 preview.https://www.tencent.com/en-us/articles/2202320.html, 2026.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents Tencent unveils Hy3 preview.https://www.tencent.com/en-us/articles/2202320.html, 2026

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T00:35:52.988276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:35:52.988276Z digest=sha256:9ad51d1d89f915059cf95543d4a6408038502c005f10fae7efe5f204cc238fae

Observation 77423fd4-60a9-4daf-8854-e3c86d5d6246 · outbound

This paper cites AppWorld: A controllable world of apps and people for benchmarking interactive coding agents.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents AppWorld: A controllable world of apps and people for benchmarking interactive coding agents

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:35:53.629240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:35:52.992293Z digest=sha256:ba301f960a64d0094a3605fe3a4d81315c45d7b0064b3fa5d726f3a5cbc80cb8

Observation 2e12bf06-e39b-43d8-8e1e-5e0da8119158 · outbound

This paper cites The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T00:35:52.996449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:35:52.996449Z digest=sha256:611c760aded172bc57678171879f979478701e2e1f44bab3f885980075e57c56

Observation ee682b72-a39b-4e0c-9b19-2a99d4a285e7 · outbound

This paper cites CodeIF-Bench: Evaluat- ing instruction-following capabilities of large language models in interactive code generation.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents CodeIF-Bench: Evaluat- ing instruction-following capabilities of large language models in interactive code generation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T00:35:53.001057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:35:53.001057Z digest=sha256:4ff5b9edabbd80e741c7da558ae440ad7590ada88ba5ec1866cf398ffc0a60f6

Observation 45288d20-abf6-4f03-a50f-de7b520a4308 · outbound

This paper cites Benchmarking complex instruction- following with multiple constraints composition.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents Benchmarking complex instruction- following with multiple constraints composition

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:35:53.615614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:35:53.005172Z digest=sha256:8bd41656ee423c6900a8d52d3b9a6575120bb6373478b6e40c3e7d689b93a1f4

Observation aae303cf-a1b7-4cc0-bbe1-371988a61791 · outbound

This paper cites LIFBench: Evaluating the instruction following performance and stability of large language models in long- context scenarios.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents LIFBench: Evaluating the instruction following performance and stability of large language models in long- context scenarios

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:35:53.602044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:35:53.009595Z digest=sha256:dcdbef413c7bf716b1090d4f734a159e305a9f517a4414ef4251f55672ed4e13

Observation 037af9e7-2a24-4442-a52b-9270db7beb63 · outbound

This paper cites OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T00:35:53.013801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:35:53.013801Z digest=sha256:a9e46c23a93fccb4733bf158b70fd6f23661f620ef3d2343b4a7312bbed275be

Observation 87759daf-12d3-4382-96a8-f99e0c5820a9 · outbound

This paper cites SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T00:35:53.018025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:35:53.018025Z digest=sha256:7ffc7aad468b872738d0dc1d0bfb3c6f0594fd8d22f7438998607209d580b591

Observation 6592ac29-096c-42b7-8f8a-2e77a9c3a174 · outbound

This paper cites Jimenez, Alex L.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents Jimenez, Alex L

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:35:53.587390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:35:53.022196Z digest=sha256:2323eb167c2c8700ffc957b446fb64ae25e82331a94886e2fa105891ff0c8ad7

Observation 7b207577-7c42-4392-ac3c-2ae67eeabfd9 · outbound

This paper cites $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T00:35:53.026282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:35:53.026282Z digest=sha256:60107b31edf2e9e44cd1e8950d8d1ca20cbc4c9c4366680598a5a21596d52a46

Observation fca0b68c-e694-4d6d-a48d-84bda6132c8f · outbound

This paper cites OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T00:35:53.030478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:35:53.030478Z digest=sha256:d083328bfdab5a671aa08d73d6d2bea5b40136cef4e096da57a8610b41090385

Observation ccd56daf-e901-4d80-a080-3f9f47707b8f · outbound

This paper cites GLM-5.1 release notes.https://docs.z.ai/release-notes/new-released, 2026.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents GLM-5.1 release notes.https://docs.z.ai/release-notes/new-released, 2026

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:35:53.572368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:35:53.035260Z digest=sha256:120fcd697e1cf2e6a8d3b108b3ec3396f38d6a0e9f86a6ded38104d4c8a746c3

Observation fa3a73aa-147a-4fc1-b001-5baadf511a37 · outbound

This paper cites CFBench: A comprehensive constraints- following benchmark for LLMs.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents CFBench: A comprehensive constraints- following benchmark for LLMs

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:35:53.558499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:35:53.039245Z digest=sha256:a5195258afbfad6931bbfb1025758a6fb03da589c7c5ede861ae88910408fccc

Observation 44cd9b76-4a7b-4d10-8b90-bf3e3b219e0c · outbound

This paper cites IHEval: Evaluating language models on following the instruction hierarchy.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents IHEval: Evaluating language models on following the instruction hierarchy

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:35:53.545016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:35:53.044353Z digest=sha256:4e720967c1fb2688ab75257b0cffd2827be8fd9878262f78fa7867e4ffbe58e9

Observation 422a8ed1-6684-4f66-b8f2-120899eaada8 · outbound

This paper cites SpecBench: Measuring Reward Hacking in Long-Horizon Coding Agents.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents SpecBench: Measuring Reward Hacking in Long-Horizon Coding Agents

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-16T00:35:53.049011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:35:53.049011Z digest=sha256:3b852c8bcb34132981976c00bd1f61a95f9ca78bc8ffa4476ea778b0abc4e53b

Observation 53a884e0-5fd7-4236-8f10-b025d740cf6a · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents Instruction-Following Evaluation for Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-16T00:35:53.053616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:35:53.053616Z digest=sha256:cb352618eec7edd5699f66af0e42f4e5e71c26a743cbac73063043cf5555de08

Observation 34ebd60a-8d73-4711-ba6a-41722ec77425 · outbound

This paper cites Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:35:53.531867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:35:53.057988Z digest=sha256:ae6fbf9a6fd0b66855c9f1b7ae3d85b0e18cd5ec196749f24e6c561927e2fa23

Observation 0c81da32-8ed6-4150-904c-04ab39407e5a · outbound

This paper cites Keep generated summaries compact,.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents Keep generated summaries compact,

Reference 50

Resolution
malformed identifier
raw_fallback, observed 2026-08-16T00:35:53.518154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:35:53.062182Z digest=sha256:bb4879b032d979e76dacbf5e789b4c45e2f7c0a72ec1f5771013fa59fbf558e5

Pith citing papers

No inbound Pith citation observations are available.