Pith. sign in

Paper Citation Record · LEDGER

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback

As of 7 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 5 inbound Pith citation observations for arXiv:2507.00699.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.00699 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:16:42.794738Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-01T08:17:10.481202Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

46 of 46 outbound references displayed

  • verified exact3
  • verified fuzzy7
  • unresolved36
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 98959035-344d-4dd5-b524-8cc2a7ddf776 · outbound

This paper cites Soen-101: Code generation by emulating software process models using large language model agents,.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback Soen-101: Code generation by emulating software process models using large language model agents,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:40.649612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:40.649612Z digest=sha256:9c992a2dc72f95e1e47d10453f4244c5d0a3b9beda3922e652b27f8c44a9e184

Observation d930fc1b-9079-4376-9294-e8d94752079b · outbound

This paper cites ROCODE: Integrating Backtracking Mechanism and Program Analysis in Large Language Models for Code Generation.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback ROCODE: Integrating Backtracking Mechanism and Program Analysis in Large Language Models for Code Generation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:40.840152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:40.840152Z digest=sha256:7503eddc3746efbf632700dd450b5c4473b965cc972814d69979a0915e19a6fb

Observation 815bdd69-f192-49d1-a45f-c0730df2b1fd · outbound

This paper cites Skcoder: A sketch-based approach for automatic code generation,.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback Skcoder: A sketch-based approach for automatic code generation,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:40.912600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:40.912600Z digest=sha256:3f24548dd1cc1bc08bd3902a606a37ca6062931e6cbc30f557854012e757565d

Observation 05f51ad5-7b73-48d2-af70-60563fdb60ed · outbound

This paper cites Enhancing Code Generation via Bidirectional Comment-Level Mutual Grounding.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback Enhancing Code Generation via Bidirectional Comment-Level Mutual Grounding

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:40.926501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:40.926501Z digest=sha256:3260900bb62a40fdee3d146a4c29b98eb3066ed7de8146e0018d26e04ee74c10

Observation e0d3c8c6-b55a-49ef-ad5e-effa4a22987b · outbound

This paper cites Fixing Large Language Models' Specification Misunderstanding for Better Code Generation.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback Fixing Large Language Models' Specification Misunderstanding for Better Code Generation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:41.068895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:41.068895Z digest=sha256:39cfd41f65fa423137ad9e0a539a08fe4bcbdb480c7ba28410f54503cc2ef1d2

Observation b61f1b99-39fe-4665-b59b-aa73c79502a7 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback Evaluating Large Language Models Trained on Code

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:41.283330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:41.283330Z digest=sha256:96af4d9d53096eda628d9a0fd297cc06124c3d2cbb798df15861e4ac6b4fcc55

Observation b47c095f-2342-44b3-9824-95727e3a39bc · outbound

This paper cites Program Synthesis with Large Language Models.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback Program Synthesis with Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:41.360449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:41.360449Z digest=sha256:30151f36da594a1d3401d57d1d18137c2dd56a885ca266139d2bcb4e4ae735f0

Observation 3bc9bdaa-f0df-4836-91cb-e64f7b267922 · outbound

This paper cites A Survey on Evaluating Large Language Models in Code Generation Tasks.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback A Survey on Evaluating Large Language Models in Code Generation Tasks

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:41.365515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:41.365515Z digest=sha256:008bfd4dd7a3d18cb2730793dd438f3858255e1a568c271f5a9c01cc3e3e8f9d

Observation 43c39c9e-b5aa-4659-bec4-0e0753fedd37 · outbound

This paper cites Instruction-following evaluation for large language models,.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback Instruction-following evaluation for large language models,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:41.453046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:41.453046Z digest=sha256:6b29c509b96c4bb397dd2d27366f9419a7fdf869d820beb0a8e44bc2e1be0be8

Observation da751ba6-c1d1-47df-985f-5a1f713dda1c · outbound

This paper cites InFoBench: Evaluating Instruction Following Ability in Large Language Models.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback InFoBench: Evaluating Instruction Following Ability in Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:41.489193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:41.489193Z digest=sha256:4c4e45ea29c71929c2fbee97bc4fb68107bb1476fb24041a55e4d6fc77417a59

Observation f1774027-b0c4-4f17-81a5-df9a12cc59c7 · outbound

This paper cites CodeIF: Benchmarking the Instruction-Following Capabilities of Large Language Models for Code Generation.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback CodeIF: Benchmarking the Instruction-Following Capabilities of Large Language Models for Code Generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:41.524744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:41.524744Z digest=sha256:37314bd816cd41a44396b88469c51102ba33cbc5a3ade0641498b9d5f8137cf1

Observation 1193c980-a560-4bd8-85e2-898adf58dbbe · outbound

This paper cites Codeif-bench: Evaluating instruction-following capabilities of large language models in interactive code generation,.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback Codeif-bench: Evaluating instruction-following capabilities of large language models in interactive code generation,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:41.559939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:41.559939Z digest=sha256:e9d0adc19fe3777df018ebedb27dbe3664010191441507ca292edf003ff62328

Observation 79701b5d-baf1-4604-9754-541bf883a75a · outbound

This paper cites A hierarchical and evolvable benchmark for fine-grained code instruction following with multi-turn feedback,.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback A hierarchical and evolvable benchmark for fine-grained code instruction following with multi-turn feedback,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:16:48.065482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:16:41.648086Z digest=sha256:6d9b50bbb4e84513bb1a856464d999e482207a501b33b074870f975ec4e4b5f5

Observation ea3e0791-063d-489c-a370-71143a52ca3b · outbound

This paper cites Beyond Functional Correctness: Investigating Coding Style Inconsistencies in Large Language Models.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback Beyond Functional Correctness: Investigating Coding Style Inconsistencies in Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:41.671896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:41.671896Z digest=sha256:5a376b32768df1911e7d39b331799fedf6e088f72e44f76da723827963cf8df9

Observation cea143ab-1230-4175-9893-9bb48f9b74f8 · outbound

This paper cites FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:41.694748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:41.694748Z digest=sha256:cd5e8494227de0470641b9bda749f9b47ea2e9071edcfe6db3cc82918ecc505b

Observation bb171d85-fe5a-43b1-837c-f0361ca7f398 · outbound

This paper cites Sharegpt,.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback Sharegpt,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:16:47.895647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:16:41.734747Z digest=sha256:e53eaf89c51aec789512d18c9f4611b080ca126b81c644e3a121f4a338c80e60

Observation f6eabef0-6a5c-4206-bbcc-a25731a628f8 · outbound

This paper cites Prompt-Based Cost-Effective Evaluation and Operation of ChatGPT as a Computer Programming Teaching Assistant.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback Prompt-Based Cost-Effective Evaluation and Operation of ChatGPT as a Computer Programming Teaching Assistant

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:16:45.463009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:16:41.784746Z digest=sha256:d197dfa62dea62dcd375970732ec96c5a2fb9f2bb94eda029b9628d17c703212

Observation 1e647956-e148-494d-8d08-dc3e467a846e · outbound

This paper cites tree-sitter/tree-sitter: v0.25.5,.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback tree-sitter/tree-sitter: v0.25.5,

Reference 22

Resolution
verified exact
doi, observed 2026-08-06T21:16:43.134792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:16:41.819239Z digest=sha256:e0d525f2dbc969d5e0ef992a3ae1effc153c1fedd216f8af44c107c64a261ce0

Observation a19eca15-2d4a-4f3a-9595-400190df6f7e · outbound

This paper cites GPT-4 Technical Report.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback GPT-4 Technical Report

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:41.874749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:41.874749Z digest=sha256:67f082a48e41145f1d617ffd20b93d8d7eee37d2ed44b85f9071c863177465e0

Observation e6c825af-d137-4486-9936-ea7107ce2302 · outbound

This paper cites ROUGE: A package for automatic evaluation of summaries,.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback ROUGE: A package for automatic evaluation of summaries,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:16:47.604770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:16:41.914739Z digest=sha256:848a49a46b3cc3287fbcc7231a768c0db27545cc7d074c6edfcb9176c01e8646

Observation 5a0c9c0b-bf7f-4272-9ae4-e7c2665ef132 · outbound

This paper cites GPT-4o System Card.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback GPT-4o System Card

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:41.937129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:41.937129Z digest=sha256:51d469812da2b22498503f5e390d22eb1f691cb73d4ea25de58e454f88f5f486

Observation 26e3e459-d30c-460a-a7ed-295320b0c912 · outbound

This paper cites Claude 3.7 sonnet and claude code,.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback Claude 3.7 sonnet and claude code,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:16:47.186490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:16:41.963659Z digest=sha256:6c1183c4be7f9d4e2b0bb4efca6f96b5a08b9aeb4da7581a1ade0d13942e8fda

Observation 84baed67-510b-43ca-b183-a206c5ce0c6f · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:41.994749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:41.994749Z digest=sha256:6fb927c6a544ec5ab915e07f37fc92086757e2d5a9802076745076d9f300dfc8

Observation 8abd2f57-2fd9-40f2-8480-f908485b89dd · outbound

This paper cites DeepSeek-V3 Technical Report.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback DeepSeek-V3 Technical Report

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:42.034748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:42.034748Z digest=sha256:01368a791f58409c648c1a644d79b95b601c9d23b804ee89d5a62de923a7a212

Observation 7a297632-837c-4105-8ad7-12f0e53f9b6d · outbound

This paper cites Qwen3 Technical Report.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback Qwen3 Technical Report

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:42.094905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:42.094905Z digest=sha256:0a13a0a3d12e0d45ebe37bf057d6bea12d176012115071945b4021eca92ede10

Observation 1927e71d-37ba-4f26-8e8d-3d54415b000e · outbound

This paper cites The Llama 3 Herd of Models.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback The Llama 3 Herd of Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:42.134745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:42.134745Z digest=sha256:7938c31f0370a706d63613a3d2365639a291d3d1b376c5386825b0ad6042cd2d

Observation 37067ce1-3533-4b5b-ac4f-c3abc6105d35 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention,.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback Efficient memory management for large language model serving with pagedattention,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:42.174750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:42.174750Z digest=sha256:6bf59a7f2097c6031590b0230622bdbdbde78dda7c06efc12d25cc16d02bc0c6

Observation 984bcaa4-5042-4ed1-8926-f2ea96675d1d · outbound

This paper cites Guesslang: A neural network to guess the programming language from code snippet,.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback Guesslang: A neural network to guess the programming language from code snippet,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:16:46.819549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:16:42.225023Z digest=sha256:10f4729b73e385d667e1f2b91366cbe8e5cefa108ff1b48afac2b735fa9a09b3

Observation 154ca5db-694d-4241-b4e1-f0df76995d2e · outbound

This paper cites DS-1000: A Natural and Reliable Benchmark for Data Science Code Generation.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback DS-1000: A Natural and Reliable Benchmark for Data Science Code Generation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:42.274834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:42.274834Z digest=sha256:9e150136b4c4f75f89edb2999f3d05bdb0977bf731d6dd75f3cbffd1a8a8a5ca

Observation a938fd60-4e67-407d-864f-5b65fe7c3f08 · outbound

This paper cites ClassEval: A Manually-Crafted Benchmark for Evaluating LLMs on Class-level Code Generation.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback ClassEval: A Manually-Crafted Benchmark for Evaluating LLMs on Class-level Code Generation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:42.308763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:42.308763Z digest=sha256:f4ac2a9eebf35f86f25c7a6a5d127fb08948f2b7f1d1c77b74603c651e9a7eeb

Observation 72cdc08a-3578-47b8-8024-0aced3a8923c · outbound

This paper cites Learning-based widget matching for migrating gui test cases,.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback Learning-based widget matching for migrating gui test cases,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:42.314114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:42.314114Z digest=sha256:e0c67b7211c65fc9cf6e1292f40688f1e14a24420f3a06cfcd098949a5dba795

Observation e6008267-b874-4237-a2fc-97efca23c204 · outbound

This paper cites DevEval: A Manually-Annotated Code Generation Benchmark Aligned with Real-World Code Repositories.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback DevEval: A Manually-Annotated Code Generation Benchmark Aligned with Real-World Code Repositories

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:42.331104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:42.331104Z digest=sha256:fdc4a23e4c2a3e052030054afcdb2e0b9105e9cb1e99fd87c4bdfdd2da4b32c2

Observation 459eb448-1417-48ee-bb16-61b7931bcb71 · outbound

This paper cites RustEvo^2: An Evolving Benchmark for API Evolution in LLM-based Rust Code Generation.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback RustEvo^2: An Evolving Benchmark for API Evolution in LLM-based Rust Code Generation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:42.334629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:42.334629Z digest=sha256:d537845a3176750c5cedbcdaf013c613958dd361003688051ea7d6e3b3087f8b

Observation b0b49b14-e3ac-4ed8-9abc-f6418ac8b6db · outbound

This paper cites Feedbackeval: A benchmark for evaluating large language models in feedback-driven code repair tasks,.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback Feedbackeval: A benchmark for evaluating large language models in feedback-driven code repair tasks,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:42.374745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:42.374745Z digest=sha256:bca45337911be9aceea415ba46c2bfe60ea7d6f8b08a580e004ebeadbb10c3fb

Observation a95a8c12-ebd9-42f2-bcf9-72e7ba49ac5e · outbound

This paper cites Benchmarking complex instruction-following with multiple constraints composition,.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback Benchmarking complex instruction-following with multiple constraints composition,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:16:46.644479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:16:42.424811Z digest=sha256:270a850158d45e913c13c037a6d04af33ffb6af5ea3d8f4aae91f0eeef30f40c

Observation 472da2bc-29fc-4846-9f68-71cf6927d5ff · outbound

This paper cites Generating Equivalent Representations of Code By A Self-Reflection Approach.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback Generating Equivalent Representations of Code By A Self-Reflection Approach

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:16:44.004826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:16:42.503377Z digest=sha256:cdb96d215510b91e17d0445b104ee79519c9e25cba090f3575da76f781176051

Observation 674a015e-81df-4cc0-8811-24c3c0e410d4 · outbound

This paper cites Codescore: Evaluating code generation by learning code execution,.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback Codescore: Evaluating code generation by learning code execution,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:42.546851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:42.546851Z digest=sha256:7d81733bf4011939bbdd0ed68909664825d37f19648a256b2409d347d3b1311e

Observation d0639506-78c1-4147-a23d-9a7d2faeb7e1 · outbound

This paper cites Benchmarking Complex Instruction-Following with Multiple Constraints Composition.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback Benchmarking Complex Instruction-Following with Multiple Constraints Composition

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:42.464751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:42.464751Z digest=sha256:1a6b2c00188edc680d8cd893bdb55cfdcb2aee753e87d369b411f2f7e4aa7774

Observation 40aa013e-3916-4b50-84d9-62400e7e56f0 · outbound

This paper cites WizardLM: Empowering large pre-trained language models to follow complex instructions.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback WizardLM: Empowering large pre-trained language models to follow complex instructions

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:42.626496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:42.626496Z digest=sha256:198bf6c954dcd831c4a594ed58baf187a53515efca852c2f8019f28eba9d4636

Observation ca44739c-0b3f-4e91-aadd-61e8e4bc0971 · outbound

This paper cites Magicoder: Empowering Code Generation with OSS-Instruct.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback Magicoder: Empowering Code Generation with OSS-Instruct

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:42.654739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:42.654739Z digest=sha256:c76204cc2cb7c29b0d31d69c113f914e80c3c3aadb859977f4fe33b1a69dd002

Observation 7d317827-9bbe-47f6-a99c-de1dfb3fb9b5 · outbound

This paper cites Self-Instruct: Aligning Language Models with Self-Generated Instructions.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback Self-Instruct: Aligning Language Models with Self-Generated Instructions

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:42.591153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:42.591153Z digest=sha256:8bedcb038ecabac4aa1b12b14253fbd564d747fd390cb14a15a4b4505f836f0c

Observation 57834f6e-a9ec-4a0c-8657-fbc15a330194 · outbound

This paper cites Genetic instruct: Scaling up synthetic generation of coding instructions for large language models,.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback Genetic instruct: Scaling up synthetic generation of coding instructions for large language models,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:16:46.608394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:16:42.750003Z digest=sha256:9151c080a186b4d989063e1e724e6593e1591d7d5f11a020361c4253161fd495

Observation 9b693b4d-0478-4898-8651-a21d4f48dba9 · outbound

This paper cites WaveCoder: Widespread And Versatile Enhancement For Code Large Language Models By Instruction Tuning.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback WaveCoder: Widespread And Versatile Enhancement For Code Large Language Models By Instruction Tuning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:42.694742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:42.694742Z digest=sha256:f9811fd4b17c07ecdd7fb73594cf085e9a637b563be75ce938017d5599ee2f4e

Observation 54853ea3-73e6-4c99-80e1-4982df7be0bd · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback Instruction-Following Evaluation for Large Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:41.484523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:41.484523Z digest=sha256:7e913bebfbd20d1c1882113c4d4d5c3809d7ab7af4d251bb03dc847bed9b1fb2

Observation 918c7b8c-6c40-4878-8711-7b74779eb127 · outbound

This paper cites SOEN-101: Code Generation by Emulating Software Process Models Using Large Language Model Agents.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback SOEN-101: Code Generation by Emulating Software Process Models Using Large Language Model Agents

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:40.724767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:40.724767Z digest=sha256:86ddafe6b296b6ea638301247dbfc1903c5a0da971dc9ad420a25ebc26beb013

Observation eabb4b14-24c4-49af-a418-bedcec7e29c1 · outbound

This paper cites Genetic Instruct: Scaling up Synthetic Generation of Coding Instructions for Large Language Models.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback Genetic Instruct: Scaling up Synthetic Generation of Coding Instructions for Large Language Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:42.794738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:42.794738Z digest=sha256:addcda9ce65f4d3849156db5bf0b491f9fce9d9d42488d84911672146677d9ba

Pith citing papers

Observation 39c86638-d467-4dad-9eba-bfbb7c127715 · inbound

Bridging Generation and Training: A Systematic Review of Quality Issues in LLMs for Code cites this paper.

Bridging Generation and Training: A Systematic Review of Quality Issues in LLMs for Code A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:21:11.028874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T17:37:51.790000Z digest=sha256:eda0438bebd84e1016e736676fd1f17b428b5f36bd5f2b862d4644ccc0c31ac7

Observation 290c59ff-805d-4373-ba60-59b11e8a9b3e · inbound

When Should Models Change Their Minds? Contextual Belief Management in Large Language Models cites this paper.

When Should Models Change Their Minds? Contextual Belief Management in Large Language Models A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:23:12.976976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-29T07:17:07.720589Z digest=sha256:04d6cc1d263e4098a14aebd8fbb5cc5bdcb59b16e5a17d673adbb07a4efff722

Observation 57e95625-d60e-4a7b-a963-890c4f78eb3d · inbound

Prompt Governance? On Governing Technologies Governed by Natural Language cites this paper.

Prompt Governance? On Governing Technologies Governed by Natural Language A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:25:32.385646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-01T08:17:10.481202Z digest=sha256:8a65287b2b893aa307522f543f3d8bacce1e198e11ccd53475de9e7fe5f25478

Observation c805b868-143f-4aaf-89c6-f46868efae58 · inbound

CodeChat-Eval: Evaluating Large Language Models in Multi-Turn Code Refinement Dialogues cites this paper.

CodeChat-Eval: Evaluating Large Language Models in Multi-Turn Code Refinement Dialogues A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-04T20:20:06.984330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-25T20:22:29.287534Z digest=sha256:7e5edf57a2c5d5558c88362168bb076e6b894c9141d640d8456205569293827e

Observation 9f50b92b-d06f-40a4-a0b3-2d24dcfef8ae · inbound

CodeChat-Eval: Evaluating Large Language Models in Multi-Turn Code Refinement Dialogues cites this paper.

CodeChat-Eval: Evaluating Large Language Models in Multi-Turn Code Refinement Dialogues A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:55:35.347673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-01T07:05:50.811872Z digest=sha256:d216bee3cdea685a3c7f8f3bc1fe0e93fd95c353f8ab6105d3fd33ac62bdb927