Pith. sign in

Paper Citation Record · LEDGER

HardTests: Synthesizing High-Quality Test Cases for LLM Coding

As of 8 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 8 inbound Pith citation observations for arXiv:2505.24098.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.24098 v1

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:39:59.087204Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T00:58:27.773874Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T20:00:08.065119Z

Reference resolution

35 of 35 outbound references displayed

  • verified exact0
  • verified fuzzy5
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d8a05b95-d4d3-4b7c-b4df-3ee993ddfcb5 · outbound

This paper cites 1 1 0\n1000000000\n1000000000.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding 1 1 0\n1000000000\n1000000000

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:40:00.263496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:39:59.087204Z digest=sha256:eb1e2a26760c96d586281b6b73858e9afcf1a487894a23cdca568e86c9edee7a

Observation e799839a-3f33-456d-93ae-86b8d820e280 · outbound

This paper cites 123 * The Python code block under each field should be independent.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding 123 * The Python code block under each field should be independent

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:40:00.825078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:39:58.827514Z digest=sha256:8de711bfc00ef58cb20659d421ff770b9b5eacb9a85ed637f6a0b16781bc3f5f

Observation 061501b8-bf53-43ab-a118-6973da65e000 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding Evaluating Large Language Models Trained on Code

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:55.616900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:55.616900Z digest=sha256:3f42a0a8f59f4b8ab2af50530192413ceff4337081b941eb8eb37671fe22fc10

Observation 1e236a05-f216-4964-b56f-286e65daac9d · outbound

This paper cites {n} { m}.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding {n} { m}

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:40:00.424892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:39:59.017463Z digest=sha256:83feba0cbd95675d0b0fe09234b4e010386d722af55eac7db02c4b0c3d4af555

Observation 2df2e97a-34d1-494c-8141-6e19422bf7fb · outbound

This paper cites Measuring Coding Challenge Competence With APPS.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding Measuring Coding Challenge Competence With APPS

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:55.896221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:55.896221Z digest=sha256:0e8840e2d6e7ac0393849b297dfa6d0d2a19ad49e15763a9c98a77d5ae77e376

Observation e776c6e1-13d8-4fcc-b270-a166d4077ee4 · outbound

This paper cites ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:56.068560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:56.068560Z digest=sha256:ca01890f7253634ee116ca999b546e8da87cc8fb1f2df1f7ac83275846ce0632

Observation 24efad3f-0eec-4e62-8e7b-b58fb412704c · outbound

This paper cites TestGenEval: A Real World Unit Test Generation and Test Completion Benchmark.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding TestGenEval: A Real World Unit Test Generation and Test Completion Benchmark

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:56.138659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:56.138659Z digest=sha256:403f22f487012f85c581845041ddfe83a8077b2bad8692d1417f1f43e0f0e4a5

Observation 82c7557f-7a31-4555-a1dc-c81aa76c3ac5 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:56.209649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:56.209649Z digest=sha256:25e31a519e5cca3a32dd020e0d5f3a8622bbd288a1e583ae8027998fbabc6757

Observation 20512dfd-441a-43d1-8ede-44926968ecd0 · outbound

This paper cites Competition-Level Code Generation with AlphaCode.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding Competition-Level Code Generation with AlphaCode

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:56.557357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:56.557357Z digest=sha256:89ca5e629ff75a4050bcac52fe369245222f93a98cf448104eb9a46347c5c2e1

Observation 6bf60f56-81a0-44af-8315-8fc8bfe65d73 · outbound

This paper cites Scattered Forest Search: Smarter Code Space Exploration with LLMs.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding Scattered Forest Search: Smarter Code Space Exploration with LLMs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:56.669431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:56.669431Z digest=sha256:4b09f87d2ad092eb3185f89d92758c1a24af52c246d890c846da0f7d82fe0827

Observation 87d06d85-d232-48cf-8b08-7e612b513762 · outbound

This paper cites Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:56.781297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:56.781297Z digest=sha256:8e5bb080b41dea31900ac57c9811513121e6505a937cac10b36a74787950e367

Observation 4fe2febf-5736-49ce-ae1f-f9eefc4f3563 · outbound

This paper cites rStar-Coder: Scaling Competitive Code Reasoning with a Large-Scale Verified Dataset.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding rStar-Coder: Scaling Competitive Code Reasoning with a Large-Scale Verified Dataset

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:56.905741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:56.905741Z digest=sha256:b6e9314b6da0c945597f94ed283c8aabcf48a8880f34683c8110e9855a3bd29d

Observation 202c5d0d-e4de-49bf-9e6e-a487bf6ee355 · outbound

This paper cites URL http://dx.doi.org/10.1145/3510454.3516829.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding URL http://dx.doi.org/10.1145/3510454.3516829

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:57.037173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:57.037173Z digest=sha256:76af85adf77cddd7506a1bcf42f8326b0bd77701020ce29b338df9ffaf445e86

Observation d89f2e36-8c6d-44cb-a848-bf6604ace09f · outbound

This paper cites OpenAI o1 System Card.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding OpenAI o1 System Card

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:57.197862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:57.197862Z digest=sha256:4cf28d95d41a40baba6b40c6466654448e3293fab1d320b6a8f202582e672f50

Observation 82330be4-dce1-4de6-8fb7-88075f7db129 · outbound

This paper cites Competitive Programming with Large Reasoning Models.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding Competitive Programming with Large Reasoning Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:57.365862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:57.365862Z digest=sha256:467dbcbd8a339960e3b7af55cfacbbe654d9aad5680fb8a98d4d3bfd2d534f44

Observation aa34b9b4-ffad-4aa3-83ca-aa3829bb7f6f · outbound

This paper cites COFFE: A Code Efficiency Benchmark for Code Generation.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding COFFE: A Code Efficiency Benchmark for Code Generation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:57.465250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:57.465250Z digest=sha256:12dcbc597846f2d1f34be2c7426a4699a24cd3871ae1060718df26e378f8f0f2

Observation ca57237a-acf1-443b-a1e0-ad4fabb214be · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:57.560048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:57.560048Z digest=sha256:9fda2caee1f170d11b81d072f822a77665a69583a7549c13c861af290c414229

Observation 0d04703e-ce4e-42be-a82d-05ff6f566335 · outbound

This paper cites Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:57.636759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:57.636759Z digest=sha256:e17d7b67dcf10188f4d942de8cb0d2e5cda21fbd20b4bd2d79e8ba4a718ad07d

Observation 9d92569e-dd0a-41c5-9707-ddcccb60ff5d · outbound

This paper cites TESTEVAL: Benchmarking Large Language Models for Test Case Generation.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding TESTEVAL: Benchmarking Large Language Models for Test Case Generation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:57.721681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:57.721681Z digest=sha256:ff360ef4958974757e9c0424ddccc8f75e00a362d908366c41a8198b5a4e3b06

Observation c61da6a1-b863-4077-b298-007329eaa1e2 · outbound

This paper cites KodCode: A Diverse, Challenging, and Verifiable Synthetic Dataset for Coding.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding KodCode: A Diverse, Challenging, and Verifiable Synthetic Dataset for Coding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:57.824234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:57.824234Z digest=sha256:dc26c18f396ec02aa5f24fc51a0404ee8b220fa456f08f2b277447608b6d82fa

Observation 8b65fe22-9ec2-4d7a-b4b4-46c5125f8fd9 · outbound

This paper cites LIMO: Less is More for Reasoning.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding LIMO: Less is More for Reasoning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:57.934398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:57.934398Z digest=sha256:3816f64d661a784f466184904d189be4af965e4c01dea33768a0f9901268c0ee

Observation 938c11e9-cd80-453e-88aa-75b304857657 · outbound

This paper cites No More Manual Tests? Evaluating and Improving ChatGPT for Unit Test Generation.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding No More Manual Tests? Evaluating and Improving ChatGPT for Unit Test Generation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:58.040594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:58.040594Z digest=sha256:789a09a87fb8f949a88c9c5816c0ca689a43fc1a2ba5dffa3cbae4ae4ec04a7a

Observation ae0285f0-db1e-4107-888c-a8124f293355 · outbound

This paper cites SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:58.116873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:58.116873Z digest=sha256:663b173d5708de671b763ae69a45be7be9045cf51a87ec08852eff2e0dbe8ab5

Observation 7fe0e6ca-7484-47d1-a1e4-74c89c72c79e · outbound

This paper cites ALGO: Synthesizing Algorithmic Programs with LLM-Generated Oracle Verifiers.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding ALGO: Synthesizing Algorithmic Programs with LLM-Generated Oracle Verifiers

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:58.211974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:58.211974Z digest=sha256:c11a12f99e5d3f78f7bb78ed50a3a216a670f63662155451a5644b10b276a95d

Observation 05460c55-7a88-42d1-a9f5-469868613ccd · outbound

This paper cites TestBench: Evaluating Class-Level Test Case Generation Capability of Large Language Models.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding TestBench: Evaluating Class-Level Test Case Generation Capability of Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:58.370843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:58.370843Z digest=sha256:6ee73357bfefa230955b21fd2a4202324d0925203e6c5bd214e01b3a1d9923ce

Observation e74403fe-084f-422b-9b8d-5d4808eebf4f · outbound

This paper cites LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:58.472965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:58.472965Z digest=sha256:b84cc02f8597662212aac457ba6ae93bef041a9ef35db6cc067222f1e1b1d1fb

Observation 5a0d58dc-98f6-4d04-abb3-7da3faa8eed2 · outbound

This paper cites an unresolved cited work.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:40:01.099920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:39:58.578150Z digest=sha256:5ce46dc60bf42b3940cb331fb2756234feee86fd42cad2ea97a819eda0af693e

Observation dcfff285-cc90-4ffd-a491-3123e6f4879e · outbound

This paper cites core logic.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding core logic

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:40:00.965186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:39:58.714694Z digest=sha256:cebc2d2ce1aa1d6fdc9dccd023d2778aed48fccc049f41dd5a185470604cb59a

Observation d69afd2e-7914-4815-a77d-59fb5b820682 · outbound

This paper cites 3\n1 10\n2 8\n3 10.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding 3\n1 10\n2 8\n3 10

Reference 1000

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:40:00.611877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:39:58.905686Z digest=sha256:784103e47fc7da66143cf42ee48da93d7c63eff20074e8e91ae28837f08a204f

Observation c47beb24-a9e0-4f8b-8614-f0da8e6c0637 · outbound

This paper cites ISBN 9781450304436.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding ISBN 9781450304436

Reference 2011

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:55.691285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:55.691285Z digest=sha256:6c8f8b6ca3c3b733612571626bb73a6e5874acf4d38a633cebf11629e9571706

Observation 5115e784-7837-4ed2-986a-be32a29beb91 · outbound

This paper cites Program Synthesis with Large Language Models.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding Program Synthesis with Large Language Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:55.538922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:55.538922Z digest=sha256:21f1974b045ba10479ea4f11498d964b2f5aed69bce708c7ce7d86b6fb03be1c

Observation 5d72edb9-f7f9-44e4-8f82-d98ba46e6223 · outbound

This paper cites TACO: Topics in Algorithmic COde generation dataset.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding TACO: Topics in Algorithmic COde generation dataset

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:56.321057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:56.321057Z digest=sha256:514218eed488873e7728818547ecf31c044db42d96e9b2b21f906dd0abfb3788

Observation ba80164b-d1a3-4ad6-b274-f52d49fe5572 · outbound

This paper cites LIMR: Less is More for RL Scaling.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding LIMR: Less is More for RL Scaling

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:56.430228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:56.430228Z digest=sha256:d0497288f0836c0854462f4602fc8340a1a9950e4228e1a7ed1d7f06e1bad685

Observation f145b96d-c545-4261-bb99-07a16a9a238b · outbound

This paper cites TDD-Bench Verified: Can LLMs Generate Tests for Issues Before They Get Resolved?.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding TDD-Bench Verified: Can LLMs Generate Tests for Issues Before They Get Resolved?

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:55.474805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:55.474805Z digest=sha256:ecbe9f0ef9e72bbcda17aeabd5705892d200dbe01acbe91562eaab296962f058

Observation 51511022-9523-4e30-8b97-5c23997f1d84 · outbound

This paper cites OpenCodeReasoning: Advancing Data Distillation for Competitive Coding.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding OpenCodeReasoning: Advancing Data Distillation for Competitive Coding

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:55.391872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:55.391872Z digest=sha256:efc110603040c01f94f3ea10b78f71936b22b62122ec639a4728775f07e08a40

Pith citing papers

Observation 57da5526-4a26-4dbb-8eed-1f45a9385867 · inbound

Efficiency of turbulence cites this paper.

Efficiency of turbulence HardTests: Synthesizing High-Quality Test Cases for LLM Coding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T00:58:27.773874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:58:27.773874Z digest=sha256:fc7baaaf3e201e13215a52581bac09c161f85875f0b9a97e738b5bb3298f4c94

Observation bb1423bc-89a7-4034-949f-64dd43891641 · inbound

Navigating the Clutter: Waypoint-Based Bi-Level Planning for Multi-Robot Systems cites this paper.

Navigating the Clutter: Waypoint-Based Bi-Level Planning for Multi-Robot Systems HardTests: Synthesizing High-Quality Test Cases for LLM Coding

Reference 118

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:06:05.487820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-09T23:27:54.704794Z digest=sha256:fc88c39675e0f5f21c71a232dc426c26180eb5d51c29f061f8902fbdce4bbff6

Observation 02a07b45-6973-4a29-84d8-0db561655f4d · inbound

VeriContest: A Competitive-Programming Benchmark for Verifiable Code Generation cites this paper.

VeriContest: A Competitive-Programming Benchmark for Verifiable Code Generation HardTests: Synthesizing High-Quality Test Cases for LLM Coding

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:01:32.448442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T01:21:42.562823Z digest=sha256:2f30810c32f08bda60e59f124e4968b26933efdb759be229fe4e0aa4d4981150

Observation 7f72180a-fb8a-48f2-93ee-2b50f2e2cfc5 · inbound

FrontierSmith: Synthesizing Open-Ended Coding Problems at Scale cites this paper.

FrontierSmith: Synthesizing Open-Ended Coding Problems at Scale HardTests: Synthesizing High-Quality Test Cases for LLM Coding

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:03:29.053765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T02:02:25.597640Z digest=sha256:f8c6d1bfbe85fdfb648808907a96c211b3df1b457c78ae9bfd38519721ecd503

Observation c6f235e1-2126-485a-a8bb-23a8c8e84267 · inbound

CP-Agent: A Calibrated Risk-Controlled Agent for Feedback-Driven Competitive Programming cites this paper.

CP-Agent: A Calibrated Risk-Controlled Agent for Feedback-Driven Competitive Programming HardTests: Synthesizing High-Quality Test Cases for LLM Coding

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:24:40.201084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T13:19:22.541146Z digest=sha256:3762b2bb833d176039d6375b004da6b4eae2054dc3e0125f421fe4290f50357d

Observation 99f18f25-3205-49c3-8e8a-86caaa7b88ff · inbound

The Generalization Spectrum: A Chromatographic Approach to Evaluating Learning Algorithms cites this paper.

The Generalization Spectrum: A Chromatographic Approach to Evaluating Learning Algorithms HardTests: Synthesizing High-Quality Test Cases for LLM Coding

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-04T20:00:08.067026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-25T20:55:15.784610Z digest=sha256:07107cd66474e0ce6a777c05892101f7869b6dea9e5c9e2ba22b195eab78004f

Observation b618dcfd-225a-4794-85de-af1428477042 · inbound

The Generalization Spectrum: A Chromatographic Approach to Evaluating Learning Algorithms cites this paper.

The Generalization Spectrum: A Chromatographic Approach to Evaluating Learning Algorithms HardTests: Synthesizing High-Quality Test Cases for LLM Coding

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:09:50.596811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T05:29:21.598397Z digest=sha256:0a76aebdb202ac821a5b1d2d86229b07401bff2e64b73c220a5d9dd075750b98

Observation 0aa7b6c7-458a-4bf8-a924-98bbb9ae9c7b · inbound

Coding Agents as Test-Suite Auditors: Finding What Official Suites Miss While Approaching What They Catch cites this paper.

Coding Agents as Test-Suite Auditors: Finding What Official Suites Miss While Approaching What They Catch HardTests: Synthesizing High-Quality Test Cases for LLM Coding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T22:23:16.109674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:23:16.109674Z digest=sha256:e654b119e820e31505e71d5d5cc3bea620dd2d4a36768ea34d699e78eae33c99