Pith. sign in

Paper Citation Record · LEDGER

HardTests: Synthesizing High-Quality Test Cases for LLM Coding

As of 19 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 8 inbound Pith citation observations for arXiv:2505.24098.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.24098 v1

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:39:59.087204Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T00:58:27.773874Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T20:00:08.065119Z

Reference resolution

35 of 35 outbound references displayed

  • verified exact0
  • verified fuzzy5
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d8a05b95-d4d3-4b7c-b4df-3ee993ddfcb5 · outbound

This paper cites 1 1 0\n1000000000\n1000000000.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding 1 1 0\n1000000000\n1000000000

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:40:00.263496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T12:39:59.087204Z digest=sha256:808648f7551bb8377d2c62868c5fca89ffe12c4c3dce79084a713b5bb4ed231c

Observation e799839a-3f33-456d-93ae-86b8d820e280 · outbound

This paper cites 123 * The Python code block under each field should be independent.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding 123 * The Python code block under each field should be independent

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:40:00.825078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T12:39:58.827514Z digest=sha256:d9b37941e94afe3e18cd243625d0d8a16ae7ee0e6ba75fef7af351f3ea8231d1

Observation 061501b8-bf53-43ab-a118-6973da65e000 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding Evaluating Large Language Models Trained on Code

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:55.616900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:55.616900Z digest=sha256:ee8c2275954424f8730d0a88d99a642ba66abf43a52b742c2394dd71a7fba656

Observation 1e236a05-f216-4964-b56f-286e65daac9d · outbound

This paper cites {n} { m}.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding {n} { m}

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:40:00.424892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T12:39:59.017463Z digest=sha256:0ee4bb22c3e5d20fc55f3a77933f2081f706bc466613cca1833368e209363c9b

Observation 2df2e97a-34d1-494c-8141-6e19422bf7fb · outbound

This paper cites Measuring Coding Challenge Competence With APPS.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding Measuring Coding Challenge Competence With APPS

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:55.896221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:55.896221Z digest=sha256:d1f1491a0f19bc2d596b17270642ba0b9e289de99a00753e0253bd06ed5fa456

Observation e776c6e1-13d8-4fcc-b270-a166d4077ee4 · outbound

This paper cites ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:56.068560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:56.068560Z digest=sha256:92a92d14d9781a20e5bfdba6010e06899f68db36c00fe8a1c1d3169511bbe5bc

Observation 24efad3f-0eec-4e62-8e7b-b58fb412704c · outbound

This paper cites TestGenEval: A Real World Unit Test Generation and Test Completion Benchmark.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding TestGenEval: A Real World Unit Test Generation and Test Completion Benchmark

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:56.138659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:56.138659Z digest=sha256:88d789046a8a80391322b6812869c938e014338a72787483d719d4ea0211c67f

Observation 82c7557f-7a31-4555-a1dc-c81aa76c3ac5 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:56.209649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:56.209649Z digest=sha256:3123b77871f5458de8cce2d6a595e6537eb8a7848f9d7f946ba096b57d69cfc8

Observation 20512dfd-441a-43d1-8ede-44926968ecd0 · outbound

This paper cites Competition-Level Code Generation with AlphaCode.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding Competition-Level Code Generation with AlphaCode

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:56.557357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:56.557357Z digest=sha256:6cb5570738d535a1722b76cc2977a8fb94f12eb31c49a5059ce9beba22872623

Observation 6bf60f56-81a0-44af-8315-8fc8bfe65d73 · outbound

This paper cites Scattered Forest Search: Smarter Code Space Exploration with LLMs.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding Scattered Forest Search: Smarter Code Space Exploration with LLMs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:56.669431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:56.669431Z digest=sha256:8dacd16fbe9f11cff9c552f1dccca00078cf0305568d50810d559c983d919e2e

Observation 87d06d85-d232-48cf-8b08-7e612b513762 · outbound

This paper cites Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:56.781297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:56.781297Z digest=sha256:7a24f84f62232af0491b7095e85040affa83acb57f5c8f99f25452c186650935

Observation 4fe2febf-5736-49ce-ae1f-f9eefc4f3563 · outbound

This paper cites rStar-Coder: Scaling Competitive Code Reasoning with a Large-Scale Verified Dataset.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding rStar-Coder: Scaling Competitive Code Reasoning with a Large-Scale Verified Dataset

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:56.905741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:56.905741Z digest=sha256:7273ed41a179b3b42e8c8ff8c851e3628941402aadfdbf67b11d47edfb1980ee

Observation 202c5d0d-e4de-49bf-9e6e-a487bf6ee355 · outbound

This paper cites URL http://dx.doi.org/10.1145/3510454.3516829.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding URL http://dx.doi.org/10.1145/3510454.3516829

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:57.037173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:57.037173Z digest=sha256:839f4b691e5d3eac00b9074a8e792cb1e3a78251b948675a027102e36cde0c2d

Observation d89f2e36-8c6d-44cb-a848-bf6604ace09f · outbound

This paper cites OpenAI o1 System Card.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding OpenAI o1 System Card

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:57.197862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:57.197862Z digest=sha256:771fff027ae31731a8e462749cedc2e61a364409a3379a5a1a285fc38e948edc

Observation 82330be4-dce1-4de6-8fb7-88075f7db129 · outbound

This paper cites Competitive Programming with Large Reasoning Models.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding Competitive Programming with Large Reasoning Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:57.365862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:57.365862Z digest=sha256:dc4c2dece575a4168979355950b981fcecf79e699a7cf8d1b7570920d5d9ea85

Observation aa34b9b4-ffad-4aa3-83ca-aa3829bb7f6f · outbound

This paper cites COFFE: A Code Efficiency Benchmark for Code Generation.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding COFFE: A Code Efficiency Benchmark for Code Generation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:57.465250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:57.465250Z digest=sha256:62088d2b3a4faa5f114c887db217393befa6f5b65a489ea0c90a93a1f61adb31

Observation ca57237a-acf1-443b-a1e0-ad4fabb214be · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:57.560048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:57.560048Z digest=sha256:9a64bab7193667a0b88d291d289b8afc1601f06e397f7f690ec4528f3fc115fc

Observation 0d04703e-ce4e-42be-a82d-05ff6f566335 · outbound

This paper cites Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:57.636759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:57.636759Z digest=sha256:b45206d843aad02280eb257ce4fc351d999690274327116b7ce46c910501aff2

Observation 9d92569e-dd0a-41c5-9707-ddcccb60ff5d · outbound

This paper cites TESTEVAL: Benchmarking Large Language Models for Test Case Generation.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding TESTEVAL: Benchmarking Large Language Models for Test Case Generation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:57.721681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:57.721681Z digest=sha256:7096ce7207ad94c87e4a808fd6710fb5102345983de54f8585b5ef7f0962cff6

Observation c61da6a1-b863-4077-b298-007329eaa1e2 · outbound

This paper cites KodCode: A Diverse, Challenging, and Verifiable Synthetic Dataset for Coding.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding KodCode: A Diverse, Challenging, and Verifiable Synthetic Dataset for Coding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:57.824234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:57.824234Z digest=sha256:360c6eba7421d8886ee0d3583d2b47fb8e920257a884c3e88ad0c93b9e6d1921

Observation 8b65fe22-9ec2-4d7a-b4b4-46c5125f8fd9 · outbound

This paper cites LIMO: Less is More for Reasoning.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding LIMO: Less is More for Reasoning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:57.934398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:57.934398Z digest=sha256:383bd076f71ede950303f604611c357b548282f346371f623f7c3bc073dbba1e

Observation 938c11e9-cd80-453e-88aa-75b304857657 · outbound

This paper cites No More Manual Tests? Evaluating and Improving ChatGPT for Unit Test Generation.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding No More Manual Tests? Evaluating and Improving ChatGPT for Unit Test Generation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:58.040594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:58.040594Z digest=sha256:6402c839a24faef7afbd3f04e5de7b845f9c7335563022fcb14651a96893b463

Observation ae0285f0-db1e-4107-888c-a8124f293355 · outbound

This paper cites SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:58.116873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:58.116873Z digest=sha256:cd54236c6bfb0cc4cc90a40b12216dda8db93f86bccca0fccb457815d3797e76

Observation 7fe0e6ca-7484-47d1-a1e4-74c89c72c79e · outbound

This paper cites ALGO: Synthesizing Algorithmic Programs with LLM-Generated Oracle Verifiers.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding ALGO: Synthesizing Algorithmic Programs with LLM-Generated Oracle Verifiers

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:58.211974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:58.211974Z digest=sha256:67b1e97c4e2333fb75ca45dee515f5b7186cd6710813f889aad468131ac65c05

Observation 05460c55-7a88-42d1-a9f5-469868613ccd · outbound

This paper cites TestBench: Evaluating Class-Level Test Case Generation Capability of Large Language Models.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding TestBench: Evaluating Class-Level Test Case Generation Capability of Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:58.370843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:58.370843Z digest=sha256:74e35f1cf5f16886305fa6aa562df4d57e4ea29a14f7341d241e6a16176c586e

Observation e74403fe-084f-422b-9b8d-5d4808eebf4f · outbound

This paper cites LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:58.472965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:58.472965Z digest=sha256:db138c5e6d1660f082aac666e3057b81390353dbd35f866623397736505ab205

Observation 5a0d58dc-98f6-4d04-abb3-7da3faa8eed2 · outbound

This paper cites an unresolved cited work.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:40:01.099920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T12:39:58.578150Z digest=sha256:fbc6b45f62f965a222955942dce762d48540d8f046ee91a3445a7e40795f6471

Observation dcfff285-cc90-4ffd-a491-3123e6f4879e · outbound

This paper cites core logic.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding core logic

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:40:00.965186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T12:39:58.714694Z digest=sha256:cc9e78c5b85c92156edf3e7aabcfc953c4baf4cdefc893241866445c8389f9d1

Observation d69afd2e-7914-4815-a77d-59fb5b820682 · outbound

This paper cites 3\n1 10\n2 8\n3 10.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding 3\n1 10\n2 8\n3 10

Reference 1000

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:40:00.611877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T12:39:58.905686Z digest=sha256:01cd0d3959e0dfcd0f665fdda9dc1682efac783fa2634335f5e9f8219a6f7098

Observation c47beb24-a9e0-4f8b-8614-f0da8e6c0637 · outbound

This paper cites ISBN 9781450304436.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding ISBN 9781450304436

Reference 2011

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:55.691285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:55.691285Z digest=sha256:ab8e6ffe3162b2cc5019a8b84514efd0178715ad15f01143bd2ecd817b9f6e50

Observation 5115e784-7837-4ed2-986a-be32a29beb91 · outbound

This paper cites Program Synthesis with Large Language Models.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding Program Synthesis with Large Language Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:55.538922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:55.538922Z digest=sha256:784af7355fd5bd2d28e8eedcc2174e7518cb0de631d0194a3d5c311669c24418

Observation 5d72edb9-f7f9-44e4-8f82-d98ba46e6223 · outbound

This paper cites TACO: Topics in Algorithmic COde generation dataset.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding TACO: Topics in Algorithmic COde generation dataset

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:56.321057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:56.321057Z digest=sha256:27e5c9a0055214782b39d0a1f8ba27cb3858bb04b714ed8693d3849553cb1ce3

Observation ba80164b-d1a3-4ad6-b274-f52d49fe5572 · outbound

This paper cites LIMR: Less is More for RL Scaling.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding LIMR: Less is More for RL Scaling

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:56.430228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:56.430228Z digest=sha256:dfbfa9e190f709dc9174b046ea52d88a4e40479a263fdbc1e623e7cfe5c1b1bf

Observation f145b96d-c545-4261-bb99-07a16a9a238b · outbound

This paper cites TDD-Bench Verified: Can LLMs Generate Tests for Issues Before They Get Resolved?.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding TDD-Bench Verified: Can LLMs Generate Tests for Issues Before They Get Resolved?

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:55.474805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:55.474805Z digest=sha256:888a6d4a13331b8e263dc426d17dffe9e373eb16b9edcf1cbb1bbbc551b48d90

Observation 51511022-9523-4e30-8b97-5c23997f1d84 · outbound

This paper cites OpenCodeReasoning: Advancing Data Distillation for Competitive Coding.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding OpenCodeReasoning: Advancing Data Distillation for Competitive Coding

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:55.391872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:55.391872Z digest=sha256:ebaffb3d9ce6abd0816eaaee5f67a612a7d82caa69999896749fcd4bbcb12dfa

Pith citing papers

Observation 57da5526-4a26-4dbb-8eed-1f45a9385867 · inbound

Efficiency of turbulence cites this paper.

Efficiency of turbulence HardTests: Synthesizing High-Quality Test Cases for LLM Coding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T00:58:27.773874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:58:27.773874Z digest=sha256:ff3baabd124374251c5588f6a5591aeb48f12ceff2e48cbf09e0a62f48d06c3e

Observation bb1423bc-89a7-4034-949f-64dd43891641 · inbound

Navigating the Clutter: Waypoint-Based Bi-Level Planning for Multi-Robot Systems cites this paper.

Navigating the Clutter: Waypoint-Based Bi-Level Planning for Multi-Robot Systems HardTests: Synthesizing High-Quality Test Cases for LLM Coding

Reference 118

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:06:05.487820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-09T23:27:54.704794Z digest=sha256:88d569269484fa5d9825dd5c573db1a4adc72ad3795929c6700999bbc73fa2a2

Observation 02a07b45-6973-4a29-84d8-0db561655f4d · inbound

VeriContest: A Competitive-Programming Benchmark for Verifiable Code Generation cites this paper.

VeriContest: A Competitive-Programming Benchmark for Verifiable Code Generation HardTests: Synthesizing High-Quality Test Cases for LLM Coding

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:01:32.448442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-12T01:21:42.562823Z digest=sha256:6e0fcd973426d03b4917edb82eafa21cabb2643a79ef721a52a4c90b89427315

Observation 7f72180a-fb8a-48f2-93ee-2b50f2e2cfc5 · inbound

FrontierSmith: Synthesizing Open-Ended Coding Problems at Scale cites this paper.

FrontierSmith: Synthesizing Open-Ended Coding Problems at Scale HardTests: Synthesizing High-Quality Test Cases for LLM Coding

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:03:29.053765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T02:02:25.597640Z digest=sha256:3f1eeece375a971bdcf22cdb2dc8f8acf6769b247b2095495ded95e999e1d435

Observation c6f235e1-2126-485a-a8bb-23a8c8e84267 · inbound

CP-Agent: A Calibrated Risk-Controlled Agent for Feedback-Driven Competitive Programming cites this paper.

CP-Agent: A Calibrated Risk-Controlled Agent for Feedback-Driven Competitive Programming HardTests: Synthesizing High-Quality Test Cases for LLM Coding

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:24:40.201084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T13:19:22.541146Z digest=sha256:4b207c6ca7b140d4aa21347c32e65fafddaa033bf83ba0dce6a0cca36dd96108

Observation 99f18f25-3205-49c3-8e8a-86caaa7b88ff · inbound

The Generalization Spectrum: A Chromatographic Approach to Evaluating Learning Algorithms cites this paper.

The Generalization Spectrum: A Chromatographic Approach to Evaluating Learning Algorithms HardTests: Synthesizing High-Quality Test Cases for LLM Coding

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-04T20:00:08.067026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-25T20:55:15.784610Z digest=sha256:8847254571b038f68cfbd8f0a80551d9ff29abf7e6e36b671849f6b9480eba97

Observation b618dcfd-225a-4794-85de-af1428477042 · inbound

The Generalization Spectrum: A Chromatographic Approach to Evaluating Learning Algorithms cites this paper.

The Generalization Spectrum: A Chromatographic Approach to Evaluating Learning Algorithms HardTests: Synthesizing High-Quality Test Cases for LLM Coding

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:09:50.596811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-26T05:29:21.598397Z digest=sha256:dd52e949460397298f05359c16662d19debc55bc9e72c3d97f4fcb203748bb95

Observation 0aa7b6c7-458a-4bf8-a924-98bbb9ae9c7b · inbound

Coding Agents as Test-Suite Auditors: Finding What Official Suites Miss While Approaching What They Catch cites this paper.

Coding Agents as Test-Suite Auditors: Finding What Official Suites Miss While Approaching What They Catch HardTests: Synthesizing High-Quality Test Cases for LLM Coding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T22:23:16.109674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:23:16.109674Z digest=sha256:a317839cfaab833320a28cc8127f2e914866d552a36677325327b778a5eda484