Pith. sign in

Paper Citation Record · LEDGER

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient

As of 10 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 1 inbound Pith citation observation for arXiv:2502.01683.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.01683 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T18:09:01.728564Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T19:39:36.819142Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-10T22:40:49.906159Z

Reference resolution

45 of 45 outbound references displayed

  • verified exact1
  • verified fuzzy4
  • unresolved40
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9cfc5650-ba39-478e-84b4-ee36304aca78 · outbound

This paper cites Claude 3.5.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient Claude 3.5

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:09:02.311909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T18:09:01.591091Z digest=sha256:63e252f82aac0f6573b39c40e09a3d276fc7d92a5ea29b560b07091ece584900

Observation 409b5d83-32f7-40d0-90ca-e91565196e3c · outbound

This paper cites an unresolved cited work.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-09T18:09:02.304253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T18:09:01.594566Z digest=sha256:6b29ae5cce1c1f804d14cd947425edc5cd6e785632e1f25353a6d09b25642835

Observation 36a20913-6729-453e-880c-f85fcff73202 · outbound

This paper cites Scaling Synthetic Data Creation with 1,000,000,000 Personas.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient Scaling Synthetic Data Creation with 1,000,000,000 Personas

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-09T18:09:01.597866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:09:01.597866Z digest=sha256:c857f9a1c9d230d5b7854ba5291302b6537295523cd7add82126ba9d12bb73f5

Observation 92662432-f6f3-4d5e-954c-b51bd990f875 · outbound

This paper cites Yu, Qiang Yang, and Xing Xie.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient Yu, Qiang Yang, and Xing Xie

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-09T18:09:01.601806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:09:01.601806Z digest=sha256:924403de647c985aa2abd0875f82ef40df266bd15bf7ac0a0fcc6dafa5f6e8ba

Observation 63181b01-6028-476d-a761-cf84cc0c8618 · outbound

This paper cites FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-09T18:09:01.604922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:09:01.604922Z digest=sha256:79871ba4c890773172174efd564b83156942cb9dae0e856f5e9b15c945263f93

Observation a34b53b7-cf57-4948-83e5-602a6b6665bf · outbound

This paper cites an unresolved cited work.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-09T18:09:01.608303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:09:01.608303Z digest=sha256:c87cc03da36d0080fe80e23d755fcf4483b5aa507fb2e6e1d827688663e1f597

Observation 0c544555-8fb4-4a4d-862e-3b86ac823563 · outbound

This paper cites an unresolved cited work.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-09T18:09:02.296828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T18:09:01.611435Z digest=sha256:8ed5a3e32e6c6cc68a5597c370ee56aa61b0558d510ee2cd9a4368e22009e9bc

Observation 98f7e8f4-cd76-4e95-acbf-8ac9a4df1e8d · outbound

This paper cites an unresolved cited work.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-09T18:09:01.614868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:09:01.614868Z digest=sha256:ebd50c4a8cf68a94092db2034f14edf28a23be67379330d62b5b44acb6d78135

Observation e87f3976-8e7a-4022-a988-58993e941b82 · outbound

This paper cites an unresolved cited work.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-09T18:09:02.282798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T18:09:01.617635Z digest=sha256:83bec170d4b1428d866201befb34bb32ffb997d687696e5bfe928687bdf240ae

Observation 883b262a-7a1d-4d85-a5a3-b806e69c743c · outbound

This paper cites an unresolved cited work.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-09T18:09:02.272809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T18:09:01.620591Z digest=sha256:f065254c57811e19902dc660884e2b02b47c6134a384a1c8bac809058a554f31

Observation 46110983-3b6c-4d9d-a724-85304659df95 · outbound

This paper cites GPT-4o System Card.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient GPT-4o System Card

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T18:09:01.624555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:09:01.624555Z digest=sha256:a60f5b747aecea861e7dc84166902b893ebabbb1a3aaef99aab3d087d51a097d

Observation 8048585d-e5f7-4b1b-97c7-e9d50480054f · outbound

This paper cites an unresolved cited work.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T18:09:01.627606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:09:01.627606Z digest=sha256:238fcb506dfa73a0c0fd146d6a6e27ec4e3f5af9172eacee5aa70f735b757605

Observation b41e42e1-dd22-4450-b693-b7db6bbc2283 · outbound

This paper cites Causal Machine Learning: A Survey and Open Problems.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient Causal Machine Learning: A Survey and Open Problems

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T18:09:01.630239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:09:01.630239Z digest=sha256:1a6e5be0833c8983d6edc6258a38b2104b6faee08a0a3844329c33a729a67327

Observation ed49413c-ef14-4cc2-9cf8-c678d0e839e4 · outbound

This paper cites S3Eval: A Synthetic, Scalable, Systematic Evaluation Suite for Large Language Models.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient S3Eval: A Synthetic, Scalable, Systematic Evaluation Suite for Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T18:09:01.632998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:09:01.632998Z digest=sha256:cb7cef3b3cbd06b834648e06fb10f0fcc2d9a6987fa072106424de86bc59548d

Observation 9dc084b1-b4bb-49ff-a534-a4ec3c054b90 · outbound

This paper cites u ttler, Mike Lewis, Wen - tau Yih, Tim Rockt \.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient u ttler, Mike Lewis, Wen - tau Yih, Tim Rockt \

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T18:09:01.636225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:09:01.636225Z digest=sha256:c71ffea8fc2e232a34a05bf1b6c077afc3743787006a699cdaed6c07987068d9

Observation a9011def-dae2-4f1d-87ff-a979333e27b6 · outbound

This paper cites PertEval: Unveiling Real Knowledge Capacity of LLMs with Knowledge-Invariant Perturbations.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient PertEval: Unveiling Real Knowledge Capacity of LLMs with Knowledge-Invariant Perturbations

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T18:09:01.639375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:09:01.639375Z digest=sha256:f41636fcd8cc1467db2688b862832a2b86a07133cecdd2797778e36908ac059b

Observation 9de263b1-2e74-46f2-9b23-f5dce2ebfe53 · outbound

This paper cites SciLitLLM: How to Adapt LLMs for Scientific Literature Understanding.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient SciLitLLM: How to Adapt LLMs for Scientific Literature Understanding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-09T18:09:01.642749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:09:01.642749Z digest=sha256:1ec933ddf06a972d0c0ea431d4c8a69392a56910e344925cb40b0e58a67432e6

Observation fa841f16-f458-4531-b321-e2c6e00cf99c · outbound

This paper cites Neuro-Symbolic Data Generation for Math Reasoning.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient Neuro-Symbolic Data Generation for Math Reasoning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-09T18:09:01.645942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:09:01.645942Z digest=sha256:183ec3cf55c499dedee4989db3c243aa800c3be792482ca459b3c2f4b783552a

Observation f34460d2-0fbb-435c-9df8-cf06c89b198e · outbound

This paper cites an unresolved cited work.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-09T18:09:02.259671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T18:09:01.649187Z digest=sha256:a7eb44c31ea6e222bf84313aadb1315e8aec9ebdaa64fc2b8a900941166dd3d7

Observation 6bae59c9-87d6-4204-992f-178418605f0d · outbound

This paper cites an unresolved cited work.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T18:09:01.651983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:09:01.651983Z digest=sha256:d5ee855640b9b87f559823db2f394bdd404d832a37fff6147bb02803b5888f4b

Observation 4b83488c-1a55-4c61-99f9-e55a04e3cea2 · outbound

This paper cites Arena Learning: Build Data Flywheel for LLMs Post-training via Simulated Chatbot Arena.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient Arena Learning: Build Data Flywheel for LLMs Post-training via Simulated Chatbot Arena

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-09T18:09:01.654904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:09:01.654904Z digest=sha256:cc6da663af97e934666338b4b2fd57b0d1cbb414c9f5223d759ce6b29231c1da

Observation 062255dc-8e6f-4159-a0b9-9b7359880775 · outbound

This paper cites an unresolved cited work.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-09T18:09:02.250661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T18:09:01.658012Z digest=sha256:873f3d1ada86dbab8185c95a18dd8647c49446e9191285c24961b82350f870ac

Observation 47e2f6de-f977-4a8f-a546-7593aeafc266 · outbound

This paper cites Efficacy of Synthetic Data as a Benchmark.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient Efficacy of Synthetic Data as a Benchmark

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T18:09:01.660763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:09:01.660763Z digest=sha256:7bbc7c7efaf49c968f584bd63153ad3e87288e73d0ab080408938154d3794f54

Observation 813d0a62-8588-41a4-9753-68695f3865a2 · outbound

This paper cites Jointly Measuring Diversity and Quality in Text Generation Models.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient Jointly Measuring Diversity and Quality in Text Generation Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T18:09:01.663911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:09:01.663911Z digest=sha256:cd710b7b4a059928bf3dade743638118923e17de5124d8e026435cf1e1ad0d6d

Observation 01342007-e778-4b7c-9999-a77c32f4bee6 · outbound

This paper cites text-embedding-ada-002.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient text-embedding-ada-002

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:09:02.241496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T18:09:01.667583Z digest=sha256:909658f7155a16e8957e45c12b2b1201911bec2b679565ed5489afeecc765742

Observation 1d3daef6-8ce7-4d38-920d-10f5a683dc7b · outbound

This paper cites Bowman, Amanda Askell, Roger Grosse, Danny Hernandez, Deep Ganguli, Evan Hubinger, Nicholas Schiefer, and Jared Kaplan.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient Bowman, Amanda Askell, Roger Grosse, Danny Hernandez, Deep Ganguli, Evan Hubinger, Nicholas Schiefer, and Jared Kaplan

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T18:09:01.670630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:09:01.670630Z digest=sha256:f07b3802d255bf17cf907a9922c77a031e9d03e2dbf3dcae72923056ff52b3f3

Observation c3ac89a6-e4f0-452f-9a4f-66abd3dcc4f5 · outbound

This paper cites an unresolved cited work.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-09T18:09:02.227588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T18:09:01.673654Z digest=sha256:3f4ec667c32c5299d2d4db67692982339ab1e7601de6e1cc3ca6f103f5f7a463

Observation 322e1967-c2d6-4d9b-9e65-c7729d065468 · outbound

This paper cites A Survey on Self-Evolution of Large Language Models.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient A Survey on Self-Evolution of Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-09T18:09:01.676414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:09:01.676414Z digest=sha256:e5dcd54c8c752eb103b075bce9eeb1fc714e41df6cea21e690792d56884f6ba6

Observation 6338cc35-3f69-4baf-81b6-61d9ce791fa7 · outbound

This paper cites Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-09T18:09:01.679147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:09:01.679147Z digest=sha256:7a05823bec9dc936d10108cf0f6d6fe1f3a05462f76b856915af01a5f9d067f1

Observation 15100bd0-b637-4948-b51c-5437ba88c379 · outbound

This paper cites an unresolved cited work.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-09T18:09:01.681746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:09:01.681746Z digest=sha256:c6688d02ae344b5bfbf222a734bbb5c1ef0ce0d6917ea67f701cbc642f9dd868

Observation 7aadc72c-cbde-4595-9497-b379dc9755cc · outbound

This paper cites A Survey on Data Synthesis and Augmentation for Large Language Models.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient A Survey on Data Synthesis and Augmentation for Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-09T18:09:01.684146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:09:01.684146Z digest=sha256:5e3ad6c6726e9f04a782978476a8f21617fdd70ba97b83cc190c2986f13e8774

Observation 802f192e-afe5-4bc8-8a87-08459d2ff2e5 · outbound

This paper cites Le, Ed H.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient Le, Ed H

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:09:02.219700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T18:09:01.686768Z digest=sha256:f97e1fe62fe34eb38658de9dd4e6322a69e639b7d1c03a984312297fdc6cfdf8

Observation 0dfa2173-c8c7-4ebc-843b-b17971b1f720 · outbound

This paper cites Smith, Daniel Khashabi, and Hannaneh Hajishirzi.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient Smith, Daniel Khashabi, and Hannaneh Hajishirzi

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-09T18:09:01.689516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:09:01.689516Z digest=sha256:ed3fd68b17038ace837ba6d8b17d3b4cfde7443d1dba7224ab7e35f940ac4ed8

Observation f201a6fa-2ab2-48c6-9b03-78a3b22f9612 · outbound

This paper cites MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-09T18:09:01.692472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:09:01.692472Z digest=sha256:59e127b8605886a05b95fc19db433711aff5c0621aa9d493958d51db945072b2

Observation a00b135b-b20e-4822-b2c6-cbe659517f71 · outbound

This paper cites an unresolved cited work.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-09T18:09:01.695833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:09:01.695833Z digest=sha256:4fc6af2ff780daf7324dcd792f9928513d5c562f4d53aa65d6c7d76e67c112da

Observation 349d573e-8f42-4ca0-8451-ab19848bfad3 · outbound

This paper cites Qwen2 Technical Report.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient Qwen2 Technical Report

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-09T18:09:01.699197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:09:01.699197Z digest=sha256:e1bf0fc5a747b00f8378032969e41c3752fd9f3dc7f73200b733b7a578475dba

Observation 23275359-7005-4eb7-865c-0dfc726f1abf · outbound

This paper cites Kwok, Zhenguo Li, Adrian Weller, and Weiyang Liu.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient Kwok, Zhenguo Li, Adrian Weller, and Weiyang Liu

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-09T18:09:01.702260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:09:01.702260Z digest=sha256:ec610bcd6457f80d992889f6c9c36a8a5c7c4c860772d34c6d2ff937ab718863

Observation 6d891ced-ae80-4651-8f13-d79f342473d2 · outbound

This paper cites Ratner, Ranjay Krishna, Jiaming Shen, and Chao Zhang.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient Ratner, Ranjay Krishna, Jiaming Shen, and Chao Zhang

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:09:02.206843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T18:09:01.705080Z digest=sha256:7a8e87ae2d478fc7aedb7f27140d2fc5ecb022169cf104703b76b2e44bcaaf14

Observation 01b89090-28f8-44de-9f76-cf7415f5dd79 · outbound

This paper cites an unresolved cited work.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-09T18:09:01.708002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:09:01.708002Z digest=sha256:7bbafda80b9cb5e63711564801490f327d7d1f88904073e55951490db4efade7

Observation 1a30f425-3980-4673-b508-1551318daebf · outbound

This paper cites Xing, Hao Zhang, Joseph E.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient Xing, Hao Zhang, Joseph E

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-09T18:09:01.711080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:09:01.711080Z digest=sha256:ca06607b36ec939988c086e0e0e1d86caf117f6900fccf1db573461d3e1d40a7

Observation 327f84c9-0dcd-4483-98f7-2268689f33f9 · outbound

This paper cites an unresolved cited work.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-09T18:09:02.191757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T18:09:01.714267Z digest=sha256:609b39ce0f13769aff25f1f60a9aab3a54af13de1c1a06f0f9091ed98a302510

Observation 1ee7a73b-7456-4d9b-8939-e8246680a985 · outbound

This paper cites Dynamic Evaluation of Large Language Models by Meta Probing Agents.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient Dynamic Evaluation of Large Language Models by Meta Probing Agents

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-09T18:09:01.717256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:09:01.717256Z digest=sha256:bc4b01ace049ac5615677194c71bb4774822f83d1e04cf13316d1ea0cc954eb2

Observation a16acfb5-74d5-4b93-9f88-ea0107d16fcb · outbound

This paper cites an unresolved cited work.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient Unresolved cited work

Reference 43

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-09T18:09:02.157754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T18:09:01.720740Z digest=sha256:b246fce6b16f8695697ac066c2e373ea1bf9e89e7d5baf88a46cd6f15e62b34f

Observation 8b6d801a-7488-4531-aaa1-6d43c4e56507 · outbound

This paper cites URL: " 'urlintro :=.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient URL: " 'urlintro :=

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-09T18:09:01.724723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:09:01.724723Z digest=sha256:c7cbdd6ef424444b5c97fbf005d1b8df12177f0751584a722947d6baa568ba26

Observation 26448a66-4b4f-4581-b83b-0d457d9303ed · outbound

This paper cites write newline.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient write newline

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-09T18:09:01.728564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:09:01.728564Z digest=sha256:d293b53aaf8fb5d330185d795b83f88356fe242b6af4ac527d8ccc2da1f009e9

Pith citing papers

Observation c925eb55-b244-4ff4-a252-25304c795a85 · inbound

Automatically Generating Hard Math Problems from Hypothesis-Driven Error Analysis cites this paper.

Automatically Generating Hard Math Problems from Hypothesis-Driven Error Analysis LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:40:49.911691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T19:39:36.819142Z digest=sha256:57614bb52db268d6b1374f64d5d015a1958cca7d8ea0ea91ac974cb75bde83f2