Pith. sign in

Paper Citation Record · LEDGER

DataPrep-Bench: Benchmarking LLMs as Training Data Preparators

As of 15 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 0 inbound Pith citation observations for arXiv:2607.20465.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.20465 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T13:43:29.483303Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

46 of 46 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved46
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bd4b03ec-4b6f-40c5-8b3e-b5487fe65a13 · outbound

This paper cites Fowlkes, Stefano Soatto, and Pietro Perona.

DataPrep-Bench: Benchmarking LLMs as Training Data Preparators Fowlkes, Stefano Soatto, and Pietro Perona

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T13:43:29.268938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:43:29.268938Z digest=sha256:bbbdef256b6863b7a70ff26cc333cbcf5ff44872d4215d929ec5aa2ea490b808

Observation cdffe553-588b-46b7-888b-cadc8844152d · outbound

This paper cites A Survey of Multimodal Large Language Model from A Data-centric Perspective.

DataPrep-Bench: Benchmarking LLMs as Training Data Preparators A Survey of Multimodal Large Language Model from A Data-centric Perspective

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T13:43:29.274883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:43:29.274883Z digest=sha256:3d98ec2233d4bf554a10c67e251a16658555bc2fd4803997debe7686475160d1

Observation fd684c46-25f3-4722-bfef-fec878732057 · outbound

This paper cites Text2sql-flow: A robust sql-aware data augmentation framework for text-to-sql.arXiv preprint arXiv:2511.10192, 2025.

DataPrep-Bench: Benchmarking LLMs as Training Data Preparators Text2sql-flow: A robust sql-aware data augmentation framework for text-to-sql.arXiv preprint arXiv:2511.10192, 2025

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T13:43:29.281395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:43:29.281395Z digest=sha256:9575e6f35b9fe2174951718e7a23d87d602bd8ac0ffe312929c7f0643552f88b

Observation a93b4c45-4b63-45cb-be31-147ea1e05d4b · outbound

This paper cites Data-juicer: A one-stop data processing system for large language models.

DataPrep-Bench: Benchmarking LLMs as Training Data Preparators Data-juicer: A one-stop data processing system for large language models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T13:43:29.286818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:43:29.286818Z digest=sha256:77e1be7dc7c10b74759d739d1199fb8ec13b61d9d6a2427c2b9f323b48c16404

Observation dcbe75e0-c1c3-41c6-a1a3-81dff4e03df5 · outbound

This paper cites Dc-bench: Dataset condensation benchmark.Advances in Neural Information Processing Systems, 35:810–822, 2022.

DataPrep-Bench: Benchmarking LLMs as Training Data Preparators Dc-bench: Dataset condensation benchmark.Advances in Neural Information Processing Systems, 35:810–822, 2022

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T13:43:29.292145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:43:29.292145Z digest=sha256:9ca09b08c5159f2422e4e55e9d0a77fe38e09a65aead5247096569455950051e

Observation 7c4aa522-9ce7-4455-859b-e95d8a464882 · outbound

This paper cites Enhancing chat language models by scaling high-quality instructional conversations, 2023.

DataPrep-Bench: Benchmarking LLMs as Training Data Preparators Enhancing chat language models by scaling high-quality instructional conversations, 2023

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T13:43:29.296952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:43:29.296952Z digest=sha256:553e02fc33ef397030da05262ae6e556ccc232211c3b31be131ce611abed42a8

Observation 603995aa-e9eb-4334-a2a1-06b30605e8f0 · outbound

This paper cites Unleashing LLM Reasoning Capability via Scalable Question Synthesis from Scratch.

DataPrep-Bench: Benchmarking LLMs as Training Data Preparators Unleashing LLM Reasoning Capability via Scalable Question Synthesis from Scratch

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T13:43:29.303491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:43:29.303491Z digest=sha256:73929f1844ec0e8d3f5f88d3996df34ab1cb4ba3729d13ef025eebb932e9ee64

Observation a22e3b31-e2d7-4070-a10c-45ae2789e857 · outbound

This paper cites MegaScience: Pushing the Frontiers of Post-Training Datasets for Science Reasoning.

DataPrep-Bench: Benchmarking LLMs as Training Data Preparators MegaScience: Pushing the Frontiers of Post-Training Datasets for Science Reasoning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T13:43:29.308285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:43:29.308285Z digest=sha256:adfa476d94bff4730bac3e48384bcdc3d5ca6a0541523795bc6b884f2f6b3fed

Observation 941e3dc9-bf3c-4716-a14f-fbaec6ad7448 · outbound

This paper cites The Vendi Score: A Diversity Evaluation Metric for Machine Learning.

DataPrep-Bench: Benchmarking LLMs as Training Data Preparators The Vendi Score: A Diversity Evaluation Metric for Machine Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T13:43:29.312742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:43:29.312742Z digest=sha256:1d49890bf87c31548b7ee84f5b90fdda8ecab56b8c6e6a0b9a59f0ec677278a3

Observation 858d25fd-5f67-4e8b-aa18-786e218ef8c7 · outbound

This paper cites Datacomp: In search of the next generation of multimodal datasets.

DataPrep-Bench: Benchmarking LLMs as Training Data Preparators Datacomp: In search of the next generation of multimodal datasets

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T13:43:29.317738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:43:29.317738Z digest=sha256:8c240bc63c8ead65234815be6b9eb56993c23ec6fcdcc08e41343f7c34635d46

Observation 96272a94-f334-4826-be5b-c904b8f16b8c · outbound

This paper cites Closing the data loop: Using opendataarena to engineer superior training datasets.arXiv preprint arXiv:2601.09733, 2025.

DataPrep-Bench: Benchmarking LLMs as Training Data Preparators Closing the data loop: Using opendataarena to engineer superior training datasets.arXiv preprint arXiv:2601.09733, 2025

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T13:43:29.321741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:43:29.321741Z digest=sha256:44d1ef098a630804a94e6ec6067969ee2df7bd6c10fb24c8bd76b395de9c07c0

Observation edec8fc0-5fe7-4784-8327-c3dd57a4efee · outbound

This paper cites The Llama 3 Herd of Models.

DataPrep-Bench: Benchmarking LLMs as Training Data Preparators The Llama 3 Herd of Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T13:43:29.326449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:43:29.326449Z digest=sha256:eee0a4ed5af7acf1e995caadfa04137e5c464d05f955e288ad597725bc9ba6fc

Observation 8260ba64-d930-43a0-9770-7534233be136 · outbound

This paper cites Textbooks Are All You Need.

DataPrep-Bench: Benchmarking LLMs as Training Data Preparators Textbooks Are All You Need

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T13:43:29.331049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:43:29.331049Z digest=sha256:9db58d67857dd8fb53d29e9a53e8ceb8240fec0ba0ba33a6790e3e5ef727bcbc

Observation fd6371c9-9507-4444-b855-2c9090b67b95 · outbound

This paper cites Lawyer llama technical report, 2023.

DataPrep-Bench: Benchmarking LLMs as Training Data Preparators Lawyer llama technical report, 2023

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T13:43:29.335854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:43:29.335854Z digest=sha256:0ab14534d460d1ae6301ab468b312e24c9036af1e74c35ded2dadcc749a3c26f

Observation 9c181c15-a749-47ab-9e2a-19cf6cb0d272 · outbound

This paper cites Mistral 7B.

DataPrep-Bench: Benchmarking LLMs as Training Data Preparators Mistral 7B

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T13:43:29.340436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:43:29.340436Z digest=sha256:be2c180d8d0f375f8e0dfea71d7feb24c1e93419d8a8e2be0c9281c50e63ca5c

Observation 56135507-05ce-4684-a697-169e33636bf7 · outbound

This paper cites Scaling Laws for Neural Language Models.

DataPrep-Bench: Benchmarking LLMs as Training Data Preparators Scaling Laws for Neural Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T13:43:29.344746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:43:29.344746Z digest=sha256:659c6b77202f8b216d4605c98e4f0c8af6367272d5f5bdcfa99e924d68997a15

Observation cc07498d-6c21-4743-bacd-6e6d90f45803 · outbound

This paper cites Miranda, Alisa Liu, Nouha Dziri, Shane Lyu, Yuling Gu, Saumya Malik, Victoria Graf, Jena D.

DataPrep-Bench: Benchmarking LLMs as Training Data Preparators Miranda, Alisa Liu, Nouha Dziri, Shane Lyu, Yuling Gu, Saumya Malik, Victoria Graf, Jena D

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T13:43:29.349168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:43:29.349168Z digest=sha256:c4af85355e26bdf0764d661c83177716427fc00d5940cc3c2a9a9ed5078f95d5

Observation 798812e9-500e-42ac-beab-8982f38a73f6 · outbound

This paper cites Datacomp-lm: In search of the next generation of training sets for language models.

DataPrep-Bench: Benchmarking LLMs as Training Data Preparators Datacomp-lm: In search of the next generation of training sets for language models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T13:43:29.353546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:43:29.353546Z digest=sha256:b5b56866b187b94c244a684ff66b65190b42d51368464ec235936cb16059d5fc

Observation 29130112-8ba6-494e-9325-3879df750380 · outbound

This paper cites Infinity Instruct: Scaling Instruction Selection and Synthesis to Enhance Language Models.

DataPrep-Bench: Benchmarking LLMs as Training Data Preparators Infinity Instruct: Scaling Instruction Selection and Synthesis to Enhance Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T13:43:29.358479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:43:29.358479Z digest=sha256:4f6031a11a18c26e6ad7d54cdebf88db4d0484ed871cd936c839e91cd78d4442

Observation 32307cbf-ed4d-4525-ba4b-b64736018cbd · outbound

This paper cites Superfiltering: Weak-to-strong data filtering for fast instruction-tuning.

DataPrep-Bench: Benchmarking LLMs as Training Data Preparators Superfiltering: Weak-to-strong data filtering for fast instruction-tuning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T13:43:29.363120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:43:29.363120Z digest=sha256:131405aba9c5fee895c01ae4bc3a06304224c92ea30ee0cc1917cbc6895f3f34

Observation 951c07c5-babc-41dd-9ebd-a3df9f7e8723 · outbound

This paper cites Dataflow: An llm-driven framework for unified data preparation and workflow automation in the era of data-centric ai.arXiv preprint arXiv:2512.16676, 2025.

DataPrep-Bench: Benchmarking LLMs as Training Data Preparators Dataflow: An llm-driven framework for unified data preparation and workflow automation in the era of data-centric ai.arXiv preprint arXiv:2512.16676, 2025

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T13:43:29.368048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:43:29.368048Z digest=sha256:25b3101e08f47e2326d6d16978bbfb752faa9e3603140d611885e3ace519b17c

Observation 0018cca2-de3d-4cf0-89a8-2fcedc1c0926 · outbound

This paper cites Data preparation for large language models.Journal of Computer Science and Technology, 2026.

DataPrep-Bench: Benchmarking LLMs as Training Data Preparators Data preparation for large language models.Journal of Computer Science and Technology, 2026

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T13:43:29.372838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:43:29.372838Z digest=sha256:46b5973e3be2bccf65508176f0d2f007272355655859e81ee6fa13265649a38f

Observation b170a5ed-a4c8-4819-8681-4183ffe4a4bf · outbound

This paper cites What makes good data for alignment? a comprehensive study of automatic data selection in instruction tuning.

DataPrep-Bench: Benchmarking LLMs as Training Data Preparators What makes good data for alignment? a comprehensive study of automatic data selection in instruction tuning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T13:43:29.377408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:43:29.377408Z digest=sha256:4a84754f053945274b5466d5b7167522f188786f80e9b5e87d39e162a629b44e

Observation 8e194211-7910-45fc-ab33-883e69dbf12e · outbound

This paper cites Dataperf: Benchmarks for data-centric ai development.

DataPrep-Bench: Benchmarking LLMs as Training Data Preparators Dataperf: Benchmarks for data-centric ai development

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T13:43:29.382203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:43:29.382203Z digest=sha256:50626df542e9b34ce4e83ebd5756a2ad2878fad309989f41a97b19bf3d0419e8

Observation e5fcee22-949f-4808-97ea-f56f89739ed4 · outbound

This paper cites GPT-4 technical report, 2023.

DataPrep-Bench: Benchmarking LLMs as Training Data Preparators GPT-4 technical report, 2023

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T13:43:29.386982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:43:29.386982Z digest=sha256:fa0da1bd04981ce27762fd84d021281e4830b2458d649c210dab4608c69b556b

Observation d1908213-b27b-41dd-8382-606338e87975 · outbound

This paper cites Towards tailored recovery of lexical diversity in literary machine translation.

DataPrep-Bench: Benchmarking LLMs as Training Data Preparators Towards tailored recovery of lexical diversity in literary machine translation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T13:43:29.391092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:43:29.391092Z digest=sha256:530b8a7a0e88c3d39dc3cca867eb85b23fd16fa7875826bab1b72dbfd3dae3fe

Observation 0428d130-e587-457f-8628-4a16b3509635 · outbound

This paper cites Fino1: On the Transferability of Reasoning-Enhanced LLMs and Reinforcement Learning to Finance.

DataPrep-Bench: Benchmarking LLMs as Training Data Preparators Fino1: On the Transferability of Reasoning-Enhanced LLMs and Reinforcement Learning to Finance

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T13:43:29.395146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:43:29.395146Z digest=sha256:f4c4878e6cfecc2533d1cbbc3f095e518813ce35f10c28b0462c6997f4ca7066

Observation 3ccdba98-9f92-4550-a0c8-458feef7a0ab · outbound

This paper cites A survey on domain adaptation theory: learning bounds and theoretical guarantees.

DataPrep-Bench: Benchmarking LLMs as Training Data Preparators A survey on domain adaptation theory: learning bounds and theoretical guarantees

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T13:43:29.400467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:43:29.400467Z digest=sha256:6618cb983960293ff97c0c2092ea79a8ca78e4a243539ec0b0acae1a499f3fd7

Observation ae485336-5f26-4516-87d6-4d769b4338dc · outbound

This paper cites Let’s verify math questions step by step.arXiv preprint arXiv:2505.13903, 2025.

DataPrep-Bench: Benchmarking LLMs as Training Data Preparators Let’s verify math questions step by step.arXiv preprint arXiv:2505.13903, 2025

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T13:43:29.404990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:43:29.404990Z digest=sha256:f900c5f02baf5b443dbaf8741d81dabeaf20fc6dacf4e0189647a2d25b5b2d66

Observation 5b144c66-67dc-4ca7-8326-4f5ae8b151e7 · outbound

This paper cites Reasonmed: A 370k multi-agent generated dataset for advancing medical reasoning,.

DataPrep-Bench: Benchmarking LLMs as Training Data Preparators Reasonmed: A 370k multi-agent generated dataset for advancing medical reasoning,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T13:43:29.409080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:43:29.409080Z digest=sha256:c6168a8072e3cea3da4633172568f60f7f356867353ef79a5cbdd6701ac40b9b

Observation 7f9efaef-e9bc-4663-b11d-f628591b0d7d · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

DataPrep-Bench: Benchmarking LLMs as Training Data Preparators LLaMA: Open and Efficient Foundation Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T13:43:29.418322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:43:29.418322Z digest=sha256:a3da9652271a5c09514a5c93a4f49c7f64b9e86d6bd410b8e720d4ed2773c48b

Observation e9eb37d7-e4c7-44c3-8ed3-777f9175fe87 · outbound

This paper cites Self-instruct: Aligning language models with self-generated instructions.

DataPrep-Bench: Benchmarking LLMs as Training Data Preparators Self-instruct: Aligning language models with self-generated instructions

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T13:43:29.422758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:43:29.422758Z digest=sha256:ac4772ff576299ae7cdc433d28d0e132c1d0cca38139c872c0813eb1a28ae930

Observation 8a4bba88-7a4f-4b5c-b9e0-987ea6a9ec2b · outbound

This paper cites QuRating: Selecting High-Quality Data for Training Language Models.

DataPrep-Bench: Benchmarking LLMs as Training Data Preparators QuRating: Selecting High-Quality Data for Training Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T13:43:29.427294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:43:29.427294Z digest=sha256:2e608b7097da1aa955c013bb50c88a7d9cb6a6b5d7e91471ba68698e132a924d

Observation d3d50ac5-3475-4a71-a7fc-109dec7d815d · outbound

This paper cites LESS: Selecting Influential Data for Targeted Instruction Tuning.

DataPrep-Bench: Benchmarking LLMs as Training Data Preparators LESS: Selecting Influential Data for Targeted Instruction Tuning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T13:43:29.432403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:43:29.432403Z digest=sha256:63a8d6dca44d0502b36a0a3a1620398263b333dabb6cb06792c2e8757b396428

Observation 31dd247f-db92-44bf-b5b7-d0f7224e1082 · outbound

This paper cites Wizardlm: Empowering large pre-trained language models to follow complex instructions, 2023.

DataPrep-Bench: Benchmarking LLMs as Training Data Preparators Wizardlm: Empowering large pre-trained language models to follow complex instructions, 2023

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T13:43:29.438152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:43:29.438152Z digest=sha256:3d49693c33b70f44b140e26dad54446e17462b3a427e2298512890fa16ec1423

Observation cd37b376-0ddf-4eb3-9bed-1cd38b27c9e0 · outbound

This paper cites Logics-stem: Empowering llm reasoning via failure-driven post-training and document knowledge enhancement, 2026.

DataPrep-Bench: Benchmarking LLMs as Training Data Preparators Logics-stem: Empowering llm reasoning via failure-driven post-training and document knowledge enhancement, 2026

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T13:43:29.443264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:43:29.443264Z digest=sha256:b3fcfa5f167bbab75124f806805c4fd1538ce9509cff852be2fb87f5cd5f181c

Observation ab7bbbe0-d135-48ed-9c48-1b67fee50a50 · outbound

This paper cites Qwen2 Technical Report.

DataPrep-Bench: Benchmarking LLMs as Training Data Preparators Qwen2 Technical Report

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T13:43:29.447430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:43:29.447430Z digest=sha256:0b8fba7b8139a3d2266c73807bca99abcb74a74d834542bfa44c6f6f3ed2741c

Observation 2c5787fa-93fb-43fc-811f-d053fa1ca382 · outbound

This paper cites Qwen2.5 Technical Report.

DataPrep-Bench: Benchmarking LLMs as Training Data Preparators Qwen2.5 Technical Report

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T13:43:29.452083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:43:29.452083Z digest=sha256:bf86c7ebdfa60fec54928be7de2037f81994f8751219030696b1be6d107a2431

Observation a1be480d-9b52-4913-8e55-650ef7ec0d5b · outbound

This paper cites OpenCSG Chinese Corpus: A Series of High-quality Chinese Datasets for LLM Training.

DataPrep-Bench: Benchmarking LLMs as Training Data Preparators OpenCSG Chinese Corpus: A Series of High-quality Chinese Datasets for LLM Training

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T13:43:29.456364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:43:29.456364Z digest=sha256:3a684f709ddb7ab019a19380a7ade27b82a9008ec632e5cfd6c519a280adab59

Observation 2a79cf5e-56d5-45fa-b806-dfec79adf9fb · outbound

This paper cites Disc-lawllm: Fine-tuning large language models for intelligent legal services, 2023.

DataPrep-Bench: Benchmarking LLMs as Training Data Preparators Disc-lawllm: Fine-tuning large language models for intelligent legal services, 2023

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T13:43:29.460718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:43:29.460718Z digest=sha256:9f06f409d974dde2b47f81a065364365c75f996a557ad75f6ca9f2bf6c84f1b9

Observation ca1f35df-7d2d-4df3-852a-e25dd474a82b · outbound

This paper cites Ultramedical: Building specialized generalists in biomedicine, 2024.

DataPrep-Bench: Benchmarking LLMs as Training Data Preparators Ultramedical: Building specialized generalists in biomedicine, 2024

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T13:43:29.464976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:43:29.464976Z digest=sha256:b95f2f3428513380ba2ba934146fe878b27c01f14871752ea424feec01a655fd

Observation 8f808fa4-a4bc-489c-80c0-886038268843 · outbound

This paper cites LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models.

DataPrep-Bench: Benchmarking LLMs as Training Data Preparators LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T13:43:29.469176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:43:29.469176Z digest=sha256:c03804054d37834a8610842c096c24b382b47cc82dba5fa37bf50b83de0cb103

Observation 4720352e-9d9c-4b37-ac62-f800bdecbbd1 · outbound

This paper cites Google-proof.

DataPrep-Bench: Benchmarking LLMs as Training Data Preparators Google-proof

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T13:43:29.473875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:43:29.473875Z digest=sha256:762064c8527fc02781ef4fac100b9f49805acf3c7174b3737fc01669a5c7acba

Observation 27cc783e-2d4e-47a6-837b-b1d93658916a · outbound

This paper cites an unresolved cited work.

DataPrep-Bench: Benchmarking LLMs as Training Data Preparators Unresolved cited work

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T13:43:29.478789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:43:29.478789Z digest=sha256:a3296148195aee95c956451927ac5db7f485b9f380160a7bbbb2aa1a0eea8913

Observation a483e706-d926-47cf-b11a-b27d902d965a · outbound

This paper cites an unresolved cited work.

DataPrep-Bench: Benchmarking LLMs as Training Data Preparators Unresolved cited work

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T13:43:29.483303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:43:29.483303Z digest=sha256:5d2e304908b53cecefdd0d0e302315e7690533f25fb1eeadf1526c4635aa6d7b

Observation 17223647-62dc-4e87-b1fc-eedfb60ea8fc · outbound

This paper cites an unresolved cited work.

DataPrep-Bench: Benchmarking LLMs as Training Data Preparators Unresolved cited work

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-02T13:43:29.413829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:43:29.413829Z digest=sha256:1ebed9df9928ef9a9f5c4ce23890f8c16ac2bf1e27894355184017e07c24575b

Pith citing papers

No inbound Pith citation observations are available.