Pith. sign in

Paper Citation Record · LEDGER

Data Swarms: Optimizable Generation of Synthetic Evaluation Data

As of 7 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 0 inbound Pith citation observations for arXiv:2506.00741.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.00741 v2

Coverage vector

measured 64 of 64 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:03:42.484850Z

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

64 of 64 outbound references displayed

  • verified exact2
  • verified fuzzy33
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d016e96c-6595-4339-b0cc-cb96a723844b · outbound

This paper cites Kgquiz: Evaluating the generalization of encoded knowledge in large language models.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Kgquiz: Evaluating the generalization of encoded knowledge in large language models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:34.567215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:34.567215Z digest=sha256:92d35e9956934a221c50eab0194d641a281101a8b1de193f934450332acc6c9d

Observation 3d0e456c-a8c9-4635-8c4b-0b332bff53bd · outbound

This paper cites AutoEval Done Right: Using Synthetic Data for Model Evaluation.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data AutoEval Done Right: Using Synthetic Data for Model Evaluation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:34.669899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:34.669899Z digest=sha256:61ade2016fe31922021005712bddcb3835bc0992b90a4771ff52a1336809c070

Observation e5d3074a-1dc2-4a2c-8ffe-8c5179553f00 · outbound

This paper cites Adaptively evaluating models with task elicitation.arXiv preprint arXiv:2503.01986, 2025.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Adaptively evaluating models with task elicitation.arXiv preprint arXiv:2503.01986, 2025

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:34.895664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:34.895664Z digest=sha256:8c8094ad3c20407f2574032e18b2ee460aec321a5c1ce813d477bdb1e8cebcac

Observation b921ed93-8e24-4bca-9e25-0bce68d43405 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Training Verifiers to Solve Math Word Problems

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:35.075359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:35.075359Z digest=sha256:6ec40529e7495eb318a4670d26476579b959470e64f2d40f7ce1bab07512d9eb

Observation 55493653-a93b-4192-b056-e954f47d10a2 · outbound

This paper cites Knowledge crosswords: Geometric knowledge reasoning with large language models.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Knowledge crosswords: Geometric knowledge reasoning with large language models

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:50.656173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:03:35.188332Z digest=sha256:d1dfa71e648dd1c7f78258474d7e325743019f055d4e0d953bc983af8c260dec

Observation d38a2290-f18f-47e1-a916-e0d9b553976b · outbound

This paper cites Self-Boosting Large Language Models with Synthetic Preference Data.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Self-Boosting Large Language Models with Synthetic Preference Data

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:35.341828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:35.341828Z digest=sha256:a2d72a8bc9a0905feb0f7c85ce5d045da0ab5428e10310e327df2f5bf7f781c6

Observation 013d73e7-0a88-4b2d-b3f3-9d4d03f5362b · outbound

This paper cites Alpacafarm: A simulation framework for methods that learn from human feedback.Advances in Neural Information Processing Systems, 36, 2024.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Alpacafarm: A simulation framework for methods that learn from human feedback.Advances in Neural Information Processing Systems, 36, 2024

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:35.502201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:35.502201Z digest=sha256:2d91117662decb5ccff1e7f12abe9c67fd2b3a0c8d776f18e7d03bbce67fea99

Observation d41b1c59-fd0e-4f48-afef-853faba12af0 · outbound

This paper cites Clas- sifying the classifier: dissecting the weight space of neural networks.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Clas- sifying the classifier: dissecting the weight space of neural networks

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:50.381847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:03:35.540623Z digest=sha256:aa96bacb4c385d0c84412a786b970bc9438dfdd9c383bcf6e4dd8e91c2b9a70d

Observation e987ecb4-b7f4-4c52-b6c9-3e905dfb908e · outbound

This paper cites Heterogeneous swarms: Jointly optimizing model roles and weights for multi-llm systems.arXiv preprint arXiv:2502.04510, 2025.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Heterogeneous swarms: Jointly optimizing model roles and weights for multi-llm systems.arXiv preprint arXiv:2502.04510, 2025

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:35.613142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:35.613142Z digest=sha256:b3fc42b4a93ba259bb2616d055271a2283bbcac816af56cc88ac5bb17635d9d5

Observation fd4fb5fa-c7bd-4b36-85ce-d6907ad443ec · outbound

This paper cites Model swarms: Collaborative search to adapt LLM experts via swarm intelligence.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Model swarms: Collaborative search to adapt LLM experts via swarm intelligence

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:50.181618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:03:35.762723Z digest=sha256:4e2c452ba40b0dcf36256921a08d0bc5475b5686236e210fada342907fa1e8d2

Observation 6301967c-1eb0-41d1-8804-d4b16fd9704a · outbound

This paper cites Promptbreeder: Self-referential self-improvement via prompt evolution.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Promptbreeder: Self-referential self-improvement via prompt evolution

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:49.915732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:03:35.916500Z digest=sha256:d1947c86304d0b5e0f4eca56177db30f277b64b493a31aedd54f2679189170ca

Observation 07f41ea0-74e2-4f30-ae48-8840954b843f · outbound

This paper cites Open llm leaderboard v2.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Open llm leaderboard v2

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:36.045712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:36.045712Z digest=sha256:fc1a4b5cf4de3264c3ace3aa8480f84241219a41a2a3d4f8e2f7948850391a15

Observation e1da0c0b-27ac-47ea-bb3c-1a6b8142652a · outbound

This paper cites Time travel in llms: Tracing data contamination in large language models.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Time travel in llms: Tracing data contamination in large language models

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:49.649629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:03:36.184289Z digest=sha256:6880452436b2bf9553fde7e4319ef5f856ee986b7a3795ca57601a6d5a2b7f79

Observation 8178c412-ca1b-45f4-8c1f-4987319c36c6 · outbound

This paper cites Automated evaluation of retrieval-augmented language models with task-specific exam generation.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Automated evaluation of retrieval-augmented language models with task-specific exam generation

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:49.382389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:03:36.332679Z digest=sha256:7c842a43255fe269a73ee4d0cbb309b4409053bf7729fec758f568eb8d4bb22e

Observation 81453eb1-4329-4599-b071-66bf8072cbe2 · outbound

This paper cites Measuring massive multitask language understanding.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Measuring massive multitask language understanding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:36.464445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:36.464445Z digest=sha256:ca73011777047d61a0ba9d37f756f54eb0d14513fbeb5f3eddac8d2d74bf6c3b

Observation 5b49f598-0ae1-451f-a5fa-5464a234093a · outbound

This paper cites Datagen: Unified synthetic dataset generation via large language models.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Datagen: Unified synthetic dataset generation via large language models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:49.134596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:03:36.603676Z digest=sha256:eaaeea26c8c3218d27a872f057da67773eb20ca045048d6dbab31aaf86a513e5

Observation 4f27019b-e3df-4b9c-8f00-91f772bb388f · outbound

This paper cites Camels in a Changing Climate: Enhancing LM Adaptation with Tulu 2.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Camels in a Changing Climate: Enhancing LM Adaptation with Tulu 2

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:36.765030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:36.765030Z digest=sha256:cc81c5dcfe0cd5b53d270b0d31eea8bafd1747832c0eb2607bd3690f988b67e2

Observation e55dc512-29eb-492a-8c5c-652e67ef1c96 · outbound

This paper cites Wildteaming at scale: From in- the-wild jailbreaks to (adversarially) safer language models.Advances in Neural Information Processing Systems, 37:47094–47165, 2024.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Wildteaming at scale: From in- the-wild jailbreaks to (adversarially) safer language models.Advances in Neural Information Processing Systems, 37:47094–47165, 2024

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:48.937768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:03:36.978535Z digest=sha256:a4cfa9b9a8c75345a848ce448c0c5dfb286cf5ecff392cc5eb5513fcfb8cc5d5

Observation 7002d98f-da27-43be-829b-d8edf4490dcd · outbound

This paper cites Teaching language models to hallucinate less with synthetic tasks.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Teaching language models to hallucinate less with synthetic tasks

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:48.645883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:03:37.116286Z digest=sha256:d768cf9763aa34ae74dc691012987dd496dccfc12145162f2ea3ff921e63bdc4

Observation 34ed43f9-c088-4ec6-9a3f-ff7c1b7f7e9b · outbound

This paper cites Au- tonomous evaluation of llms for truth maintenance and reasoning tasks.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Au- tonomous evaluation of llms for truth maintenance and reasoning tasks

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:48.458711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:03:37.251988Z digest=sha256:977da06363c5105d744348d08e9cb5b21e93c5013ebadf6f51b030dbda3c3165

Observation 736059c6-e70f-43bd-9a8a-6dc29d8e1b03 · outbound

This paper cites Realtime qa: What’s the answer right now? Advances in neural information processing systems, 36:49025–49043, 2023.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Realtime qa: What’s the answer right now? Advances in neural information processing systems, 36:49025–49043, 2023

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:37.403279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:37.403279Z digest=sha256:f562a7fd8349329b3a7d5c4e3ea9ce4cb411de62611e175af990fe80e3c221ae

Observation e30f5e7f-8934-4ad5-8a4c-84496bd00fd0 · outbound

This paper cites Particle swarm optimization.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Particle swarm optimization

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:37.581123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:37.581123Z digest=sha256:6c537147e8c5ac762e180508a4b52c1f03b709219e3dfbe07ebf9f7b42689ee2

Observation e7b23351-9ef6-4c55-8752-f02f950bcf67 · outbound

This paper cites Openassistant conversations-democratizing large language model alignment.Advances in Neural Information Processing Systems, 36:47669–47681, 2023.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Openassistant conversations-democratizing large language model alignment.Advances in Neural Information Processing Systems, 36:47669–47681, 2023

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:48.194529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:03:37.767108Z digest=sha256:425ec8de2988cfd71f65a18ba7b84259407f3027d12f34f827f20ff969ee168b

Observation e40cc949-e6b2-43c0-b393-e6d8927bf100 · outbound

This paper cites Eliciting Language Model Behaviors with Investigator Agents.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Eliciting Language Model Behaviors with Investigator Agents

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:37.921485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:37.921485Z digest=sha256:3e59a73c9c940368c1585d882de9d48e0dccaad36639dba37b4e1c842fa19918

Observation f3b21320-0d42-4975-ba23-a63db57a5890 · outbound

This paper cites Autobencher: Towards declarative benchmark construction.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Autobencher: Towards declarative benchmark construction

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:38.098878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:38.098878Z digest=sha256:cb95e85f60edd7306def0539666309ebfa526f03d139e1a6f830c275fc39a889

Observation d244d9dd-3773-4826-ad02-5471a91be604 · outbound

This paper cites Gen- dataagent: On-the-fly dataset augmentation with synthetic data.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Gen- dataagent: On-the-fly dataset augmentation with synthetic data

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:47.979881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:03:38.227102Z digest=sha256:053d4986c360c59e2ed7d7bd1f75515a9d2426e4fd80d3d57eb968598a45f890

Observation 8da932b7-31ef-4c50-8c38-e62277f8c458 · outbound

This paper cites Hemm: Holistic evaluation of multimodal foundation models.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Hemm: Holistic evaluation of multimodal foundation models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:47.710055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:03:38.373411Z digest=sha256:666a951839dfda9554ed7705e5388fb7039057d8b2fbe5f9742088ee671960ee

Observation a89be644-4aab-44c0-92d4-34e813ff1c33 · outbound

This paper cites Holistic evaluation of language models.Transactions on Machine Learning Research, 2022.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Holistic evaluation of language models.Transactions on Machine Learning Research, 2022

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:47.397511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:03:38.517906Z digest=sha256:797b085af1eabe6f0f56a8e55eb63c82912b92fb00c63591db3339fcfe16a450

Observation 50e6751b-e4e9-4fd8-b1ca-b844c8542f33 · outbound

This paper cites Truthfulqa: Measuring how models mimic human falsehoods.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Truthfulqa: Measuring how models mimic human falsehoods

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:38.666791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:38.666791Z digest=sha256:9b1e4ed0d85de96071c687cfb2b1c9afbe5ffcb387a05091f6b5cc86b66b0ef9

Observation 34406dc8-cbb5-4a5b-a9e9-46bc8580c585 · outbound

This paper cites Best practices and lessons learned on synthetic data.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Best practices and lessons learned on synthetic data

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:47.205529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:03:38.854417Z digest=sha256:9f846084ec9c7bbe8e8ca00643bdbd651cab597df5888224e46909f8dfa51053

Observation 080387ab-ccae-423b-adc7-7451384d0e5d · outbound

This paper cites Fictitious Synthetic Data Can Improve LLM Factuality via Prerequisite Learning.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Fictitious Synthetic Data Can Improve LLM Factuality via Prerequisite Learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:39.014964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:39.014964Z digest=sha256:2e71af83e2257885382ab08f5435b5c985fbeaefc5504e094e63eaff89064ae6

Observation ec6e65cc-1c8c-4b78-b252-8c1ea0f7441f · outbound

This paper cites Adaptive labeling for efficient out-of-distribution model evaluation.Advances in Neural Information Processing Systems, 37:70981–71003, 2024.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Adaptive labeling for efficient out-of-distribution model evaluation.Advances in Neural Information Processing Systems, 37:70981–71003, 2024

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:46.853533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:03:39.157235Z digest=sha256:3371bb5f603c1b34ad9a08e3508d275ae8f5279fb71c4141a58d42226ab10129

Observation 93bb72f4-c315-4cbb-9533-827ac636cf70 · outbound

This paper cites Unearthing Skill-Level Insights for Understanding Trade-Offs of Foundation Models.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Unearthing Skill-Level Insights for Understanding Trade-Offs of Foundation Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:39.313256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:39.313256Z digest=sha256:c146a9c13858cb3bcf4ebffc87adf6cfcc754e397afa7c87552862968dde257a

Observation 4c594c94-5c1a-42e8-8e81-dcd7a06e0923 · outbound

This paper cites Enhancing reason- ing capabilities of llms via principled synthetic logic corpus.Advances in Neural Information Processing Systems, 37:73572–73604, 2024.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Enhancing reason- ing capabilities of llms via principled synthetic logic corpus.Advances in Neural Information Processing Systems, 37:73572–73604, 2024

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:46.425652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:03:39.451503Z digest=sha256:a9d5668afc28c040726ba1c1e541cc5527530112fe035d05f5201877d385a700

Observation 68588ff6-2b66-43e7-8058-4dc7072c5046 · outbound

This paper cites Caps: Collaborative and private synthetic data generation from distributed sources.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Caps: Collaborative and private synthetic data generation from distributed sources

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:46.206363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:03:39.570494Z digest=sha256:c5cffa614cfc4385fc3a1ac1803348ae9628bb7a5e10e19471d5b141bf840ba2

Observation c185db6c-5de2-4728-b983-20224c9e5697 · outbound

This paper cites Humanity's Last Exam.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Humanity's Last Exam

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:39.704349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:39.704349Z digest=sha256:6c6ee6e1768f6e4fc184a0e6859aae51a3a32759ca3e85e32bf4031367486fd0

Observation 1440dcac-f8d7-4683-9148-2dba3f03f29a · outbound

This paper cites Quantifying language models’ sensitivity to spurious features in prompt design or: How i learned to start worrying about prompt formatting.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Quantifying language models’ sensitivity to spurious features in prompt design or: How i learned to start worrying about prompt formatting

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:39.843361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:39.843361Z digest=sha256:b275e26669bce99d562b1869e91a7d91c7873f72cef4d521ca350f282ef9413c

Observation 4b304ac3-1d89-486f-a811-dd6974a27604 · outbound

This paper cites Explore Theory of Mind: Program-guided adversarial data generation for theory of mind reasoning.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Explore Theory of Mind: Program-guided adversarial data generation for theory of mind reasoning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:39.949056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:39.949056Z digest=sha256:6625904d224eae5a813b505b2bf28e29fea7f765cca2febffb49bcf1bbe6d1b4

Observation a82e3d99-1f46-439c-93e6-25414521ab33 · outbound

This paper cites Rl on incorrect synthetic data scales the efficiency of llm math reasoning by eight-fold.Advances in Neural Information Processing Systems, 37:43000–43031, 2024.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Rl on incorrect synthetic data scales the efficiency of llm math reasoning by eight-fold.Advances in Neural Information Processing Systems, 37:43000–43031, 2024

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:40.030301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:40.030301Z digest=sha256:9a4536f2d716d1a74a0a5f494067dce98ba2f51bb2b17e79d4df7bf805a00f04

Observation 2d635b6f-aa87-4f42-b0eb-f251f2f356b5 · outbound

This paper cites Is your benchmark truly adversarial? AdvScore: Evaluating Human-Grounded Adversarialness.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Is your benchmark truly adversarial? AdvScore: Evaluating Human-Grounded Adversarialness

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:03:42.954865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:03:40.130964Z digest=sha256:f2e6bd6433c9b94c99b9303f22555ae51a75462dc7e17af0bbef6b41692c32c1

Observation a4ceafbe-6bbf-488b-b76a-aeffe7a8c783 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Gemma 2: Improving Open Language Models at a Practical Size

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:40.248109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:40.248109Z digest=sha256:417b3110d3e27c08dd318c3dad9891134ae07b973d6936b97fdb365bf4e065cc

Observation 83c22784-02cf-4af9-8e02-6683f78e080e · outbound

This paper cites Measuring general intelligence with generated games, 2025.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Measuring general intelligence with generated games, 2025

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:40.350204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:40.350204Z digest=sha256:528af3f54956554fe1d1dd947eafcd5eebe77d53ebea533a847f19388f7211a0

Observation 287aab16-f8da-4da7-aaa2-a1a53c8a67fe · outbound

This paper cites Cuts: Customizable tabular synthetic data generation.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Cuts: Customizable tabular synthetic data generation

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:45.994917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:03:40.455814Z digest=sha256:1b9b8fe5cd7886e06b1b9f3bad81da4218fcd915473458c4f78ad13c5e7d3f7a

Observation fa9c1b34-5dae-480e-8558-4dea15fe50b3 · outbound

This paper cites The Power of LLM-Generated Synthetic Data for Stance Detection in Online Political Discussions.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data The Power of LLM-Generated Synthetic Data for Stance Detection in Online Political Discussions

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:03:42.726856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:03:40.514127Z digest=sha256:5ad3989de03d50c952e88bb6cc5a48afa5c7e05533f76bc091e113db21594941

Observation 12ecc956-b073-4191-9415-7a6d530c350f · outbound

This paper cites Glue: A multi-task benchmark and analysis platform for natural language understanding.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Glue: A multi-task benchmark and analysis platform for natural language understanding

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:45.805852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:03:40.582255Z digest=sha256:dc00324926d4b0fdd64ce84ede979737b9d67f0164a4608784994218d69e0eea

Observation e5220f4a-7129-497a-b745-4f8170da5201 · outbound

This paper cites Can language models solve graph problems in natural language?Advances in Neural Information Processing Systems, 36:30840–30861, 2023.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Can language models solve graph problems in natural language?Advances in Neural Information Processing Systems, 36:30840–30861, 2023

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:45.621047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:03:40.660003Z digest=sha256:55755116974dd561856975c611e4c4cff59415363578bcad7c6ca604a7575b90

Observation b7cc40a9-d296-4b25-ac97-2c099f57c4c4 · outbound

This paper cites Pandalm: An automatic evaluation benchmark for llm instruction tuning optimization.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Pandalm: An automatic evaluation benchmark for llm instruction tuning optimization

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:45.358576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:03:40.729147Z digest=sha256:19d22aedd9d1d4cb521c1a0d216a784e19a705fc7e17041e1d3d5c6047873b8a

Observation 1064b364-68fa-45f1-9753-a3f9176832e9 · outbound

This paper cites Self-instruct: Aligning language models with self-generated in- structions.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Self-instruct: Aligning language models with self-generated in- structions

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:40.827284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:40.827284Z digest=sha256:173ce9464f1a626bae7b52b0153ba9dde2aa1167d279ce5a5f31cd50de3ba807

Observation f672e44b-eda7-4f67-8007-da3b6ac579dc · outbound

This paper cites Pre-training with synthetic data helps offline reinforcement learning.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Pre-training with synthetic data helps offline reinforcement learning

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:45.214965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:03:40.925782Z digest=sha256:81c38437029a868b1eee995e6959339f4c770006c26ff5990afff253766d08c5

Observation 2ae4d87d-0c52-43f4-8558-51801c4ca646 · outbound

This paper cites Rocketeval: Efficient automated llm evaluation via grading checklist.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Rocketeval: Efficient automated llm evaluation via grading checklist

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:45.017392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:03:40.994839Z digest=sha256:830bd5147c8e66252f59c2fbd9306962d9cb9c394c4acdbe8a69fc207c2e1485

Observation a5a81aa4-3368-492c-8936-94765473778d · outbound

This paper cites Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:41.104348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:41.104348Z digest=sha256:b1628699a62d90c3615852d14111f2bd0c059cf6db5be40fd1b4d635e7435d64

Observation 5a54daff-d92d-42da-95e8-c21d034ee6d0 · outbound

This paper cites Differentially private synthetic data via foundation model apis 2: Text.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Differentially private synthetic data via foundation model apis 2: Text

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:44.837302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:03:41.220080Z digest=sha256:8906a688957dccb732bb7cee737f19264765bdaa1e8888079f570d143e2593cb

Observation d570d6b4-12d0-49a0-b1d7-6ca0dab74574 · outbound

This paper cites Automating dataset updates towards reliable and timely evaluation of large language models.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Automating dataset updates towards reliable and timely evaluation of large language models

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:44.651043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:03:41.331758Z digest=sha256:cdb185e6c7973f21700c68a10d77c46d047bcd0bc94844d8fd2b9872a728af55

Observation 22db04ed-bee3-4fd8-a2eb-5c6de61b3b6e · outbound

This paper cites xfinder: Large language models as automated evaluators for reliable evaluation.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data xfinder: Large language models as automated evaluators for reliable evaluation

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:44.481634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:03:41.402978Z digest=sha256:b1ead2f404a7b59c690f00fec692380b1791a87545b2bd7ceb04c229b31ed685

Observation 36e675ce-81da-4eaf-a022-33e188051eeb · outbound

This paper cites Hellaswag: Can a machine really finish your sentence? InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4791–4800, 2019.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Hellaswag: Can a machine really finish your sentence? InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4791–4800, 2019

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:41.541623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:41.541623Z digest=sha256:ec3f6340e69f25adcd3f96fc8505bce053143d6da16bb122659e5e61c9c66f9c

Observation e5df9015-eeba-493d-a863-5c1d497aa6a3 · outbound

This paper cites EvalTree: Profiling Language Model Weaknesses via Hierarchical Capability Trees.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data EvalTree: Profiling Language Model Weaknesses via Hierarchical Capability Trees

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:41.647357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:41.647357Z digest=sha256:1637cd44616207c5fcd04b1b79d1bcf134ecd8281ad5283784d7448d471202ff

Observation 421897cc-4d5d-4cfe-90cd-ce34ec6505a1 · outbound

This paper cites Bidirectional LMs are Better Knowledge Memorizers? A Benchmark for Real-world Knowledge Injection.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Bidirectional LMs are Better Knowledge Memorizers? A Benchmark for Real-world Knowledge Injection

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:41.749831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:41.749831Z digest=sha256:7fd1e51db040d378ad0df2d2b56314c2723e717945faaf6536934d7dc949017f

Observation 7fff0ab4-4543-462f-9808-82a749400c7f · outbound

This paper cites Wildchat: 1m chatgpt interaction logs in the wild.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Wildchat: 1m chatgpt interaction logs in the wild

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:44.304019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:03:41.852941Z digest=sha256:2bcce9353edae3cb8c2028b38df1acd69bf6cc0a8ed10ae1c2dd95881dd50015

Observation 523b8033-85f6-4e45-b746-d0638c7ad679 · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in Neural Information Processing Systems, 36:46595–46623, 2023.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in Neural Information Processing Systems, 36:46595–46623, 2023

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:41.967706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:41.967706Z digest=sha256:65f070b851f8ffecb34d64d8be83fa5e41d5039395bda3c97f6d18252f428071

Observation c2b40608-3786-4bce-bc64-3ffb03f02354 · outbound

This paper cites Lima: Less is more for alignment.Advances in Neural Information Processing Systems, 36:55006–55021, 2023.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Lima: Less is more for alignment.Advances in Neural Information Processing Systems, 36:55006–55021, 2023

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:44.147466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:03:42.068582Z digest=sha256:8a0821c97cfcf83939df75535d90e073b5f105f582c9364d79f54a410252f23c

Observation 3f59b8ae-51c8-4f18-bccd-539b6abfc1b2 · outbound

This paper cites Sotopia: Interactive evaluation for social intelligence in language agents.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Sotopia: Interactive evaluation for social intelligence in language agents

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:43.981981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:03:42.161531Z digest=sha256:2c531bb3188fb03a3a6ea6228dad5d68df87811253faefb8283b18a3e64d5b02

Observation 49d8ee22-76eb-4750-a158-8d9a5de1bc00 · outbound

This paper cites Dyval: Dynamic evaluation of large language models for reasoning tasks.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Dyval: Dynamic evaluation of large language models for reasoning tasks

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:43.812704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:03:42.262166Z digest=sha256:d3bc8600559c2b6b2e64378c7583083db0a93560d80303f99b57941b659d51c6

Observation 170741c2-449c-4e06-b467-adda07e6e7b4 · outbound

This paper cites Dynamic evaluation of large language models by meta probing agents.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Dynamic evaluation of large language models by meta probing agents

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:43.645456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:03:42.372390Z digest=sha256:b4689b39b9d509ee4402ef386fe85104cc2d518cd4e1c67874da54277c345255

Observation 09208314-afdc-414f-8d4a-dc9ae8757df8 · outbound

This paper cites Top Secret.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Top Secret

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:43.487172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:03:42.484850Z digest=sha256:57320d79087103ab95ea24dfc5934150c69d4a926477c46f652fc7188bcd072b

Pith citing papers

No inbound Pith citation observations are available.