Pith. sign in

Paper Citation Record · LEDGER

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions

As of 9 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 0 inbound Pith citation observations for arXiv:2506.16679.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.16679 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:42:44.693479Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

45 of 45 outbound references displayed

  • verified exact0
  • verified fuzzy19
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b29b55f1-1eae-4431-8287-c32b58950ab5 · outbound

This paper cites eDiff-I: Text-to-Image Diffusion Models with an Ensemble of Expert Denoisers.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions eDiff-I: Text-to-Image Diffusion Models with an Ensemble of Expert Denoisers

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:39.205574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:39.205574Z digest=sha256:746d70c93576dffc45787fb00d9ea6e4c9e8e67f5d5f5eae40ff952806d46ede

Observation b7ccddc5-04e7-49b8-b370-194bce0d1623 · outbound

This paper cites Improving image generation with better captions, 2023.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions Improving image generation with better captions, 2023

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:39.318251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:39.318251Z digest=sha256:a15547c64ab16cfd8803aaf945542f6f04480a079ae627d361cd016f8e4a27c8

Observation e9c07c11-e186-4cf8-810b-83d1fd08c5fa · outbound

This paper cites Easily Ac- cessible Text-to-Image Generation Amplifies Demographic Stereotypes at Large Scale.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions Easily Ac- cessible Text-to-Image Generation Amplifies Demographic Stereotypes at Large Scale

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:39.446715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:39.446715Z digest=sha256:3b6b61d70ab87bdbd832e9261a0578fd64d81e8eb770e20309f5fa28660ac234

Observation 558f82cb-ca6b-4513-bf14-22cd904bf406 · outbound

This paper cites Se- mantics derived automatically from language corpora con- tain human-like biases.Science, 356(6334):183–186, 2017.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions Se- mantics derived automatically from language corpora con- tain human-like biases.Science, 356(6334):183–186, 2017

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:39.584832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:39.584832Z digest=sha256:1e44deb8d65498d52ded581ac224423b3c2d9bcd1ffdf1f33bc183f7942ca8f3

Observation 2155edb3-d6f1-4785-8840-a235f3b9fe30 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:39.794485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:39.794485Z digest=sha256:4a9704673f2fdd03c2e646e0d8b800cd7052a6b7123c3c463490aa51c5d1737e

Observation 6624f839-a4f3-4b3a-8d99-8774cb2628d5 · outbound

This paper cites Lmdeploy: A toolkit for com- pressing, deploying, and serving llm.https://github.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions Lmdeploy: A toolkit for com- pressing, deploying, and serving llm.https://github

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:39.904443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:39.904443Z digest=sha256:cb777549a45903c2701d7bc95e86fd62c920fa11beb83b061266b4fe4756f5b7

Observation 6a27f1dc-6ac3-4360-9fe3-16dc6ee37eab · outbound

This paper cites DeepFloyd-IF-I-XL-v1.0: DeepFloyd’s Image Generation Model, 2023.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions DeepFloyd-IF-I-XL-v1.0: DeepFloyd’s Image Generation Model, 2023

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:40.102939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:40.102939Z digest=sha256:ea18f398623963412f416eaf31640737210dc805e05a41b3f08d4ea1f3048f55

Observation d0220b11-71d3-4c05-8397-1d2acf8f482f · outbound

This paper cites Scaling rectified flow trans- formers for high-resolution image synthesis.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions Scaling rectified flow trans- formers for high-resolution image synthesis

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:40.255286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:40.255286Z digest=sha256:f51f58f11cb6de1a585aceff027f365306c47fa6681f8a5a32209b5eb17aa62c

Observation 347fc60b-73ce-4c28-b156-dd2933804507 · outbound

This paper cites Auditing and instructing text-to-image gener- ation models on fairness.AI and Ethics, 2024.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions Auditing and instructing text-to-image gener- ation models on fairness.AI and Ethics, 2024

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:40.405376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:40.405376Z digest=sha256:6713ce4a11c3ee09422143f05cb8331cffa2f1f5aa3ec6822182af90fce8a3d6

Observation a1c12695-5946-44e0-8b8d-c2ddea93614f · outbound

This paper cites Multilingual text-to-image generation magnifies gender stereotypes and prompt engineering may not help you, 2024.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions Multilingual text-to-image generation magnifies gender stereotypes and prompt engineering may not help you, 2024

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:40.495626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:40.495626Z digest=sha256:d9c6e442ad523f10ea42cadd07ff25e8523e78a5260bd2f4b6c2213f553b2d80

Observation 22dec8b3-0204-4ef4-89e8-31bd5a06cafd · outbound

This paper cites Datacomp: In search of the next generation of multimodal datasets.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions Datacomp: In search of the next generation of multimodal datasets

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:40.624921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:40.624921Z digest=sha256:1987da81d94a95ea2cb38259343a6c71a401ca470b09f5d389f8b8f0ddea9017

Observation 61b54fcd-4a5f-434d-944c-ae818c146561 · outbound

This paper cites CommonCanvas: An Open Diffusion Model Trained with Creative-Commons Images.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions CommonCanvas: An Open Diffusion Model Trained with Creative-Commons Images

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:40.734956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:40.734956Z digest=sha256:ed1e67fcb6621c624e53e78859800387f5841371fcddf81685509e45bfb8fe31

Observation 54200460-ec09-4521-bc29-7146b74c3f7b · outbound

This paper cites Hal- lusionbench: An advanced diagnostic suite for entangled language hallucination and visual illusion in large vision- language models.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions Hal- lusionbench: An advanced diagnostic suite for entangled language hallucination and visual illusion in large vision- language models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:40.846001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:40.846001Z digest=sha256:949843d6cd28990b7064f27b791a21fb99e372bbbf296e3a13ff352ec3bc0cf7

Observation 4b7918e8-bed1-449e-b904-6cf61361700b · outbound

This paper cites Benchmarking of deep architectures for segmentation of medical images.Trans.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions Benchmarking of deep architectures for segmentation of medical images.Trans

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:40.966618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:40.966618Z digest=sha256:b757ce33e669c706f838bb780d0ec837cdb55e5cb03d697335104dd17d38fa19

Observation fc7b575b-4bc4-4a28-8285-15bc81136e46 · outbound

This paper cites CLIPScore: a reference-free evaluation met- ric for image captioning.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions CLIPScore: a reference-free evaluation met- ric for image captioning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:41.194882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:41.194882Z digest=sha256:30313e71d6c0761dbda75e8b0c9e1b6f750cff6329c2c6252c1f8ef10e2745df

Observation c571d370-70d8-47c8-b310-bf761d777dbf · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilib- rium.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions Gans trained by a two time-scale update rule converge to a local nash equilib- rium

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:41.290350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:41.290350Z digest=sha256:c7f90e470264a0560c3e34afee07231669848ff21f86332cd7131047488787ca

Observation b095a204-ce0b-497f-80f7-b008bf11c865 · outbound

This paper cites Classifier-Free Diffusion Guidance.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions Classifier-Free Diffusion Guidance

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:41.417383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:41.417383Z digest=sha256:e128f8c57932006b7b568921500b33dbfb1fb7b95d35ef044426cac8e60ebdd9

Observation 81b22947-d3cf-4904-9aca-c6978eecb2f8 · outbound

This paper cites Fairface: Face at- tribute dataset for balanced race, gender, and age for bias measurement and mitigation.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions Fairface: Face at- tribute dataset for balanced race, gender, and age for bias measurement and mitigation

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:51.905389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:42:41.602235Z digest=sha256:60f605db0890a991b8436971da67eca9183b3703225814308a326d40634817ca

Observation 58cfc80b-e862-46cc-a7cb-63029ed7ea07 · outbound

This paper cites Transformers are minimax optimal nonparametric in-context learners.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions Transformers are minimax optimal nonparametric in-context learners

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:51.555707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:42:41.764753Z digest=sha256:b7bd05fa6b609f9685554784d944914e430d000c97d56293da3da1bcbfe8b4e8

Observation 0a02f6ce-6ae5-48ba-a14f-a37192251143 · outbound

This paper cites Pick-a-pic: An open dataset of user preferences for text-to-image generation.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions Pick-a-pic: An open dataset of user preferences for text-to-image generation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:51.178767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:42:41.934753Z digest=sha256:099f8b9059ae5ea19c029acda37fd231eeb2efb8e9b18324bbbb6a38dbed24ff

Observation 08b35d18-f633-4871-9784-f9114dbe2b77 · outbound

This paper cites Genai-bench: Evaluat- ing and improving compositional text-to-visual generation,.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions Genai-bench: Evaluat- ing and improving compositional text-to-visual generation,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:50.974970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:42:42.064986Z digest=sha256:a578c20ef3ec1d2e0a77d310156cfacfb6216ee9c1891e22cf25d97a8012c660

Observation da717749-599d-4916-9ecb-f93dbace5d54 · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:42.178165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:42.178165Z digest=sha256:1f5bb29c39231db78aa04c73fade11a28ac8f642d17a36718873ed0ff9d28957

Observation 6e38cb9e-9dfb-4235-9150-5a65b0a9f6fb · outbound

This paper cites Manmatha, Ashwin Swaminathan, Zhuowen Tu, Stefano Ermon, and Stefano Soatto.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions Manmatha, Ashwin Swaminathan, Zhuowen Tu, Stefano Ermon, and Stefano Soatto

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:50.455089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:42:42.375184Z digest=sha256:6b483667f999441d40ac26a43a70c311957a2526c26fb2a09ff05a3afebc250d

Observation 1b97e62d-a13d-48ba-a05f-7af49f06fea8 · outbound

This paper cites an unresolved cited work.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:42:50.064635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:42:42.455111Z digest=sha256:a1ea1336a90f66c2b1cb3d450659fffa6cea2d5525caf0448a967549ac07e668

Observation 908c64c8-763d-4626-b7b0-54265600e425 · outbound

This paper cites MVPTR: multi- level semantic alignment for vision-language pre-training via multi-stage learning.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions MVPTR: multi- level semantic alignment for vision-language pre-training via multi-stage learning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:49.814499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:42:42.544752Z digest=sha256:303b3a3104ae45ef0f95302c8c2d6c883985dc145465ca5b93214bbc7d001215

Observation d98e2388-d1ae-4244-977f-c4bbd7a40305 · outbound

This paper cites Playground v3: Im- proving text-to-image alignment with deep-fusion large lan- guage models, 2024.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions Playground v3: Im- proving text-to-image alignment with deep-fusion large lan- guage models, 2024

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:49.498510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:42:42.629743Z digest=sha256:84111951ea62a5e167fd05efa8d8bc5bc3db07727a2e7392c387175913334b0c

Observation 990f61d5-5c46-4105-ba7e-0679f2c7a42d · outbound

This paper cites Decoupled weight decay regularization.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions Decoupled weight decay regularization

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:49.261266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:42:42.715356Z digest=sha256:a5d7e04d96b8c5abd04ab4dad381a4ffec1cd5a50969c299ed768ff1cd2e3066

Observation 7b80c2bb-db38-4371-8bc0-d7b1e7857db7 · outbound

This paper cites DataDecide: How to Predict Best Pretraining Data with Small Experiments.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions DataDecide: How to Predict Best Pretraining Data with Small Experiments

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:42.799200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:42.799200Z digest=sha256:018d7996a805f5777e4e9c4ef23748eee5846578348c78db1e83591b9952a31c

Observation 5583e7a2-d656-4ab7-9bce-9fadce231b7f · outbound

This paper cites On aliased resizing and surprising subtleties in GAN evaluation.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions On aliased resizing and surprising subtleties in GAN evaluation

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:49.045221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:42:42.975038Z digest=sha256:f9d1a61056317ac7a8a2833c325138f75fdb3fb4d7c09bb7338c02b5b7910837

Observation a0f40c02-ffab-472f-9a2c-777a7742a149 · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:43.044829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:43.044829Z digest=sha256:01e4e3019144dc552eb87e5e95fb5a08aad5e8ac9bd4216329db1d3ce1e116ed

Observation 99aa93f8-c68d-42aa-87c6-b454a4a6f723 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions High-resolution image synthesis with latent diffusion models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:48.820391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:42:43.133552Z digest=sha256:cd21e8568383ae9e4fa4e57843e3f6d8ff897b339e512e2e85d6f9e1d2cd01a4

Observation 547bb8b8-6e95-42d6-a8f6-2dd81afe9eec · outbound

This paper cites Sara Mahdavi, Rapha Gontijo Lopes, Tim Salimans, Jonathan Ho, David J.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions Sara Mahdavi, Rapha Gontijo Lopes, Tim Salimans, Jonathan Ho, David J

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:48.515166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:42:43.199837Z digest=sha256:2e99cc3c1a5342f280b87087606efbe69cb810992cfca91f420db8ad369172ea

Observation 4dd22a3e-4cda-4a02-9453-0aa14946d9a1 · outbound

This paper cites Fast High-Resolution Image Synthesis with Latent Adversarial Diffusion Distillation.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions Fast High-Resolution Image Synthesis with Latent Adversarial Diffusion Distillation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:43.346847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:43.346847Z digest=sha256:d6214b97458a006fe87aba0cf327d521c674651b4231c4c5435d9643de4bc566

Observation 901e554f-1db5-4b46-a477-311ee7b789d0 · outbound

This paper cites Safe latent diffusion: Mitigating inappro- priate degeneration in diffusion models.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions Safe latent diffusion: Mitigating inappro- priate degeneration in diffusion models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:48.294749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:42:43.454840Z digest=sha256:2385847d3712de966657369f868ae7703e6204687fefd58f9c6acb89b89180fc

Observation fb5fa092-cb34-4479-b4f0-4d7e0a19c085 · outbound

This paper cites LAION- 400M: open dataset of clip-filtered 400 million image-text pairs.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions LAION- 400M: open dataset of clip-filtered 400 million image-text pairs

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:48.074836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:42:43.546387Z digest=sha256:6a2ccf3b38e95199e0c33827002b9cb533e508d8f25758032d1f60dbbea0a25e

Observation e9f375a3-88be-4dab-bf85-3746ad49ef47 · outbound

This paper cites Laion-5b: An open large-scale dataset for training next gen- eration image-text models.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions Laion-5b: An open large-scale dataset for training next gen- eration image-text models

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:47.803600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:42:43.727063Z digest=sha256:e52604ee7238c3bd5e47254c772f6ca7aa6f218238285e47af1107625343dfe6

Observation 99029762-b815-4202-903e-33770a0a3c76 · outbound

This paper cites The bias amplification paradox in text-to-image generation.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions The bias amplification paradox in text-to-image generation

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:47.536040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:42:43.854962Z digest=sha256:3764eff17bd0c63b2435f1bde8c7010471541c6cafdb03d9e512d3bb488325b2

Observation 716dcf68-43c4-4a15-97de-064c5d401727 · outbound

This paper cites Scaling autoregressive models for content-rich text-to-image generation.Transactions on Machine Learn- ing Research, 2022.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions Scaling autoregressive models for content-rich text-to-image generation.Transactions on Machine Learn- ing Research, 2022

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:47.312561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:42:43.941301Z digest=sha256:6ae85d6f12fd41730d870ecf00c348de1d736bb23e221b2c5c0bcf522ec80906

Observation 517e503f-be60-4ce8-938e-aaf294774c0c · outbound

This paper cites Sigmoid loss for language image pre-training.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions Sigmoid loss for language image pre-training

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:47.121589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:42:44.054838Z digest=sha256:931f4f4fbb41bb065d919370aeae74be84e74078d954bcbd4e522d2b28a6f61c

Observation c3150f97-83ce-48d9-837e-d5497f7ee3cc · outbound

This paper cites The unreasonable effectiveness of deep features as a perceptual metric.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions The unreasonable effectiveness of deep features as a perceptual metric

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:46.932291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:42:44.165791Z digest=sha256:a6f5ff104fc9172706f8c625d0ed1d8621a0dccb935ffb920ba6ec7f3e08111d

Observation 0aef22c2-3279-4cba-a8ce-a278e4b1b84e · outbound

This paper cites SGLang: Efficient Execution of Structured Language Model Programs.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions SGLang: Efficient Execution of Structured Language Model Programs

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:44.284962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:44.284962Z digest=sha256:6a2680018ae286a4a343ff686dc150392f0637c69e20d97560c1eca8eb33ed07

Observation e9e566db-2c49-4d4b-bd6b-68a5572cd0ec · outbound

This paper cites an unresolved cited work.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:42:46.644753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:42:44.394752Z digest=sha256:94d7f3b99f1afb00109c40f1730fca4258aae6790d4ba76b44b73b243bec3d4b

Observation ee55de2c-6c89-4dc3-8284-a4fb32739b76 · outbound

This paper cites an unresolved cited work.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:42:46.457415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:42:44.514755Z digest=sha256:1a0a5d829536531080d7dbe245920f124b1d41a06ee0897d5a9a0b3eac57b780

Observation 96d665c8-a8f2-4835-bba7-91f98df2851c · outbound

This paper cites an unresolved cited work.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:42:46.138421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:42:44.616284Z digest=sha256:0f18864be4232a0d4ba5c967e6209cecb70d7192bb1d22a62e6faa73b6415f03

Observation 9fa2384b-0944-42cf-99c9-b5e9f07da3c9 · outbound

This paper cites psychologist.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions psychologist

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:45.864900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:42:44.693479Z digest=sha256:115b515a423541faefd0c0ac7a26686ee9b2c574a65bd1d2d4ce5b9e1ccd45c3

Pith citing papers

No inbound Pith citation observations are available.