Pith. sign in

Paper Citation Record · LEDGER

Improving Generative Pre-Training: An In-depth Study of Masked Image Modeling and Denoising Models

As of 20 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 1 inbound Pith citation observation for arXiv:2412.19104.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.19104 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T01:02:56.130644Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T21:44:32.969001Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-10T21:44:33.449281Z

Reference resolution

45 of 45 outbound references displayed

  • verified exact0
  • verified fuzzy17
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 16711921-5997-4531-a563-f35eaefae475 · outbound

This paper cites Beit: Bert pre-training of image transformers.

Improving Generative Pre-Training: An In-depth Study of Masked Image Modeling and Denoising Models Beit: Bert pre-training of image transformers

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:02:56.657516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T01:02:55.954431Z digest=sha256:476ec8e66653a42d2be0b61bf104587bc8e71dc77f9e05690c326f62929a6ada

Observation 5087204a-062a-409a-b718-f7f3958eb2a1 · outbound

This paper cites End-to- end object detection with transformers.

Improving Generative Pre-Training: An In-depth Study of Masked Image Modeling and Denoising Models End-to- end object detection with transformers

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T01:02:55.959353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:02:55.959353Z digest=sha256:fa220df3bbc1719495595b75ace8ef876302c369a56e9ecfd2af4f927e1ac7b9

Observation 32281388-d623-4ce1-a379-50e416210459 · outbound

This paper cites Pre-trained image processing transformer.

Improving Generative Pre-Training: An In-depth Study of Masked Image Modeling and Denoising Models Pre-trained image processing transformer

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:02:56.636932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T01:02:55.963595Z digest=sha256:1cde8d619e3e55a667c2d6bbeb793fe91da9491711ab37654e98aa94686aa512

Observation 2701d573-baa7-4bc1-bdf2-a5e256a7d246 · outbound

This paper cites A simple framework for contrastive learning of visual representations.

Improving Generative Pre-Training: An In-depth Study of Masked Image Modeling and Denoising Models A simple framework for contrastive learning of visual representations

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:02:56.622645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T01:02:55.968274Z digest=sha256:c7613d21454e5b4fd56d0041758a7f54c5d6f4a01325731a704a0f9525f4eb62

Observation 1396c092-74d4-48ce-aff1-55db76de6ef9 · outbound

This paper cites Context autoencoder for self- supervised representation learning.

Improving Generative Pre-Training: An In-depth Study of Masked Image Modeling and Denoising Models Context autoencoder for self- supervised representation learning

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:02:56.610732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T01:02:55.972787Z digest=sha256:8af6970f6d799b16ccdbe2868e68d012f444d6d085acdcc17b7baa4e1bb3fda9

Observation d7d51be8-6c58-424d-809a-3ea35b81b605 · outbound

This paper cites Deconstructing Denoising Diffusion Models for Self-Supervised Learning.

Improving Generative Pre-Training: An In-depth Study of Masked Image Modeling and Denoising Models Deconstructing Denoising Diffusion Models for Self-Supervised Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T01:02:55.977251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:02:55.977251Z digest=sha256:49c92fa541a943e04273b1440f5e990c565cc664f7a867aba34269e22cf5e9f0

Observation b5cd8573-1a21-4279-8acb-f726d32b164e · outbound

This paper cites Emerging Property of Masked Token for Effective Pre-training.

Improving Generative Pre-Training: An In-depth Study of Masked Image Modeling and Denoising Models Emerging Property of Masked Token for Effective Pre-training

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T01:02:55.982018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:02:55.982018Z digest=sha256:9b6450d5c964d18fbbe9e6f325766ea0cec6fdf2524b0e030c7ec3bd7e2951ea

Observation 61993712-0102-4c2a-a66e-f00572367c43 · outbound

This paper cites Salience-based adaptive masking: revisit- ing token dynamics for enhanced pre-training.

Improving Generative Pre-Training: An In-depth Study of Masked Image Modeling and Denoising Models Salience-based adaptive masking: revisit- ing token dynamics for enhanced pre-training

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T01:02:55.986055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:02:55.986055Z digest=sha256:2d7703e060a59dbc8046df5b4a9ab072e756f2da40d1a0ac8cb9277c0acc9946

Observation 98186132-9faa-425b-9ccc-300f3065584a · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

Improving Generative Pre-Training: An In-depth Study of Masked Image Modeling and Denoising Models Imagenet: A large-scale hierarchical image database

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:02:56.592510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T01:02:55.990047Z digest=sha256:10126dc8e3204423a73272304081a04f75879ab121281f808618ccabb8a2e95a

Observation c85341db-e4bf-430b-9759-e2d35b77f496 · outbound

This paper cites Bootstrapped masked autoencoders for vision bert pretraining.

Improving Generative Pre-Training: An In-depth Study of Masked Image Modeling and Denoising Models Bootstrapped masked autoencoders for vision bert pretraining

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:02:56.580387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T01:02:55.993721Z digest=sha256:678540a9baf8f16402939484e4c548a8532db8fd7476188d45cdb482a8659ce4

Observation c4df558c-a581-40cd-b195-a7ad2a4ebc4c · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Improving Generative Pre-Training: An In-depth Study of Masked Image Modeling and Denoising Models An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T01:02:55.997475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:02:55.997475Z digest=sha256:8af0b74fb2193d22ed9e9e008d19222a8a8670abf41b4f0eb2535bba2ea80d0d

Observation 2bd72a33-0b08-491e-8170-f5dd86aa2b06 · outbound

This paper cites Bootstrap your own latent-a new approach to self-supervised learning.

Improving Generative Pre-Training: An In-depth Study of Masked Image Modeling and Denoising Models Bootstrap your own latent-a new approach to self-supervised learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T01:02:56.001518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:02:56.001518Z digest=sha256:e976883d5b19b1b09cce02ca65e73308f579c925e9594804af68c3a77a8a57a9

Observation bc3a8e35-3d52-4ab4-9de4-f16633a7ee7b · outbound

This paper cites Momentum contrast for unsupervised visual rep- resentation learning.

Improving Generative Pre-Training: An In-depth Study of Masked Image Modeling and Denoising Models Momentum contrast for unsupervised visual rep- resentation learning

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:02:56.555082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T01:02:56.005316Z digest=sha256:abd85376ffd0f026a5d3ba87eb2eea36d865e957e5f8eb73a12ed7940dccfcd0

Observation 86a9b906-3589-472f-8b84-a44b2e2780b5 · outbound

This paper cites Masked autoencoders are scalable vision learners.

Improving Generative Pre-Training: An In-depth Study of Masked Image Modeling and Denoising Models Masked autoencoders are scalable vision learners

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:02:56.542971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T01:02:56.008566Z digest=sha256:308cc5e64fe208ccc7d6a7bb179576fa8b31d68a211907620e9b19e138b1bb84

Observation 4c025cd7-2c8b-41aa-ae69-cc942aacf501 · outbound

This paper cites Unsupervised keypoints from pretrained diffusion models.

Improving Generative Pre-Training: An In-depth Study of Masked Image Modeling and Denoising Models Unsupervised keypoints from pretrained diffusion models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:02:56.530117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T01:02:56.011935Z digest=sha256:a9d4395a8120afc34fc97eb30e9adb315f146a37768c79ed21e678598bc28010

Observation b592c68a-e5d3-41c6-bbc6-8c102808af91 · outbound

This paper cites Unsupervised semantic correspondence using stable diffu- sion.

Improving Generative Pre-Training: An In-depth Study of Masked Image Modeling and Denoising Models Unsupervised semantic correspondence using stable diffu- sion

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:02:56.515903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T01:02:56.015246Z digest=sha256:400b1a8b6997200fbca414eb900b6dc13e9bd2545c1b367e3c7b0534b40cc811

Observation 291338fb-aadd-448b-a116-f7d26e922394 · outbound

This paper cites Denoising dif- fusion probabilistic models.

Improving Generative Pre-Training: An In-depth Study of Masked Image Modeling and Denoising Models Denoising dif- fusion probabilistic models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T01:02:56.018534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:02:56.018534Z digest=sha256:c9687148fe8eb6e687bf236665cc82f4d4ffb16483c039f3246fbc0c95aecbbe

Observation b658b922-b7e5-443e-9cf8-226ffe8eb353 · outbound

This paper cites Simpler Diffusion (SiD2): 1.5 FID on ImageNet512 with pixel-space diffusion.

Improving Generative Pre-Training: An In-depth Study of Masked Image Modeling and Denoising Models Simpler Diffusion (SiD2): 1.5 FID on ImageNet512 with pixel-space diffusion

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T01:02:56.021927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:02:56.021927Z digest=sha256:402293b6efd1d0e7e527c8b3fc7cb69e857d0a74e7f1f4a23264e248c3a640cb

Observation 23cb11a5-a493-41a2-a8c4-70b7c6d91670 · outbound

This paper cites Elucidating the design space of diffusion-based generative models.

Improving Generative Pre-Training: An In-depth Study of Masked Image Modeling and Denoising Models Elucidating the design space of diffusion-based generative models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T01:02:56.026556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:02:56.026556Z digest=sha256:de1672b6d84d26bbaed9a12ffe8aec550ac5641b10b0aade61e5d5f9c7c0eae6

Observation be2c7a95-e086-42a5-b3d6-133558da4569 · outbound

This paper cites 3d object representations for fine-grained categorization.

Improving Generative Pre-Training: An In-depth Study of Masked Image Modeling and Denoising Models 3d object representations for fine-grained categorization

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:02:56.486692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T01:02:56.030867Z digest=sha256:a2f21c8d5b30000f1fcd73d55d669448897dd6a88936099640855105c2f6d359

Observation ae5d13f3-306d-4e1d-a894-85f8d185d5df · outbound

This paper cites Microsoft coco: Common objects in context.

Improving Generative Pre-Training: An In-depth Study of Masked Image Modeling and Denoising Models Microsoft coco: Common objects in context

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:02:56.471380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T01:02:56.035619Z digest=sha256:12a336a66ae73ab97b4334df290c6b5106a7a609e8166d1a842d0dee41aa4d40

Observation 94fbd934-087b-4a80-bf16-4acef01956f1 · outbound

This paper cites Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow.

Improving Generative Pre-Training: An In-depth Study of Masked Image Modeling and Denoising Models Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T01:02:56.039956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:02:56.039956Z digest=sha256:783b712bc8777be969065c7c251c8a2e2d1cc1cec81a7a4e03c774b1b037ebd3

Observation f1a503a7-af4d-4f3f-9d0a-1a8998f08085 · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

Improving Generative Pre-Training: An In-depth Study of Masked Image Modeling and Denoising Models Swin transformer: Hierarchical vision transformer using shifted windows

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T01:02:56.043992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:02:56.043992Z digest=sha256:0465ddd46fc916becbdaee7ec2c74a98fab4277244441e7149d47dbfe9b230b7

Observation 5c7668bf-1b5c-4d68-a530-4310e4b1c7c0 · outbound

This paper cites Dpm-solver: A fast ode solver for diffusion probabilistic model sampling in around 10 steps.

Improving Generative Pre-Training: An In-depth Study of Masked Image Modeling and Denoising Models Dpm-solver: A fast ode solver for diffusion probabilistic model sampling in around 10 steps

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T01:02:56.047752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:02:56.047752Z digest=sha256:ab7477c13e97a4a71ac72bc9df977ee11d8cd40948fde469032d960f25d0c612

Observation 7f8750ac-287b-4b24-b12c-c32d67ae9c1f · outbound

This paper cites Diffusion hyperfeatures: Searching through time and space for semantic correspondence.

Improving Generative Pre-Training: An In-depth Study of Masked Image Modeling and Denoising Models Diffusion hyperfeatures: Searching through time and space for semantic correspondence

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T01:02:56.051861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:02:56.051861Z digest=sha256:86e2dc4a424665ade881876e152fd02094492fc05e2af5aeb342ab33b696d077

Observation f9eacf2f-b69f-4cc6-9322-ce74d2f1a1b5 · outbound

This paper cites Fine-Grained Visual Classification of Aircraft.

Improving Generative Pre-Training: An In-depth Study of Masked Image Modeling and Denoising Models Fine-Grained Visual Classification of Aircraft

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T01:02:56.056244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:02:56.056244Z digest=sha256:f20e09e17340534415a0cb2f98f67a68802bf1d055875103ee255fea454264ba

Observation 612eb5ba-aade-47ff-9f1a-7232b6554df0 · outbound

This paper cites Improved denoising diffusion probabilistic models.

Improving Generative Pre-Training: An In-depth Study of Masked Image Modeling and Denoising Models Improved denoising diffusion probabilistic models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T01:02:56.060588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:02:56.060588Z digest=sha256:183ae7ac188844d3f3ab1fcc65402f48ca6266246b8e7f38dcfd8bbd6051b6c7

Observation 18bec5c4-2ecb-458b-8f3f-36eb1d70a5ad · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Improving Generative Pre-Training: An In-depth Study of Masked Image Modeling and Denoising Models Learning transferable visual models from natural language supervi- sion

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T01:02:56.064746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:02:56.064746Z digest=sha256:3a081be0a1655b548b859a5bf75bcc995011fe12a98f13863ce038a394a50420

Observation 3ec8e3e9-8654-46e0-9131-a11da3de7bde · outbound

This paper cites Zero-shot text-to-image generation.

Improving Generative Pre-Training: An In-depth Study of Masked Image Modeling and Denoising Models Zero-shot text-to-image generation

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:02:56.401192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T01:02:56.068865Z digest=sha256:6ed5404eb8dfefd86929690eeef681dec784f9b7c9fb68e2e71511d1c1f54619

Observation d0e1d641-e3ae-421f-935a-9f0e56efedbd · outbound

This paper cites Vi- sion transformers for dense prediction.

Improving Generative Pre-Training: An In-depth Study of Masked Image Modeling and Denoising Models Vi- sion transformers for dense prediction

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:02:56.385852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T01:02:56.072689Z digest=sha256:d413b3fcf4ea07ad5d5ba8067d2a2045550f609cb21974f0195f902f142a6e28

Observation dba26b97-fd97-4ae6-8257-6e7627cd1a2a · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Improving Generative Pre-Training: An In-depth Study of Masked Image Modeling and Denoising Models High-resolution image synthesis with latent diffusion models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T01:02:56.076514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:02:56.076514Z digest=sha256:8834ed4c44a1c6d23394f26811d2678858af629a8f809200402e9f80bb4ceacc

Observation 51f75e41-5357-4a8c-8600-e8ecf3d88f91 · outbound

This paper cites Photorealistic text-to-image diffusion models with deep language understanding.

Improving Generative Pre-Training: An In-depth Study of Masked Image Modeling and Denoising Models Photorealistic text-to-image diffusion models with deep language understanding

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T01:02:56.080252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:02:56.080252Z digest=sha256:c4b310decc53a849eb8651f5d7b94f17f59b5d80be629604d58c8aa23d09eeaa

Observation 57f56fab-e262-40a6-bff8-b6a7f45af10d · outbound

This paper cites LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs.

Improving Generative Pre-Training: An In-depth Study of Masked Image Modeling and Denoising Models LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T01:02:56.083835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:02:56.083835Z digest=sha256:03aa4b09e0e14c38058ee0e52bf442814da229936e9e54fa1248edbdf684852a

Observation 36742fdd-214e-455c-a1d9-bb9a22c92ee1 · outbound

This paper cites Denoising Diffusion Implicit Models.

Improving Generative Pre-Training: An In-depth Study of Masked Image Modeling and Denoising Models Denoising Diffusion Implicit Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T01:02:56.087361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:02:56.087361Z digest=sha256:7735f62300bca771f90e915152ecfffd502234357ecf517fd5e573a59afc04a1

Observation 91c850b1-cf7c-44a6-a098-f159ba2dc59a · outbound

This paper cites Building a bird recognition app and large scale dataset with citizen scientists: The fine print in fine-grained dataset collection.

Improving Generative Pre-Training: An In-depth Study of Masked Image Modeling and Denoising Models Building a bird recognition app and large scale dataset with citizen scientists: The fine print in fine-grained dataset collection

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T01:02:56.090773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:02:56.090773Z digest=sha256:c12ac8925240494c1426a7cf2a0ba2cf8093d28acd1011c2a7ad296ffc0ed8f3

Observation 268ff3d6-50e0-463a-9f0b-b1a741311cdd · outbound

This paper cites The iNaturalist Species Classification and Detection Dataset.

Improving Generative Pre-Training: An In-depth Study of Masked Image Modeling and Denoising Models The iNaturalist Species Classification and Detection Dataset

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T01:02:56.094114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:02:56.094114Z digest=sha256:f0cc46b80f24f2d483a8a78d3e3815670a333cf76ff7c3bb32dc189fabacd55b

Observation 68825b24-dc32-460c-a38f-6accbec73dfa · outbound

This paper cites The inaturalist species classification and de- tection dataset.

Improving Generative Pre-Training: An In-depth Study of Masked Image Modeling and Denoising Models The inaturalist species classification and de- tection dataset

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T01:02:56.097953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:02:56.097953Z digest=sha256:5ccd6169cc634b7cde4103e04253e2a49510f153238c1b5c370f80ab0fe3f209

Observation f51c4bde-3480-4fdb-a04a-1c68904bdb9b · outbound

This paper cites The caltech-ucsd birds-200-2011 dataset.

Improving Generative Pre-Training: An In-depth Study of Masked Image Modeling and Denoising Models The caltech-ucsd birds-200-2011 dataset

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:02:56.341954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T01:02:56.102152Z digest=sha256:6d2bba1d796d2d07ca7cf9c8004c132f5b280e844b07bebb1b8a9ab15fa01679

Observation f0e5fe92-855a-4733-980d-f7e7da21bf5c · outbound

This paper cites Diffusion models as masked autoencoders.

Improving Generative Pre-Training: An In-depth Study of Masked Image Modeling and Denoising Models Diffusion models as masked autoencoders

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:02:56.327716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T01:02:56.105987Z digest=sha256:f65f9b17e8d90d9d5b2407b98b4c4eca64a54b7cca2c5c741ebd6e6ff2b7166f

Observation 381dce87-cf43-4677-8762-d4e2ab1c558a · outbound

This paper cites Simmim: A simple framework for masked image modeling.

Improving Generative Pre-Training: An In-depth Study of Masked Image Modeling and Denoising Models Simmim: A simple framework for masked image modeling

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:02:56.314775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T01:02:56.109871Z digest=sha256:ae7ae759b92c9955bfbd85cb57c0ffce5194a34c3612a08c7e1341c310a9d1b0

Observation b41bdae4-4d42-4748-97e6-a984aa472455 · outbound

This paper cites Masked Image Modeling with Denoising Contrast.

Improving Generative Pre-Training: An In-depth Study of Masked Image Modeling and Denoising Models Masked Image Modeling with Denoising Contrast

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T01:02:56.113714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:02:56.113714Z digest=sha256:5743cc28341fe2d79c418d58b04387c75993077ebd24f2cbc6ef7bbccc09dc21

Observation 79d7d66e-c587-4e38-8cd1-03f8cdf3894c · outbound

This paper cites Fast Training of Diffusion Models with Masked Transformers.

Improving Generative Pre-Training: An In-depth Study of Masked Image Modeling and Denoising Models Fast Training of Diffusion Models with Masked Transformers

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T01:02:56.117688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:02:56.117688Z digest=sha256:ce251d8ae150ab55b891ade968d37d384469deb57d0a412157dfca3a13b84384

Observation 5ae084a7-dc13-400f-b74c-87767a6ab6d4 · outbound

This paper cites Rethinking semantic segmen- tation from a sequence-to-sequence perspective with trans- formers.

Improving Generative Pre-Training: An In-depth Study of Masked Image Modeling and Denoising Models Rethinking semantic segmen- tation from a sequence-to-sequence perspective with trans- formers

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T01:02:56.122105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:02:56.122105Z digest=sha256:a2c1ef96e0f90d42deb5a871ea39294023ca1f613657bc8f830003e0b3fe59d0

Observation dd94211a-ca3e-4b40-9297-b6e0e4920042 · outbound

This paper cites Scene parsing through ade20k dataset.

Improving Generative Pre-Training: An In-depth Study of Masked Image Modeling and Denoising Models Scene parsing through ade20k dataset

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T01:02:56.126490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:02:56.126490Z digest=sha256:028c5f390e391125329bfd8c5e84f56e60635979b2ce0837b1391d7264db27bd

Observation 1232b076-f48f-4cc3-a6b9-2e678fbfb910 · outbound

This paper cites Deformable DETR: Deformable Transformers for End-to-End Object Detection.

Improving Generative Pre-Training: An In-depth Study of Masked Image Modeling and Denoising Models Deformable DETR: Deformable Transformers for End-to-End Object Detection

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T01:02:56.130644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:02:56.130644Z digest=sha256:6faf7c5e3dd96087517c20841ce767036152391c767f8ba728cd20abd4ff26b7

Pith citing papers

Observation b261d85d-9469-407c-83b5-0640b749b916 · inbound

TADFormer : Task-Adaptive Dynamic Transformer for Efficient Multi-Task Learning cites this paper.

TADFormer : Task-Adaptive Dynamic Transformer for Efficient Multi-Task Learning Improving Generative Pre-Training: An In-depth Study of Masked Image Modeling and Denoising Models

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-10T21:44:33.453634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T21:44:32.969001Z digest=sha256:79ee82141a09ca254013d7a500da88d3dccec78bb6a8c8b1285941887f12abaa