Pith. sign in

Paper Citation Record · LEDGER

E-MD3C: Taming Masked Diffusion Transformers for Efficient Zero-Shot Object Customization

As of 10 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 0 inbound Pith citation observations for arXiv:2502.09164.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.09164 v1

Coverage vector

measured 14 of 14 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T22:29:07.222964Z

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

14 of 14 outbound references displayed

  • verified exact2
  • verified fuzzy3
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7b57e698-9dab-40a5-ab65-93c744415817 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

E-MD3C: Taming Masked Diffusion Transformers for Efficient Zero-Shot Object Customization An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T22:29:07.168345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:29:07.168345Z digest=sha256:c20b3657cc054a71d4b43448587f4bb95deb6dfb1094155d6705234865259ebc

Observation c981bd9c-4c00-47aa-b3fa-6a8a49770912 · outbound

This paper cites Table 4: Parameters and Configs.

E-MD3C: Taming Masked Diffusion Transformers for Efficient Zero-Shot Object Customization Table 4: Parameters and Configs

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T22:29:07.452595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T22:29:07.213613Z digest=sha256:3c9e0a1a6f4cd5c5f95c5231210c0aac5ec5077a535673727c95481bd7b65920

Observation 4fde4436-0347-4095-b9d9-8304312d294e · outbound

This paper cites Classifier-Free Diffusion Guidance.

E-MD3C: Taming Masked Diffusion Transformers for Efficient Zero-Shot Object Customization Classifier-Free Diffusion Guidance

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T22:29:07.179565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:29:07.179565Z digest=sha256:5dec65b3d8e64500e5dd6fd11bb2a435db082b3a00d16fb4843146f8d6aeda44

Observation 95f03259-266a-43c0-ad38-043273007b9e · outbound

This paper cites Quality-aware Masked Diffusion Transformer for Enhanced Music Generation.

E-MD3C: Taming Masked Diffusion Transformers for Efficient Zero-Shot Object Customization Quality-aware Masked Diffusion Transformer for Enhanced Music Generation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T22:29:07.184390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:29:07.184390Z digest=sha256:aac8416d2317cf5475f1c5161b415a9b906a67829134c4a5526c9fb9c011ea35

Observation acefdf97-c10b-44ab-b4c0-26a16dab06a0 · outbound

This paper cites MDT-A2G: Exploring Masked Diffusion Transformers for Co-Speech Gesture Generation.

E-MD3C: Taming Masked Diffusion Transformers for Efficient Zero-Shot Object Customization MDT-A2G: Exploring Masked Diffusion Transformers for Co-Speech Gesture Generation

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-07T22:29:07.317167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T22:29:07.189731Z digest=sha256:869187e800356a29ad29683f2db539c50c0e60436379ad868c50eeaa63b13160

Observation 1dc0d800-1ea2-40fe-9add-beb7e385dcb2 · outbound

This paper cites X., Sun, J., Zhu, Y ., Kweon, I.

E-MD3C: Taming Masked Diffusion Transformers for Efficient Zero-Shot Object Customization X., Sun, J., Zhu, Y ., Kweon, I

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T22:29:07.469005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T22:29:07.195032Z digest=sha256:e524acd7182c8af96e19862bb6a46bc5094ac2b838346b5bb016d14f72f3a95b

Observation a5310589-ce07-41bd-87cb-440e456a07c7 · outbound

This paper cites Denoising Diffusion Implicit Models.

E-MD3C: Taming Masked Diffusion Transformers for Efficient Zero-Shot Object Customization Denoising Diffusion Implicit Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T22:29:07.204102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:29:07.204102Z digest=sha256:44ebd09a6cebd519612fce3c40b7001142a964745100e1e5468ad0a7389b8a1e

Observation 7de29935-77c8-4786-950c-1c5f863b867f · outbound

This paper cites an unresolved cited work.

E-MD3C: Taming Masked Diffusion Transformers for Efficient Zero-Shot Object Customization Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-07T22:29:07.437582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T22:29:07.218433Z digest=sha256:a448665be9a7c628776572f81fb37f5875f59b6040465723dcc5d8094027da01

Observation c11005bf-ab28-4721-9672-39015a79a6c7 · outbound

This paper cites We mainly use DINOv2, but the other options may be worth trying.

E-MD3C: Taming Masked Diffusion Transformers for Efficient Zero-Shot Object Customization We mainly use DINOv2, but the other options may be worth trying

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T22:29:07.422646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T22:29:07.222964Z digest=sha256:5331266553eb01f8783446aba598d176d821b74cb4c22476099be2c900b87fe3

Observation 2a61c3e7-74b7-492f-9360-5c487a3446ad · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

E-MD3C: Taming Masked Diffusion Transformers for Efficient Zero-Shot Object Customization BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T22:29:07.162667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:29:07.162667Z digest=sha256:09139d478689e7271cf7dce4c210193deb0a5119fd01353647c7f45c84bbaa39

Observation 75741d7b-c34c-41a6-a371-b8e7c276f273 · outbound

This paper cites Muse: Text-To-Image Generation via Masked Generative Transformers.

E-MD3C: Taming Masked Diffusion Transformers for Efficient Zero-Shot Object Customization Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T22:29:07.156820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:29:07.156820Z digest=sha256:86bbd97d9f3d2df4faef08f68f5eedb6e25c4775d6e4431d127e73c2d47a9fea

Observation 71607520-7fe8-4a01-8ac7-a1891f5b8704 · outbound

This paper cites On the Pros and Cons of Momentum Encoder in Self-Supervised Visual Representation Learning.

E-MD3C: Taming Masked Diffusion Transformers for Efficient Zero-Shot Object Customization On the Pros and Cons of Momentum Encoder in Self-Supervised Visual Representation Learning

Reference 2023

Resolution
verified exact
local_arxiv, observed 2026-08-07T22:29:07.294414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T22:29:07.199485Z digest=sha256:198e97e98480b414a84a7faf08c8a45ebd7166dbf78842d2791d76009bd8763a

Observation e99ac24d-f270-483a-aba1-91854f94b5bd · outbound

This paper cites An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion.

E-MD3C: Taming Masked Diffusion Transformers for Efficient Zero-Shot Object Customization An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T22:29:07.174165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:29:07.174165Z digest=sha256:793c0724207f00174ff55ae4bc1d49ccb62d0b416ea9b88f4ff6aab1e46e055d

Observation 38d3543f-9b06-46c9-9e90-0819d6ef26b8 · outbound

This paper cites YouTube-VOS: A Large-Scale Video Object Segmentation Benchmark.

E-MD3C: Taming Masked Diffusion Transformers for Efficient Zero-Shot Object Customization YouTube-VOS: A Large-Scale Video Object Segmentation Benchmark

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T22:29:07.208847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:29:07.208847Z digest=sha256:8b8007fd49f496d1f50a04966eaf0f6a7e0bcfc0e22a1567662600a7e411d7c4

Pith citing papers

No inbound Pith citation observations are available.