Pith. sign in

Paper Citation Record · LEDGER

EraserDiT: Fast Video Inpainting with Diffusion Transformer Model

As of 10 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 2 inbound Pith citation observations for arXiv:2506.12853.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.12853 v2

Coverage vector

measured 17 of 17 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:42:01.742085Z

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-29T08:06:27.191812Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T08:13:15.764565Z

Reference resolution

17 of 17 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3483ca62-74a1-4e0d-84ba-3c21b377a598 · outbound

This paper cites write newline.

EraserDiT: Fast Video Inpainting with Diffusion Transformer Model write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T00:42:00.389964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:42:00.389964Z digest=sha256:4c375eb00cdd63b8ed0510778c542ae72544dd57513bf93695d529002239dd68

Observation 53f16d1d-dc0d-4c48-abef-6605b054644b · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

EraserDiT: Fast Video Inpainting with Diffusion Transformer Model An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T00:42:00.453980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:42:00.453980Z digest=sha256:5bff9ea45286363f3c084ed9222fbc4f2c8a47f248f4307bb784d86a353726f0

Observation 67c50ec6-a683-4cec-82d1-e7701c31d11e · outbound

This paper cites Scaling rectified flow transformers for high-resolution image synthesis.

EraserDiT: Fast Video Inpainting with Diffusion Transformer Model Scaling rectified flow transformers for high-resolution image synthesis

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T00:42:00.539849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:42:00.539849Z digest=sha256:2230d3e00908bd44d1c802e44df08faae52a955724b2b214344435f25aa3ecce

Observation 3e7f663e-c012-4bf3-9005-89fe071663b0 · outbound

This paper cites AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning.

EraserDiT: Fast Video Inpainting with Diffusion Transformer Model AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T00:42:00.591297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:42:00.591297Z digest=sha256:37bf4bb7b60e84f59867c179f46498caf90a9814cdff46a4f184aa05f7cf02a4

Observation ee97c058-2ea8-454b-a0ba-4012299d007b · outbound

This paper cites LTX-Video: Realtime Video Latent Diffusion.

EraserDiT: Fast Video Inpainting with Diffusion Transformer Model LTX-Video: Realtime Video Latent Diffusion

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T00:42:00.641831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:42:00.641831Z digest=sha256:7a4cfa31490c3abf8996878d7d50a515d2de9544ae342faab54652e2ee7273bf

Observation 8caa6034-e3c3-437b-b8fb-59371aa0ed6b · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

EraserDiT: Fast Video Inpainting with Diffusion Transformer Model HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T00:42:00.695623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:42:00.695623Z digest=sha256:4cb206bea8ed531f6cd323c6a5a30c6968379138b877fe7556f78290869cdc8d

Observation caec44e0-8720-4554-94b9-469e9107e8b2 · outbound

This paper cites DiffuEraser: A Diffusion Model for Video Inpainting.

EraserDiT: Fast Video Inpainting with Diffusion Transformer Model DiffuEraser: A Diffusion Model for Video Inpainting

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T00:42:00.744948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:42:00.744948Z digest=sha256:2ed4f5730c08bad3ac39b50f94064b77eaae1a0a4fae15e73b64e0b9d2b2189e

Observation bb869301-bdcf-47b5-bf86-f559623a30b7 · outbound

This paper cites Flow Matching for Generative Modeling.

EraserDiT: Fast Video Inpainting with Diffusion Transformer Model Flow Matching for Generative Modeling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T00:42:00.808194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:42:00.808194Z digest=sha256:db3b9718ade2ee6974d693a42741bea6163e2f2f33290577715998e329d0c1d5

Observation e518d056-4dfd-4cd7-adaa-13d0761ebfff · outbound

This paper cites Fuseformer: Fusing fine-grained information in transformers for video inpainting.

EraserDiT: Fast Video Inpainting with Diffusion Transformer Model Fuseformer: Fusing fine-grained information in transformers for video inpainting

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:42:02.904955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T00:42:00.901619Z digest=sha256:c020a1b4fa7db0a160a1d96ebc0037b3075e2a79e22ebafff68a967804064f4d

Observation 54dc4eef-3242-4a4a-bc00-e7799952ecbd · outbound

This paper cites A benchmark dataset and evaluation methodology for video object segmentation.

EraserDiT: Fast Video Inpainting with Diffusion Transformer Model A benchmark dataset and evaluation methodology for video object segmentation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T00:42:00.948590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:42:00.948590Z digest=sha256:ca21012bce9d23c8a8538fd62b1aacd3ed350837356c692e4ce863e46a375572

Observation 91ae787f-d58e-4b4f-b9fe-9062b07118ce · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

EraserDiT: Fast Video Inpainting with Diffusion Transformer Model Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T00:42:01.013801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:42:01.013801Z digest=sha256:0461c34ea8db14d15796725b4440c5e4f99b158834cfd65144d6407a98d96908

Observation 5d89be9f-e883-4ec5-82b5-ac9fe4c4b544 · outbound

This paper cites Sam 2: Segment anything in images and videos, 2024.

EraserDiT: Fast Video Inpainting with Diffusion Transformer Model Sam 2: Segment anything in images and videos, 2024

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T00:42:01.099668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:42:01.099668Z digest=sha256:a4430a0ec8e02857a536dbe6e68198e185f680049e6ac92019b99fb92fe0678f

Observation 75211975-9bbd-4e64-b978-3649f0ec11d7 · outbound

This paper cites Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks.

EraserDiT: Fast Video Inpainting with Diffusion Transformer Model Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T00:42:01.243679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:42:01.243679Z digest=sha256:419e3181077cab5ab6176c19f54da6549e9e1f64a525c8a9ea1e3043ccccb1a6

Observation 109b83bc-d7bc-42a0-a6de-a15e43cc009d · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

EraserDiT: Fast Video Inpainting with Diffusion Transformer Model CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T00:42:01.340005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:42:01.340005Z digest=sha256:6b8b546d723e814c4e5cbee658ec525583de860357f0ab2809d4db304e9d1185

Observation 2fef8977-4eda-493b-9afd-2474a34a432e · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

EraserDiT: Fast Video Inpainting with Diffusion Transformer Model MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T00:42:01.390837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:42:01.390837Z digest=sha256:8a7e85b17ba66b704e965bedb7b6ae314dea796f705d3359fe9088e15c769ff3

Observation a2dd6f99-90b5-4012-99ed-15d4a0f124de · outbound

This paper cites Learning joint spatial-temporal transformations for video inpainting.

EraserDiT: Fast Video Inpainting with Diffusion Transformer Model Learning joint spatial-temporal transformations for video inpainting

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:42:02.553991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T00:42:01.526322Z digest=sha256:a2dd503e41cae4969d76b2b400b2b3d8e2caf6cd28daeda12519d4b8cb97a809

Observation b668a2d2-770f-4aa4-8e00-51b4780f253c · outbound

This paper cites Propainter: Improving propagation and transformer for video inpainting.

EraserDiT: Fast Video Inpainting with Diffusion Transformer Model Propainter: Improving propagation and transformer for video inpainting

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:42:02.238566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T00:42:01.742085Z digest=sha256:5eb496d6c5c516099e608593ab22c5363b5b3789dfc76996c601935a4788222d

Pith citing papers

Observation 97933a7e-95b9-450b-a998-2f41a4d87ec0 · inbound

CLEAR: Context-Aware Learning with End-to-End Mask-Free Inference for Adaptive Video Subtitle Removal cites this paper.

CLEAR: Context-Aware Learning with End-to-End Mask-Free Inference for Adaptive Video Subtitle Removal EraserDiT: Fast Video Inpainting with Diffusion Transformer Model

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-15T00:39:35.679421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T00:39:15.790131Z digest=sha256:014c5bb0d1c20dfb28ace3d66f9ac2e662f5b92942f06135bbbc300db127be83

Observation c6e7c9ce-3d05-4f72-b280-60117090d07c · inbound

GenEraser: Generalizable Video Object Removal via Balanced Text-Mask Guidance and Decoupled Locator-Preserver cites this paper.

GenEraser: Generalizable Video Object Removal via Balanced Text-Mask Guidance and Decoupled Locator-Preserver EraserDiT: Fast Video Inpainting with Diffusion Transformer Model

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T08:13:15.766085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T08:06:27.191812Z digest=sha256:41e8e72940b687ddc590f925f01dbec8a2268e79c84e369793b32beae418db52