Pith. sign in

Paper Citation Record · LEDGER

EraserDiT: Fast Video Inpainting with Diffusion Transformer Model

As of 7 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 2 inbound Pith citation observations for arXiv:2506.12853.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.12853 v2

Coverage vector

measured 17 of 17 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:42:01.742085Z

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-29T08:06:27.191812Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T08:13:15.764565Z

Reference resolution

17 of 17 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3483ca62-74a1-4e0d-84ba-3c21b377a598 · outbound

This paper cites write newline.

EraserDiT: Fast Video Inpainting with Diffusion Transformer Model write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T00:42:00.389964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:42:00.389964Z digest=sha256:20141d5e7291c9a83549b263fca55c50c84ee19bc020d9e3fc4c9ca223f70dea

Observation 53f16d1d-dc0d-4c48-abef-6605b054644b · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

EraserDiT: Fast Video Inpainting with Diffusion Transformer Model An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T00:42:00.453980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:42:00.453980Z digest=sha256:d9ecb4c3729ddc49bcf9c34461a3e6650a98573e1ca8cbc743ca46a32baa8561

Observation 67c50ec6-a683-4cec-82d1-e7701c31d11e · outbound

This paper cites Scaling rectified flow transformers for high-resolution image synthesis.

EraserDiT: Fast Video Inpainting with Diffusion Transformer Model Scaling rectified flow transformers for high-resolution image synthesis

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T00:42:00.539849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:42:00.539849Z digest=sha256:ed8760787fcb509f400529eab7ab8f815ec30eb55b34896b565efb61cc5fdd48

Observation 3e7f663e-c012-4bf3-9005-89fe071663b0 · outbound

This paper cites AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning.

EraserDiT: Fast Video Inpainting with Diffusion Transformer Model AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T00:42:00.591297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:42:00.591297Z digest=sha256:ace137b66ba078e8dc738f542652cc9a1042d4e31f025549b6f20092cde11666

Observation ee97c058-2ea8-454b-a0ba-4012299d007b · outbound

This paper cites LTX-Video: Realtime Video Latent Diffusion.

EraserDiT: Fast Video Inpainting with Diffusion Transformer Model LTX-Video: Realtime Video Latent Diffusion

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T00:42:00.641831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:42:00.641831Z digest=sha256:8595000f46e2a3aab169106ed353a2be1dad636c3195513de75f838514d0083d

Observation 8caa6034-e3c3-437b-b8fb-59371aa0ed6b · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

EraserDiT: Fast Video Inpainting with Diffusion Transformer Model HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T00:42:00.695623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:42:00.695623Z digest=sha256:e2799888e2b96ba09f0a27483c9482cf30222e258798811baad23ef11d73fcb6

Observation caec44e0-8720-4554-94b9-469e9107e8b2 · outbound

This paper cites DiffuEraser: A Diffusion Model for Video Inpainting.

EraserDiT: Fast Video Inpainting with Diffusion Transformer Model DiffuEraser: A Diffusion Model for Video Inpainting

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T00:42:00.744948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:42:00.744948Z digest=sha256:389b2952f497d304105c1371bfb7f2e3c252aeeba779d33ec28e46c7152b0fba

Observation bb869301-bdcf-47b5-bf86-f559623a30b7 · outbound

This paper cites Flow Matching for Generative Modeling.

EraserDiT: Fast Video Inpainting with Diffusion Transformer Model Flow Matching for Generative Modeling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T00:42:00.808194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:42:00.808194Z digest=sha256:1a6077ffe26f2927dc18a6e22e0f1e4da978135c9e00edf30b6368f9643b65ee

Observation e518d056-4dfd-4cd7-adaa-13d0761ebfff · outbound

This paper cites Fuseformer: Fusing fine-grained information in transformers for video inpainting.

EraserDiT: Fast Video Inpainting with Diffusion Transformer Model Fuseformer: Fusing fine-grained information in transformers for video inpainting

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:42:02.904955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:42:00.901619Z digest=sha256:6a564483fcbe61638dbffd3c13c1199f4bb5e9de2606e1d3e952bc6f2a10f245

Observation 54dc4eef-3242-4a4a-bc00-e7799952ecbd · outbound

This paper cites A benchmark dataset and evaluation methodology for video object segmentation.

EraserDiT: Fast Video Inpainting with Diffusion Transformer Model A benchmark dataset and evaluation methodology for video object segmentation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T00:42:00.948590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:42:00.948590Z digest=sha256:34ada4f85e9f77b93f21007670ee9a01e56ae348657a34a7cac2aa1406ad3267

Observation 91ae787f-d58e-4b4f-b9fe-9062b07118ce · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

EraserDiT: Fast Video Inpainting with Diffusion Transformer Model Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T00:42:01.013801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:42:01.013801Z digest=sha256:7536eac8d263f05fdd3f8f1e7ea2e4b68e1501f2135ac4419610608298448f8d

Observation 5d89be9f-e883-4ec5-82b5-ac9fe4c4b544 · outbound

This paper cites Sam 2: Segment anything in images and videos, 2024.

EraserDiT: Fast Video Inpainting with Diffusion Transformer Model Sam 2: Segment anything in images and videos, 2024

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T00:42:01.099668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:42:01.099668Z digest=sha256:1c0ed6aa34c068f4c2230678489854be8551683651aee619470cf6b26058448c

Observation 75211975-9bbd-4e64-b978-3649f0ec11d7 · outbound

This paper cites Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks.

EraserDiT: Fast Video Inpainting with Diffusion Transformer Model Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T00:42:01.243679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:42:01.243679Z digest=sha256:593dab77ba13b14bf5a07ce2928d117ba1cd3089d5156867438b3ccaf41508bb

Observation 109b83bc-d7bc-42a0-a6de-a15e43cc009d · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

EraserDiT: Fast Video Inpainting with Diffusion Transformer Model CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T00:42:01.340005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:42:01.340005Z digest=sha256:607a536de88e241edeb92e51f49f71b095cd8b008c114fb8bf63e759eed213b8

Observation 2fef8977-4eda-493b-9afd-2474a34a432e · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

EraserDiT: Fast Video Inpainting with Diffusion Transformer Model MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T00:42:01.390837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:42:01.390837Z digest=sha256:30388facb9dd8098c6c79cc936dadb3e8f410773a96225dcb3db70219e594a7d

Observation a2dd6f99-90b5-4012-99ed-15d4a0f124de · outbound

This paper cites Learning joint spatial-temporal transformations for video inpainting.

EraserDiT: Fast Video Inpainting with Diffusion Transformer Model Learning joint spatial-temporal transformations for video inpainting

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:42:02.553991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:42:01.526322Z digest=sha256:89304ca1ea8d64bfc7e1b0919643ffb53d4cff92300dbf6a1751369e5c513df9

Observation b668a2d2-770f-4aa4-8e00-51b4780f253c · outbound

This paper cites Propainter: Improving propagation and transformer for video inpainting.

EraserDiT: Fast Video Inpainting with Diffusion Transformer Model Propainter: Improving propagation and transformer for video inpainting

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:42:02.238566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:42:01.742085Z digest=sha256:d13f27820acd001076de937aef73bb2c25c380e8b9f861e3e2ddc57f84701f48

Pith citing papers

Observation 97933a7e-95b9-450b-a998-2f41a4d87ec0 · inbound

CLEAR: Context-Aware Learning with End-to-End Mask-Free Inference for Adaptive Video Subtitle Removal cites this paper.

CLEAR: Context-Aware Learning with End-to-End Mask-Free Inference for Adaptive Video Subtitle Removal EraserDiT: Fast Video Inpainting with Diffusion Transformer Model

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-15T00:39:35.679421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T00:39:15.790131Z digest=sha256:5736bf760b264e747d8baef5f13a205758d6892e34f039b9ad93cb2e255890c2

Observation c6e7c9ce-3d05-4f72-b280-60117090d07c · inbound

GenEraser: Generalizable Video Object Removal via Balanced Text-Mask Guidance and Decoupled Locator-Preserver cites this paper.

GenEraser: Generalizable Video Object Removal via Balanced Text-Mask Guidance and Decoupled Locator-Preserver EraserDiT: Fast Video Inpainting with Diffusion Transformer Model

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T08:13:15.766085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T08:06:27.191812Z digest=sha256:551cfc48c8c670be9b14faf164b9a56666a2316584537b7795aa5d33ede10077