Pith. sign in

Paper Citation Record · LEDGER

MADI: Masking-Augmented Diffusion with Inference-Time Scaling for Visual Editing

As of 18 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 0 inbound Pith citation observations for arXiv:2507.13401.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.13401 v1

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T16:48:34.038440Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

30 of 30 outbound references displayed

  • verified exact0
  • verified fuzzy7
  • unresolved22
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 40aa7432-0898-4b1f-b3ec-b7918a3ee569 · outbound

This paper cites an unresolved cited work.

MADI: Masking-Augmented Diffusion with Inference-Time Scaling for Visual Editing Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T16:48:29.723374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:48:29.723374Z digest=sha256:48bd93b15bad9e2ef5a435b57ba19e4b3c4770079c46cf6a1cfbddada1e44ae7

Observation 07764351-64d3-4d07-9710-39ec82cf5df2 · outbound

This paper cites Muse: Text-To-Image Generation via Masked Generative Transformers.

MADI: Masking-Augmented Diffusion with Inference-Time Scaling for Visual Editing Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T16:48:29.960911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:48:29.960911Z digest=sha256:bf96ddd47602a474bc455a927cae2d291e2b47026f9b90f56045bb61df705207

Observation 4f0d0e10-8936-4265-8514-462f8edd60b0 · outbound

This paper cites Maskgit: Masked generative image transformer.

MADI: Masking-Augmented Diffusion with Inference-Time Scaling for Visual Editing Maskgit: Masked generative image transformer

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:48:36.487936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T16:48:30.075783Z digest=sha256:d6ed25f9f241b6926e741443b01ee7922bb5fa87332a82641b231d15b5e448ad

Observation 1f70df94-920b-4d6f-ba6d-ad9e89f71182 · outbound

This paper cites Bert: Pre-training of deep bidi- rectional transformers for language understanding.

MADI: Masking-Augmented Diffusion with Inference-Time Scaling for Visual Editing Bert: Pre-training of deep bidi- rectional transformers for language understanding

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T16:48:30.199603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:48:30.199603Z digest=sha256:2963ed6ec9434732ce554c1d07545b486459f779656a78048ebcecf9bf33d5a1

Observation 3bc04cb4-535b-4f8b-bc66-7c3d66b9c0df · outbound

This paper cites Got: Unleashing reasoning capability of multimodal large language model for visual generation and editing, 2025.

MADI: Masking-Augmented Diffusion with Inference-Time Scaling for Visual Editing Got: Unleashing reasoning capability of multimodal large language model for visual generation and editing, 2025

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:48:36.261778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T16:48:30.337360Z digest=sha256:8c2c1e2978231721f75ec84ad7a4979e741e71a75bddab6cde82bb428255dab7

Observation 3f997ebe-b0a6-493b-bacd-890a5167ea63 · outbound

This paper cites Geneval: An object-focused framework for evaluating text-to-image alignment, 2023.

MADI: Masking-Augmented Diffusion with Inference-Time Scaling for Visual Editing Geneval: An object-focused framework for evaluating text-to-image alignment, 2023

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T16:48:30.466238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:48:30.466238Z digest=sha256:d9dbd9e7087fd6b047656669ccadb6c740862b85410532f0684e8a0427a5e037

Observation a06ac850-4280-4fab-9f07-12874da7d55f · outbound

This paper cites Think before you speak: Training Language Models With Pause Tokens.

MADI: Masking-Augmented Diffusion with Inference-Time Scaling for Visual Editing Think before you speak: Training Language Models With Pause Tokens

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T16:48:30.625805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:48:30.625805Z digest=sha256:f7b6885ef0f92e617e15c053d0b0217f74199c8fabdec3a45fd1777d5200a34d

Observation 2ea94b04-339f-4b24-b1a7-8745c2666fef · outbound

This paper cites Masked autoencoders are scalable vision learners.

MADI: Masking-Augmented Diffusion with Inference-Time Scaling for Visual Editing Masked autoencoders are scalable vision learners

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T16:48:30.762790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:48:30.762790Z digest=sha256:d816dfc1110b8f628241cccbb83988a5e248a238901234f23a523dc22919256c

Observation b6b3c055-c2ef-4540-89f1-a02462553536 · outbound

This paper cites Denoising diffusion probabilistic models, 2020.

MADI: Masking-Augmented Diffusion with Inference-Time Scaling for Visual Editing Denoising diffusion probabilistic models, 2020

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T16:48:30.895115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:48:30.895115Z digest=sha256:7e959433cbca28c2dbfc3be4652dad2b4e9f3e6d3f4cf14f38073c9e1191730a

Observation d48de95d-3f88-46fe-9146-6531635ffd86 · outbound

This paper cites [MASK] is All You Need.

MADI: Masking-Augmented Diffusion with Inference-Time Scaling for Visual Editing [MASK] is All You Need

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T16:48:31.037163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:48:31.037163Z digest=sha256:0762a496c18b07ccd581bfe46898292b4bd71fd006d74f90d4d33325e4b87ca2

Observation 79961ac2-5dae-4ca5-a46b-1789267e4633 · outbound

This paper cites Learning action and reasoning-centric image editing from videos and simulations, 2024.

MADI: Masking-Augmented Diffusion with Inference-Time Scaling for Visual Editing Learning action and reasoning-centric image editing from videos and simulations, 2024

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:48:35.946212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T16:48:31.317425Z digest=sha256:6f07018d4022d8d68599c0e1d835dc43a4291be48a13a4f1cccd4373c57490ec

Observation a4587575-9a8b-40b7-8d54-b84b02fe8bb1 · outbound

This paper cites IDEA-Bench: How Far are Generative Models from Professional Designing?.

MADI: Masking-Augmented Diffusion with Inference-Time Scaling for Visual Editing IDEA-Bench: How Far are Generative Models from Professional Designing?

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T16:48:31.473458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:48:31.473458Z digest=sha256:322e458d6c740acbdc0d4bf0568a86d8c1f27c3f1d4343b713774f6b56727137

Observation 1ef9e1ca-9511-49b7-9078-440483fe6ff7 · outbound

This paper cites Show your work: Scratchpads for intermediate computation with language models.

MADI: Masking-Augmented Diffusion with Inference-Time Scaling for Visual Editing Show your work: Scratchpads for intermediate computation with language models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T16:48:31.622705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:48:31.622705Z digest=sha256:1dd1e9c4eefdf80f6d8a656c848752f68c29f934b45f78b7849ffaab78493b17

Observation 16ce3eac-24a5-4948-8e64-b59f1362e2f1 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

MADI: Masking-Augmented Diffusion with Inference-Time Scaling for Visual Editing DINOv2: Learning Robust Visual Features without Supervision

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T16:48:31.785039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:48:31.785039Z digest=sha256:17509597c68631c261f8321bf334518c4df7faccba1b68ac8c07e1f337b1b037

Observation bfb6dff6-9d3d-4afe-9f2a-69bc43bf3ab6 · outbound

This paper cites Cogcom: A visual language model with chain-of-manipulations reasoning.

MADI: Masking-Augmented Diffusion with Inference-Time Scaling for Visual Editing Cogcom: A visual language model with chain-of-manipulations reasoning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T16:48:31.920475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:48:31.920475Z digest=sha256:ea552553cfb59d83ab56fdba3b38dacf0eb4ce3bb6bf736863d20af14b57d08c

Observation d1757b3d-0c9a-4f5d-89e7-9679c5766510 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

MADI: Masking-Augmented Diffusion with Inference-Time Scaling for Visual Editing High-resolution image synthesis with latent diffusion models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T16:48:32.050994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:48:32.050994Z digest=sha256:d300d0f674173c0eb84160ae081eb6eac0436a6fc619fe284e546fdc0fe1c866

Observation 7e4e583e-a1c9-46e0-ad27-16f7a2b88cde · outbound

This paper cites Emu edit: Precise image editing via recognition and generation tasks.

MADI: Masking-Augmented Diffusion with Inference-Time Scaling for Visual Editing Emu edit: Precise image editing via recognition and generation tasks

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T16:48:32.208981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:48:32.208981Z digest=sha256:c76e9a91e24b682cd58530dc1b5938116b8d4140f7a1f39badc95be1c5017c73

Observation 1f67c3cd-25fd-407c-824d-030d14b09bf6 · outbound

This paper cites Seededit: Align image re-generation to image editing, 2024.

MADI: Masking-Augmented Diffusion with Inference-Time Scaling for Visual Editing Seededit: Align image re-generation to image editing, 2024

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:48:35.606020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T16:48:32.376334Z digest=sha256:5a3d5139f770660ae3cad2a4b156334c318833cd8c73cc2fb79b9df3fd2b77a7

Observation f0b5e86f-e786-4a8e-b815-749978105f92 · outbound

This paper cites Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole.

MADI: Masking-Augmented Diffusion with Inference-Time Scaling for Visual Editing Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:48:35.283366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T16:48:32.531370Z digest=sha256:b870cf64887d33dd65aa686011e0424c295a92a45940e0f351c139a41da1540c

Observation 17aed3b1-0717-4500-9363-f91343110c5d · outbound

This paper cites an unresolved cited work.

MADI: Masking-Augmented Diffusion with Inference-Time Scaling for Visual Editing Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:48:34.948628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T16:48:32.709600Z digest=sha256:52bc2cb7cf47a13e85bdec5d8db86984f32e5b9a1a7189baae7c73e9a0926898

Observation 420c2d08-bc12-487c-b64b-2a1f0898d681 · outbound

This paper cites Omniedit: Building image editing generalist models through specialist supervision.

MADI: Masking-Augmented Diffusion with Inference-Time Scaling for Visual Editing Omniedit: Building image editing generalist models through specialist supervision

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T16:48:32.860648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:48:32.860648Z digest=sha256:15205e65f8022464839b21bf4d1d1e1ffd0a5223503dd42758d9d305473a8f55

Observation 69d05fb5-b143-487f-b40b-12e909dcec6d · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

MADI: Masking-Augmented Diffusion with Inference-Time Scaling for Visual Editing Chain-of-thought prompting elicits reasoning in large language models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T16:48:33.014953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:48:33.014953Z digest=sha256:c2005818741948a6fb4a4f3dc53f4cd82b778cbeea6d5ca3e5a7b696b8f4298e

Observation 6955a465-4660-4db5-a29b-62daf9ac3edc · outbound

This paper cites OmniGen: Unified Image Generation.

MADI: Masking-Augmented Diffusion with Inference-Time Scaling for Visual Editing OmniGen: Unified Image Generation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T16:48:33.123960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:48:33.123960Z digest=sha256:a57748451b7e9dd247038b7ad541237573ed35d6c26b9bd2d11ae1c52d4c80b9

Observation 2a0b67ba-be06-4d2a-8a44-7e5bda064be9 · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

MADI: Masking-Augmented Diffusion with Inference-Time Scaling for Visual Editing Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T16:48:33.253084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:48:33.253084Z digest=sha256:a144c3d467324aa52d2151a82708d0b63ab04572b2cd358eb616e5ccbe9fa0d7

Observation 88691822-0144-4c17-bedb-9eba35c6cf76 · outbound

This paper cites Llava-cot: Let vision language models reason step-by-step, 2025.

MADI: Masking-Augmented Diffusion with Inference-Time Scaling for Visual Editing Llava-cot: Let vision language models reason step-by-step, 2025

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:48:34.671198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T16:48:33.380627Z digest=sha256:0a2135d0141ede5946092148aeb58660be30ff83f17f7091c0347db676361a19

Observation 483ca2e6-6335-4900-9e6c-adaba0f497af · outbound

This paper cites Complex-Edit: Cot-like instruction generation for complexity-controllable image editing benchmark, 2025.

MADI: Masking-Augmented Diffusion with Inference-Time Scaling for Visual Editing Complex-Edit: Cot-like instruction generation for complexity-controllable image editing benchmark, 2025

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:48:34.416256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T16:48:33.548602Z digest=sha256:64296b1fc7a2484bc4b43bbfc6b78bbabd2437d24fe9477788b24e1df7ef0c34

Observation cfe9c930-2380-48b4-8efe-c962a317a2b0 · outbound

This paper cites A tale of two features: Stable diffusion complements dino for zero-shot semantic correspondence.

MADI: Masking-Augmented Diffusion with Inference-Time Scaling for Visual Editing A tale of two features: Stable diffusion complements dino for zero-shot semantic correspondence

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T16:48:33.661386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:48:33.661386Z digest=sha256:1e63d456325eac4b1e370a3c2c965d3714b48fe3ae645ea4c0641e0cb82e621a

Observation d6782d8f-6fa3-448d-9f95-a07b0a0c41b0 · outbound

This paper cites Magicbrush: A manually annotated dataset for instruction-guided image editing.

MADI: Masking-Augmented Diffusion with Inference-Time Scaling for Visual Editing Magicbrush: A manually annotated dataset for instruction-guided image editing

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T16:48:33.801958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:48:33.801958Z digest=sha256:1e3f95da006e2302bc9fe4d81ecc8f46d0e13f70fce9a1272474fa8c2fc5fe72

Observation a49bcbc2-53c8-4629-abe3-83e89f15f9a0 · outbound

This paper cites Ultraedit: Instruction-based fine-grained image editing at scale.

MADI: Masking-Augmented Diffusion with Inference-Time Scaling for Visual Editing Ultraedit: Instruction-based fine-grained image editing at scale

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T16:48:33.923207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:48:33.923207Z digest=sha256:a0cdba2d704479ae90743d38085218c31211b2413c2cdad6c7124dddaa42cf09

Observation 54f2b4ff-c8c9-4a47-9964-0a21b2e591d9 · outbound

This paper cites Fast Training of Diffusion Models with Masked Transformers.

MADI: Masking-Augmented Diffusion with Inference-Time Scaling for Visual Editing Fast Training of Diffusion Models with Masked Transformers

Reference 30

Resolution
malformed identifier
no resolver link, observed 2026-08-06T16:48:34.038440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:48:34.038440Z digest=sha256:c5da9b6b6855bfb2f3ee37def41109be3d06b36fbe8ae631678de2e86416b441

Pith citing papers

No inbound Pith citation observations are available.