Pith. sign in

Paper Citation Record · LEDGER

A Survey of Multimodal-Guided Image Editing with Text-to-Image Diffusion Models

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 20 inbound Pith citation observations for arXiv:2406.14555.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.14555 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:53:12.692444Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T14:48:32.608389Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 3b454903-1db3-4f09-86aa-f565ce66a6a7 · inbound

TrojanEdit: Multimodal Backdoor Attack Against Image Editing Model cites this paper.

TrojanEdit: Multimodal Backdoor Attack Against Image Editing Model A Survey of Multimodal-Guided Image Editing with Text-to-Image Diffusion Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T15:06:14.629218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:06:14.629218Z digest=sha256:3700bf1b032c7f6b7c5c20231b29ae265f0cf819206173abde2f48e5fe29687e

Observation a2866124-2653-436d-8b23-6533b1f0e1cc · inbound

Automatic Evaluation for Text-to-image Generation: Task-decomposed Framework, Distilled Training, and Meta-evaluation Benchmark cites this paper.

Automatic Evaluation for Text-to-image Generation: Task-decomposed Framework, Distilled Training, and Meta-evaluation Benchmark A Survey of Multimodal-Guided Image Editing with Text-to-Image Diffusion Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T14:19:08.555991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:19:08.555991Z digest=sha256:0dafaa16eb40af89c883be763255abe1d7553c2e0b485a5d83114b41c84ef991

Observation df2b15bc-6ce3-4c3e-b404-cd3a3b408ac6 · inbound

SPAgent: Adaptive Task Decomposition and Model Selection for General Video Generation and Editing cites this paper.

SPAgent: Adaptive Task Decomposition and Model Selection for General Video Generation and Editing A Survey of Multimodal-Guided Image Editing with Text-to-Image Diffusion Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T10:44:14.782930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:44:14.782930Z digest=sha256:82594f92b35146175bce5753fadde21a2b53ba14f8f8e21d971f49c802267a70

Observation 0eb85e0c-4e60-4fec-9d77-24ec06be32b3 · inbound

BadPatch: Diffusion-Based Generation of Physical Adversarial Patches cites this paper.

BadPatch: Diffusion-Based Generation of Physical Adversarial Patches A Survey of Multimodal-Guided Image Editing with Text-to-Image Diffusion Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T04:26:55.119335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:26:55.119335Z digest=sha256:7615db0a55126e80d821a519eb00415611851bdcb0ac673ff035909ee9199010

Observation e1874453-7955-4a51-b45f-d8baa20ebf60 · inbound

AsymRnR: Video Diffusion Transformers Acceleration with Asymmetric Reduction and Restoration cites this paper.

AsymRnR: Video Diffusion Transformers Acceleration with Asymmetric Reduction and Restoration A Survey of Multimodal-Guided Image Editing with Text-to-Image Diffusion Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T14:44:17.579527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:44:17.579527Z digest=sha256:c4c112cb52a58082bf8ee03f4a2837d4ff8d7fc4bb6c55a39aa857f0c7473bf9

Observation 88cee15e-435d-4c33-bf70-304f6e7d0221 · inbound

Mapping the Mind of an Instruction-based Image Editing using SMILE cites this paper.

Mapping the Mind of an Instruction-based Image Editing using SMILE A Survey of Multimodal-Guided Image Editing with Text-to-Image Diffusion Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:23.106767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:23.106767Z digest=sha256:0dfd3648cfd4a8a422b1bdd61e8e744de073db29bb358029a53805b7d2a282f6

Observation 9766ea11-053d-4910-ae3d-92454d611ade · inbound

Free-Form Motion Control: Controlling the 6D Poses of Camera and Objects in Video Generation cites this paper.

Free-Form Motion Control: Controlling the 6D Poses of Camera and Objects in Video Generation A Survey of Multimodal-Guided Image Editing with Text-to-Image Diffusion Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T22:31:47.641270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:31:47.641270Z digest=sha256:596f159c28a850ddf0685cdf8ecfa8f4430456bea6bb51a8041a63ed56444cf7

Observation 0e345701-4149-431f-a655-25e323bdb32e · inbound

A Survey on Pre-Trained Diffusion Model Distillations cites this paper.

A Survey on Pre-Trained Diffusion Model Distillations A Survey of Multimodal-Guided Image Editing with Text-to-Image Diffusion Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T05:26:57.014104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T05:26:57.014104Z digest=sha256:89612eb366eac5a592f28ce07fa4cb4d91cca36d5d2233beb98f6b49a45390be

Observation ab8fc5a4-095a-4db2-9e23-ffc234927d19 · inbound

Text to Image Generation and Editing: A Survey cites this paper.

Text to Image Generation and Editing: A Survey A Survey of Multimodal-Guided Image Editing with Text-to-Image Diffusion Models

Reference 238

Resolution
unresolved
no resolver link, observed 2026-08-16T00:53:12.692444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:53:12.692444Z digest=sha256:561f604076387a886b1c85a8e64ad1d780dfdeda56054507c7a868cb11a5481d

Observation 878b40ac-f11b-43b7-a642-64d481dac74d · inbound

KRIS-Bench: Benchmarking Next-Level Intelligent Image Editing Models cites this paper.

KRIS-Bench: Benchmarking Next-Level Intelligent Image Editing Models A Survey of Multimodal-Guided Image Editing with Text-to-Image Diffusion Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:18.230361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:18.230361Z digest=sha256:8d4856f43edb4bba5de1dd5b3c63fdf9897c2099bbde3b07aac0253b0d57fedc

Observation 6c2fbbb3-ed15-49b7-9928-5a3138001b26 · inbound

MMAFFBen: A Multilingual and Multimodal Affective Analysis Benchmark for Evaluating LLMs and VLMs cites this paper.

MMAFFBen: A Multilingual and Multimodal Affective Analysis Benchmark for Evaluating LLMs and VLMs A Survey of Multimodal-Guided Image Editing with Text-to-Image Diffusion Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T12:27:34.712223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:27:34.712223Z digest=sha256:23712327dc8f2593be3118ca95f2d6616ccf01827bbca547b63bb156514674a0

Observation 40b743f7-bc70-444f-8b57-053c8fc6fc83 · inbound

AnyI2V: Animating Any Conditional Image with Motion Control cites this paper.

AnyI2V: Animating Any Conditional Image with Motion Control A Survey of Multimodal-Guided Image Editing with Text-to-Image Diffusion Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:27.267209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:27.267209Z digest=sha256:4dcb935ecf6b95287e2478d0fc614b4b2d7d3fb77d588a8b54a43e514c92ee4a

Observation ba30208d-2536-458a-92f1-067bd37636db · inbound

Mastering Regional 3DGS: Locating, Initializing, and Editing with Diverse 2D Priors cites this paper.

Mastering Regional 3DGS: Locating, Initializing, and Editing with Diverse 2D Priors A Survey of Multimodal-Guided Image Editing with Text-to-Image Diffusion Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T19:34:25.575695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:34:25.575695Z digest=sha256:13cc02a55e04bdf478936cad3998b9a9575f4321e01b6d2664583370371f69cf

Observation 8165f4f7-8e44-44bb-922b-a229563d7772 · inbound

CharaConsist: Fine-Grained Consistent Character Generation cites this paper.

CharaConsist: Fine-Grained Consistent Character Generation A Survey of Multimodal-Guided Image Editing with Text-to-Image Diffusion Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T17:10:56.151197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:10:56.151197Z digest=sha256:ea97e6b3ec996e640406d32744377d01ec586688256971137f1e0deabfb8d5c5

Observation e6af47d5-4af3-440e-8548-3eb7ebda4aa4 · inbound

Report of the 5th PVUW Challenge: Towards More Diverse Modalities in Pixel-Level Understanding cites this paper.

Report of the 5th PVUW Challenge: Towards More Diverse Modalities in Pixel-Level Understanding A Survey of Multimodal-Guided Image Editing with Text-to-Image Diffusion Models

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:31:15.614068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-07T16:52:26.262129Z digest=sha256:e9efb94ecd1a71ea047dffd122612d161c83207ae762fdacffa86944946f1745

Observation 59122a93-fd2f-47a9-9bbf-ac778dbaa19e · inbound

3DReflecNet: A Large-Scale Dataset for 3D Reconstruction of Reflective, Transparent, and Low-Texture Objects cites this paper.

3DReflecNet: A Large-Scale Dataset for 3D Reconstruction of Reflective, Transparent, and Low-Texture Objects A Survey of Multimodal-Guided Image Editing with Text-to-Image Diffusion Models

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:31:29.034804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-12T04:10:35.320940Z digest=sha256:bf46fd36bb08d1857c24741c956d865f427b441482de5956a2c0dc055bebda51

Observation 892abf8a-cee7-46cd-9aea-c74da73112c2 · inbound

CPC-VAR:Continual Personalized and Compositional Generation in Visual Autoregressive Models cites this paper.

CPC-VAR:Continual Personalized and Compositional Generation in Visual Autoregressive Models A Survey of Multimodal-Guided Image Editing with Text-to-Image Diffusion Models

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:03:06.239593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T07:01:01.427273Z digest=sha256:cfd921c6dc9c1d69c32ee76631e27d2019f956d35ae9f53e84f8bce87ba70282

Observation 4396eee7-f712-44f0-95e4-134bbbd06b38 · inbound

Rethinking Scribble-Guided Image Editing: Generalization, Instruction Adherence, and Multi-Tasking cites this paper.

Rethinking Scribble-Guided Image Editing: Generalization, Instruction Adherence, and Multi-Tasking A Survey of Multimodal-Guided Image Editing with Text-to-Image Diffusion Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-06-29T23:14:02.005615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T23:06:29.200541Z digest=sha256:3814b000d9d1e223e12ad2173265aee532fa341b993293f3fc66a438515f4cbb

Observation 0b2d1af0-f569-4eb2-8d1a-0abb21da0e73 · inbound

Seek to Segment: Active Perception for Panoramic Referring Segmentation cites this paper.

Seek to Segment: Active Perception for Panoramic Referring Segmentation A Survey of Multimodal-Guided Image Editing with Text-to-Image Diffusion Models

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:48:32.609856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-03T14:39:22.617747Z digest=sha256:96af09415badadfada0f9820aac6d14877938aedb7d4e9bfd7ac9a91cd5524c4

Observation 434f46b2-f6e4-45c2-bbba-52c4b15aa7ec · inbound

Appearance Pointers -- Multimodal Region Control of Diffusion Transformers cites this paper.

Appearance Pointers -- Multimodal Region Control of Diffusion Transformers A Survey of Multimodal-Guided Image Editing with Text-to-Image Diffusion Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T12:45:18.568969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:45:18.568969Z digest=sha256:0e79a9b4f4a6252514ca0b6576fe89e8eec8e3e9f409cb765e96611474a685ab