Pith. sign in

Paper Citation Record · LEDGER

A Survey of Multimodal-Guided Image Editing with Text-to-Image Diffusion Models

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2406.14555.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.14555 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T05:26:57.014104Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T14:48:32.608389Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0e345701-4149-431f-a655-25e323bdb32e · inbound

A Survey on Pre-Trained Diffusion Model Distillations cites this paper.

A Survey on Pre-Trained Diffusion Model Distillations A Survey of Multimodal-Guided Image Editing with Text-to-Image Diffusion Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T05:26:57.014104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T05:26:57.014104Z digest=sha256:1032190e3940b8981f660afc1103c2ef94732e3f1715d117104e7e0e9552d523

Observation 878b40ac-f11b-43b7-a642-64d481dac74d · inbound

KRIS-Bench: Benchmarking Next-Level Intelligent Image Editing Models cites this paper.

KRIS-Bench: Benchmarking Next-Level Intelligent Image Editing Models A Survey of Multimodal-Guided Image Editing with Text-to-Image Diffusion Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:18.230361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:18.230361Z digest=sha256:0d60107ccb7bc477d2fba6238431da0e3e933272149c5849ce1d5ad809adffdd

Observation 6c2fbbb3-ed15-49b7-9928-5a3138001b26 · inbound

MMAFFBen: A Multilingual and Multimodal Affective Analysis Benchmark for Evaluating LLMs and VLMs cites this paper.

MMAFFBen: A Multilingual and Multimodal Affective Analysis Benchmark for Evaluating LLMs and VLMs A Survey of Multimodal-Guided Image Editing with Text-to-Image Diffusion Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T12:27:34.712223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:27:34.712223Z digest=sha256:4a75404f7629aee040daae7d7ab94d33f83a1326208d0d4b50659b1917302008

Observation 40b743f7-bc70-444f-8b57-053c8fc6fc83 · inbound

AnyI2V: Animating Any Conditional Image with Motion Control cites this paper.

AnyI2V: Animating Any Conditional Image with Motion Control A Survey of Multimodal-Guided Image Editing with Text-to-Image Diffusion Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:27.267209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:27.267209Z digest=sha256:7be3eba77bcedf8af7c67d5ac93c812eb552c83ef163060501c800b01465cd58

Observation ba30208d-2536-458a-92f1-067bd37636db · inbound

Mastering Regional 3DGS: Locating, Initializing, and Editing with Diverse 2D Priors cites this paper.

Mastering Regional 3DGS: Locating, Initializing, and Editing with Diverse 2D Priors A Survey of Multimodal-Guided Image Editing with Text-to-Image Diffusion Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T19:34:25.575695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:34:25.575695Z digest=sha256:e1951502db4ebd53a3a06c880f0b46e23761f15436dd1a2d4c62f8a9f98e6c81

Observation 8165f4f7-8e44-44bb-922b-a229563d7772 · inbound

CharaConsist: Fine-Grained Consistent Character Generation cites this paper.

CharaConsist: Fine-Grained Consistent Character Generation A Survey of Multimodal-Guided Image Editing with Text-to-Image Diffusion Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T17:10:56.151197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:10:56.151197Z digest=sha256:933cb0c11a387e19311abe996dcd7617217edca99975c7371f1fe489c6d1cb69

Observation e6af47d5-4af3-440e-8548-3eb7ebda4aa4 · inbound

Report of the 5th PVUW Challenge: Towards More Diverse Modalities in Pixel-Level Understanding cites this paper.

Report of the 5th PVUW Challenge: Towards More Diverse Modalities in Pixel-Level Understanding A Survey of Multimodal-Guided Image Editing with Text-to-Image Diffusion Models

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:31:15.614068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-07T16:52:26.262129Z digest=sha256:8cdbe66806b732eaccbd2fb4000b20e70cf4900d1069acdb7b88b0ea19b86447

Observation 59122a93-fd2f-47a9-9bbf-ac778dbaa19e · inbound

3DReflecNet: A Large-Scale Dataset for 3D Reconstruction of Reflective, Transparent, and Low-Texture Objects cites this paper.

3DReflecNet: A Large-Scale Dataset for 3D Reconstruction of Reflective, Transparent, and Low-Texture Objects A Survey of Multimodal-Guided Image Editing with Text-to-Image Diffusion Models

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:31:29.034804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T04:10:35.320940Z digest=sha256:50c8232b5396261be4c31a57fef8b77e518134c293f4f2177224e6606c988d0a

Observation 892abf8a-cee7-46cd-9aea-c74da73112c2 · inbound

CPC-VAR:Continual Personalized and Compositional Generation in Visual Autoregressive Models cites this paper.

CPC-VAR:Continual Personalized and Compositional Generation in Visual Autoregressive Models A Survey of Multimodal-Guided Image Editing with Text-to-Image Diffusion Models

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:03:06.239593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T07:01:01.427273Z digest=sha256:f34fafa2fe8c265ee6429c14a7809cc2146d529491a8be180436b20f127b05e8

Observation 4396eee7-f712-44f0-95e4-134bbbd06b38 · inbound

Rethinking Scribble-Guided Image Editing: Generalization, Instruction Adherence, and Multi-Tasking cites this paper.

Rethinking Scribble-Guided Image Editing: Generalization, Instruction Adherence, and Multi-Tasking A Survey of Multimodal-Guided Image Editing with Text-to-Image Diffusion Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-06-29T23:14:02.005615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T23:06:29.200541Z digest=sha256:39e307193289dd753a6083adb992b0340960590375fe620124777e07c70e25d1

Observation 0b2d1af0-f569-4eb2-8d1a-0abb21da0e73 · inbound

Seek to Segment: Active Perception for Panoramic Referring Segmentation cites this paper.

Seek to Segment: Active Perception for Panoramic Referring Segmentation A Survey of Multimodal-Guided Image Editing with Text-to-Image Diffusion Models

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:48:32.609856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-03T14:39:22.617747Z digest=sha256:ec10af1b0259c5027331c11e3d97a579d5032b45ad36b355f826e6c4687ab291

Observation 434f46b2-f6e4-45c2-bbba-52c4b15aa7ec · inbound

Appearance Pointers -- Multimodal Region Control of Diffusion Transformers cites this paper.

Appearance Pointers -- Multimodal Region Control of Diffusion Transformers A Survey of Multimodal-Guided Image Editing with Text-to-Image Diffusion Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T12:45:18.568969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:45:18.568969Z digest=sha256:fe4f380572c8c3674167aba16f8128b1a697a15b1de32c250a0da9b1c43da710