Pith. sign in

Paper Citation Record · LEDGER

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection

As of 8 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 1 inbound Pith citation observation for arXiv:2505.19149.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.19149 v1

Coverage vector

measured 68 of 68 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:23:11.635467Z

measured 69 of 69 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T01:56:46.514177Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T12:36:57.287687Z

Reference resolution

68 of 68 outbound references displayed

  • verified exact2
  • verified fuzzy15
  • unresolved51
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2e6166ce-f527-4459-b734-4e4f87ae8b6e · outbound

This paper cites Blended diffusion for text-driven editing of natural images.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Blended diffusion for text-driven editing of natural images

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:05.290015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:05.290015Z digest=sha256:8b8489f83edaf3354786f9e50a8ecfb34f110135166b7e923b9a4ef42129ac86

Observation f7b4572f-e89d-450a-bb12-a64c87fed13e · outbound

This paper cites HumanEdit: A High-Quality Human-Rewarded Dataset for Instruction-based Image Editing.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection HumanEdit: A High-Quality Human-Rewarded Dataset for Instruction-based Image Editing

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:05.407350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:05.407350Z digest=sha256:3f57756b9c12ba36c49b4b7bb12cf18c42cf6873021385bcb5cabc9fb6135bee

Observation b6e93817-f659-499d-9034-4f6d70423030 · outbound

This paper cites Instructpix2pix: Learning to follow image editing instructions.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Instructpix2pix: Learning to follow image editing instructions

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:05.484280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:05.484280Z digest=sha256:ef16c534bbe13d1f07e7dfb8801cfa18fa0f4de72d7c3790f4f23775f7d72e1a

Observation d1b14510-6aab-402a-ac90-5245746e5a62 · outbound

This paper cites Training-free layout control with cross-attention guidance.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Training-free layout control with cross-attention guidance

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:05.552410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:05.552410Z digest=sha256:b5540743a3d6824c7966d8774e383fdf5fcc372b82a3b6ebea1a35a2a77de4ae

Observation 52d020b0-924a-4840-82e2-70cde09142ac · outbound

This paper cites Zero-shot Image Editing with Reference Imitation.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Zero-shot Image Editing with Reference Imitation

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:23:12.316270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:23:05.622338Z digest=sha256:b12df8fbc251a7aa7e080c2afa1a42261b4cafd8174ba10ca502ff433ee6b4e0

Observation ad20f985-a1de-444b-8ac2-69340bf30b0c · outbound

This paper cites Anydoor: Zero-shot object-level image customization.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Anydoor: Zero-shot object-level image customization

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:05.729789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:05.729789Z digest=sha256:c292d2e028361651cdd3b0e52f434824d5135eacf18c2cbb9e5beaa2647924ae

Observation d692a309-d3a0-4e9a-873e-1454d442229b · outbound

This paper cites Region-Aware Text-to-Image Generation via Hard Binding and Soft Refinement.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Region-Aware Text-to-Image Generation via Hard Binding and Soft Refinement

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:05.812424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:05.812424Z digest=sha256:1a0eab167cb58e695cb6feff739f77db2b122ce48376d11a539e827b46893dd6

Observation bb31ca85-84df-4703-b8fa-12cece736a53 · outbound

This paper cites Swiftbrush v2: Make your one-step diffusion model better than its teacher.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Swiftbrush v2: Make your one-step diffusion model better than its teacher

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:23:13.375314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:23:05.893851Z digest=sha256:df7d04baac7e477de5726a7b9c71f2bca67cec6c6d65bf914ffc65a68f56b1b2

Observation beb1240c-aab3-4b4b-9585-24117ff3a691 · outbound

This paper cites Turboedit: Text- based image editing using few-step diffusion models.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Turboedit: Text- based image editing using few-step diffusion models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:23:13.364833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:23:05.995294Z digest=sha256:ddb0a01057c49f76e7e74adb704166e26dd51d1d07da28e6a09d410854b8a422

Observation cc29225c-ccd1-442b-a46c-3a0f1384b12e · outbound

This paper cites Textcrafter: Accurately rendering multiple texts in complex visual scenes.arXiv preprint arXiv:2503.23461, 2025.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Textcrafter: Accurately rendering multiple texts in complex visual scenes.arXiv preprint arXiv:2503.23461, 2025

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:06.129326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:06.129326Z digest=sha256:ce7c3e4df21906132fc0becc291b80b1353c6926926a401bdc1b0069a3c39d9e

Observation 3be98bb8-76f9-43c3-ac61-cbe3ad2df277 · outbound

This paper cites Complex multistep image-editing dataset, 2025.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Complex multistep image-editing dataset, 2025

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:23:13.353470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:23:06.203916Z digest=sha256:61a5bc1dc534f4c08cfb556f1ca2a48b4cf66e1dfeb370d17200d2420994b907

Observation bc80ed0c-6205-467f-a165-7c0712f4bede · outbound

This paper cites Guiding Instruction-based Image Editing via Multimodal Large Language Models.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Guiding Instruction-based Image Editing via Multimodal Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:06.307126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:06.307126Z digest=sha256:8b12cec0afce96740b7814e49cbcbf421d5e4b3675ebf88052219b5987f3a631

Observation 6179cd9c-4a9f-4af8-9af2-4803e939fce2 · outbound

This paper cites Renoise: Real image inversion through iterative noising.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Renoise: Real image inversion through iterative noising

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:06.408275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:06.408275Z digest=sha256:8404b1406f9e7e54a0095b2a952ce09c3205fd4f63dbb612137b6f07ea8956d0

Observation 4fb036e1-7338-4489-b762-6b96481cbebf · outbound

This paper cites Generative adversarial nets.Advances in neural information processing systems, 27, 2014.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Generative adversarial nets.Advances in neural information processing systems, 27, 2014

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:06.481834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:06.481834Z digest=sha256:c128384eb2a77383004e79960510883adb09318d66dfac23985c1b0c84f0ab5a

Observation 18a4508d-02b4-4113-8c94-b06b31da0179 · outbound

This paper cites Multi-Reward as Condition for Instruction-based Image Editing.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Multi-Reward as Condition for Instruction-based Image Editing

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:06.580357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:06.580357Z digest=sha256:a32f4d37dfc69570ff2d83259349f2df1fa3205ad1b2630ca73e059e52701633

Observation e89e5aa7-2461-4333-884d-72f237573a52 · outbound

This paper cites FreeEdit: Mask-free Reference-based Image Editing with Multi-modal Instruction.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection FreeEdit: Mask-free Reference-based Image Editing with Multi-modal Instruction

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:06.656115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:06.656115Z digest=sha256:fe958f1d68c5fd1904c7e7add97fe9048dc04cc7eaeb311e9c62a973a3c116d6

Observation e34b6a98-679d-4106-b753-83c1d60c63ed · outbound

This paper cites Prompt-to-Prompt Image Editing with Cross Attention Control.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Prompt-to-Prompt Image Editing with Cross Attention Control

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:06.733485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:06.733485Z digest=sha256:b4cba8fe3ddf1199b9ca8625b0c200c3111838d85f64510d8e3cf093e276da53

Observation bd5d451a-9dc0-4e6b-894f-3e0d9697428b · outbound

This paper cites Denoising diffusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Denoising diffusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:06.825599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:06.825599Z digest=sha256:1458a354a0ae57c215728a7793c175510e6e451d603225c665df40f056d5d855

Observation 0b967774-43c1-4d18-b6c2-5b40dc7ce02b · outbound

This paper cites Smartedit: Exploring complex instruction- based image editing with multimodal large language models.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Smartedit: Exploring complex instruction- based image editing with multimodal large language models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:06.886247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:06.886247Z digest=sha256:d04afc4edc534c621cce79a74889c95f31b243b0594e0fc48091ac72656e08d8

Observation fc4172b0-2a0d-4d7a-9a6a-f7500c1b3deb · outbound

This paper cites Brushnet: A plug-and-play image inpainting model with decomposed dual-branch diffusion.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Brushnet: A plug-and-play image inpainting model with decomposed dual-branch diffusion

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:23:13.323963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:23:06.997606Z digest=sha256:ed7c848f36dd1451bf99fba3f77f32d55ce32f80c91d6cec45e7009b593d09f1

Observation 35adf55b-9118-4577-a2c7-70df5da76cb4 · outbound

This paper cites Image Inpainting Models are Effective Tools for Instruction-guided Image Editing.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Image Inpainting Models are Effective Tools for Instruction-guided Image Editing

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:23:11.973222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:23:07.084405Z digest=sha256:035f2f1df69665f851afd6935be90ea2841afed1f1539eb4a94d3436d392f13c

Observation 42379728-bd94-4049-a1bd-ee305d5d2994 · outbound

This paper cites A style-based generator architecture for generative adversarial networks.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection A style-based generator architecture for generative adversarial networks

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:07.148075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:07.148075Z digest=sha256:27c6afc61412adb75ea7314d3987fb6a104a26c17a2c638d4f5131cd68ad83b4

Observation b5cfbf33-76e4-4636-8adc-368730060563 · outbound

This paper cites Analyzing and improving the image quality of stylegan.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Analyzing and improving the image quality of stylegan

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:07.213831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:07.213831Z digest=sha256:bed9122390098046b94362dca43f214def8d618a993f89b510d2e3909f30034a

Observation 1bff09a2-e945-41c2-b4f7-5b5baacb48fd · outbound

This paper cites Auto-encoding variational bayes, 2013.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Auto-encoding variational bayes, 2013

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:07.318655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:07.318655Z digest=sha256:964aca4c2ce1cf4a581e04f8cfd12839c492232cd84747098fc303558aaf492f

Observation 7cbfeb69-1b04-4dba-b626-80087ca8934b · outbound

This paper cites Generating images with multimodal language models.Advances in Neural Information Processing Systems, 36:21487–21506, 2023.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Generating images with multimodal language models.Advances in Neural Information Processing Systems, 36:21487–21506, 2023

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:07.418627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:07.418627Z digest=sha256:aad37c39a38dfd1be05233875be040e248a3b09bfbee6527c5f34794bc0971ee

Observation d5ff60fe-417b-4454-8ebd-7287f5e77646 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection LLaVA-OneVision: Easy Visual Task Transfer

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:07.470068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:07.470068Z digest=sha256:4ad322735625961dbacfdae69b65e3798e9e94747d554a3b6ec221c7e2e654da

Observation 6c8b85b2-4ba2-4ead-9039-fa2ec3d732cb · outbound

This paper cites Q-Insight: Understanding Image Quality via Visual Reinforcement Learning.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Q-Insight: Understanding Image Quality via Visual Reinforcement Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:07.552002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:07.552002Z digest=sha256:3af18d9ea7a74a68ecb64bfcff25d9268bdfdf77aaafc097d97162bccc8899ce

Observation 4eb9912e-f721-4025-97c0-40031ef509ed · outbound

This paper cites Resvr: Joint rescaling and viewport rendering of omnidirectional images.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Resvr: Joint rescaling and viewport rendering of omnidirectional images

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:23:13.292457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:23:07.636743Z digest=sha256:10cf4d31be42bbc5fa0df7be0ee59a63444fb2cffd6e3a5c7f61a6c06e3d41da

Observation a4e84820-72e2-4453-b3ea-d26e37bf8cac · outbound

This paper cites OmniDrag: Enabling Motion Control for Omnidirectional Image-to-Video Generation.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection OmniDrag: Enabling Motion Control for Omnidirectional Image-to-Video Generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:07.710615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:07.710615Z digest=sha256:75043d17538a478c9ab65476a9983786cb996cdc13a6146d8899992837263bbe

Observation 1604f9f3-00db-4a10-b30c-a40fd6fab71e · outbound

This paper cites BrushEdit: All-In-One Image Inpainting and Editing.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection BrushEdit: All-In-One Image Inpainting and Editing

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:07.789904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:07.789904Z digest=sha256:1681698c84e81e4528b9fe08606d525cea6ede191a6ab48d501688c8c52447e2

Observation aff9180f-e37b-4128-92cc-dfa09b296a79 · outbound

This paper cites Adversarial supervision makes layout-to-image diffusion models thrive.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Adversarial supervision makes layout-to-image diffusion models thrive

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:23:13.281552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:23:07.849807Z digest=sha256:e5fe80f5ad49615ad0ee9c75be7245433adf7590e0df0421263e9f7f5644ea31

Observation 918830ec-ccf9-4edd-aa77-01b5b8bc6c3d · outbound

This paper cites Drag your noise: Interactive point-based editing via diffusion semantic propagation.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Drag your noise: Interactive point-based editing via diffusion semantic propagation

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:23:13.269834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:23:07.911521Z digest=sha256:59b26c9f627a510710847bc515086b13dc4345addbb5adbd6d131fdb2eb08518

Observation 0216285b-ef45-4c89-9d0a-cc1b30b8f39b · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:07.989663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:07.989663Z digest=sha256:5324ff2ae8ade31f92cfe3a9b00d3132649f9d82af5506d31fa905a6ff4b63cd

Observation f93afe6a-2e55-44bd-b5be-7f4a4feb3d03 · outbound

This paper cites Step1X-Edit: A Practical Framework for General Image Editing.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Step1X-Edit: A Practical Framework for General Image Editing

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:08.052387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:08.052387Z digest=sha256:98895306671ef2485f196bbffa233b8724125e22e1ebb9c380f88ad11dbc8464

Observation b4db43ad-cf9d-489f-b737-f0535a7cf1c4 · outbound

This paper cites Magicquill: An intelligent interactive image editing system.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Magicquill: An intelligent interactive image editing system

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:23:13.252940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:23:08.111111Z digest=sha256:f52fd2ef3f0734db7a842b7e8554ddbcf56672b74c773d72353203b430ccb2f4

Observation 26de2773-2e35-41dc-80b3-1fe414a377ba · outbound

This paper cites Diffeditor: Boosting accuracy and flexibility on diffusion-based image editing.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Diffeditor: Boosting accuracy and flexibility on diffusion-based image editing

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:08.156958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:08.156958Z digest=sha256:ea1636c68df9ab2d8a38cdf15f847855805dd8592291c42f0cb2108c9a6cd0c6

Observation a1e50af1-e713-44a2-90d2-cf6413222ae1 · outbound

This paper cites Dragondiffusion: Enabling drag-style manipulation on diffusion models.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Dragondiffusion: Enabling drag-style manipulation on diffusion models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:23:13.237351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:23:08.216468Z digest=sha256:aabc788ab965c2cde4396ed74aa683dada7c977a1e61e5404985c80010a74fac

Observation bc87defe-6221-4c52-85a4-5b9f57067a60 · outbound

This paper cites T2i-adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection T2i-adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:08.307630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:08.307630Z digest=sha256:d3ada4f9ba309e7d0fce105d59eb20e441137a0f75a9d8a5e26294bd0b40fdd7

Observation 709b76fe-869d-46c9-ad80-238c731b80df · outbound

This paper cites Handiffuser: Text-to-image generation with realistic hand appearances.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Handiffuser: Text-to-image generation with realistic hand appearances

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:23:13.221725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:23:08.367757Z digest=sha256:deebba8cebb0359fec372f45854ff06b97dd346a99fcd367b3ab4453fc0138b4

Observation 8935cc25-1e62-46ae-8909-beac8916a29e · outbound

This paper cites Transfer between Modalities with MetaQueries.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Transfer between Modalities with MetaQueries

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:08.422122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:08.422122Z digest=sha256:1c8e2b5c1a0bc7dbc23fbbb2767c31d6ce4304df750b9f112ea17c472a277423

Observation 261e470d-a3b1-47ce-b1f6-9b32998d414a · outbound

This paper cites Drag your gan: Interactive point-based manipulation on the generative image manifold.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Drag your gan: Interactive point-based manipulation on the generative image manifold

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:08.494940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:08.494940Z digest=sha256:5d7e18efe87aae4b1cee4d14b6036840bdf81094d5ce393260e7b670d05b947e

Observation 248cb43e-66ce-4cf6-8f70-34f87b5d6722 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:08.555573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:08.555573Z digest=sha256:33c5c6f1aa0bb366edb98a9bfc050ee85f51050224dc3d6c6b8f5aed47d622f8

Observation 1fe2ce5e-4ff4-40db-bbb9-4a029952f774 · outbound

This paper cites Learning transferable visual models from natural language supervision.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Learning transferable visual models from natural language supervision

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:08.617044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:08.617044Z digest=sha256:2d1c335f2113438a86e152ae849eb88081942fdd7c7bd92b3f3dd3db8cce11a7

Observation 1266e1bc-ecbe-42e0-a468-dc2f0145448f · outbound

This paper cites High- resolution image synthesis with latent diffusion models.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection High- resolution image synthesis with latent diffusion models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:08.712174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:08.712174Z digest=sha256:e927b7e30e6b41a0115eb8917053da5cf2fa343045ec7c70b271140ef95fb32c

Observation c1320835-73ca-48a9-a69c-5a5bfcbcd9c2 · outbound

This paper cites Photorealistic text-to-image diffusion models with deep language understanding.Advances in neural information processing systems, 35:36479–36494, 2022.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Photorealistic text-to-image diffusion models with deep language understanding.Advances in neural information processing systems, 35:36479–36494, 2022

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:08.840800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:08.840800Z digest=sha256:8771ccbd8fd5be3529fc3e76ec82b3f6844aa8f6971b01ed79aa228a16a36ef6

Observation 9e154348-4e85-4f2d-ae28-df9b0607a3ce · outbound

This paper cites Emu edit: Precise image editing via recognition and generation tasks.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Emu edit: Precise image editing via recognition and generation tasks

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:23:13.180735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:23:08.922371Z digest=sha256:246283b99879169831d2b3302821670f7fd5b5d5f03098b01f8c2ed49eff1f40

Observation 3dbd24ee-55f1-4e00-b648-d9568b50ec81 · outbound

This paper cites Dragdiffusion: Harnessing diffusion models for interactive point-based image editing.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Dragdiffusion: Harnessing diffusion models for interactive point-based image editing

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:09.001818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:09.001818Z digest=sha256:b1dab8baa81a847e24684fe3258425904b5363c227852009c96e53cdb2b83f92

Observation 74a839d2-7d8f-40ea-b90e-8fc5727eed50 · outbound

This paper cites Insert Anything: Image Insertion via In-Context Editing in DiT.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Insert Anything: Image Insertion via In-Context Editing in DiT

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:09.117476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:09.117476Z digest=sha256:912a1cf621636b7af3d6d0c69262127927c8d5eb9d60fcdfa6053fad7cc06722

Observation 38c37376-6b06-426a-9519-dd98316ca5ec · outbound

This paper cites Invertible Consistency Distillation for Text-Guided Image Editing in Around 7 Steps.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Invertible Consistency Distillation for Text-Guided Image Editing in Around 7 Steps

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:09.205192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:09.205192Z digest=sha256:ddcaa5e98340b657dcf07ffa18bd4bf13f6b18603f68b61a27ce76955506ea2a

Observation 1a6114cb-3d3d-428e-8ad6-c3e28434033a · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:09.297602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:09.297602Z digest=sha256:1c611ea16ea5901507f0d92e42e5cae3cc8914dce6b7b3e942c68353a750aacd

Observation e986d506-1f5c-4c35-b9c7-24a90e475e81 · outbound

This paper cites Gpt-4o system card, 2024.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Gpt-4o system card, 2024

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:09.385434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:09.385434Z digest=sha256:e177aff76a82adbdad7d69d28b8f49a6b2ed12c31f0edc54cbf0781e1f42c6f5

Observation f90c1bc6-8a38-4195-af3e-38350038b4a6 · outbound

This paper cites MetaMorph: Multimodal Understanding and Generation via Instruction Tuning.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection MetaMorph: Multimodal Understanding and Generation via Instruction Tuning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:09.580072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:09.580072Z digest=sha256:9f070858fbd94fc36cfefd036a63953fa6a6c9bff016e07fe6e0ad3b1fa0c058

Observation 40030a32-4388-47ac-be64-d4212c2b8a14 · outbound

This paper cites Plug-and-play diffusion features for text-driven image-to-image translation.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Plug-and-play diffusion features for text-driven image-to-image translation

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:09.686407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:09.686407Z digest=sha256:5f73324f42269fda23a2feb53287490c803374f7b6d639b749cc8e4491f4b23b

Observation 94e5fa03-dff7-4e81-bb65-58b861e3ee4b · outbound

This paper cites FlexEdit: Marrying Free-Shape Masks to VLLM for Flexible Image Editing.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection FlexEdit: Marrying Free-Shape Masks to VLLM for Flexible Image Editing

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:09.768033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:09.768033Z digest=sha256:1b19a3f24a3d1b799d2e8b991cf140be5fd31cdfd22edb161343aef877f828e9

Observation 850fd77c-bdfc-4d80-bbda-a51ff526850d · outbound

This paper cites 360dvd: Controllable panorama video generation with 360-degree video diffusion model.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection 360dvd: Controllable panorama video generation with 360-degree video diffusion model

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:09.890616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:09.890616Z digest=sha256:8e93964ea568acecacff5923f12f6c32956de9f33fc3ace3a84de469dcd40fa7

Observation 3206f2f9-55c5-4c1a-ba13-1b381ed42fe6 · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Emu3: Next-Token Prediction is All You Need

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:10.024028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:10.024028Z digest=sha256:f3b7124a6c6078e85ca66d726f4cbd5eede780573645f3e43d41802377041fd8

Observation d922f159-7107-48d9-afba-24819e1eeba1 · outbound

This paper cites In- stancediffusion: Instance-level control for image generation.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection In- stancediffusion: Instance-level control for image generation

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:23:13.144422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:23:10.133279Z digest=sha256:1546b117655f31ce14b7e485a2338969db1fd98df08f55769d464c138773ce47

Observation 6e286aec-fcff-49d6-9e79-82842c48ebc6 · outbound

This paper cites Genartist: Multimodal llm as an agent for unified image generation and editing.Advances in Neural Information Processing Systems, 37:128374–128395, 2024.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Genartist: Multimodal llm as an agent for unified image generation and editing.Advances in Neural Information Processing Systems, 37:128374–128395, 2024

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:23:13.051692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:23:10.221516Z digest=sha256:daa4c02f4ad6aed4cefbb9824e0714e41f901612792d14ed55e68cd85ffdce37

Observation 1516427e-e951-4437-bea2-2cad37912ca3 · outbound

This paper cites Image quality assessment: from error visibility to structural similarity.IEEE transactions on image processing, 13(4):600– 612, 2004.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Image quality assessment: from error visibility to structural similarity.IEEE transactions on image processing, 13(4):600– 612, 2004

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:10.370380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:10.370380Z digest=sha256:e24ed67a19951011a22ed6a7de05c4e9d6a905a977d4b734492dfbae7e27f53e

Observation 57fa69e6-d237-49bf-b1a2-53afc3db4e95 · outbound

This paper cites Visionllm v2: An end-to-end generalist multimodal large language model for hundreds of vision-language tasks.Advances in Neural Information Processing Systems, 37:69925–69975, 2024.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Visionllm v2: An end-to-end generalist multimodal large language model for hundreds of vision-language tasks.Advances in Neural Information Processing Systems, 37:69925–69975, 2024

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:10.504956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:10.504956Z digest=sha256:d11b3529aaf3871e36282710d1b2b56340d78b3a3fb0f2fd277275bf9d63942a

Observation 9a5c87af-ddff-4ea1-9bcd-82103e720673 · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:10.608892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:10.608892Z digest=sha256:1af3c1ce5138958679e730385b3f3cf269009d90f34745388221217d71dd3b07

Observation d328811f-6996-4525-912b-34f4991f13f7 · outbound

This paper cites IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:10.781055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:10.781055Z digest=sha256:65cd39014412f0c6f777ee6fa882521aca3ce3958df0b90facac50b4ead0c16f

Observation 07d7c50d-9c33-495c-96bb-081da015947a · outbound

This paper cites AnyEdit: Mastering Unified High-Quality Image Editing for Any Idea.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection AnyEdit: Mastering Unified High-Quality Image Editing for Any Idea

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:10.975139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:10.975139Z digest=sha256:9fb4c3b49d964d20fd5fdf1a60db06c7b988c2f214637dcb88ad3c3473b3c82f

Observation d1a5f95d-8415-4ef0-9669-7b883dcc3d10 · outbound

This paper cites Magicbrush: A manually annotated dataset for instruction-guided image editing.Advances in Neural Information Processing Systems, 36:31428–31449, 2023.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Magicbrush: A manually annotated dataset for instruction-guided image editing.Advances in Neural Information Processing Systems, 36:31428–31449, 2023

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:11.145234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:11.145234Z digest=sha256:fdbb20f02b08b9b982883104027d2500f07bb6c8c6c8c2beae4812dc27868d9b

Observation 3abffb1d-1c68-472d-97e3-a1aa62985ec2 · outbound

This paper cites Adding conditional control to text-to-image diffusion models.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Adding conditional control to text-to-image diffusion models

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:23:12.809196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:23:11.296196Z digest=sha256:f1030e5ac545ed52c9827569512a2f13b13779b9072c95b169abdfd311657e36

Observation f58d4b08-9b8a-4631-973f-5bfe97a10850 · outbound

This paper cites The unrea- sonable effectiveness of deep features as a perceptual metric.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection The unrea- sonable effectiveness of deep features as a perceptual metric

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:11.402858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:11.402858Z digest=sha256:06728aef6a0d79904b825eae98a89ddf27280fc709880fb16690eaac9e507ea1

Observation c353a799-e1be-4c50-8eb6-d00667549de2 · outbound

This paper cites GoodDrag: Towards Good Practices for Drag Editing with Diffusion Models.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection GoodDrag: Towards Good Practices for Drag Editing with Diffusion Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:11.514069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:11.514069Z digest=sha256:238d71212474a71362245f14a6acc7dece336a64114c6c82e2b81cd58b60406d

Observation c66d4d9f-6be9-41e9-8238-2925bfbf133b · outbound

This paper cites Ultraedit: Instruction-based fine-grained image editing at scale.Advances in Neural Information Processing Systems, 37:3058–3093, 2024.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Ultraedit: Instruction-based fine-grained image editing at scale.Advances in Neural Information Processing Systems, 37:3058–3093, 2024

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:23:12.600148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:23:11.635467Z digest=sha256:8b5b5e6cee95b344d068d31e89b158fe74233b61f69a9eebe017cb5096c5e95a

Pith citing papers

Observation 0025a0dc-b722-4309-ab78-5eb73859901f · inbound

TextWand: A Unified Framework for Scene Text Editing cites this paper.

TextWand: A Unified Framework for Scene Text Editing MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:36:57.289066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T01:56:46.514177Z digest=sha256:1fcd92a0078c1a94ca0135d413ff507b40066b851658a57e62eff9f8e029b44d