Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:23:11.635467Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 1 inbound Pith citation observation for arXiv:2505.19149.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:23:11.635467Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-28T01:56:46.514177Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-02T12:36:57.287687Z
68 of 68 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 2e6166ce-f527-4459-b734-4e4f87ae8b6e · outbound
MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Blended diffusion for text-driven editing of natural images
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7b4572f-e89d-450a-bb12-a64c87fed13e · outbound
MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection HumanEdit: A High-Quality Human-Rewarded Dataset for Instruction-based Image Editing
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6e93817-f659-499d-9034-4f6d70423030 · outbound
MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Instructpix2pix: Learning to follow image editing instructions
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1b14510-6aab-402a-ac90-5245746e5a62 · outbound
MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Training-free layout control with cross-attention guidance
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52d020b0-924a-4840-82e2-70cde09142ac · outbound
MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Zero-shot Image Editing with Reference Imitation
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ad20f985-a1de-444b-8ac2-69340bf30b0c · outbound
MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Anydoor: Zero-shot object-level image customization
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d692a309-d3a0-4e9a-873e-1454d442229b · outbound
MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Region-Aware Text-to-Image Generation via Hard Binding and Soft Refinement
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb31ca85-84df-4703-b8fa-12cece736a53 · outbound
MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Swiftbrush v2: Make your one-step diffusion model better than its teacher
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation beb1240c-aab3-4b4b-9585-24117ff3a691 · outbound
MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Turboedit: Text- based image editing using few-step diffusion models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cc29225c-ccd1-442b-a46c-3a0f1384b12e · outbound
MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Textcrafter: Accurately rendering multiple texts in complex visual scenes.arXiv preprint arXiv:2503.23461, 2025
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3be98bb8-76f9-43c3-ac61-cbe3ad2df277 · outbound
MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Complex multistep image-editing dataset, 2025
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation bc80ed0c-6205-467f-a165-7c0712f4bede · outbound
MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Guiding Instruction-based Image Editing via Multimodal Large Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6179cd9c-4a9f-4af8-9af2-4803e939fce2 · outbound
MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Renoise: Real image inversion through iterative noising
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4fb036e1-7338-4489-b762-6b96481cbebf · outbound
MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Generative adversarial nets.Advances in neural information processing systems, 27, 2014
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18a4508d-02b4-4113-8c94-b06b31da0179 · outbound
MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Multi-Reward as Condition for Instruction-based Image Editing
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e89e5aa7-2461-4333-884d-72f237573a52 · outbound
MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection FreeEdit: Mask-free Reference-based Image Editing with Multi-modal Instruction
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e34b6a98-679d-4106-b753-83c1d60c63ed · outbound
MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Prompt-to-Prompt Image Editing with Cross Attention Control
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd5d451a-9dc0-4e6b-894f-3e0d9697428b · outbound
MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Denoising diffusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b967774-43c1-4d18-b6c2-5b40dc7ce02b · outbound
MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Smartedit: Exploring complex instruction- based image editing with multimodal large language models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc4172b0-2a0d-4d7a-9a6a-f7500c1b3deb · outbound
MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Brushnet: A plug-and-play image inpainting model with decomposed dual-branch diffusion
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 35adf55b-9118-4577-a2c7-70df5da76cb4 · outbound
MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Image Inpainting Models are Effective Tools for Instruction-guided Image Editing
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 42379728-bd94-4049-a1bd-ee305d5d2994 · outbound
MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection A style-based generator architecture for generative adversarial networks
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5cfbf33-76e4-4636-8adc-368730060563 · outbound
MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Analyzing and improving the image quality of stylegan
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1bff09a2-e945-41c2-b4f7-5b5baacb48fd · outbound
MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Auto-encoding variational bayes, 2013
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7cbfeb69-1b04-4dba-b626-80087ca8934b · outbound
MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Generating images with multimodal language models.Advances in Neural Information Processing Systems, 36:21487–21506, 2023
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5ff60fe-417b-4454-8ebd-7287f5e77646 · outbound
MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection LLaVA-OneVision: Easy Visual Task Transfer
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c8b85b2-4ba2-4ead-9039-fa2ec3d732cb · outbound
MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Q-Insight: Understanding Image Quality via Visual Reinforcement Learning
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4eb9912e-f721-4025-97c0-40031ef509ed · outbound
MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Resvr: Joint rescaling and viewport rendering of omnidirectional images
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a4e84820-72e2-4453-b3ea-d26e37bf8cac · outbound
MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection OmniDrag: Enabling Motion Control for Omnidirectional Image-to-Video Generation
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1604f9f3-00db-4a10-b30c-a40fd6fab71e · outbound
MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection BrushEdit: All-In-One Image Inpainting and Editing
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aff9180f-e37b-4128-92cc-dfa09b296a79 · outbound
MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Adversarial supervision makes layout-to-image diffusion models thrive
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 918830ec-ccf9-4edd-aa77-01b5b8bc6c3d · outbound
MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Drag your noise: Interactive point-based editing via diffusion semantic propagation
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0216285b-ef45-4c89-9d0a-cc1b30b8f39b · outbound
MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f93afe6a-2e55-44bd-b5be-7f4a4feb3d03 · outbound
MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Step1X-Edit: A Practical Framework for General Image Editing
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4db43ad-cf9d-489f-b737-f0535a7cf1c4 · outbound
MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Magicquill: An intelligent interactive image editing system
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 26de2773-2e35-41dc-80b3-1fe414a377ba · outbound
MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Diffeditor: Boosting accuracy and flexibility on diffusion-based image editing
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1e50af1-e713-44a2-90d2-cf6413222ae1 · outbound
MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Dragondiffusion: Enabling drag-style manipulation on diffusion models
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation bc87defe-6221-4c52-85a4-5b9f57067a60 · outbound
MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection T2i-adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 709b76fe-869d-46c9-ad80-238c731b80df · outbound
MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Handiffuser: Text-to-image generation with realistic hand appearances
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8935cc25-1e62-46ae-8909-beac8916a29e · outbound
MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Transfer between Modalities with MetaQueries
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 261e470d-a3b1-47ce-b1f6-9b32998d414a · outbound
MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Drag your gan: Interactive point-based manipulation on the generative image manifold
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 248cb43e-66ce-4cf6-8f70-34f87b5d6722 · outbound
MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1fe2ce5e-4ff4-40db-bbb9-4a029952f774 · outbound
MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Learning transferable visual models from natural language supervision
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1266e1bc-ecbe-42e0-a468-dc2f0145448f · outbound
MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection High- resolution image synthesis with latent diffusion models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1320835-73ca-48a9-a69c-5a5bfcbcd9c2 · outbound
MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Photorealistic text-to-image diffusion models with deep language understanding.Advances in neural information processing systems, 35:36479–36494, 2022
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e154348-4e85-4f2d-ae28-df9b0607a3ce · outbound
MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Emu edit: Precise image editing via recognition and generation tasks
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3dbd24ee-55f1-4e00-b648-d9568b50ec81 · outbound
MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Dragdiffusion: Harnessing diffusion models for interactive point-based image editing
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74a839d2-7d8f-40ea-b90e-8fc5727eed50 · outbound
MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Insert Anything: Image Insertion via In-Context Editing in DiT
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38c37376-6b06-426a-9519-dd98316ca5ec · outbound
MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Invertible Consistency Distillation for Text-Guided Image Editing in Around 7 Steps
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a6114cb-3d3d-428e-8ad6-c3e28434033a · outbound
MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Chameleon: Mixed-Modal Early-Fusion Foundation Models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e986d506-1f5c-4c35-b9c7-24a90e475e81 · outbound
MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Gpt-4o system card, 2024
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f90c1bc6-8a38-4195-af3e-38350038b4a6 · outbound
MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection MetaMorph: Multimodal Understanding and Generation via Instruction Tuning
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40030a32-4388-47ac-be64-d4212c2b8a14 · outbound
MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Plug-and-play diffusion features for text-driven image-to-image translation
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94e5fa03-dff7-4e81-bb65-58b861e3ee4b · outbound
MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection FlexEdit: Marrying Free-Shape Masks to VLLM for Flexible Image Editing
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 850fd77c-bdfc-4d80-bbda-a51ff526850d · outbound
MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection 360dvd: Controllable panorama video generation with 360-degree video diffusion model
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3206f2f9-55c5-4c1a-ba13-1b381ed42fe6 · outbound
MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Emu3: Next-Token Prediction is All You Need
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d922f159-7107-48d9-afba-24819e1eeba1 · outbound
MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection In- stancediffusion: Instance-level control for image generation
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6e286aec-fcff-49d6-9e79-82842c48ebc6 · outbound
MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Genartist: Multimodal llm as an agent for unified image generation and editing.Advances in Neural Information Processing Systems, 37:128374–128395, 2024
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1516427e-e951-4437-bea2-2cad37912ca3 · outbound
MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Image quality assessment: from error visibility to structural similarity.IEEE transactions on image processing, 13(4):600– 612, 2004
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57fa69e6-d237-49bf-b1a2-53afc3db4e95 · outbound
MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Visionllm v2: An end-to-end generalist multimodal large language model for hundreds of vision-language tasks.Advances in Neural Information Processing Systems, 37:69925–69975, 2024
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a5c87af-ddff-4ea1-9bcd-82103e720673 · outbound
MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Show-o: One Single Transformer to Unify Multimodal Understanding and Generation
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d328811f-6996-4525-912b-34f4991f13f7 · outbound
MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07d7c50d-9c33-495c-96bb-081da015947a · outbound
MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection AnyEdit: Mastering Unified High-Quality Image Editing for Any Idea
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1a5f95d-8415-4ef0-9669-7b883dcc3d10 · outbound
MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Magicbrush: A manually annotated dataset for instruction-guided image editing.Advances in Neural Information Processing Systems, 36:31428–31449, 2023
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3abffb1d-1c68-472d-97e3-a1aa62985ec2 · outbound
MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Adding conditional control to text-to-image diffusion models
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f58d4b08-9b8a-4631-973f-5bfe97a10850 · outbound
MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection The unrea- sonable effectiveness of deep features as a perceptual metric
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c353a799-e1be-4c50-8eb6-d00667549de2 · outbound
MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection GoodDrag: Towards Good Practices for Drag Editing with Diffusion Models
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c66d4d9f-6be9-41e9-8238-2925bfbf133b · outbound
MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Ultraedit: Instruction-based fine-grained image editing at scale.Advances in Neural Information Processing Systems, 37:3058–3093, 2024
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0025a0dc-b722-4309-ab78-5eb73859901f · inbound
TextWand: A Unified Framework for Scene Text Editing MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.