Pith. sign in

Paper Citation Record · LEDGER

InstructPix2Pix: Learning to Follow Image Editing Instructions

As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 66 inbound Pith citation observations for arXiv:2211.09800.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2211.09800 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 66 of 66 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:35:31.628468Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T10:19:47.353241Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 34651357-07bd-470f-85cd-dc9b41cfc461 · inbound

Adding Conditional Control to Text-to-Image Diffusion Models cites this paper.

Adding Conditional Control to Text-to-Image Diffusion Models InstructPix2Pix: Learning to Follow Image Editing Instructions

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:43:11.001734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-16T22:43:10.880338Z digest=sha256:5f2454f30cd72d6c24a7a4e1bd75ef71e957e53c4076ea4f4f1db0a88b39b934

Observation 922c80f7-4e45-4891-b88f-d5a711fbae3e · inbound

Scaling Robot Learning with Semantically Imagined Experience cites this paper.

Scaling Robot Learning with Semantically Imagined Experience InstructPix2Pix: Learning to Follow Image Editing Instructions

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-17T18:59:10.533140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-17T18:59:10.352342Z digest=sha256:664a3c9928161a35ce919339532d37c361a8797761f35538e1bdd672b5dc8f92

Observation fa1e7c72-3bf2-44cf-afa8-309a3a5c3638 · inbound

Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models cites this paper.

Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models InstructPix2Pix: Learning to Follow Image Editing Instructions

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:50:24.120824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-13T22:50:24.053411Z digest=sha256:4153736d59c533727e833ce6f5d5ea09f5f89f9d4eb0a2820cf66c175efaada0

Observation 11808f46-2e0f-44c1-ab10-41bbf9d5f819 · inbound

Visual Instruction Tuning cites this paper.

Visual Instruction Tuning InstructPix2Pix: Learning to Follow Image Editing Instructions

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:22:03.765463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-11T08:22:03.403362Z digest=sha256:22d99c89d5ea1aa3448016be9f2a444f0a9010af59579cf4050e7a423e0f6509

Observation 49a2736b-01a0-4c07-9fce-86ac1e7c455e · inbound

Scaling Properties of Diffusion Models for Perceptual Tasks cites this paper.

Scaling Properties of Diffusion Models for Perceptual Tasks InstructPix2Pix: Learning to Follow Image Editing Instructions

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T22:02:10.986162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T22:02:10.986162Z digest=sha256:6f000300287b92b1a1d78341a47c9e0ee23f403236e9942694c8b61091a1fb5c

Observation 87a08015-fa52-4bcc-b0d4-e156fc1b40b2 · inbound

Test-time Conditional Text-to-Image Synthesis Using Diffusion Models cites this paper.

Test-time Conditional Text-to-Image Synthesis Using Diffusion Models InstructPix2Pix: Learning to Follow Image Editing Instructions

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:50.216713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:21:50.216713Z digest=sha256:23cf7b5a400a271c3b9749587e0e2c6ad1e2f9a4d6fa346c10841b4c414f03cb

Observation 10e527bd-ec3e-4cfe-876c-59d1c3e5d242 · inbound

Generating Compositional Scenes via Text-to-image RGBA Instance Generation cites this paper.

Generating Compositional Scenes via Text-to-image RGBA Instance Generation InstructPix2Pix: Learning to Follow Image Editing Instructions

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T19:15:04.376912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:15:04.376912Z digest=sha256:994dfc5a84b4c6381947bdb9d592a1c0c3a6e113304f59b3b68b77c1252e9c96

Observation 4fa9bb79-0659-4e63-b02b-5ad9b44968bd · inbound

Medical Video Generation for Disease Progression Simulation cites this paper.

Medical Video Generation for Disease Progression Simulation InstructPix2Pix: Learning to Follow Image Editing Instructions

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T18:11:36.548692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:11:36.548692Z digest=sha256:527ceb81674ea2e47ce162445642c9a373ebbe22e5b24e4e10c6b78afa2fcfe8

Observation c43ad6e2-5f43-4618-aa74-d675d58c2d94 · inbound

DynamicAvatars: Accurate Dynamic Facial Avatars Reconstruction and Precise Editing with Diffusion Models cites this paper.

DynamicAvatars: Accurate Dynamic Facial Avatars Reconstruction and Precise Editing with Diffusion Models InstructPix2Pix: Learning to Follow Image Editing Instructions

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T14:02:24.267946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:02:24.267946Z digest=sha256:1f53621f9c2da419f798b21fe3bd7a839a4ef8fab1fdd5ab2e7c3e01d030971d

Observation 434f9eea-d3f4-4f4b-a7c0-2c1979eaef30 · inbound

Generative Image Layer Decomposition with Visual Effects cites this paper.

Generative Image Layer Decomposition with Visual Effects InstructPix2Pix: Learning to Follow Image Editing Instructions

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T11:50:01.253553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:50:01.253553Z digest=sha256:32ba2d3917cd6051e50810cb74b6cf22f626cbe8989f61a4f38238d5fd91a3be

Observation 82d60bff-21e2-439c-a7d3-e7b271963f40 · inbound

DreamDance: Animating Human Images by Enriching 3D Geometry Cues from 2D Poses cites this paper.

DreamDance: Animating Human Images by Enriching 3D Geometry Cues from 2D Poses InstructPix2Pix: Learning to Follow Image Editing Instructions

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T05:29:09.411240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:29:09.411240Z digest=sha256:cfa9da0c5eb16c21635dd7cae7ffff7cf4202c61a66cfc8dd446e1ebc2c3af91

Observation ac375ce6-948c-480a-b0be-7e486dc5038e · inbound

Panoptic Diffusion Models: co-generation of images and segmentation maps cites this paper.

Panoptic Diffusion Models: co-generation of images and segmentation maps InstructPix2Pix: Learning to Follow Image Editing Instructions

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T23:02:16.067052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:02:16.067052Z digest=sha256:4a130c49c74822336aea39b9121e82eac106fc7512eca3d2daa294b4fa8d0278

Observation db3a6920-fcbc-4a6f-a2c8-c679f3d23b69 · inbound

Do We Need to Design Specific Diffusion Models for Different Tasks? Try ONE-PIC cites this paper.

Do We Need to Design Specific Diffusion Models for Different Tasks? Try ONE-PIC InstructPix2Pix: Learning to Follow Image Editing Instructions

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T20:35:35.170866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:35:35.170866Z digest=sha256:fffaac8c504b5c06f44d4c66dea9fc307773fb3e190458a4f3891aa4b07447c1

Observation 15a9f3fc-9913-456b-9b02-345d978cf95b · inbound

MoViE: Mobile Diffusion for Video Editing cites this paper.

MoViE: Mobile Diffusion for Video Editing InstructPix2Pix: Learning to Follow Image Editing Instructions

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T19:34:26.356105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:34:26.356105Z digest=sha256:2d26bd5fcf5b06643d0084d5131faab653eae37467d055b2450e7cd57132476f

Observation dff65b37-337a-4aed-ae0f-3ebf9808426e · inbound

TryOffAnyone: Tiled Cloth Generation from a Dressed Person cites this paper.

TryOffAnyone: Tiled Cloth Generation from a Dressed Person InstructPix2Pix: Learning to Follow Image Editing Instructions

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T17:48:01.601880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:48:01.601880Z digest=sha256:c59d1acfb75e5a534bcffe24a398cbfdde6668e905d01f6bd90cf8c918acb8f0

Observation f35f8025-a3a7-4da2-b35d-adaf8d018768 · inbound

Learning Complex Non-Rigid Image Edits from Multimodal Conditioning cites this paper.

Learning Complex Non-Rigid Image Edits from Multimodal Conditioning InstructPix2Pix: Learning to Follow Image Editing Instructions

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T16:16:14.631506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:16:14.631506Z digest=sha256:94032ca313bdea8be1faa998765929fa47eec5b2d23757eda17c2de24d98a05a

Observation 17b58c6a-19b7-46cb-bbd7-041225f24725 · inbound

Relational Programming with Foundation Models cites this paper.

Relational Programming with Foundation Models InstructPix2Pix: Learning to Follow Image Editing Instructions

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T12:13:21.378477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:13:21.378477Z digest=sha256:c3b37835b0ea0a0368c5a4e2ff9930eca2b942ac05259dcc96be15c61fd9924a

Observation 927a235e-6e8a-4d3a-bd6c-b0743d3740d5 · inbound

DiffuEraser: A Diffusion Model for Video Inpainting cites this paper.

DiffuEraser: A Diffusion Model for Video Inpainting InstructPix2Pix: Learning to Follow Image Editing Instructions

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T19:25:45.092982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:25:45.092982Z digest=sha256:0cecb08ee810d2f21b650fb0e2723162f9d3ece5965014a8160e5f89404375b3

Observation e15659c6-0fe5-466e-a797-2181377d1e4c · inbound

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions cites this paper.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions InstructPix2Pix: Learning to Follow Image Editing Instructions

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T17:30:47.782994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:30:47.782994Z digest=sha256:35376102dfcaac325a73b1b50e2fe19a04a5f86f1207c908f5adb7c18beb9cbc

Observation c6aeb455-3477-4235-b048-7c12595c190b · inbound

Hunyuan3D 2.0: Scaling Diffusion Models for High Resolution Textured 3D Assets Generation cites this paper.

Hunyuan3D 2.0: Scaling Diffusion Models for High Resolution Textured 3D Assets Generation InstructPix2Pix: Learning to Follow Image Editing Instructions

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-23T04:55:25.035559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-23T04:53:58.465445Z digest=sha256:222a53a6c503eac311fdfa75c732a4facd8775c9f4582e7766044d04eedbc14d

Observation bd8ef176-e7e2-47f5-ac8b-04c75b823e8c · inbound

Generative AI for Vision: A Comprehensive Study of Frameworks and Applications cites this paper.

Generative AI for Vision: A Comprehensive Study of Frameworks and Applications InstructPix2Pix: Learning to Follow Image Editing Instructions

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T00:58:34.705487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:58:34.705487Z digest=sha256:3577f9c6a5c79b0a032d5be516542bb1f18cbb1c36f7d150be7ec6496316ed4f

Observation 609d7847-06fe-45fd-81f0-198bdca9e3db · inbound

ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features cites this paper.

ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features InstructPix2Pix: Learning to Follow Image Editing Instructions

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T22:51:30.349907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:51:30.349907Z digest=sha256:c81396fb260b78b5c54894ef77fa96deb642a09264d488b18a2ba3636ae7dcb3

Observation 386e9fb3-b3f9-4fce-82af-47e5f9ba8d05 · inbound

Dynamic watermarks in images generated by diffusion models cites this paper.

Dynamic watermarks in images generated by diffusion models InstructPix2Pix: Learning to Follow Image Editing Instructions

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T23:18:33.138406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:18:33.138406Z digest=sha256:3c0060db2763375697c6c920792ff56ad7d76f8b94fcab67ab368700abe296c3

Observation 1bb1be5e-c29d-4e0f-a6d7-f475657a2bcd · inbound

DCT-Shield: A Robust Frequency Domain Defense against Malicious Image Editing cites this paper.

DCT-Shield: A Robust Frequency Domain Defense against Malicious Image Editing InstructPix2Pix: Learning to Follow Image Editing Instructions

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T10:35:31.628468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:35:31.628468Z digest=sha256:0465aee86233b7e125d69d692db9aea264470f8edee4c363a85a1b9df8a65644

Observation d62cbe5c-fae3-47a9-bc43-d094f7718c17 · inbound

Multi-Modal Language Models as Text-to-Image Model Evaluators cites this paper.

Multi-Modal Language Models as Text-to-Image Model Evaluators InstructPix2Pix: Learning to Follow Image Editing Instructions

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:11.180247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:11.180247Z digest=sha256:73b3e3af540af358b2f5f82a8fefec7d84bf8dc3f55571c57e4a7dc579e5cf88

Observation 1fc12b40-2d61-4f11-ba21-05dbdbe4eec2 · inbound

Object detection in adverse weather conditions for autonomous vehicles using Instruct Pix2Pix cites this paper.

Object detection in adverse weather conditions for autonomous vehicles using Instruct Pix2Pix InstructPix2Pix: Learning to Follow Image Editing Instructions

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T22:03:31.519796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:03:31.519796Z digest=sha256:7075ae8a4354f1833d64b21b9f767c953a433661849dc32d24f7cd58e0914d0e

Observation 4b93ca34-00f3-4f49-aa6f-cb0bacb2aa0b · inbound

IMAGE-ALCHEMY: Advancing subject fidelity in personalised text-to-image generation cites this paper.

IMAGE-ALCHEMY: Advancing subject fidelity in personalised text-to-image generation InstructPix2Pix: Learning to Follow Image Editing Instructions

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-15T21:06:44.365413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:06:44.365413Z digest=sha256:ee21519a4109dd8101786f4586587b0677c3b14343cd3b7f7b1de6c6586ecb55

Observation a1c9d41b-2786-45f5-9d06-965b0b4c424b · inbound

OmniStyle: Filtering High Quality Style Transfer Data at Scale cites this paper.

OmniStyle: Filtering High Quality Style Transfer Data at Scale InstructPix2Pix: Learning to Follow Image Editing Instructions

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:12.277342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:12.277342Z digest=sha256:544d342400ee175e69e8d1dde12d07b6486fa7a9ba6a369592a49fd9562a4a77

Observation 3bee0394-b482-426b-abc3-b620aa31cd07 · inbound

ConText-CIR: Learning from Concepts in Text for Composed Image Retrieval cites this paper.

ConText-CIR: Learning from Concepts in Text for Composed Image Retrieval InstructPix2Pix: Learning to Follow Image Editing Instructions

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:13.438855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:13.438855Z digest=sha256:af8a56309d172d8f4de892c6b71f74b19d05dc861e9caddf79ce0b51987919e8

Observation fdc78d5b-3e9c-4473-89d0-cfb761724a05 · inbound

D-Fusion: Direct Preference Optimization for Aligning Diffusion Models with Visually Consistent Samples cites this paper.

D-Fusion: Direct Preference Optimization for Aligning Diffusion Models with Visually Consistent Samples InstructPix2Pix: Learning to Follow Image Editing Instructions

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T13:23:01.364254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:23:01.364254Z digest=sha256:a74bc4b4a506e287303dc28fc7ac851ab51358c31ed413e01d50a64ee23610d4

Observation cccf28c1-60a8-4def-a0fb-1ab155fb55d4 · inbound

Multiverse Through Deepfakes: The MultiFakeVerse Dataset of Person-Centric Visual and Conceptual Manipulations cites this paper.

Multiverse Through Deepfakes: The MultiFakeVerse Dataset of Person-Centric Visual and Conceptual Manipulations InstructPix2Pix: Learning to Follow Image Editing Instructions

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:33.037487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:33.037487Z digest=sha256:7785f2b3672d1787dc3dc076590acac8113843387c253cea5383e88c1769e52c

Observation 61af728a-1a10-497a-ac43-c29e965d6ebf · inbound

EditInspector: A Benchmark for Evaluation of Text-Guided Image Edits cites this paper.

EditInspector: A Benchmark for Evaluation of Text-Guided Image Edits InstructPix2Pix: Learning to Follow Image Editing Instructions

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:36.403646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:36.403646Z digest=sha256:8b68ee21de009ff35e2d6e217db5438e6565e5e2bdac6bc21df2eee2b7ea1a3c

Observation baed8f9a-6e54-423d-8192-56728cce2a6f · inbound

Enhance Multimodal Consistency and Coherence for Text-Image Plan Generation cites this paper.

Enhance Multimodal Consistency and Coherence for Text-Image Plan Generation InstructPix2Pix: Learning to Follow Image Editing Instructions

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T04:14:28.819047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:14:28.819047Z digest=sha256:f230a43f64efae3de75f93064b4d60f5640dbadd3c1a1285ef49d3d2c154b677

Observation f3bbb458-e222-4b58-93c6-596f1e009382 · inbound

Balancing Preservation and Modification: A Region and Semantic Aware Metric for Instruction-Based Image Editing cites this paper.

Balancing Preservation and Modification: A Region and Semantic Aware Metric for Instruction-Based Image Editing InstructPix2Pix: Learning to Follow Image Editing Instructions

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T00:43:14.170464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:43:14.170464Z digest=sha256:e01b458229d18aefe9a5d200bb314fc5de178f794310f3971dac5b5997dd1147

Observation ea8fbe9a-87ea-411f-90f5-9c4c64d68460 · inbound

VectorEdits: A Dataset and Benchmark for Instruction-Based Editing of Vector Graphics cites this paper.

VectorEdits: A Dataset and Benchmark for Instruction-Based Editing of Vector Graphics InstructPix2Pix: Learning to Follow Image Editing Instructions

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:36.246517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:48:36.246517Z digest=sha256:702c0883590311afd29ae35c88af13591f26ea998a7336d7c96581bea8691622

Observation 1a5bff8e-7310-44af-8f0a-d530163efac2 · inbound

IdeaBlocks: Expressing and Reusing Divergent Intents for Graphic Design Exploration using Generative AI cites this paper.

IdeaBlocks: Expressing and Reusing Divergent Intents for Graphic Design Exploration using Generative AI InstructPix2Pix: Learning to Follow Image Editing Instructions

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-19T02:11:59.520541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-19T02:07:33.016608Z digest=sha256:8b0809dc962772f70dce30fb73cc49a0a8c1295e1e9ac59cbf0b62fbeffe4367

Observation 08e3e62b-3e29-48f2-8d61-66cffb66356e · inbound

DiffIER: Optimizing Diffusion Models with Iterative Error Reduction cites this paper.

DiffIER: Optimizing Diffusion Models with Iterative Error Reduction InstructPix2Pix: Learning to Follow Image Editing Instructions

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-05T19:03:04.735019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:03:04.735019Z digest=sha256:750ccb16e520b158e621108d74ef96b6992e97132c9473f50851643f4e461146

Observation 8cb9e917-1ede-49b4-9232-e5366f38c831 · inbound

Exploring Diffusion Models for Generative Forecasting of Financial Charts cites this paper.

Exploring Diffusion Models for Generative Forecasting of Financial Charts InstructPix2Pix: Learning to Follow Image Editing Instructions

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T16:42:42.470069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:42:42.470069Z digest=sha256:0fb8a70fb7323f5ac1ea788d218ec230aa91b9cff0a95b3df912395e3247c840

Observation 4bdf3bb4-eb11-4b9a-a585-9812b8da4bd2 · inbound

Plot'n Polish: Zero-shot Story Visualization and Disentangled Editing with Text-to-Image Diffusion Models cites this paper.

Plot'n Polish: Zero-shot Story Visualization and Disentangled Editing with Text-to-Image Diffusion Models InstructPix2Pix: Learning to Follow Image Editing Instructions

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:15.273719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:15.273719Z digest=sha256:da043c3a023ef97b41d71879e62cd4b14b53e3131971fbc96b8734812bf826ab

Observation 8f01c66d-4a3d-4e75-bf91-16c3b69fc7ed · inbound

Delta Rectified Flow Sampling for Text-to-Image Editing cites this paper.

Delta Rectified Flow Sampling for Text-to-Image Editing InstructPix2Pix: Learning to Follow Image Editing Instructions

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:02:48.587947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-18T19:02:18.176561Z digest=sha256:ea94663e13629a4fcfddb7cad388d7f27069bccaccd08f02f705d9a6ed3b2b34

Observation 06513b82-e5ca-455f-851a-2fe0d7f0ea78 · inbound

Cultural Counterfactuals: Evaluating Cultural Biases in Large Vision-Language Models with Counterfactual Examples cites this paper.

Cultural Counterfactuals: Evaluating Cultural Biases in Large Vision-Language Models with Counterfactual Examples InstructPix2Pix: Learning to Follow Image Editing Instructions

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T19:26:23.720833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T19:26:23.720833Z digest=sha256:4102e9d0fd98321ca6a8229b65cee61918ec6a392b3710b1b7d3e10b7d8b9dbd

Observation edddbb6d-b6b1-42bd-af50-7f243d42f846 · inbound

Edit in 2D, Verify in 3D: Reinforcement Learning for Multi-view Consistent Scene Editing cites this paper.

Edit in 2D, Verify in 3D: Reinforcement Learning for Multi-view Consistent Scene Editing InstructPix2Pix: Learning to Follow Image Editing Instructions

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T19:14:04.763185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:14:04.763185Z digest=sha256:1a6a7b933b698cb830854c9ba0f310f5c51087b75978a4a5bf5853dc0c936d3f

Observation 355b8205-b0dc-438d-89df-505769fd47cb · inbound

Pinterest Canvas: Large-Scale Image Generation at Pinterest cites this paper.

Pinterest Canvas: Large-Scale Image Generation at Pinterest InstructPix2Pix: Learning to Follow Image Editing Instructions

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-15T13:50:51.666642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T13:50:51.666642Z digest=sha256:9332d1ca299facd611332bcc7cff95b960955dc8a6291f6636d78e1b46247f89

Observation a6fb18b7-afa4-4db4-8329-5f3cf95c5d2c · inbound

Setting-Matched and Semantics-Scaled Benchmarking of One-Step Generative Models Against Multistep Diffusion and Flow Models cites this paper.

Setting-Matched and Semantics-Scaled Benchmarking of One-Step Generative Models Against Multistep Diffusion and Flow Models InstructPix2Pix: Learning to Follow Image Editing Instructions

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T12:20:00.845639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-15T12:15:38.186914Z digest=sha256:1e9f02e8c9e6bf13bee2fb0dd7cda6244388e4235827f0c4bcc320ed9a8740ee

Observation 5a2df389-9536-4f65-9144-3fb1f3fe585d · inbound

From Video to Control: A Survey of Learning Manipulation Interfaces from Temporal Visual Data cites this paper.

From Video to Control: A Survey of Learning Manipulation Interfaces from Temporal Visual Data InstructPix2Pix: Learning to Follow Image Editing Instructions

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T17:03:01.182136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-13T17:02:18.358675Z digest=sha256:f0013a581b15dcf43a2254e3cf908df632e852ea2379ba3902406fdfb5aa1acd

Observation 5c975759-e2d1-4b93-acea-1314cb1582fe · inbound

PostureObjectstitch: Anomaly Image Generation Considering Assembly Relationships in Industrial Scenarios cites this paper.

PostureObjectstitch: Anomaly Image Generation Considering Assembly Relationships in Industrial Scenarios InstructPix2Pix: Learning to Follow Image Editing Instructions

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:20:30.532026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T14:16:40.509571Z digest=sha256:599a2062cd2b100add7f9ba8ee63732cb3438a8d56c55c53cf7fd3120cd599ad

Observation c03836f6-970c-464a-90b5-3d94a2152bd8 · inbound

PhysEdit: Physically-Consistent Region-Aware Image Editing via Adaptive Spatio-Temporal Reasoning cites this paper.

PhysEdit: Physically-Consistent Region-Aware Image Editing via Adaptive Spatio-Temporal Reasoning InstructPix2Pix: Learning to Follow Image Editing Instructions

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:36:10.241697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-09T19:33:13.609552Z digest=sha256:c56d578d837d140371c0a68ad2bca3bc1fc5ca48f4226e2c4780f0a6b8c0a018

Observation 05d3ca51-54b0-4702-82d0-3e083d7410fe · inbound

Stylistic Attribute Control in Latent Diffusion Models cites this paper.

Stylistic Attribute Control in Latent Diffusion Models InstructPix2Pix: Learning to Follow Image Editing Instructions

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:15:38.115740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-08T18:43:21.423012Z digest=sha256:4946cf5e0f843224ab2dd05b0211ee11d0ee7f9b3c8bdad3e4f5d0e38c48f528

Observation b9fab31e-863d-4804-afa5-3b8c8244141d · inbound

Edit-GRPO: A Locality-Preserving Policy Optimization Framework for Image Editing cites this paper.

Edit-GRPO: A Locality-Preserving Policy Optimization Framework for Image Editing InstructPix2Pix: Learning to Follow Image Editing Instructions

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-19T20:47:45.919637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-19T20:44:54.356658Z digest=sha256:daaf379c870a56b18d11b98108f0995dcf2cf36d0186b75554f3a95c3f84b8ec

Observation 86a3163b-14b5-4e60-91a9-8dad8fbd0a02 · inbound

Functionalization via Structure Completion and Motion Rectification cites this paper.

Functionalization via Structure Completion and Motion Rectification InstructPix2Pix: Learning to Follow Image Editing Instructions

Reference 103

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T12:28:16.854252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-20T12:25:07.157086Z digest=sha256:f69ec57b831ef8f6aab3fdc862214e25ad52f6dedbfb1a0ca39e0d5e7db38ff2

Observation 8ce526d6-a5da-497b-a961-57b2ea8a45b8 · inbound

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation cites this paper.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation InstructPix2Pix: Learning to Follow Image Editing Instructions

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:26:20.779221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:87bf39f74843ac56032a2bb61b09783cd3115141ca6d0373ccacd6b7f82a6a79

Observation 5ebbcc7c-0bda-4033-8050-a93dad14a216 · inbound

SimInsert: Seamless Video Object Insertion via Regional Sparse Attention Fusion cites this paper.

SimInsert: Seamless Video Object Insertion via Regional Sparse Attention Fusion InstructPix2Pix: Learning to Follow Image Editing Instructions

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:50:21.388354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-25T04:46:50.749223Z digest=sha256:065c06c8119c794df9d98db4ed6f6a80b55fdc73b0aee4a6ef81139d32437e0f

Observation 77169000-6a4a-4a34-a9c1-f9bb71bc9ffc · inbound

Toward Native Multimodal Modeling: A Roadmap cites this paper.

Toward Native Multimodal Modeling: A Roadmap InstructPix2Pix: Learning to Follow Image Editing Instructions

Reference 102

Resolution
verified exact
arxiv_id, observed 2026-06-29T23:04:01.954409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-29T22:58:38.610609Z digest=sha256:a605f5eec3a78074f1313969b4d11f4313f64abfa5ee357bf5cf3ea6163cfc8a

Observation 863259a4-4e01-465e-b3e2-2d932d79820b · inbound

ROGLE: Robust Global-Local Alignment with Automated Region Supervision for Text-Based Person Search cites this paper.

ROGLE: Robust Global-Local Alignment with Automated Region Supervision for Text-Based Person Search InstructPix2Pix: Learning to Follow Image Editing Instructions

Reference 135

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T22:36:16.968065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-06-28T15:18:17.427707Z digest=sha256:ad89080b1da3d48112cdbbb8bb048aeadd5359aed98e5adf81de109b11fba267

Observation 04f877a9-7018-40cc-bdcd-7835cf3eec3c · inbound

ROGLE: Robust Global-Local Alignment with Automated Region Supervision for Text-Based Person Search cites this paper.

ROGLE: Robust Global-Local Alignment with Automated Region Supervision for Text-Based Person Search InstructPix2Pix: Learning to Follow Image Editing Instructions

Reference 135

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T23:17:28.938252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-07-02T23:17:03.456746Z digest=sha256:904621fc1886ad02b6b9bf0fd615ff6ea7c793a9d8efaba45572e174eeb187a7

Observation c2042e52-ee71-4825-838d-6b24c348afd3 · inbound

Towards Characterizing Scientific Image Utility and Upgradability cites this paper.

Towards Characterizing Scientific Image Utility and Upgradability InstructPix2Pix: Learning to Follow Image Editing Instructions

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:06:27.858570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-28T11:10:37.254428Z digest=sha256:030b68c2f4a23884686569efe660fded2cca70f3591104e67d47b8e8fc1e2db6

Observation 89529731-c2a7-4cea-a778-0bc9f75f5248 · inbound

Organizational Control Layer: Governance Infrastructure at the Execution Boundary of LLM Agent Systems cites this paper.

Organizational Control Layer: Governance Infrastructure at the Execution Boundary of LLM Agent Systems InstructPix2Pix: Learning to Follow Image Editing Instructions

Reference 114

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T11:16:53.842480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-06-28T04:17:18.483765Z digest=sha256:82263766552af0e01d80aa5b5ea97e04b3d2f9659f6907774bbc7cba4ef1e533

Observation c9a827fd-2e15-493f-98fd-f320ebffea8f · inbound

V2V-Bench: A Comprehensive Benchmark for Video-to-Video Generation Evaluation cites this paper.

V2V-Bench: A Comprehensive Benchmark for Video-to-Video Generation Evaluation InstructPix2Pix: Learning to Follow Image Editing Instructions

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:16:56.708468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-28T02:20:05.748127Z digest=sha256:c7f76e48434b8d1b24720fd548eeb3a632255b96089d6e191684c19fa60aa075

Observation a289de02-4cc8-4a9d-897a-e471dd055dae · inbound

A Systematic Study of Behavioral Cloning for Scientific Data Annotation cites this paper.

A Systematic Study of Behavioral Cloning for Scientific Data Annotation InstructPix2Pix: Learning to Follow Image Editing Instructions

Reference 252

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T16:23:39.126376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-06-29T16:23:08.402194Z digest=sha256:cc4675357226dcfd9f626abbe08f02701d1cf4ad0d257bac4ec1fe3d45c535ae

Observation b8a358a4-3344-4a8c-93da-c1a1a31630ee · inbound

RS-Gen: A Multi-Stage Agentic Framework for Reasoning and Search-Augmented Image Generation cites this paper.

RS-Gen: A Multi-Stage Agentic Framework for Reasoning and Search-Augmented Image Generation InstructPix2Pix: Learning to Follow Image Editing Instructions

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:19:47.354767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-26T08:59:40.467458Z digest=sha256:f9c18818a3a79211499e4042d7a3d33a7fdeb66d933dc3e6341e14c20f096437

Observation d7f77fce-cb2f-4147-8067-3f8ee17353b0 · inbound

Xema: Efficient Diffusion Serving through Fine-Grained Memory Management and Auto-Configuration cites this paper.

Xema: Efficient Diffusion Serving through Fine-Grained Memory Management and Auto-Configuration InstructPix2Pix: Learning to Follow Image Editing Instructions

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-14T06:49:04.349040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T06:49:04.349040Z digest=sha256:c800e47c14ea5bbdb9b02a0ae558717fc8a1f221b4c1a771ab9a35012a248d61

Observation 3f2f0e2d-9c25-493c-8a7a-0a695b88da38 · inbound

Think, Plan, Paint: Layout-Aware Reasoning for Controllable Image Generation in Unified Models cites this paper.

Think, Plan, Paint: Layout-Aware Reasoning for Controllable Image Generation in Unified Models InstructPix2Pix: Learning to Follow Image Editing Instructions

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T21:07:57.019322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:07:57.019322Z digest=sha256:e5615be7b331f945458a83c51b96b3590a362cc6c10342368b578d037069476e

Observation c1b52947-7354-4cd5-b1bc-5a32b213b692 · inbound

Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing cites this paper.

Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing InstructPix2Pix: Learning to Follow Image Editing Instructions

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T13:38:58.541944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:38:58.541944Z digest=sha256:c8f5550b46eb09a20fa37fc207f1153c8252ef873cbd60dfbca23f4c3c4463b9

Observation d86f7635-6d70-449c-b424-9f0d501077c6 · inbound

OSVE: One Step Video Editing with One Step Diffusion Models cites this paper.

OSVE: One Step Video Editing with One Step Diffusion Models InstructPix2Pix: Learning to Follow Image Editing Instructions

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T11:29:13.382881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:29:13.382881Z digest=sha256:02948251390542e49cee61baa4a9bab91298da7b0f09f047753c44d11261a9f8

Observation b6e65be5-c2a9-4b52-8761-8bbd738cea59 · inbound

Localize, Don't Beautify: Client-Side Control of Image-Editing APIs for Cosmetic Surgery Previews cites this paper.

Localize, Don't Beautify: Client-Side Control of Image-Editing APIs for Cosmetic Surgery Previews InstructPix2Pix: Learning to Follow Image Editing Instructions

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T15:02:45.733977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:02:45.733977Z digest=sha256:39539a1609ed71a9ec26056ce4562032bced7863cb60c5ef0d78a7d382638ad2

Observation 78ddb145-ebb8-495e-b56c-1fe9612cbb23 · inbound

SI-Edit: Toward Sketch-Instruction Guided Local Image Editing with Pixel-Level Precision cites this paper.

SI-Edit: Toward Sketch-Instruction Guided Local Image Editing with Pixel-Level Precision InstructPix2Pix: Learning to Follow Image Editing Instructions

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T23:41:46.478815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:41:46.478815Z digest=sha256:46935ccbb5616ee1a04b3e15873effd4d6d98cc69add086f3cdfbbd89dcb4cd7