Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T20:05:16.256044Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 1 inbound Pith citation observation for arXiv:2502.05165.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T20:05:16.256044Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T23:35:15.341406Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-06T23:35:15.923842Z
58 of 58 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation a09b5255-306e-4e4f-874b-fa53aa34b126 · outbound
Multitwine: Multi-Object Compositing with Text and Layout Control Cross-image attention for zero- shot appearance transfer
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9b0ac045-08d8-484c-90f3-58f2e18265de · outbound
Multitwine: Multi-Object Compositing with Text and Layout Control Flamingo: a visual language model for few-shot learning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1305d9a-3c8c-4819-b371-1232a3b83acc · outbound
Multitwine: Multi-Object Compositing with Text and Layout Control Vip- llava: Making large multimodal models understand arbitrary visual prompts
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e06f5dcb-2bfb-4d6e-95c5-b1b6f20e0033 · outbound
Multitwine: Multi-Object Compositing with Text and Layout Control Masactrl: Tuning-free mu- tual self-attention control for consistent image synthesis and editing
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 0b29a90d-0dba-41e3-8e3e-9d8987dc54ce · outbound
Multitwine: Multi-Object Compositing with Text and Layout Control Re-Imagen: Retrieval-Augmented Text-to-Image Generator
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc45db59-27d1-4316-8531-c3c37be8189a · outbound
Multitwine: Multi-Object Compositing with Text and Layout Control Subject-driven text-to-image generation via apprenticeship learning
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e9684e9a-4df2-466f-bb84-2292e57918ea · outbound
Multitwine: Multi-Object Compositing with Text and Layout Control AnyDoor: Zero-shot Object-level Image Customization
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c81b1664-3105-45b1-a436-1affe30cc0f3 · outbound
Multitwine: Multi-Object Compositing with Text and Layout Control Be Yourself: Bounded Attention for Multi-Subject Text-to-Image Generation
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d1a345a-e7fb-47e4-a468-6d2ca1e9b7ac · outbound
Multitwine: Multi-Object Compositing with Text and Layout Control DreamLLM: Synergistic Multimodal Comprehension and Creation
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd0abeb5-e9a7-4178-94c4-1f5c602025b3 · outbound
Multitwine: Multi-Object Compositing with Text and Layout Control PaLM-E: An Embodied Multimodal Language Model
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 837e58cb-59c3-4b1b-889b-a222bf3bd5a2 · outbound
Multitwine: Multi-Object Compositing with Text and Layout Control Scaling recti- fied flow transformers for high-resolution image synthesis
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 341e031a-cf5f-497a-9792-6bcec4aacf70 · outbound
Multitwine: Multi-Object Compositing with Text and Layout Control An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2aa2c571-9ff0-4ef6-92f2-19c13569d0d1 · outbound
Multitwine: Multi-Object Compositing with Text and Layout Control Prompt-to-Prompt Image Editing with Cross Attention Control
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27d371a3-5e0a-4a91-93c2-0ed07b72e59f · outbound
Multitwine: Multi-Object Compositing with Text and Layout Control Clipscore: A reference-free evaluation met- ric for image captioning
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 05e0f83e-8f4a-4001-8fd9-3a8c887f304c · outbound
Multitwine: Multi-Object Compositing with Text and Layout Control Taming Encoder for Zero Fine-tuning Image Customization with Text-to-Image Diffusion Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c2e602d-e6af-4798-b16f-8c8a4187f114 · outbound
Multitwine: Multi-Object Compositing with Text and Layout Control Rendering synthetic objects into legacy photographs
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a365751e-6298-45bf-abae-01df37cb318c · outbound
Multitwine: Multi-Object Compositing with Text and Layout Control 3d object manipulation in a single photograph using stock 3d models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation fc442340-45ae-4eaf-aa15-629daae51b1b · outbound
Multitwine: Multi-Object Compositing with Text and Layout Control Gen- erating images with multimodal language models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67e4255b-c2fd-400f-a0da-e82010b9a3b4 · outbound
Multitwine: Multi-Object Compositing with Text and Layout Control Open images v5 text annotation and yet another mask text spotter
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation bb5a39c4-c7a1-4822-925e-b6c12b0d0de9 · outbound
Multitwine: Multi-Object Compositing with Text and Layout Control Putting people in their place: Affordance-aware hu- man insertion into scenes
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5c546ea4-3ffb-4918-a06e-95704f0b3072 · outbound
Multitwine: Multi-Object Compositing with Text and Layout Control Photo clip art
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation fb1d63c8-216a-4531-bc3c-b1b7af3a797f · outbound
Multitwine: Multi-Object Compositing with Text and Layout Control Blip-diffusion: Pre- trained subject representation for controllable text-to-image generation and editing
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ba07cc92-e700-4c65-b17c-d117e8ad268a · outbound
Multitwine: Multi-Object Compositing with Text and Layout Control UNIMO-G: Unified Image Generation through Multimodal Conditional Diffusion
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44f52a1d-a541-41f1-b616-4545299ffd2e · outbound
Multitwine: Multi-Object Compositing with Text and Layout Control Improved baselines with visual instruction tuning, 2023
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 27ca87c1-5ba6-48a3-a799-d75cc0711370 · outbound
Multitwine: Multi-Object Compositing with Text and Layout Control Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 304baa49-21d6-44ff-9d49-97752c29c1cd · outbound
Multitwine: Multi-Object Compositing with Text and Layout Control Tf-icon: Diffusion-based training-free cross-domain image composi- tion
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 55d4a265-1e4a-4841-a59c-7c2027df956a · outbound
Multitwine: Multi-Object Compositing with Text and Layout Control DINOv2: Learning Robust Visual Features without Supervision
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ee35b5d-4d11-4b72-8216-e62c5eb15e9c · outbound
Multitwine: Multi-Object Compositing with Text and Layout Control Kosmos-G: Generating Images in Context with Multimodal Large Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05aa7bc1-d977-40cc-874c-3723d47fdd57 · outbound
Multitwine: Multi-Object Compositing with Text and Layout Control https://pixabay.com/, 2024
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 2f083b3d-2002-4b00-b2e3-be4473e9468f · outbound
Multitwine: Multi-Object Compositing with Text and Layout Control SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04b7e7bd-be0f-4737-a9bc-320e07c1cee5 · outbound
Multitwine: Multi-Object Compositing with Text and Layout Control High-Quality Entity Segmentation
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ddc5b5da-451c-4048-8daa-bb80f0dc28b5 · outbound
Multitwine: Multi-Object Compositing with Text and Layout Control Learning transferable visual models from natural language supervi- sion
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbf8423c-7629-4a99-b439-08159bd32e96 · outbound
Multitwine: Multi-Object Compositing with Text and Layout Control High-resolution image synthesis with latent diffusion models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf7d524c-438c-4870-87c3-c5bf63b1fdfb · outbound
Multitwine: Multi-Object Compositing with Text and Layout Control Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d800c370-6ce3-41b4-8b68-a767ed7aaa08 · outbound
Multitwine: Multi-Object Compositing with Text and Layout Control Diffuse to Choose: Enriching Image Conditioned Inpainting in Latent Diffusion Models for Virtual Try-All
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 355e135b-1342-4b10-aede-3384f133622d · outbound
Multitwine: Multi-Object Compositing with Text and Layout Control Video visual relation detection
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 28cf2d92-51e0-480e-8eff-dd3d8276fdd4 · outbound
Multitwine: Multi-Object Compositing with Text and Layout Control Annotating objects and relations in user- generated videos
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3304e5c2-fee3-44c9-bbde-0da93e303fdb · outbound
Multitwine: Multi-Object Compositing with Text and Layout Control InstantBooth: Personalized Text-to-Image Generation without Test-Time Finetuning
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69b55bb5-e0e3-454d-85c4-d8487227bf77 · outbound
Multitwine: Multi-Object Compositing with Text and Layout Control ObjectStitch: Generative Object Compositing
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 607a3426-bca4-4f7d-ad04-7fcc67d41958 · outbound
Multitwine: Multi-Object Compositing with Text and Layout Control Imprint: Generative object compositing by learning identity-preserving representation
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 483e5ba2-0012-4a11-8aa4-8364c98eb43b · outbound
Multitwine: Multi-Object Compositing with Text and Layout Control Emu: Generative Pretraining in Multimodality
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5d21661-b75f-462b-bbde-c39f0e36d0be · outbound
Multitwine: Multi-Object Compositing with Text and Layout Control Generative multimodal mod- els are in-context learners
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 58db0e90-c271-4b15-aec1-08024664cc3f · outbound
Multitwine: Multi-Object Compositing with Text and Layout Control Thinking Outside the BBox: Unconstrained Generative Object Compositing
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 762dcc6a-bb1d-4db7-9f9b-e93c91d3ba83 · outbound
Multitwine: Multi-Object Compositing with Text and Layout Control CogVLM: Visual Expert for Pretrained Language Models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1e141ef-52ce-4b42-be37-33b9a6a7d42b · outbound
Multitwine: Multi-Object Compositing with Text and Layout Control Elite: Encoding visual con- cepts into textual embeddings for customized text-to-image generation
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbae3f95-8114-431e-bcab-e2b2cf83b116 · outbound
Multitwine: Multi-Object Compositing with Text and Layout Control Fastcomposer: Tuning-free multi- subject image generation with localized attention
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 2f9f0806-d127-4b70-bb66-5b3301e8a885 · outbound
Multitwine: Multi-Object Compositing with Text and Layout Control GroundingBooth: Grounding Text-to-Image Customization
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b051d4b7-f7c2-40d4-9b8d-aab3aaf5e34e · outbound
Multitwine: Multi-Object Compositing with Text and Layout Control Paint by example: Exemplar-based image editing with diffusion mod- els
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f61748aa-6e5d-4c4e-b0bb-5bedaee888c9 · outbound
Multitwine: Multi-Object Compositing with Text and Layout Control CustomNet: Zero-shot Object Customization with Variable-Viewpoints in Text-to-Image Diffusion Models
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f381e241-b1f2-44ae-bfa9-92faa49e1e97 · outbound
Multitwine: Multi-Object Compositing with Text and Layout Control ControlCom: Controllable Image Composition using Diffusion Model
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation caa0283f-05fd-4101-ad25-87b5b5c818df · outbound
Multitwine: Multi-Object Compositing with Text and Layout Control LoCo: Locally Constrained Training-Free Layout-to-Image Synthesis
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a560b2d-248f-499f-a8f0-9d9bf0d2d100 · outbound
Multitwine: Multi-Object Compositing with Text and Layout Control Can you provide a grammatically correct one-line caption for the relation <object A> <relation> <object B> in the image?
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8a454c94-da25-417d-9c2f-aa9ef7a0911a · outbound
Multitwine: Multi-Object Compositing with Text and Layout Control Figure 4
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 6594bd52-d1fc-4891-8733-bd6287635ded · outbound
Multitwine: Multi-Object Compositing with Text and Layout Control used to extract them
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e4a496a8-e693-4bfe-919b-41542389de35 · outbound
Multitwine: Multi-Object Compositing with Text and Layout Control Background images are sourced from Pixabay [29], while objects are from Pixabay [29], MultiBench [23], and DreamBooth [34]
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 6a58adb2-d0ad-42c8-ad3e-6fb483ece5b6 · outbound
Multitwine: Multi-Object Compositing with Text and Layout Control Further details on user studies can be found in Section 3.3
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8cbbc66a-c750-4349-91f2-6ecbaa65528c · outbound
Multitwine: Multi-Object Compositing with Text and Layout Control Without multi-view data ( i.e., video data, manually collected data), the model struggles to prop- erly repose and combine objects to align with the textual description
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 6d49043d-648e-4ee1-a8ae-04f25ddf9c8c · outbound
Multitwine: Multi-Object Compositing with Text and Layout Control Unresolved cited work
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 58b4d835-8f33-4c81-bfa6-ae638a5cc0b7 · inbound
BlenderFusion: 3D-Grounded Visual Editing and Generative Compositing Multitwine: Multi-Object Compositing with Text and Layout Control
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.