Pith. sign in

Paper Citation Record · LEDGER

Multitwine: Multi-Object Compositing with Text and Layout Control

As of 9 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 1 inbound Pith citation observation for arXiv:2502.05165.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.05165 v1

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T20:05:16.256044Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:35:15.341406Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T23:35:15.923842Z

Reference resolution

58 of 58 outbound references displayed

  • verified exact0
  • verified fuzzy27
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a09b5255-306e-4e4f-874b-fa53aa34b126 · outbound

This paper cites Cross-image attention for zero- shot appearance transfer.

Multitwine: Multi-Object Compositing with Text and Layout Control Cross-image attention for zero- shot appearance transfer

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:05:17.161126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:05:16.003397Z digest=sha256:3d5b5e83222e359880b798e9b59088f38c2890e59228b73e3dc2c329a6c48211

Observation 9b0ac045-08d8-484c-90f3-58f2e18265de · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Multitwine: Multi-Object Compositing with Text and Layout Control Flamingo: a visual language model for few-shot learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.009424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.009424Z digest=sha256:15be98e872bc759f4a3c7bed6dd8fd0ef4587f207bb3dae24db1e1795cda46ae

Observation e1305d9a-3c8c-4819-b371-1232a3b83acc · outbound

This paper cites Vip- llava: Making large multimodal models understand arbitrary visual prompts.

Multitwine: Multi-Object Compositing with Text and Layout Control Vip- llava: Making large multimodal models understand arbitrary visual prompts

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:05:17.134681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:05:16.014215Z digest=sha256:6e2948d7bf41c26e4aefa6b8e4f848344a2789f1375e01b2497f374d901981bc

Observation e06f5dcb-2bfb-4d6e-95c5-b1b6f20e0033 · outbound

This paper cites Masactrl: Tuning-free mu- tual self-attention control for consistent image synthesis and editing.

Multitwine: Multi-Object Compositing with Text and Layout Control Masactrl: Tuning-free mu- tual self-attention control for consistent image synthesis and editing

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:05:17.119564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:05:16.018799Z digest=sha256:3b1bc93c82aedf32fbb4d6ef19aa4827d0827ead0b10497bc87d6ecf3ae02f24

Observation 0b29a90d-0dba-41e3-8e3e-9d8987dc54ce · outbound

This paper cites Re-Imagen: Retrieval-Augmented Text-to-Image Generator.

Multitwine: Multi-Object Compositing with Text and Layout Control Re-Imagen: Retrieval-Augmented Text-to-Image Generator

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.023288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.023288Z digest=sha256:ac827b1101d0877bae465d0b75f215109dc556d89fafe628df631b124e7a0aee

Observation fc45db59-27d1-4316-8531-c3c37be8189a · outbound

This paper cites Subject-driven text-to-image generation via apprenticeship learning.

Multitwine: Multi-Object Compositing with Text and Layout Control Subject-driven text-to-image generation via apprenticeship learning

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:05:17.103540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:05:16.028321Z digest=sha256:4bcd9974ae4f20aae30216391f2ccb9563894810c36361482203215f60998da3

Observation e9684e9a-4df2-466f-bb84-2292e57918ea · outbound

This paper cites AnyDoor: Zero-shot Object-level Image Customization.

Multitwine: Multi-Object Compositing with Text and Layout Control AnyDoor: Zero-shot Object-level Image Customization

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.032810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.032810Z digest=sha256:6d9f3feac01ee9cd81051efcc36c2d492f6b4836a45698448f0b6f06e4d03a28

Observation c81b1664-3105-45b1-a436-1affe30cc0f3 · outbound

This paper cites Be Yourself: Bounded Attention for Multi-Subject Text-to-Image Generation.

Multitwine: Multi-Object Compositing with Text and Layout Control Be Yourself: Bounded Attention for Multi-Subject Text-to-Image Generation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.038213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.038213Z digest=sha256:f7496fbe08d9d4e563c3aaba9c54786a60af9e40c7899073a40f92188677b47d

Observation 1d1a345a-e7fb-47e4-a468-6d2ca1e9b7ac · outbound

This paper cites DreamLLM: Synergistic Multimodal Comprehension and Creation.

Multitwine: Multi-Object Compositing with Text and Layout Control DreamLLM: Synergistic Multimodal Comprehension and Creation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.042916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.042916Z digest=sha256:6a705501573123793b59f829eafd732318ff0e15639de245695d04b790d072bb

Observation dd0abeb5-e9a7-4178-94c4-1f5c602025b3 · outbound

This paper cites PaLM-E: An Embodied Multimodal Language Model.

Multitwine: Multi-Object Compositing with Text and Layout Control PaLM-E: An Embodied Multimodal Language Model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.047459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.047459Z digest=sha256:b9a7bff752ac93a3df650986cfc8466479472ab6c32099d6c4d48bda062a2b66

Observation 837e58cb-59c3-4b1b-889b-a222bf3bd5a2 · outbound

This paper cites Scaling recti- fied flow transformers for high-resolution image synthesis.

Multitwine: Multi-Object Compositing with Text and Layout Control Scaling recti- fied flow transformers for high-resolution image synthesis

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:05:17.087545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:05:16.051903Z digest=sha256:c8d8d020a9887005cc758f8dea23638745a29f7e85f5798a5fd390ccdfb62a6b

Observation 341e031a-cf5f-497a-9792-6bcec4aacf70 · outbound

This paper cites An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion.

Multitwine: Multi-Object Compositing with Text and Layout Control An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.056096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.056096Z digest=sha256:e0bb5a6cc74d30792813f4f3f9aeabab4515a9ee5f882aa42c95d5b7c44e6da2

Observation 2aa2c571-9ff0-4ef6-92f2-19c13569d0d1 · outbound

This paper cites Prompt-to-Prompt Image Editing with Cross Attention Control.

Multitwine: Multi-Object Compositing with Text and Layout Control Prompt-to-Prompt Image Editing with Cross Attention Control

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.060539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.060539Z digest=sha256:f02c324dd9caf969bc741bd607e9a1c4799c0ce7836df323639386c23fee43c4

Observation 27d371a3-5e0a-4a91-93c2-0ed07b72e59f · outbound

This paper cites Clipscore: A reference-free evaluation met- ric for image captioning.

Multitwine: Multi-Object Compositing with Text and Layout Control Clipscore: A reference-free evaluation met- ric for image captioning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:05:17.071654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:05:16.064651Z digest=sha256:1411e3bbce4a637f581dc715a2dd5772b175537aeed19c732956e5e6cf9fa494

Observation 05e0f83e-8f4a-4001-8fd9-3a8c887f304c · outbound

This paper cites Taming Encoder for Zero Fine-tuning Image Customization with Text-to-Image Diffusion Models.

Multitwine: Multi-Object Compositing with Text and Layout Control Taming Encoder for Zero Fine-tuning Image Customization with Text-to-Image Diffusion Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.068679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.068679Z digest=sha256:0db4bcbea216be7aaacf27a289d3f9b23fce1e97f06aa8613871af4db6e3ada8

Observation 5c2e602d-e6af-4798-b16f-8c8a4187f114 · outbound

This paper cites Rendering synthetic objects into legacy photographs.

Multitwine: Multi-Object Compositing with Text and Layout Control Rendering synthetic objects into legacy photographs

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:05:17.055767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:05:16.073123Z digest=sha256:b38df49db1f3c1a1a8cb30398b80227a9d617454c9fcfd33a3ed9ea3972a95ed

Observation a365751e-6298-45bf-abae-01df37cb318c · outbound

This paper cites 3d object manipulation in a single photograph using stock 3d models.

Multitwine: Multi-Object Compositing with Text and Layout Control 3d object manipulation in a single photograph using stock 3d models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:05:17.040988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:05:16.077249Z digest=sha256:a088c142d5512cde983959dce91325934028fc0580c827e278db7af87c5ae259

Observation fc442340-45ae-4eaf-aa15-629daae51b1b · outbound

This paper cites Gen- erating images with multimodal language models.

Multitwine: Multi-Object Compositing with Text and Layout Control Gen- erating images with multimodal language models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.081647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.081647Z digest=sha256:104270cf9bbd79b3af8fa5b0f85b97a72f8092789bfe53795e780fefe3c43500

Observation 67e4255b-c2fd-400f-a0da-e82010b9a3b4 · outbound

This paper cites Open images v5 text annotation and yet another mask text spotter.

Multitwine: Multi-Object Compositing with Text and Layout Control Open images v5 text annotation and yet another mask text spotter

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:05:17.016080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:05:16.086200Z digest=sha256:82b7017a1fbfeb6ab6f8aebb51f0ee4e762dc1634b2477c151944087102baaf5

Observation bb5a39c4-c7a1-4822-925e-b6c12b0d0de9 · outbound

This paper cites Putting people in their place: Affordance-aware hu- man insertion into scenes.

Multitwine: Multi-Object Compositing with Text and Layout Control Putting people in their place: Affordance-aware hu- man insertion into scenes

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:05:17.000101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:05:16.090712Z digest=sha256:cd3b5d8871685abf098d27ac380bed1b3cd126da9d3bed76beee41a15661e27e

Observation 5c546ea4-3ffb-4918-a06e-95704f0b3072 · outbound

This paper cites Photo clip art.

Multitwine: Multi-Object Compositing with Text and Layout Control Photo clip art

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:05:16.984562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:05:16.095184Z digest=sha256:cfb1018ad20afa6a644aba38ae79e3a9943e9671b42d29dbb11234cb9ed14279

Observation fb1d63c8-216a-4531-bc3c-b1b7af3a797f · outbound

This paper cites Blip-diffusion: Pre- trained subject representation for controllable text-to-image generation and editing.

Multitwine: Multi-Object Compositing with Text and Layout Control Blip-diffusion: Pre- trained subject representation for controllable text-to-image generation and editing

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:05:16.970266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:05:16.099723Z digest=sha256:d442a4a3cd6b80cddb219ecefe49f8a0349a9561d7626232d9b6244ec51d9b74

Observation ba07cc92-e700-4c65-b17c-d117e8ad268a · outbound

This paper cites UNIMO-G: Unified Image Generation through Multimodal Conditional Diffusion.

Multitwine: Multi-Object Compositing with Text and Layout Control UNIMO-G: Unified Image Generation through Multimodal Conditional Diffusion

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.103942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.103942Z digest=sha256:3634432df16cbc3cdd386f2c5bbb50b60ea7c38df848e8102a85e4ca0998cf6a

Observation 44f52a1d-a541-41f1-b616-4545299ffd2e · outbound

This paper cites Improved baselines with visual instruction tuning, 2023.

Multitwine: Multi-Object Compositing with Text and Layout Control Improved baselines with visual instruction tuning, 2023

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:05:16.955457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:05:16.108408Z digest=sha256:e60e21ee238e9ff4ea23696f37a26423348654077a3dd453f91154cdd9ab0e33

Observation 27ca87c1-5ba6-48a3-a799-d75cc0711370 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

Multitwine: Multi-Object Compositing with Text and Layout Control Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.112663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.112663Z digest=sha256:c048064eb670b25dae9482bd6504ab00a286ca5d22833a3486867b314e6fd406

Observation 304baa49-21d6-44ff-9d49-97752c29c1cd · outbound

This paper cites Tf-icon: Diffusion-based training-free cross-domain image composi- tion.

Multitwine: Multi-Object Compositing with Text and Layout Control Tf-icon: Diffusion-based training-free cross-domain image composi- tion

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:05:16.940940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:05:16.117452Z digest=sha256:c13a4d1e681fc23b2d21a0eb034fc22600606e34625b1bc1acadce578c152f16

Observation 55d4a265-1e4a-4841-a59c-7c2027df956a · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Multitwine: Multi-Object Compositing with Text and Layout Control DINOv2: Learning Robust Visual Features without Supervision

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.121656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.121656Z digest=sha256:fdae7f225d30f5e816b90c73d63fcfa077f62813958c857ddbe5e511c00dbd35

Observation 7ee35b5d-4d11-4b72-8216-e62c5eb15e9c · outbound

This paper cites Kosmos-G: Generating Images in Context with Multimodal Large Language Models.

Multitwine: Multi-Object Compositing with Text and Layout Control Kosmos-G: Generating Images in Context with Multimodal Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.125929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.125929Z digest=sha256:8ec37991bc38fd7960c6f88345e055f7bb79112dabaf4f72d7482b64933e0f6f

Observation 05aa7bc1-d977-40cc-874c-3723d47fdd57 · outbound

This paper cites https://pixabay.com/, 2024.

Multitwine: Multi-Object Compositing with Text and Layout Control https://pixabay.com/, 2024

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:05:16.925460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:05:16.130241Z digest=sha256:00a88440401543baf99f20acf584d95467de7f3ba5be234ae56ff2bda1d2e8ae

Observation 2f083b3d-2002-4b00-b2e3-be4473e9468f · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Multitwine: Multi-Object Compositing with Text and Layout Control SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.134464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.134464Z digest=sha256:c13697db7b70b367ff22e10407002eecf2701985c8a059cf90b1c3a13c67e7e6

Observation 04b7e7bd-be0f-4737-a9bc-320e07c1cee5 · outbound

This paper cites High-Quality Entity Segmentation.

Multitwine: Multi-Object Compositing with Text and Layout Control High-Quality Entity Segmentation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.138710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.138710Z digest=sha256:0fe1d8e5c1760697b25c7644dec0007448ec0cc144a6be3d3012221554a6de43

Observation ddc5b5da-451c-4048-8daa-bb80f0dc28b5 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Multitwine: Multi-Object Compositing with Text and Layout Control Learning transferable visual models from natural language supervi- sion

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.143110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.143110Z digest=sha256:dd5dac4783ec2a0d2093f5670d936ac3ca52ddf2f5816c8f82c26a1b989ee0b9

Observation fbf8423c-7629-4a99-b439-08159bd32e96 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Multitwine: Multi-Object Compositing with Text and Layout Control High-resolution image synthesis with latent diffusion models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.147319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.147319Z digest=sha256:64738e0fb4e85c67966400efe9d786cca0774b3d5385dbdb37598653b423a974

Observation bf7d524c-438c-4870-87c3-c5bf63b1fdfb · outbound

This paper cites Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation.

Multitwine: Multi-Object Compositing with Text and Layout Control Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:05:16.888041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:05:16.151779Z digest=sha256:4723a49d69b0a28589b17ea64cc66d454465e57b5709850923433da989a787b9

Observation d800c370-6ce3-41b4-8b68-a767ed7aaa08 · outbound

This paper cites Diffuse to Choose: Enriching Image Conditioned Inpainting in Latent Diffusion Models for Virtual Try-All.

Multitwine: Multi-Object Compositing with Text and Layout Control Diffuse to Choose: Enriching Image Conditioned Inpainting in Latent Diffusion Models for Virtual Try-All

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.156007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.156007Z digest=sha256:1d64bb43b2ab55a879197734644d9aef7bb9645bc23b4c049b850c008d40d24e

Observation 355e135b-1342-4b10-aede-3384f133622d · outbound

This paper cites Video visual relation detection.

Multitwine: Multi-Object Compositing with Text and Layout Control Video visual relation detection

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:05:16.871595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:05:16.160475Z digest=sha256:af173884d607a6c398555b2793bdc7669b81f6375c61abe749d3ac79309296b9

Observation 28cf2d92-51e0-480e-8eff-dd3d8276fdd4 · outbound

This paper cites Annotating objects and relations in user- generated videos.

Multitwine: Multi-Object Compositing with Text and Layout Control Annotating objects and relations in user- generated videos

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:05:16.856146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:05:16.164892Z digest=sha256:ee82599ca26b2c6f95da3943a16e419a3b29245239c02cd2f6099fc29b12ad4a

Observation 3304e5c2-fee3-44c9-bbde-0da93e303fdb · outbound

This paper cites InstantBooth: Personalized Text-to-Image Generation without Test-Time Finetuning.

Multitwine: Multi-Object Compositing with Text and Layout Control InstantBooth: Personalized Text-to-Image Generation without Test-Time Finetuning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.169140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.169140Z digest=sha256:e96ae021d932bdad2c5c1f14b3fcce605062036a5f0dc9fb3eaf14cb57d0df02

Observation 69b55bb5-e0e3-454d-85c4-d8487227bf77 · outbound

This paper cites ObjectStitch: Generative Object Compositing.

Multitwine: Multi-Object Compositing with Text and Layout Control ObjectStitch: Generative Object Compositing

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.173596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.173596Z digest=sha256:85a43d2289bc1f918f69e2f04b7cfae8b125602ee9d92faa273ae0ce202fb84b

Observation 607a3426-bca4-4f7d-ad04-7fcc67d41958 · outbound

This paper cites Imprint: Generative object compositing by learning identity-preserving representation.

Multitwine: Multi-Object Compositing with Text and Layout Control Imprint: Generative object compositing by learning identity-preserving representation

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:05:16.840319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:05:16.177752Z digest=sha256:d337984c90a6f066072433a1a7ec5fc197c7b4adde8be3529e04078c94668942

Observation 483e5ba2-0012-4a11-8aa4-8364c98eb43b · outbound

This paper cites Emu: Generative Pretraining in Multimodality.

Multitwine: Multi-Object Compositing with Text and Layout Control Emu: Generative Pretraining in Multimodality

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.182015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.182015Z digest=sha256:c43f8829b347dd9634b13d106cee582c6fc96277c650ba3cf5c4a6bffd116527

Observation d5d21661-b75f-462b-bbde-c39f0e36d0be · outbound

This paper cites Generative multimodal mod- els are in-context learners.

Multitwine: Multi-Object Compositing with Text and Layout Control Generative multimodal mod- els are in-context learners

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:05:16.825417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:05:16.186695Z digest=sha256:be99545dbf4c732ec184d4c7715afabc3a390e40fe05c32dd6309c259eb6b8f9

Observation 58db0e90-c271-4b15-aec1-08024664cc3f · outbound

This paper cites Thinking Outside the BBox: Unconstrained Generative Object Compositing.

Multitwine: Multi-Object Compositing with Text and Layout Control Thinking Outside the BBox: Unconstrained Generative Object Compositing

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.190538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.190538Z digest=sha256:173e854fb6df9a4c1ed21c1329b8476210a389e80cae920faeb1e3c5adaa58ae

Observation 762dcc6a-bb1d-4db7-9f9b-e93c91d3ba83 · outbound

This paper cites CogVLM: Visual Expert for Pretrained Language Models.

Multitwine: Multi-Object Compositing with Text and Layout Control CogVLM: Visual Expert for Pretrained Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.194518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.194518Z digest=sha256:b894c4523a1b514e884ea35fbd346d60140d32dcc4c8313a7fa0a8c40feb05fa

Observation c1e141ef-52ce-4b42-be37-33b9a6a7d42b · outbound

This paper cites Elite: Encoding visual con- cepts into textual embeddings for customized text-to-image generation.

Multitwine: Multi-Object Compositing with Text and Layout Control Elite: Encoding visual con- cepts into textual embeddings for customized text-to-image generation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.198714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.198714Z digest=sha256:ffa79d45a706c33b051bd0faa63f2ab1423105797d0a8951d7affb861fd2eeb2

Observation dbae3f95-8114-431e-bcab-e2b2cf83b116 · outbound

This paper cites Fastcomposer: Tuning-free multi- subject image generation with localized attention.

Multitwine: Multi-Object Compositing with Text and Layout Control Fastcomposer: Tuning-free multi- subject image generation with localized attention

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:05:16.799738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:05:16.202729Z digest=sha256:6a39a173edd9ce4e690d94b8a5d8d8dda2e87d6706f5145f14c9f1772c80a587

Observation 2f9f0806-d127-4b70-bb66-5b3301e8a885 · outbound

This paper cites GroundingBooth: Grounding Text-to-Image Customization.

Multitwine: Multi-Object Compositing with Text and Layout Control GroundingBooth: Grounding Text-to-Image Customization

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.206770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.206770Z digest=sha256:fc001f7c3c8e48e6cceadd032ad5aad01527dadd8a9b55b86fd7286c9a782c60

Observation b051d4b7-f7c2-40d4-9b8d-aab3aaf5e34e · outbound

This paper cites Paint by example: Exemplar-based image editing with diffusion mod- els.

Multitwine: Multi-Object Compositing with Text and Layout Control Paint by example: Exemplar-based image editing with diffusion mod- els

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.211720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.211720Z digest=sha256:8a4ce97ee8cd7906066c47fbcc5962af7c6b9de7ab352746441468d9f6ed5cfe

Observation f61748aa-6e5d-4c4e-b0bb-5bedaee888c9 · outbound

This paper cites CustomNet: Zero-shot Object Customization with Variable-Viewpoints in Text-to-Image Diffusion Models.

Multitwine: Multi-Object Compositing with Text and Layout Control CustomNet: Zero-shot Object Customization with Variable-Viewpoints in Text-to-Image Diffusion Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.216485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.216485Z digest=sha256:f1d34f04baa388ad2b1d4a1b36bbaabc3967dab92b3e355e2c678fecb977f4f0

Observation f381e241-b1f2-44ae-bfa9-92faa49e1e97 · outbound

This paper cites ControlCom: Controllable Image Composition using Diffusion Model.

Multitwine: Multi-Object Compositing with Text and Layout Control ControlCom: Controllable Image Composition using Diffusion Model

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.220898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.220898Z digest=sha256:e4479e63fabdad19d383585837c3a0596536a3ba615a176929136e53c4fa23b6

Observation caa0283f-05fd-4101-ad25-87b5b5c818df · outbound

This paper cites LoCo: Locally Constrained Training-Free Layout-to-Image Synthesis.

Multitwine: Multi-Object Compositing with Text and Layout Control LoCo: Locally Constrained Training-Free Layout-to-Image Synthesis

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.225176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.225176Z digest=sha256:74c5a69f762fb2bf9de55594df7084b4c20ec130475061c2a783042183c33418

Observation 1a560b2d-248f-499f-a8f0-9d9bf0d2d100 · outbound

This paper cites Can you provide a grammatically correct one-line caption for the relation <object A> <relation> <object B> in the image?.

Multitwine: Multi-Object Compositing with Text and Layout Control Can you provide a grammatically correct one-line caption for the relation <object A> <relation> <object B> in the image?

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:05:16.774028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:05:16.229501Z digest=sha256:931addbe536f42b6734b485a55c2feb93bf537abc7c3d118bcebb79d75e2d6a2

Observation 8a454c94-da25-417d-9c2f-aa9ef7a0911a · outbound

This paper cites Figure 4.

Multitwine: Multi-Object Compositing with Text and Layout Control Figure 4

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:05:16.758388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:05:16.234124Z digest=sha256:9eb3fc44c72fc5764f66920a5d9594028a1fd2bb189d187a3b36a66256c0f229

Observation 6594bd52-d1fc-4891-8733-bd6287635ded · outbound

This paper cites used to extract them.

Multitwine: Multi-Object Compositing with Text and Layout Control used to extract them

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:05:16.740942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:05:16.238462Z digest=sha256:dd1240b2dfe2e4909c89bebde98e88f70046e1797d1bc1f0231a75d5634cd16c

Observation e4a496a8-e693-4bfe-919b-41542389de35 · outbound

This paper cites Background images are sourced from Pixabay [29], while objects are from Pixabay [29], MultiBench [23], and DreamBooth [34].

Multitwine: Multi-Object Compositing with Text and Layout Control Background images are sourced from Pixabay [29], while objects are from Pixabay [29], MultiBench [23], and DreamBooth [34]

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:05:16.725567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:05:16.242546Z digest=sha256:5b292c0a7d29d356042c6aed62cf196a3c62454d447d2d726791c915da15427d

Observation 6a58adb2-d0ad-42c8-ad3e-6fb483ece5b6 · outbound

This paper cites Further details on user studies can be found in Section 3.3.

Multitwine: Multi-Object Compositing with Text and Layout Control Further details on user studies can be found in Section 3.3

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:05:16.710737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:05:16.246690Z digest=sha256:32e6952ecc6c3bde19186be9e506c7af7e8ad73ec80b1d3a1b1d5da485399df0

Observation 8cbbc66a-c750-4349-91f2-6ecbaa65528c · outbound

This paper cites Without multi-view data ( i.e., video data, manually collected data), the model struggles to prop- erly repose and combine objects to align with the textual description.

Multitwine: Multi-Object Compositing with Text and Layout Control Without multi-view data ( i.e., video data, manually collected data), the model struggles to prop- erly repose and combine objects to align with the textual description

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:05:16.694958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:05:16.251155Z digest=sha256:23826c8deb4045eed7f8a483409e5cbdb4e03cdd7a6acb81edeb37d0903f7598

Observation 6d49043d-648e-4ee1-a8ae-04f25ddf9c8c · outbound

This paper cites an unresolved cited work.

Multitwine: Multi-Object Compositing with Text and Layout Control Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-08T20:05:16.677728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:05:16.256044Z digest=sha256:9549a7c1f7ad52ebb5367b2c288902942252135a9c88a13a9144e09f6cc2682e

Pith citing papers

Observation 58b4d835-8f33-4c81-bfa6-ae638a5cc0b7 · inbound

BlenderFusion: 3D-Grounded Visual Editing and Generative Compositing cites this paper.

BlenderFusion: 3D-Grounded Visual Editing and Generative Compositing Multitwine: Multi-Object Compositing with Text and Layout Control

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-08-06T23:35:15.948136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:35:15.341406Z digest=sha256:42b1e8ef0784cf83f7dad1c29212c8c21205d887e639c3a06b43de89533aa2e6