Pith. sign in

Paper Citation Record · LEDGER

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions

As of 11 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 0 inbound Pith citation observations for arXiv:2501.12173.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.12173 v1

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T17:30:48.007551Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

54 of 54 outbound references displayed

  • verified exact1
  • verified fuzzy27
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d490b31e-3d0d-4c43-b68c-25bcb4026d10 · outbound

This paper cites Spatext: Spatio-textual representation for con- trollable image generation.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions Spatext: Spatio-textual representation for con- trollable image generation

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:30:48.686884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T17:30:47.768865Z digest=sha256:a780fe6dd7ee69a38a740022a387a33dd5a2dc1583178b91d10b6dca05165bc8

Observation df8c0798-d017-47ee-be16-10ae77b31b18 · outbound

This paper cites MultiDiffusion: Fusing Diffusion Paths for Controlled Image Generation.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions MultiDiffusion: Fusing Diffusion Paths for Controlled Image Generation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T17:30:47.773975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:30:47.773975Z digest=sha256:f073f8a2af3f9a9fb840ae042001eeeeeed04ee374b8f5a3c75b2481a4cb352f

Observation ba481bdf-bd9b-4bab-8d27-4b19380350ed · outbound

This paper cites Sutherland, Michael Arbel, and Arthur Gretton.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions Sutherland, Michael Arbel, and Arthur Gretton

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:30:48.673792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T17:30:47.778737Z digest=sha256:76d7e1575a7938cf888dfac99d7ce106879d1042af0d4eddf3c1821fcfab00c9

Observation e15659c6-0fe5-466e-a797-2181377d1e4c · outbound

This paper cites InstructPix2Pix: Learning to Follow Image Editing Instructions.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions InstructPix2Pix: Learning to Follow Image Editing Instructions

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T17:30:47.782994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:30:47.782994Z digest=sha256:6c27eda30d8e1a4b54d202ef7f01fe3dae6e032e1835f9793bdf977179215c1c

Observation a7e13581-a7e8-4a81-a5d6-46a2e4255ddd · outbound

This paper cites PhotoVerse: Tuning-Free Image Customization with Text-to-Image Diffusion Models.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions PhotoVerse: Tuning-Free Image Customization with Text-to-Image Diffusion Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T17:30:47.787809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:30:47.787809Z digest=sha256:60a74efebfb93e1a00601d7e0ee5eb83f71e9c24702a51d08422379db51960b1

Observation b51779b0-dd86-4f76-8c11-dd7cd6fe8f80 · outbound

This paper cites Training-free layout control with cross-attention guidance.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions Training-free layout control with cross-attention guidance

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T17:30:47.792688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:30:47.792688Z digest=sha256:284a9d5ddbac234e3e08c37a3444557e4c4d7f98254bc6801dd403a045185db5

Observation 64544040-90d9-42c6-b024-18658ba570df · outbound

This paper cites AnyDoor: Zero-shot Object-level Image Customization.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions AnyDoor: Zero-shot Object-level Image Customization

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T17:30:47.798320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:30:47.798320Z digest=sha256:1f1bd2e3cd59c67a88c329be5d4001cc8fec0ae96c23bc64da38d9488adfe916

Observation 84e5c1ab-d65e-4699-993f-ed758ebd7b9d · outbound

This paper cites LayoutDiffuse: Adapting Foundational Diffusion Models for Layout-to-Image Generation.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions LayoutDiffuse: Adapting Foundational Diffusion Models for Layout-to-Image Generation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T17:30:47.802873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:30:47.802873Z digest=sha256:0ce087214b7b9841427e5ed47f6e60adeca861f8cb55170bbc4ab0f04c9f693a

Observation 74322a34-e077-4956-a9bb-5c5d8e95b8ae · outbound

This paper cites KPE: Keypoint Pose Encoding for Transformer-based Image Generation.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions KPE: Keypoint Pose Encoding for Transformer-based Image Generation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T17:30:47.807837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:30:47.807837Z digest=sha256:4e6f441ce450d042be31580ad206f20526c8816a2c0fc6167e98e1409ca7b74b

Observation 7463dbd3-605e-4864-920a-c7f8f865d4b9 · outbound

This paper cites Viton-hd: High-resolution virtual try-on via misalignment-aware normalization.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions Viton-hd: High-resolution virtual try-on via misalignment-aware normalization

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:30:48.653415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T17:30:47.812542Z digest=sha256:c8238f139a357ebbfb14cc6554a80c9e9ef67c84973bfc259404c74772ddd5c9

Observation a75b5cc1-77df-4446-a201-24c77098d514 · outbound

This paper cites Measures of the amount of ecologic association between species.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions Measures of the amount of ecologic association between species

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:30:48.641094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T17:30:47.816586Z digest=sha256:971251f9c02668ff94992c801fefe58eb906ca694b93b5da92f4e069b08d667c

Observation f4ce4cb4-362e-4645-b07b-e210759a69ec · outbound

This paper cites Stylegan-human: A data-centric odyssey of human genera- tion.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions Stylegan-human: A data-centric odyssey of human genera- tion

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:30:48.628108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T17:30:47.820786Z digest=sha256:8da4f7299d55e58b4e5b8382445559b3caadf1ee14c15987beaa572148cf0edf

Observation c72b606b-1c4d-4089-bd16-dd1af06003e9 · outbound

This paper cites An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T17:30:47.825923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:30:47.825923Z digest=sha256:db2f497b2cb896aa4f39ba5e61f5c8be45b7199baf39acfb8b58b3b7af378e5f

Observation ba458290-f62a-4d56-9c3a-441c41d42976 · outbound

This paper cites A versatile benchmark for de- tection, pose estimation, segmentation and re-identification of clothing images.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions A versatile benchmark for de- tection, pose estimation, segmentation and re-identification of clothing images

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:30:48.615396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T17:30:47.830402Z digest=sha256:474ad7df394cab84844af7e680a06df967712a019183c73934b0c6412e3dbebb

Observation 6536aca5-b052-495a-9982-daf7022e9f41 · outbound

This paper cites Context- aware layout to image generation with enhanced object ap- pearance.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions Context- aware layout to image generation with enhanced object ap- pearance

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:30:48.602017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T17:30:47.835299Z digest=sha256:35e21a81396c5e29cafa3fef91a53ec45ea4645077508cb7c6c161057164d15f

Observation 75ae43e7-c555-42b4-af0a-fe35d4fb7642 · outbound

This paper cites Denoising dif- fusion probabilistic models.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions Denoising dif- fusion probabilistic models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T17:30:47.839765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:30:47.839765Z digest=sha256:1006e702ef7d55b8369cc634b218cb4f050a54bdfc79273b0e772735c9d0dc72

Observation e0be945d-b4ad-488d-954a-f4da9bd4ba9b · outbound

This paper cites CogVLM2: Visual Language Models for Image and Video Understanding.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions CogVLM2: Visual Language Models for Image and Video Understanding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T17:30:47.844735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:30:47.844735Z digest=sha256:d9da31e868e92f25c1c1b7fd7e828967585c76e03221862ebc0b6970753cd8b7

Observation 66cc599d-d0ce-48dd-bff7-158d4bc5a32d · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions LoRA: Low-Rank Adaptation of Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T17:30:47.849204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:30:47.849204Z digest=sha256:647e32a30dfc09d2659696e5906af29c32cdf9e154e5c98833e9b853204553b0

Observation d177a0fb-e45b-4063-aacc-13e6ea135bf0 · outbound

This paper cites ´Etude comparative de la distribution florale dans une portion des alpes et des jura.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions ´Etude comparative de la distribution florale dans une portion des alpes et des jura

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:30:48.580058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T17:30:47.854040Z digest=sha256:434788e800ca03f5a68c746cb015bf1ab9e151bd98f212817f88533b98f12485

Observation ee5ccada-a3ec-471c-9e24-3d1f12f62f34 · outbound

This paper cites Text2human: Text-driven controllable human image generation.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions Text2human: Text-driven controllable human image generation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:30:48.567316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T17:30:47.858407Z digest=sha256:32ec92f40fdb94d3217522ff6201d24b82d346a8805d5c42096aece836bf652f

Observation fd99474c-d18a-49c5-8832-fbd29d4872e3 · outbound

This paper cites Dense text-to-image generation with attention modulation.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions Dense text-to-image generation with attention modulation

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:30:48.554819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T17:30:47.862430Z digest=sha256:cacde5cdea15533f76c2f21012a1fcd967d793490426939a7cf273053d083075

Observation e86d74a4-fa59-48f2-adf9-23d908bed295 · outbound

This paper cites Auto-Encoding Variational Bayes.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions Auto-Encoding Variational Bayes

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T17:30:47.866740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:30:47.866740Z digest=sha256:de1bffafe531c022b6027bf755ac46cd68cab565695c2044fe010ba02a386cf1

Observation c50343be-f86d-46d3-9e95-fb61a07c58c5 · outbound

This paper cites Segment Anything.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions Segment Anything

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T17:30:47.871094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:30:47.871094Z digest=sha256:d5185c590e45b23d595e35f4254cb60abb1d8a9a6dddf6fdaa984440ebdd85d5

Observation 9a33e0e4-b2cd-445b-ad39-14aa23fdc43c · outbound

This paper cites Multi-concept customization of text-to-image diffusion.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions Multi-concept customization of text-to-image diffusion

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T17:30:47.875119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:30:47.875119Z digest=sha256:82c21e8993a999467b241c2a618a521d89f950d8e4691531316f44f37ce06ce1

Observation b658aa63-f2b8-4934-8adc-0ad7dbfcf72c · outbound

This paper cites Self- correction for human parsing.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions Self- correction for human parsing

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T17:30:47.879911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:30:47.879911Z digest=sha256:50fceeb79e41f5090e077849abfd286b32c5b1598f727e02caaae88b22b1209b

Observation 94595758-8b25-4f86-8660-f8c8f87ae733 · outbound

This paper cites Gligen: Open-set grounded text-to-image generation.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions Gligen: Open-set grounded text-to-image generation

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:30:48.527010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T17:30:47.883877Z digest=sha256:0330d86f2c71993054098b7c8eaa2a52ee1649ffc66a83399be5d381787c9f38

Observation 71814155-4120-4460-bfe9-62f5873f8a01 · outbound

This paper cites Image synthesis from layout with locality- aware mask adaption.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions Image synthesis from layout with locality- aware mask adaption

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:30:48.514255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T17:30:47.889188Z digest=sha256:6f3666b46a238cd9ed4b68d433d3cab794b084c21d57f50dbd44817d0ea1db67

Observation d292a7d4-5b9c-433c-ab03-2970a6ca3eb8 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T17:30:47.894131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:30:47.894131Z digest=sha256:31d99aba7debfe033e391eaee127dad2164666ce6caa14f61a9e4239f6410905

Observation 1438a259-7aa5-4b42-8dcb-6d800e6bf84b · outbound

This paper cites Dress code: High- resolution multi-category virtual try-on, 2022.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions Dress code: High- resolution multi-category virtual try-on, 2022

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:30:48.501571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T17:30:47.901085Z digest=sha256:c676465ce52bc2a4496d60ff367663a68308894d41b441b35b3c14db53f50f49

Observation f07dde4c-a53b-47c4-832f-dc8412757e0f · outbound

This paper cites $\lambda$-ECLIPSE: Multi-Concept Personalized Text-to-Image Diffusion Models by Leveraging CLIP Latent Space.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions $\lambda$-ECLIPSE: Multi-Concept Personalized Text-to-Image Diffusion Models by Leveraging CLIP Latent Space

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T17:30:47.906551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:30:47.906551Z digest=sha256:d860f4c6084e7050c1deb8dfeb0df3cab5238aeb3e1d256f3e72d2e618b3233d

Observation 6c705960-f287-4452-90e0-032760beeb40 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions Learning transferable visual models from natural language supervi- sion

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T17:30:47.911049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:30:47.911049Z digest=sha256:18c1b40bf5b540d49ad8d010b35007921bf15dd8cd9d9617a389d47e4e70fbca

Observation 50d1761e-a282-46e0-bc46-e26e004913e6 · outbound

This paper cites High-resolution image syn- thesis with latent diffusion models, 2021.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions High-resolution image syn- thesis with latent diffusion models, 2021

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:30:48.479983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T17:30:47.915124Z digest=sha256:138a4b5ca7b235aa3e807697b5ed55d2cf3b2d0badfc6e56529973fcb1a010d3

Observation f680884f-6307-44e3-80dc-cb74615c614b · outbound

This paper cites Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T17:30:47.918826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:30:47.918826Z digest=sha256:22870b3b784a281d0ea80ea268c43d8082c9850b29f1148d9cfbfbdcb7fba133

Observation 30885056-b095-4de1-8e94-778a40c7062f · outbound

This paper cites Humangan: A generative model of hu- man images.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions Humangan: A generative model of hu- man images

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:30:48.459736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T17:30:47.924213Z digest=sha256:d782f323c897e5731b7965fbb85af4738eb00a5d8a5e76c137ae4f75693f1feb

Observation f95a8c05-5756-442a-8ad7-55e739b573ce · outbound

This paper cites pytorch-fid: FID Score for PyTorch.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions pytorch-fid: FID Score for PyTorch

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T17:30:47.928078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:30:47.928078Z digest=sha256:9186a1b601c96b322ffd1f721dd5e86cb0246f7a2059bff9bbdf6846659add34

Observation 1f523453-f07d-4eea-bfd7-c5e5ae280c68 · outbound

This paper cites In- stantbooth: Personalized text-to-image generation without test-time finetuning.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions In- stantbooth: Personalized text-to-image generation without test-time finetuning

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:30:48.438335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T17:30:47.932515Z digest=sha256:dfa38b9a1c485c7714e56752197fe2e4a0141a3e12db3df088ebee94f48a3368

Observation cc25dca7-73fc-4be4-82c6-f8d7fcdf2ad4 · outbound

This paper cites Image synthesis from reconfig- urable layout and style.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions Image synthesis from reconfig- urable layout and style

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:30:48.425427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T17:30:47.936451Z digest=sha256:5194c9373f02d2804e81e1a55eb334f44bab2f38693800f6fcfa48ec2f0b57d0

Observation df2be0fc-bf8b-4631-9efb-cdf5184e4130 · outbound

This paper cites Object-centric image genera- tion from layouts.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions Object-centric image genera- tion from layouts

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:30:48.412497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T17:30:47.940545Z digest=sha256:9ae542593ca091e921870380511cfad99c24c12826854edce9cbea3134d082e8

Observation 236f6967-d908-4754-a255-3195ece747eb · outbound

This paper cites Interactive image synthesis with panoptic layout generation.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions Interactive image synthesis with panoptic layout generation

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:30:48.398448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T17:30:47.944723Z digest=sha256:1ae25117cd5c3e0334cbed2ba5b8a305338ce0f60ca334bb183ba7e3389c4a10

Observation 99752318-c0cb-4658-902b-360198564b6e · outbound

This paper cites InstantID: Zero-shot Identity-Preserving Generation in Seconds.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions InstantID: Zero-shot Identity-Preserving Generation in Seconds

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T17:30:47.949730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:30:47.949730Z digest=sha256:ea38be05f5fe18fe2a5e41a4ea54c44c83a260ab9764ae83619689a6db246555

Observation 485acd33-d9bb-4334-87ce-05158caaa609 · outbound

This paper cites Instancediffusion: Instance-level control for image generation, 2024.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions Instancediffusion: Instance-level control for image generation, 2024

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:30:48.385731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T17:30:47.954162Z digest=sha256:246cc2fdeae8f89e7f7172ac88858f3fa4dea7c13cb9f8d55fece787fce18b8c

Observation a2f20f46-d721-4072-aa23-3776efa3ea22 · outbound

This paper cites Image quality assessment: from error visibility to structural similarity.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions Image quality assessment: from error visibility to structural similarity

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:30:48.372765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T17:30:47.958179Z digest=sha256:5ac6de1bbeb11945d4594e9cf44c7b9bb4dc6625bb15fcc136e726e0a384ed3b

Observation 43066a6f-3058-4fa9-8fab-dc8d4c3bc853 · outbound

This paper cites ELITE: Encoding Visual Concepts into Textual Embeddings for Customized Text-to-Image Generation.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions ELITE: Encoding Visual Concepts into Textual Embeddings for Customized Text-to-Image Generation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T17:30:47.961855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:30:47.961855Z digest=sha256:79c45dcee55a27be480b9fb8de3fb76cccc02f873a39fd4906962811f0238ccd

Observation 414dc079-dcff-402a-b380-7eb015485c60 · outbound

This paper cites Fastcomposer: Tuning-free multi- subject image generation with localized attention.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions Fastcomposer: Tuning-free multi- subject image generation with localized attention

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:30:48.359668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T17:30:47.966274Z digest=sha256:003551be4a9b256f526b1016051b2517b325ae80578f83456c06f6bf7cacb3eb

Observation 04342a8f-1d53-4922-b7e2-4683bc520249 · outbound

This paper cites R&B: Region and Boundary Aware Zero-shot Grounded Text-to-image Generation.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions R&B: Region and Boundary Aware Zero-shot Grounded Text-to-image Generation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T17:30:47.970015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:30:47.970015Z digest=sha256:3ef33691d8dd484cb82a322c496dd58a799b2a9b078f041f106e02925ffd9538

Observation 46f84a73-4883-4025-8cd0-92a06168b2fc · outbound

This paper cites Boxdiff: Text-to-image synthesis with training-free box-constrained diffusion.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions Boxdiff: Text-to-image synthesis with training-free box-constrained diffusion

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:30:48.346658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T17:30:47.974418Z digest=sha256:50ccfe5bf03f22d3f9226a5ddfb34931a63b4c756e24ddace382e63f4fc41db6

Observation 78a9c7b3-b863-49bf-9550-172e70599253 · outbound

This paper cites Reco: Region-controlled text-to-image genera- tion.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions Reco: Region-controlled text-to-image genera- tion

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T17:30:47.978379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:30:47.978379Z digest=sha256:cf374ec0adc60e53701766e87715580e98c3db84cc0095249035235baeb7a997

Observation 038fce44-f971-4687-a31e-755c0091c412 · outbound

This paper cites IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T17:30:47.982276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:30:47.982276Z digest=sha256:7bbd807fda370886f5e63cceca3ae443796615345df3a0028ff5a365324994a0

Observation e5f886bb-68c3-462f-b0e4-f26e54b30344 · outbound

This paper cites Customnet: Zero-shot object customization with variable-viewpoints in text-to-image dif- fusion models, 2023.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions Customnet: Zero-shot object customization with variable-viewpoints in text-to-image dif- fusion models, 2023

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:30:48.326113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T17:30:47.986643Z digest=sha256:15a7cc1a51845ff823af9c121c37df1b31a54b332ae89923500b6184228424a1

Observation ba63e5ea-e2e2-4041-96e5-31f3b0664d9b · outbound

This paper cites HumanDiffusion: a Coarse-to-Fine Alignment Diffusion Framework for Controllable Text-Driven Person Image Generation.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions HumanDiffusion: a Coarse-to-Fine Alignment Diffusion Framework for Controllable Text-Driven Person Image Generation

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-08-10T17:30:48.049415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T17:30:47.991487Z digest=sha256:08b5d2dec093723400073ca926a174c7b383cb3d4678c593decdb9ad1953ea9c

Observation 784dc3f6-544e-41cb-874d-bae8c2f5aef7 · outbound

This paper cites The unreasonable effectiveness of deep features as a perceptual metric.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions The unreasonable effectiveness of deep features as a perceptual metric

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T17:30:47.995603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:30:47.995603Z digest=sha256:0469228a1650d22543af773bba79859eaaded08c9848b3ad283304580fadd6b4

Observation 544697bd-b9d8-430b-a76c-8fa90abb9735 · outbound

This paper cites Layoutdiffusion: Controllable diffu- sion model for layout-to-image generation.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions Layoutdiffusion: Controllable diffu- sion model for layout-to-image generation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:30:48.305842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T17:30:47.999705Z digest=sha256:44a488bc507802b0af66218df7d7e6d9e921b708caafd8bb4be9ee82272f77e4

Observation 70dcb08b-9447-4c06-9930-be3b2f0b5f9b · outbound

This paper cites clip-score: CLIP Score for Py- Torch.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions clip-score: CLIP Score for Py- Torch

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:30:48.292780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T17:30:48.003474Z digest=sha256:62d5bf0c7d2b66ed5aed0709eb084f147d6f73cd9f2c10523f4838f507d611de

Observation f573d113-9743-4574-a9d8-2b343db30785 · outbound

This paper cites Migc: Multi-instance generation controller for text-to-image synthesis, 2024.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions Migc: Multi-instance generation controller for text-to-image synthesis, 2024

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:30:48.280060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T17:30:48.007551Z digest=sha256:c9427ff7ffe7191e80704e9efc1502efd1f065b62fb33660d12ec9daad97fb7d

Pith citing papers

No inbound Pith citation observations are available.