Pith. sign in

REVIEW 1 cited by

MagiCapture: High-Resolution Multi-Concept Portrait Customization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.06895 v2 pith:K4KX3PJU submitted 2023-09-13 cs.CV cs.GRcs.LG

classification cs.CVcs.GRcs.LG
keywords imagesportraitmagicapturesubjectaddressconceptsgeneratehigh-resolution
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large-scale text-to-image models including Stable Diffusion are capable of generating high-fidelity photorealistic portrait images. There is an active research area dedicated to personalizing these models, aiming to synthesize specific subjects or styles using provided sets of reference images. However, despite the plausible results from these personalization methods, they tend to produce images that often fall short of realism and are not yet on a commercially viable level. This is particularly noticeable in portrait image generation, where any unnatural artifact in human faces is easily discernible due to our inherent human bias. To address this, we introduce MagiCapture, a personalization method for integrating subject and style concepts to generate high-resolution portrait images using just a few subject and style references. For instance, given a handful of random selfies, our fine-tuned model can generate high-quality portrait images in specific styles, such as passport or profile photos. The main challenge with this task is the absence of ground truth for the composed concepts, leading to a reduction in the quality of the final output and an identity shift of the source subject. To address these issues, we present a novel Attention Refocusing loss coupled with auxiliary priors, both of which facilitate robust learning within this weakly supervised learning setting. Our pipeline also includes additional post-processing steps to ensure the creation of highly realistic outputs. MagiCapture outperforms other baselines in both quantitative and qualitative evaluations and can also be generalized to other non-human objects.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Improving Personalized Image Generation through Social Context Feedback

    cs.CV 2025-07 reject novelty 6.0 of 10

    A feedback fine-tuning method adds detector-based pose, interaction, identity, and gaze losses, time-gated by noise level, to SSR-Encoder, reporting modest gains on HICO-DET, GazeFollow, and Concept101.

Pith tools