Pith. sign in

REVIEW 14 cited by

UniReal: Universal Image Generation and Editing via Learning Real-world Dynamics

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.07774 v2 pith:7APMQMQT submitted 2024-12-10 cs.CV

classification cs.CV
keywords generationtaskseditingimageunirealcapabilityconsistencydesigned
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce UniReal, a unified framework designed to address various image generation and editing tasks. Existing solutions often vary by tasks, yet share fundamental principles: preserving consistency between inputs and outputs while capturing visual variations. Inspired by recent video generation models that effectively balance consistency and variation across frames, we propose a unifying approach that treats image-level tasks as discontinuous video generation. Specifically, we treat varying numbers of input and output images as frames, enabling seamless support for tasks such as image generation, editing, customization, composition, etc. Although designed for image-level tasks, we leverage videos as a scalable source for universal supervision. UniReal learns world dynamics from large-scale videos, demonstrating advanced capability in handling shadows, reflections, pose variation, and object interaction, while also exhibiting emergent capability for novel applications.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MultiRef: Controllable Image Generation with Multiple Visual References

    cs.CV 2025-08 conditional novelty 7.0 of 10

    MultiRef-bench shows that current image generators that accept multiple visual references still fail to combine them reliably, with the best tested model OmniGen reaching only 66.6% synthetic and 79.0% real-world alig...

  2. DataClaw0: Agentic Tailoring Multimodal Data from Raw Streams

    cs.LG 2026-06 unverdicted novelty 6.0 of 10

    DataClaw0 introduces an agentic data-tailoring paradigm, a 9B model trained on a synthetically generated dataset, and a new benchmark, claiming improved downstream adaptation in video generation, VQA, and GUI navigati...

  3. ReMoT: Reinforcement Learning with Motion Contrast Triplets

    cs.CV 2026-02 conditional novelty 6.0 of 10

    Training a 4B vision-language model on rule-generated motion-contrast triplets with GRPO lifts spatio-temporal QA accuracy by about 17 points on the authors' own benchmark and by smaller margins on standard benchmarks.

  4. UMO: Scaling Multi-Identity Consistency for Image Customization via Matching Reward

    cs.CV 2025-09 conditional novelty 6.0 of 10

    A reinforcement-learning reward based on bipartite face matching improves multi-identity consistency and reduces identity confusion in image customization models.

  5. HOComp: Interaction-Aware Human-Object Composition

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A diffusion-transformer method that composes a foreground object into a human image with MLLM-chosen interaction regions, pose keypoint supervision, and appearance/background consistency losses, plus a new paired dataset.

  6. SeqTex: Generate Mesh Textures in Video Sequence

    cs.CV 2025-07 conditional novelty 6.0 of 10

    SeqTex adapts a pretrained video diffusion model to directly generate complete UV texture maps by jointly predicting four multi-view images and the UV map as a five-frame sequence.

  7. UNIC: Unified In-Context Video Editing

    cs.CV 2025-06 conditional novelty 6.0 of 10

    One diffusion transformer handles ID insert, swap, delete, stylization, propagation, and re-camera control in a single model using in-context token concatenation with task-aware positional encoding and bias.

  8. Image Editing As Programs with Diffusion Models

    cs.CV 2025-06 conditional novelty 6.0 of 10

    IEAP decomposes complex editing instructions into atomic operations executed sequentially on a diffusion transformer, and reports state-of-the-art results on MagicBrush and AnyEdit.

  9. OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A benchmark and five-million-clip dataset for evaluating and training subject-to-video generation models, with three new metrics for subject consistency, naturalness, and text alignment.

  10. Jodi: Unification of Visual Generation and Understanding via Joint Modeling

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A single diffusion transformer with role-switch training performs joint generation, controllable generation, and multi-label perception across image and seven label domains.

  11. Pixie: Fast and Generalizable Supervised Learning of 3D Physics from Pixels

    cs.CV 2025-08 reject novelty 5.0 of 10

    A supervised 3D U-Net predicts per-voxel material fields from CLIP feature grids, enabling fast MPM-based animation, but the reported evidence depends on pseudo-labels and a VLM judge from the same model family as the...

  12. XVerse: Consistent Multi-Subject Control of Identity and Semantic Attributes via DiT Modulation

    cs.CV 2025-06 conditional novelty 5.0 of 10

    XVerse learns token-specific offsets that modify the text-stream modulation of a diffusion transformer, enabling multi-subject identity and attribute control in image generation.

  13. Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning

    cs.CV 2025-06 reject novelty 5.0 of 10

    A new micro-edit dataset and fine-tuning recipe appear to help multimodal LLMs notice small visual changes, but the central 'feature consistency loss' claim is not present in the method.

  14. ByteMorph: Benchmarking Instruction-Guided Image Editing with Non-Rigid Motions

    cs.CV 2025-06 conditional novelty 5.0 of 10

    A released 6.4 million pair dataset and 613 sample benchmark for instruction-guided image editing of non-rigid motions, plus a Flux.1-dev based baseline that outperforms open-source methods on the new benchmark.

Pith tools