Pith. sign in

REVIEW 7 cited by

LLMs Meet Multimodal Generation and Editing: A Survey

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.19334 v2 pith:KH6WYOAT submitted 2024-05-29 cs.AI cs.CLcs.CVcs.MMcs.SD

classification cs.AIcs.CLcs.CVcs.MMcs.SD
keywords multimodalgenerationllmsmodelsgenerativeadvancementsdiscussediting
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

With the recent advancement in large language models (LLMs), there is a growing interest in combining LLMs with multimodal learning. Previous surveys of multimodal large language models (MLLMs) mainly focus on multimodal understanding. This survey elaborates on multimodal generation and editing across various domains, comprising image, video, 3D, and audio. Specifically, we summarize the notable advancements with milestone works in these fields and categorize these studies into LLM-based and CLIP/T5-based methods. Then, we summarize the various roles of LLMs in multimodal generation and exhaustively investigate the critical technical components behind these methods and the multimodal datasets utilized in these studies. Additionally, we dig into tool-augmented multimodal agents that can leverage existing generative models for human-computer interaction. Lastly, we discuss the advancements in the generative AI safety field, investigate emerging applications, and discuss future prospects. Our work provides a systematic and insightful overview of multimodal generation and processing, which is expected to advance the development of Artificial Intelligence for Generative Content (AIGC) and world models. A curated list of all related papers can be found at https://github.com/YingqingHe/Awesome-LLMs-meet-Multimodal-Generation

Discussion (0). Sign in to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. GraphVid: Interactive Graph-Controllable Video Generation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    GraphVid controls video generation with user-editable interaction scene graphs, reporting FID/FVD improvements over trajectory- and text-physics baselines using 0.6B trainable parameters.

  2. Measuring Agents in Production

    cs.CY 2025-12 conditional novelty 6.0 of 10

    Most deployed AI agents are small, human-checked workflows built by prompting off-the-shelf models, with reliability as the top challenge.

  3. MaaSO: SLO-aware Orchestration of Heterogeneous Model Instances for MaaS

    cs.DC 2025-09 conditional novelty 6.0 of 10

    MaaSO assigns different parallelism strategies and batch sizes to LLM instances and routes requests by deadline, improving simulated SLO attainment by 15 to 30 percent.

  4. Rethinking Layered Graphic Design Generation with a Top-Down Approach

    cs.CV 2025-07 conditional novelty 6.0 of 10

    Accordion decomposes AI-generated raster designs into editable background, object, and vectorized text layers using a VLM-driven top-down planning pipeline.

  5. Fake it till You Make it: Reward Modeling as Discriminative Prediction

    cs.CV 2025-06 conditional novelty 6.0 of 10

    GAN-RM trains a CLIP-based discriminator to distinguish a few hundred preference proxy images from model outputs, then uses it for Best-of-N selection, SFT, and DPO.

  6. EditInspector: A Benchmark for Evaluation of Text-Guided Image Edits

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A new human-labeled benchmark shows leading vision-language models are unreliable at judging image edits, and the authors' methods improve artifact detection and difference captioning.

  7. Event-Priori-Based Vision-Language Model for Efficient Visual Understanding

    cs.CV 2025-06 conditional novelty 6.0 of 10

    EP-VLM uses event-camera motion data to sparsify image patches before a vision-language model processes them, cutting FLOPs by about half with a small accuracy drop.

Pith tools