REVIEW 4 cited by
AI-Generated Content (AIGC) for Various Data Modalities: A Survey
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
AI-generated content (AIGC) methods aim to produce text, images, videos, 3D assets, and other media using AI algorithms. Due to its wide range of applications and the potential of recent works, AIGC developments -- especially in Machine Learning (ML) and Deep Learning (DL) -- have been attracting significant attention, and this survey focuses on comprehensively reviewing such advancements in ML/DL. AIGC methods have been developed for various data modalities, such as image, video, text, 3D shape, 3D scene, 3D human avatar, 3D motion, and audio -- each presenting unique characteristics and challenges. Furthermore, there have been significant developments in cross-modality AIGC methods, where generative methods receive conditioning input in one modality and produce outputs in another. Examples include going from various modalities to image, video, 3D, and audio. This paper provides a comprehensive review of AIGC methods across different data modalities, including both single-modality and cross-modality methods, highlighting the various challenges, representative works, and recent technical directions in each setting. We also survey the representative datasets throughout the modalities, and present comparative results for various modalities. Moreover, we discuss the typical applications of AIGC methods in various domains, challenges, and future research directions.
Forward citations
Cited by 4 Pith papers
-
Visual Prompting for One-shot Controllable Video Editing without Inversion
A one-shot video editing method that uses a 2x2 visual prompt grid, modified consistency sampling, and Stein Variational Gradient Descent to propagate first-frame edits without DDIM inversion.
-
Generative Ghost: Investigating Ranking Bias Hidden in AI-Generated Videos
Text-video retrieval models systematically rank AI-generated videos above semantically matched real videos, driven by both visual and temporal cues and amplified by AI content in training data.
-
Guided Learning: Lubricating End-to-End Modeling for Multi-stage Decision-making
Adding intermediate guide losses to an end-to-end portfolio model improves backtested Sharpe and Calmar ratios versus stage-wise and unguided end-to-end baselines.
-
Compact Visual Data Representation for Green Multimedia -- A Human Visual System Perspective
A survey of compact visual data representation for green multimedia, structured around human visual system principles.
Discussion (0). Continue with ORCID to comment.