Art Arena evaluates how artistic styles from training data leak into AI-generated images without explicit prompts, revealing asymmetric blending due to differences in representational strength and interaction dynamics across models like Stable Diffusion.
hub
A Neural Algorithm of Artistic Style
19 Pith papers cite this work. Polarity classification is still indexing.
abstract
In fine art, especially painting, humans have mastered the skill to create unique visual experiences through composing a complex interplay between the content and style of an image. Thus far the algorithmic basis of this process is unknown and there exists no artificial system with similar capabilities. However, in other key areas of visual perception such as object and face recognition near-human performance was recently demonstrated by a class of biologically inspired vision models called Deep Neural Networks. Here we introduce an artificial system based on a Deep Neural Network that creates artistic images of high perceptual quality. The system uses neural representations to separate and recombine content and style of arbitrary images, providing a neural algorithm for the creation of artistic images. Moreover, in light of the striking similarities between performance-optimised artificial neural networks and biological vision, our work offers a path forward to an algorithmic understanding of how humans create and perceive artistic imagery.
hub tools
citation-role summary
citation-polarity summary
representative citing papers
The paper introduces a Markov kernel framework for exhaustively classifying corruptions in supervised learning and derives loss corrections for label, attribute, and joint cases by comparing clean and corrupted Bayes risks.
OSOG is a fast, fully differentiable physics-informed synthetic data engine for micro-optical environments that supports zero-shot transfer to real Lysozyme micrographs, inverse rendering, and sub-linear scaling synthesis.
UniTrans pretrains a bank of translator experts and learns combination coefficients from modality mappings in a scene-invariant latent space to enable zero-shot any-to-any feature translation for heterogeneous collaborative perception.
Proposes TinyUSFM-uLPIPS and TinyUSFM-NRQ metrics that show better alignment with segmentation task performance and expert preference than PSNR or VGG-LPIPS in ultrasound imaging.
Systematic benchmark reveals recent complex INR methods for continuous image super-resolution offer only marginal gains, with performance tied to training setups, auxiliary losses improving textures, and scaling laws holding.
VBench-2.0 is a benchmark suite that automatically evaluates video generative models on five dimensions of intrinsic faithfulness: Human Fidelity, Controllability, Creativity, Physics, and Commonsense using VLMs, LLMs, and anomaly detection methods.
Frontier LLMs' self-declared language support is unstable and over-optimistic, verified behavior is task-dependent, and language mismatch alone degrades collaborative agent performance.
WILD SAM combines denoised pseudo-labels from real adverse-weather images with simulation-based training to improve object detection AP by up to 13% on the Four Seasons dataset for rain and snow.
Gram-MMD is a texture-aware realism metric that computes MMD on upper-triangular Gram matrices from backbone activations, providing complementary information to semantic distributional metrics.
Hist2Style introduces a lightweight bilateral-grid network conditioned on histogram embeddings for distilling large-model stylization into real-time, structure-preserving, user-controllable photorealistic edits.
The paper introduces clean-model-based metrics that stratify test samples by vulnerability to targeted poisoning, enabling worst-case attack evaluation and vulnerability-aware defenses.
DMT uses identity and makeup encoders in a GAN to enable controllable makeup transfer from references and sampling of new styles from a prior distribution.
Style transfer-based domain adaptation improves cross-site calcification classification AUC by 0.04 on two external mammography datasets.
Universal aesthetic alignment in image models biases outputs toward conventional beauty and penalizes anti-aesthetic prompts even when they match explicit user instructions.
ArtNet and PhotoNet enable one-pass fast universal style transfer with fewer artifacts, better detail preservation, and 3-100x speedup over prior AE-based methods.
A layered image-processing pipeline with illumination transfer enables real-time facial makeup application from a single reference image while handling dark makeup and air-bangs.
DeepTEGINN is a deep learning toolbox combining image processing and graph theory to automate graph extraction from brain tissue images as an alternative to manual tracing.
The paper reviews limits in AI vision for robotics and describes work-in-progress on bridging sim-to-real domain gaps by linking real and synthetic training data.
citing papers explorer
-
The Silent Brush: Evaluating Artistic Style Leakage in AI Art Generation
Art Arena evaluates how artistic styles from training data leak into AI-generated images without explicit prompts, revealing asymmetric blending due to differences in representational strength and interaction dynamics across models like Stable Diffusion.
-
Corruptions of Supervised Learning Problems: Typology and Mitigations
The paper introduces a Markov kernel framework for exhaustively classifying corruptions in supervised learning and derives loss corrections for label, attribute, and joint cases by comparing clean and corrupted Bayes risks.
-
OSOG: A Differentiable, Physics-Informed Synthetic Data Engine for Micro-Optical Environments
OSOG is a fast, fully differentiable physics-informed synthetic data engine for micro-optical environments that supports zero-shot transfer to real Lysozyme micrographs, inverse rendering, and sub-linear scaling synthesis.
-
One Model to Translate Them All: Universal Any-to-Any Translation for Heterogeneous Collaborative Perception
UniTrans pretrains a bank of translator experts and learns combination coefficients from modality mappings in a scene-invariant latent space to enable zero-shot any-to-any feature translation for heterogeneous collaborative perception.
-
Defining Robust Ultrasound Quality Metrics via an Ultrasound Foundation Model
Proposes TinyUSFM-uLPIPS and TinyUSFM-NRQ metrics that show better alignment with segmentation task performance and expert preference than PSNR or VGG-LPIPS in ultrasound imaging.
-
Implicit Neural Representation-Based Continuous Single Image Super-Resolution: An Empirical Benchmark
Systematic benchmark reveals recent complex INR methods for continuous image super-resolution offer only marginal gains, with performance tied to training setups, auxiliary losses improving textures, and scaling laws holding.
-
VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness
VBench-2.0 is a benchmark suite that automatically evaluates video generative models on five dimensions of intrinsic faithfulness: Human Fidelity, Controllability, Creativity, Physics, and Commonsense using VLMs, LLMs, and anomaly detection methods.
-
Lost in the Tower of Babel: The Adverse Effects of Incidental Multilingualism in LLMs
Frontier LLMs' self-declared language support is unstable and over-optimistic, verified behavior is task-dependent, and language mismatch alone degrades collaborative agent performance.
-
WILD SAM: A Simulated-and-Real Data Augmentation for Autonomous Driving Perception under Challenging Weather
WILD SAM combines denoised pseudo-labels from real adverse-weather images with simulation-based training to improve object detection AP by up to 13% on the Four Seasons dataset for rain and snow.
-
Gram-MMD: A Texture-Aware Metric for Image Realism Assessment
Gram-MMD is a texture-aware realism metric that computes MMD on upper-triangular Gram matrices from backbone activations, providing complementary information to semantic distributional metrics.
-
Hist2Style: Histogram-Guided Stylization with Bilateral Grids
Hist2Style introduces a lightweight bilateral-grid network conditioned on histogram embeddings for distilling large-model stylization into real-time, structure-preserving, user-controllable photorealistic edits.
-
Are Targeted Data Poisoning Attacks as Effective as We Think?
The paper introduces clean-model-based metrics that stratify test samples by vulnerability to targeted poisoning, enabling worst-case attack evaluation and vulnerability-aware defenses.
-
Disentangled Makeup Transfer with Generative Adversarial Network
DMT uses identity and makeup encoders in a GAN to enable controllable makeup transfer from references and sampling of new styles from a prior distribution.
-
Unsupervised Domain Adaptation for Calcification Classification in Mammography Across Multi-Site Datasets
Style transfer-based domain adaptation improves cross-site calcification classification AUC by 0.04 on two external mammography datasets.
-
Position: Universal Aesthetic Alignment Narrows Artistic Expression
Universal aesthetic alignment in image models biases outputs toward conventional beauty and penalizes anti-aesthetic prompts even when they match explicit user instructions.
-
Fast Universal Style Transfer for Artistic and Photorealistic Rendering
ArtNet and PhotoNet enable one-pass fast universal style transfer with fewer artifacts, better detail preservation, and 3-100x speedup over prior AE-based methods.
-
Facial Makeup Transfer Combining Illumination Transfer
A layered image-processing pipeline with illumination transfer enables real-time facial makeup application from a single reference image while handling dark makeup and air-bangs.
-
DeepTEGINN: Deep Learning Based Tools to Extract Graphs from Images of Neural Networks
DeepTEGINN is a deep learning toolbox combining image processing and graph theory to automate graph extraction from brain tissue images as an alternative to manual tracing.
-
Efficiently Linking Real Scenes with Synthetic Data Generation for AI-based Cognitive Robotics and Computer Vision Applications
The paper reviews limits in AI vision for robotics and describes work-in-progress on bridging sim-to-real domain gaps by linking real and synthetic training data.