REVIEW 4 major objections 4 minor 1 cited by
Fine-grained emotion control in AI image generation is achievable with learnable prototypes
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-03 00:01 UTC pith:POYBWXJZ
load-bearing objection A clever, novel emotion-prototype generation system whose headline numbers are partly measuring its own objective; worth engaging, but the evaluation needs independent emotion assessment before the claims will convince me. the 4 major comments →
EmoSpace: Immersive Affective Image Generation Guided by Fine-Grained Emotion Prototypes
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central claim is that emotions in visual content are better represented as a set of learnable prototypes in a shared vision-language embedding space than as fixed categories or points on a valence-arousal plane. EmoSpace learns a bank of such prototypes from an emotional image dataset through a composite loss, letting them merge and split during training to adaptively cover the emotion distribution. At generation time, a user's free-text emotion description is mapped to a weighted combination of the most similar prototypes, which then guide a diffusion model via attention reweighting, iterative prompt refinement that adds sub-emotion descriptors, and temporal blending that shifts from co
What carries the argument
The prototype bank with merge-split dynamics, together with the mapper that transfers vision-language embeddings into the diffusion model's prompt space. The hierarchical emotion representation (basic categories plus a large set of dynamic prototypes) is the central object; it creates a 'structured-yet-continuous' emotion space that is interpretable and controllable. Multi-prototype guidance computes a weighted combination of similar prototypes for conditioning; attention reweighting injects this into the model's cross-attention layers; temporal blending schedules its influence during denoising; and iterative prompt refinement uses a language model to enrich the prompt with sub-emotion words
Load-bearing premise
The paper relies on the assumption that vision-language embedding similarity between a generated image and an emotion-description prompt faithfully measures fine-grained emotional alignment; the prototypes, mapper, and prompt refinement are all optimized to maximize this specific similarity, so if it doesn't track human perception, the quantitative gains become partly self-referential.
What would settle it
A large, pre-registered human study in which independent raters judge whether images generated by EmoSpace versus baseline methods match a set of fine-grained emotion descriptions. If humans show no significant preference for EmoSpace, or if the embedding-similarity scores used in the paper fail to predict human judgments across models, the central claim would be refuted. Alternatively, an ablation that freezes the prototype bank to a fixed set of categorical labels while keeping all other components would reveal whether the dynamic prototypes actually drive the claimed advantage.
If this is right
- Content creators can steer image generation with free-form emotion descriptions rather than fixed labels, making emotional control more intuitive and expressive.
- The same emotion-conditioning mechanism works across standard text-to-image, panorama, style transfer, and outpainting, suggesting a unified approach to affective content generation.
- If the reported gains hold, prototype-based emotion spaces could replace categorical and dimensional models for downstream generation tasks.
- The VR user study suggests that immersive presentation itself shifts emotional perception, so emotion-aware generation for VR may need to account for viewing medium.
Where Pith is reading between the lines
- The dynamic prototype bank could generalize to other abstract attributes (e.g., cultural tone, aesthetic mood) beyond emotion; the same merge-split mechanism might learn arbitrary semantic spaces from paired image-text data.
- Replacing the vision-language-similarity evaluation with forced-choice human judgments would test whether the quantitative lead reflects genuine emotional fidelity rather than optimization of the same metric used for conditioning.
- The VR finding that emotion category selection shifts toward positive emotions suggests that content tuned for desktop may not read identically in VR; a practical extension would be to condition generation on display mode.
- Because prototypes are interpretable embeddings, they could be probed to reveal which emotional concepts are being activated, enabling cross-cultural adaptation of emotion control.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes EmoSpace, a framework for fine-grained emotion-controlled image generation. It learns a bank of dynamic CLIP-space emotion prototypes via vision-language alignment, then steers a diffusion model using multi-prototype guidance, temporal blending, attention reweighting, and GPT-assisted iterative prompt refinement. The framework is extended to outpainting, stylized generation, and panorama generation. The paper reports quantitative comparisons with EmoGen, EmotiCrafter, and SDXL (Tables II-III), ablations (Table V), a small human ranking study (Sec. V-D), and a VR vs. desktop user study (Sec. VI) with the claim that EmoSpace improves emotional alignment and aesthetic quality while immersive presentation changes emotion perception.
Significance. Fine-grained emotionally controllable generative models for VR are a useful and timely topic, and the paper is ambitious in scope. The strengths are the prototype-based emotion representation, the integration of multiple control pathways, the breadth of downstream tasks, and the inclusion of a VR user study. If the central claim is supported by independent evidence, EmoSpace would be a practical contribution to affective computing and immersive content generation. However, the current evidence is not yet convincing: the automated emotion metrics are computed in the same CLIP space the method explicitly optimizes, baselines differ in base model, and the human evaluation is small and unreleased. The result therefore has moderate-to-high potential but needs additional validation.
major comments (4)
- [Sec. V-A / Sec. IV-B / Sec. IV-C2] Table II's headline emotion-accuracy gains are computed with CLIP cosine similarity in CLIP-H-14 space (Emotion-Category and Emotion-Fine-Grained, Sec. V-A). This is the same embedding space where the prototype bank is learned (Sec. IV-B1: d_v=d_t=1024), where multi-prototype guidance retrieves prototypes by cosine similarity, and where iterative prompt refinement explicitly selects candidates maximizing <phi_CLIP(c), p_emo> (Sec. IV-C2). The MLP mapper is also trained to maximize cosine similarity. Thus the reported improvements are partly self-referential: they measure how well the system optimizes its training/selection objective. The only independent evidence is the ranking study in Sec. V-D (28 users, 104 images), but the images and protocol are not released, and the VR user study in Sec. VI is not a head-to-head method comparison. Please add a human emotion-alignment evaluation on
- [Tables II and III] The quantitative comparison is confounded by base-model differences. EmoGen uses SD 1.5, while EmoSpace, EmotiCrafter, and SDXL use SDXL (Sec. V-A). The very large CLIP-Prompt gap (17.49 for EmoGen vs. 31.99 for SDXL) and aesthetic gap (6.13 vs. 16.25) are likely inherited from the base model rather than reflecting emotion-control ability. Table III equalizes the prompts but not the base models, so SD 1.5 baselines remain disadvantaged. Please either adapt EmoGen to SDXL (if feasible), provide an SD 1.5 version of EmoSpace, or include a same-base SDXL baseline for categorical/dimensional emotion conditioning. Without this, Table II's cross-method rankings for emotion and aesthetics are not interpretable.
- [Sec. IV-B, IV-C and supplementary material] Several mechanisms central to the method are only described by reference to a supplementary file that is not included. Specifically, the merge/split criteria and update rules for the dynamic prototype bank, the exact definitions of L_contrast, L_diversity, and L_dist, the three-phase temporal blending schedule, and the GPT-based iterative prompt refinement details (number of iterations, threshold, prompt template, temperature) are all deferred. In addition, the abstract states 256 prototypes while Sec. IV-B1 sets K=1024. These omissions prevent full assessment and replication of the central representation and steering contributions. Please include the supplementary material or move the detailed formulations into the main text, and resolve the prototype-count inconsistency.
- [Sec. VI-B] The VR user study uses independent-samples t-tests on a within-subject design: each participant experienced both VR and desktop conditions, so the observations are paired. The reported values (T1 t=-1.742, p=0.082; T2 t=-0.495, p=0.621; T3 t=-1.161, p=0.246) are also inconsistent with the means (VR > desktop in all cases), suggesting the sign convention or the test may be erroneous. Please use paired-samples tests and report effect sizes. This does not affect the text-to-image comparison in Sec. V, but it is load-bearing for the VR-perception claims (H1 and H2) in Sec. VI.
minor comments (4)
- [Sec. V-A] The CLIP-Prompt score for EmoSpace is computed against the original prompt, while the generation actually uses the refined prompt. The paper notes this, but it should also report alignment to the final refined prompt so that the text-image alignment measure is well defined.
- [Sec. V-B] The statement that EmoSpace is 'the first framework to support diverse emotional generation tasks' should be qualified. Prior systems support outpainting, style LoRAs, and panorama generation; the novelty is the unified emotion-conditioning mechanism, not the task support per se.
- [Fig. 8] The panels in Fig. 8 lack clear axis labels and the caption repeats numeric results from the text. Please make the figure self-contained (e.g., label 'Mean Rating' and 'Selection Distribution' directly on the axes) and avoid duplicating statistics in the caption.
- [Sec. VI-C] The 'Limitations and Future Work' subsection addresses only hardware comfort and user-requested features. There is no limitations paragraph for the proposed method or the evaluation (e.g., dependency on CLIP similarity, small human study, lack of release). Adding one would improve scientific transparency.
Circularity Check
Head-table emotion-accuracy gains rest on the same CLIP cosine similarity EmoSpace is explicitly trained/selected to maximize; independent human studies provide partial but not complete mitigation.
specific steps
-
fitted input called prediction
[Sec. V-A (Evaluation Metrics); Sec. IV-B2 (Training Objectives); Sec. IV-C2 (Iterative Prompt Refinement)]
""Emotion Accuracy: We employ CLIP-based metrics to evaluate emotion alignment at both categorical (Emotion-Category) and fine-grained (Emotion-Fine-Grained) levels using cosine similarity in CLIP feature space (higher is better)." ... "the mapper is trained using a composite loss that combines cosine similarity maximization and Euclidean distance minimization" ... "then selects the dominant emotion c* that maximizes CLIP-space similarity ⟨φ_CLIP(c_i), p_emo⟩""
The headline emotion-accuracy metrics are CLIP cosine similarities, while the MLP mapper is explicitly trained with cosine-similarity maximization, the iterative prompt refinement selects candidates by CLIP-space similarity, and the prototypes/embeddings live in CLIP-H-14 space (Sec. IV-B1). Thus EmoSpace is directly optimized for the same quantity used as the independent evaluation metric. The Emo-Cat and Emo-FG gains in Tables II and IV are at least partly forced by construction rather than being an external test of human-perceived emotion. The human ranking study in Sec. V-D gives some independent support, but the main quantitative head-to-head superiority claim rests on this self-referential metric.
full rationale
The paper is largely a systems/application paper: prototype learning, multi-prototype guidance, attention reweighting, and VR-specific extensions are evaluated with ablations, qualitative comparisons, and user studies. The central quantitative claim of 'superior performance in emotion accuracy' is where circularity enters. Sec. V-A defines Emotion-Category and Emotion-Fine-Grained as CLIP cosine similarity between generated images and emotion text. The same cosine similarity appears as the training objective for the MLP mapper (Sec. IV-B2) and as the selection criterion in iterative prompt refinement (Sec. IV-C2), with all prototypes and embeddings in CLIP-H-14 space. Consequently, the reported metric gains are partly self-referential: the method explicitly maximizes the measure used to declare it the winner. This is not a fatal flaw because the paper also reports an independent 28-participant ranking study (Sec. V-D) and a VR pairwise preference test (Sec. VI-B) that favor EmoSpace; those are genuine, if small and unreleased, human evaluations. No load-bearing self-citation chain was found: the one overlapping-author reference ([55]) is not used to justify the central derivation, and there is no imported uniqueness theorem or ansatz-by-citation. The score of 6 reflects partial circularity in the quantitative emotion-accuracy evidence, not in the whole framework. The small/unreleased human studies are a validity concern but not themselves a circularity step.
Axiom & Free-Parameter Ledger
free parameters (8)
- Number of emotion prototypes K =
1024 (abstract says 256)
- Merge threshold tau_m =
0.3
- Split threshold tau_s =
0.7
- Loss weights alpha_L, beta_L, gamma_L, delta_L =
1.0, 0.5, 0.1, 0.1
- Attention injection intensity alpha_attn =
1.5
- Softmax temperature tau_temp =
0.1
- Top-k for positive/negative prototypes =
not stated
- Three-phase temporal blending schedule =
not specified
axioms (6)
- domain assumption CLIP-H-14 embedding space is a sufficient representation for fine-grained emotion semantics and for evaluating emotion alignment.
- domain assumption EmoSet-118K's categorical labels plus BLIP-2 captions provide enough emotional supervision to learn generalizable prototypes.
- domain assumption A lightweight MLP can map CLIP emotion vectors into SDXL text-embedding space without losing emotional content.
- domain assumption Diffusion denoising exhibits a three-phase hierarchical structure that justifies the temporal blending schedule.
- domain assumption GPT-4o-mini refinement increases rather than distorts emotion-prompt correspondence.
- domain assumption Prototype learning from object classification transfers to abstract emotion semantics.
invented entities (1)
-
Dynamic emotion prototype bank with merge/split
no independent evidence
read the original abstract
Immersive affective content generation aims to create visually compelling VR imagery with controllable emotional nuance, yet existing methods typically rely on coarse labels or prompt-only control. Although modern diffusion transformers (DiTs) such as FLUX improve visual fidelity, they are not designed to incorporate structured affective representations. We present EmoSpace, an immersive affective content generation framework guided by fine-grained emotion prototypes, transforming free-form text and emotion descriptions into fine-grained affective imagery through three coordinated components. First, to represent sub-emotion variation beyond conventional categorical models, EmoSpace learns a hierarchical bank of 256 prototypes with input-conditioned adaptation through vision-language alignment. Second, Prototype-Conditioned Steering converts these prototypes into DiT-compatible generation signals through multi-pathway injection and temporal blending, while Iterative Prompt Refinement enriches prompts with prototype-aligned sub-emotion descriptors. Third, Affect-Grounded Modulation coordinates emotion conditioning with controllable LoRAs for panoramic, stylized, and multi-conditional generation. Through quantitative and qualitative evaluations, EmoSpace improves fine-grained emotional alignment while maintaining high aesthetic quality. Our user study shows that EmoSpace outputs are perceived as more emotionally aligned than baseline results and more suitable for immersive scene design. Additionally, we find that immersive presentation alters emotional perception and increases emotional engagement. Together, these findings inform the design of emotion-aware generative systems for immersive media, with potential applications including education, immersive storytelling, and artistic creation. We will release our code and models to facilitate future research along this line.
Figures
Forward citations
Cited by 1 Pith paper
-
EmoScene: A Dual-space Dataset for Controllable Affective Image Generation
EmoScene contributes 1.2M images annotated with discrete emotions, continuous VAD scores, perceptual attributes, and captions, plus a cross-attention modulation that shifts generated images toward requested affective targets.
Reference graph
Works this paper leans on
-
[1]
Feeling Virtually Present Makes Me Happier: The Influence of Immersion, Sense of Presence, and Video Contents on Positive Emotion Induction,
K. Pavic, L. Chaby, T. Gricourt, and D. Vergilino-Perez, “Feeling Virtually Present Makes Me Happier: The Influence of Immersion, Sense of Presence, and Video Contents on Positive Emotion Induction,” Cyberpsychology, Behavior, and Social Networking, vol. 26, no. 4, pp. 238–245, 2023
2023
-
[2]
Virtual Reality for Emotion Elicitation–a Review,
R. Somarathna, T. Bednarz, and G. Mohammadi, “Virtual Reality for Emotion Elicitation–a Review,”IEEE Transactions on Affective Com- puting, vol. 14, no. 4, pp. 2626–2645, 2022
2022
-
[3]
A Survey on Affective and Cognitive VR,
T. Luong, A. Lecuyer, N. Martin, and F. Argelaguet, “A Survey on Affective and Cognitive VR,”IEEE Transactions on Visualization and Computer Graphics, vol. 28, no. 12, pp. 5154–5171, 2021
2021
-
[4]
Emotional qualities of VR space,
A. Naz, R. Kopper, R. P. McMahan, and M. Nadin, “Emotional qualities of VR space,” in2017 IEEE virtual reality (VR). IEEE, 2017, pp. 3–11
2017
-
[5]
Traces in Virtual Environments: A Framework and Exploration to Conceptualize the Design of Social Vir- tual Environments,
L. Hirsch, C. George, and A. Butz, “Traces in Virtual Environments: A Framework and Exploration to Conceptualize the Design of Social Vir- tual Environments,”IEEE Transactions on Visualization and Computer Graphics, vol. 28, no. 11, pp. 3874–3884, 2022
2022
-
[6]
Affective com- puting in virtual reality: emotion recognition from brain and heartbeat dynamics using wearable sensors,
J. Mar ´ın-Morales, J. L. Higuera-Trujillo, A. Greco, J. Guixeres, C. Llinares, E. P. Scilingo, M. Alca ˜niz, and G. Valenza, “Affective com- puting in virtual reality: emotion recognition from brain and heartbeat dynamics using wearable sensors,”Scientific reports, vol. 8, no. 1, p. 13657, 2018
2018
-
[7]
Revolutionizing Virtual Reality With Generative AI: An In-Depth Review,
M. E. A. Hashim, W. A. W. Mustafa, N. S. Prameswari, M. M. Ghani, and H. F. Hanafi, “Revolutionizing Virtual Reality With Generative AI: An In-Depth Review,”Journal of Advanced Research in Computing and Applications, vol. 30, no. 1, pp. 19–30, 2023. 13
2023
-
[8]
Generative AI Meets Virtual Reality: A Comprehensive Survey on Applications, Challenges, and Future Direction,
F. Rahimi, A. Sadeghi-Niaraki, and S.-M. Choi, “Generative AI Meets Virtual Reality: A Comprehensive Survey on Applications, Challenges, and Future Direction,”IEEE Access, 2025
2025
-
[9]
Effects of Virtual Environment Platforms on Emotional Responses,
K. Kim, M. Z. Rosenthal, D. J. Zielinski, and R. Brady, “Effects of Virtual Environment Platforms on Emotional Responses,”Computer Methods and Programs in Biomedicine, vol. 113, no. 3, pp. 882–893, 2014
2014
-
[10]
The Effect of Immersion on Emotional Responses to Film Viewing in a Virtual Environment,
A. Kim, M. Chang, Y . Choi, S. Jeon, and K. Lee, “The Effect of Immersion on Emotional Responses to Film Viewing in a Virtual Environment,” in2018 IEEE Conference on Virtual Reality and 3D User Interfaces (VR). IEEE, 2018, pp. 601–602
2018
-
[11]
Effects of Immersion in a Simulated Natural Environment on Stress Reduction and Emotional Arousal: A Systematic Review and Meta-Analysis,
H. Li, Y . Ding, B. Zhao, Y . Xu, and W. Wei, “Effects of Immersion in a Simulated Natural Environment on Stress Reduction and Emotional Arousal: A Systematic Review and Meta-Analysis,”Frontiers in psy- chology, vol. 13, p. 1058177, 2023
2023
-
[12]
Basic emotions,
P. Ekman, T. Dalgleish, and M. Power, “Basic emotions,”San Francisco, USA, 1999
1999
-
[13]
A Circumplex Model of Affect,
J. A. Russell, “A Circumplex Model of Affect,”Journal of personality and social psychology, vol. 39, no. 6, p. 1161, 1980
1980
-
[14]
Emogen: Emotional Image Content Generation with Text-to-Image Diffusion Models,
J. Yang, J. Feng, and H. Huang, “Emogen: Emotional Image Content Generation with Text-to-Image Diffusion Models,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 6358–6368
2024
-
[15]
EmotiCrafter: Text-to-Emotional-Image Generation Based on Valence-Arousal Model,
S. Dang, Y . He, L. Ling, Z. Qian, N. Zhao, and N. Cao, “EmotiCrafter: Text-to-Emotional-Image Generation Based on Valence-Arousal Model,” arXiv preprint arXiv:2501.05710, 2025
arXiv 2025
-
[16]
A systematic review on affective computing: Emotion models, databases, and recent advances,
Y . Wang, W. Song, W. Tao, A. Liotta, D. Yang, X. Li, S. Gao, Y . Sun, W. Ge, W. Zhanget al., “A systematic review on affective computing: Emotion models, databases, and recent advances,”Information Fusion, vol. 83, pp. 19–52, 2022
2022
-
[17]
Emotional Category Data on Images from the International Affective Picture System,
J. A. Mikels, B. L. Fredrickson, G. R. Larkin, C. M. Lindberg, S. J. Maglio, and P. A. Reuter-Lorenz, “Emotional Category Data on Images from the International Affective Picture System,”Behavior research methods, vol. 37, no. 4, pp. 626–630, 2005
2005
-
[18]
Pleasure-Arousal-Dominance: A General Framework for Describing and Measuring Individual Differences in Temperament,
A. Mehrabian, “Pleasure-Arousal-Dominance: A General Framework for Describing and Measuring Individual Differences in Temperament,” Current psychology, vol. 14, no. 4, pp. 261–292, 1996
1996
-
[19]
Robust Image Sentiment Analysis Using Progressively Trained and Domain Transferred Deep Networks,
Q. You, J. Luo, H. Jin, and J. Yang, “Robust Image Sentiment Analysis Using Progressively Trained and Domain Transferred Deep Networks,” inProceedings of the AAAI Conference on Artificial Intelligence, vol. 29, no. 1, 2015
2015
-
[20]
MA VEN: Multi-Modal Attention for Valence-Arousal Emotion Network,
V . Ahire, K. Shah, M. Khan, N. Pakhale, L. Sookha, M. Ganaie, and A. Dhall, “MA VEN: Multi-Modal Attention for Valence-Arousal Emotion Network,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2025, pp. 5789–5799
2025
-
[21]
EmoVit: Revolutionizing Emotion Insights with Visual Instruction Tuning,
H. Xie, C.-J. Peng, Y .-W. Tseng, H.-J. Chen, C.-F. Hsu, H.-H. Shuai, and W.-H. Cheng, “EmoVit: Revolutionizing Emotion Insights with Visual Instruction Tuning,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 26 596–26 605
2024
-
[22]
Emotion-LLaMA: Multimodal Emotion Recognition and Reasoning with Instruction Tuning,
Z. Cheng, Z.-Q. Cheng, J.-Y . He, K. Wang, Y . Lin, Z. Lian, X. Peng, and A. Hauptmann, “Emotion-LLaMA: Multimodal Emotion Recognition and Reasoning with Instruction Tuning,”Advances in Neural Informa- tion Processing Systems, vol. 37, pp. 110 805–110 853, 2024
2024
-
[23]
GPT-4V with Emotion: A Zero-Shot Benchmark for Generalized Emotion Recognition,
Z. Lian, L. Sun, H. Sun, K. Chen, Z. Wen, H. Gu, B. Liu, and J. Tao, “GPT-4V with Emotion: A Zero-Shot Benchmark for Generalized Emotion Recognition,”Information Fusion, vol. 108, p. 102367, 2024
2024
-
[24]
Y . Zhou, Z. Zhang, J. Cao, J. Jia, Y . Jiang, F. Wen, X. Liu, X. Min, and G. Zhai, “MEMO-Bench: A Multiple Benchmark for Text-to-Image and Multimodal Large Language Models on Human Emotion Analysis,” arXiv preprint arXiv:2411.11235, 2024
Pith/arXiv arXiv 2024
-
[25]
Make Me Happier: Evoking Emotions Through Image Diffusion Models,
Q. Lin, J. Zhang, Y . S. Ong, and M. Zhang, “Make Me Happier: Evoking Emotions Through Image Diffusion Models,”arXiv preprint arXiv:2403.08255, 2024
Pith/arXiv arXiv 2024
-
[26]
Sal-Guide Diffusion: Saliency Maps Guide Emotional Image Generation through Adapter,
X. Lin, S. Zhong, Y . Liu, and G. Chen, “Sal-Guide Diffusion: Saliency Maps Guide Emotional Image Generation through Adapter,” in2024 IEEE International Conference on Multimedia and Expo (ICME). IEEE, 2024, pp. 1–6
2024
-
[27]
TGCA-PVT: Topic-Guided Context-Aware Pyramid Vision Transformer for Sticker Emotion Recognition,
J. Chen, W. Wang, Y . Hu, J. Chen, H. Liu, and X. Hu, “TGCA-PVT: Topic-Guided Context-Aware Pyramid Vision Transformer for Sticker Emotion Recognition,” inProceedings of the 32nd ACM International Conference on Multimedia, 2024, pp. 9709–9718
2024
-
[28]
Knowledge-Aligned Counterfactual-Enhancement Diffusion Perception for Unsupervised Cross-Domain Visual Emotion Recognition,
W. Yin, Y . Wang, G. Duan, D. Zhang, X. Hu, Y .-F. Li, and T. He, “Knowledge-Aligned Counterfactual-Enhancement Diffusion Perception for Unsupervised Cross-Domain Visual Emotion Recognition,” inPro- ceedings of the Computer Vision and Pattern Recognition Conference, 2025, pp. 3888–3898
2025
-
[29]
EmoEdit: Evoking Emotions Through Image Manipulation,
J. Yang, J. Feng, W. Luo, D. Lischinski, D. Cohen-Or, and H. Huang, “EmoEdit: Evoking Emotions Through Image Manipulation,” inPro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2025, pp. 24 690–24 699
2025
-
[30]
Image Re-Emotionalizing,
M. Xu, B. Ni, J. Tang, and S. Yan, “Image Re-Emotionalizing,” inThe Era of Interactive Media. Springer, 2012, pp. 3–14
2012
-
[31]
Affective Image Editing: Shaping Emotional Factors via Text Descriptions,
P. Zhang, S. Weng, C. Zhu, B. Tang, Z. Jia, S. Li, and B. Shi, “Affective Image Editing: Shaping Emotional Factors via Text Descriptions,”arXiv preprint arXiv:2505.18699, 2025
arXiv 2025
-
[32]
Bridge Then Begin Anew: Generating Target-Relevant Intermediate Model for Source-Free Visual Emotion Adaptation,
J. Zhu, S. Zhao, J. Jiang, W. Tang, Z. Xu, T. Han, P. Xu, and H. Yao, “Bridge Then Begin Anew: Generating Target-Relevant Intermediate Model for Source-Free Visual Emotion Adaptation,” inProceedings of the AAAI Conference on Artificial Intelligence, vol. 39, no. 2, 2025, pp. 1674–1682
2025
-
[33]
It Is Okay to Not Be Okay: Overcoming Emotional Bias in Affective Image Captioning by Contrastive Data Collection,
Y . Mohamed, F. F. Khan, K. Haydarov, and M. Elhoseiny, “It Is Okay to Not Be Okay: Overcoming Emotional Bias in Affective Image Captioning by Contrastive Data Collection,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 21 263–21 272
2022
-
[34]
J. D. Lomas, W. van der Maden, S. Bandyopadhyay, G. Lion, N. Patel, G. Jain, Y . Litowsky, H. Xue, and P. Desmet, “Improved Emotional Alignment of AI and Humans: Human Ratings of Emotions Expressed by Stable Diffusion v1, DALL-E 2, and DALL-E 3,”arXiv preprint arXiv:2405.18510, 2024
Pith/arXiv arXiv 2024
-
[35]
All One Needs to Know About Metaverse: A Complete Survey on Technological Singularity, Virtual Ecosystem, and Research Agenda,
L.-H. Lee, T. Braud, P. Y . Zhou, L. Wang, D. Xu, Z. Lin, A. Kumar, C. Bermejo, P. Huiet al., “All One Needs to Know About Metaverse: A Complete Survey on Technological Singularity, Virtual Ecosystem, and Research Agenda,”Foundations and Trends in Human-Computer Interaction, vol. 18, no. 2–3, pp. 100–337, 2024
2024
-
[36]
PanoWan: Lifting Diffusion Video Generation Models to 360° with Latitude/Longitude-aware Mechanisms,
Y . Xia, S. Weng, S. Yang, J. Liu, C. Zhu, M. Teng, Z. Jia, H. Jiang, and B. Shi, “PanoWan: Lifting Diffusion Video Generation Models to 360° with Latitude/Longitude-aware Mechanisms,”arXiv preprint arXiv:2505.22016, 2025
Pith/arXiv arXiv 2025
-
[37]
DreamCube: 3D Panorama Generation via Multi-plane Synchronization,
Y . Huang, Y . Zhou, J. Wang, K. Huang, and X. Liu, “DreamCube: 3D Panorama Generation via Multi-plane Synchronization,”arXiv preprint arXiv:2506.17206, 2025
Pith/arXiv arXiv 2025
-
[38]
LayerPano3D: Layered 3D Panorama for Hyper-Immersive Scene Gen- eration,
S. Yang, J. Tan, M. Zhang, T. Wu, G. Wetzstein, Z. Liu, and D. Lin, “LayerPano3D: Layered 3D Panorama for Hyper-Immersive Scene Gen- eration,” inProceedings of the Special Interest Group on Computer Graphics and Interactive Techniques Conference Conference Papers, 2025, pp. 1–10
2025
-
[39]
DynamicScaler: Seamless and Scalable Video Generation for Panoramic Scenes,
J. Liu, S. Lin, Y . Li, and M.-H. Yang, “DynamicScaler: Seamless and Scalable Video Generation for Panoramic Scenes,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2025, pp. 6144–6153
2025
-
[40]
Taming Stable Diffusion for Text to 360 Panorama Image Generation,
C. Zhang, Q. Wu, C. C. Gambardella, X. Huang, D. Phung, W. Ouyang, and J. Cai, “Taming Stable Diffusion for Text to 360 Panorama Image Generation,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 6347–6357
2024
-
[41]
Kiss3DGen: Repurposing Image Diffusion Models for 3D Asset Generation,
J. Lin, X. Yang, M. Chen, Y . Xu, D. Yan, L. Wu, X. Xu, L. Xu, S. Zhang, and Y .-C. Chen, “Kiss3DGen: Repurposing Image Diffusion Models for 3D Asset Generation,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2025, pp. 5870–5880
2025
-
[42]
SceneDreamer: Unbounded 3D Scene Generation from 2D Image Collections,
Z. Chen, G. Wang, and Z. Liu, “SceneDreamer: Unbounded 3D Scene Generation from 2D Image Collections,”IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 45, no. 12, pp. 15 562–15 576, 2023
2023
-
[43]
Matrix-3D: Omnidirectional Explorable 3D World Generation,
Z. Yang, W. Ge, Y . Li, J. Chen, H. Li, M. An, F. Kang, H. Xue, B. Xu, Y . Yinet al., “Matrix-3D: Omnidirectional Explorable 3D World Generation,”arXiv preprint arXiv:2508.08086, 2025
Pith/arXiv arXiv 2025
-
[44]
Infinite Photorealistic Worlds Using Procedural Generation,
A. Raistrick, L. Lipson, Z. Ma, L. Mei, M. Wang, Y . Zuo, K. Kayan, H. Wen, B. Han, Y . Wanget al., “Infinite Photorealistic Worlds Using Procedural Generation,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 12 630–12 641
2023
-
[45]
DI-PCG: Diffusion- Based Efficient Inverse Procedural Content Generation for High-Quality 3D Asset Creation,
W. Zhao, Y .-P. Cao, J. Xu, Y . Dong, and Y . Shan, “DI-PCG: Diffusion- Based Efficient Inverse Procedural Content Generation for High-Quality 3D Asset Creation,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2025, pp. 11 061–11 072
2025
-
[46]
CEAP-360VR: A Continuous Physiological and Behavioral Emotion Annotation Dataset for 360 VR Videos,
T. Xue, A. El Ali, T. Zhang, G. Ding, and P. Cesar, “CEAP-360VR: A Continuous Physiological and Behavioral Emotion Annotation Dataset for 360 VR Videos,”IEEE Transactions on Multimedia, vol. 25, pp. 243–255, 2021
2021
-
[47]
An Immer- sive and Interactive VR Dataset to Elicit Emotions,
W. Jiang, M. Windl, B. Tag, Z. Sarsenbayeva, and S. Mayer, “An Immer- sive and Interactive VR Dataset to Elicit Emotions,”IEEE Transactions on Visualization and Computer Graphics, 2024
2024
-
[48]
High-Resolution Image Synthesis with Latent Diffusion Models,
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “High-Resolution Image Synthesis with Latent Diffusion Models,” in 14 Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 10 684–10 695
2022
-
[49]
An Image Is Worth One Word: Personalizing Text-to-Image Generation Using Textual Inversion,
R. Gal, Y . Alaluf, Y . Atzmon, O. Patashnik, A. H. Bermano, G. Chechik, and D. Cohen-Or, “An Image Is Worth One Word: Personalizing Text-to-Image Generation Using Textual Inversion,”arXiv preprint arXiv:2208.01618, 2022
Pith/arXiv arXiv 2022
-
[50]
LoRA: Low-Rank Adaptation of Large Language Models,
E. J. Hu, Y . Shen, P. Wallis, Z. Allen-Zhu, Y . Li, S. Wang, L. Wang, and W. Chen, “LoRA: Low-Rank Adaptation of Large Language Models,” inICLR. OpenReview.net, 2022. [Online]. Available: http://dblp.uni-trier.de/db/conf/iclr/iclr2022.html#HuSW ALWWC22
2022
-
[51]
The Nature of Emotions: Human Emotions Have Deep Evolutionary Roots, a Fact That May Explain Their Complexity and Provide Tools for Clinical Practice,
R. Plutchik, “The Nature of Emotions: Human Emotions Have Deep Evolutionary Roots, a Fact That May Explain Their Complexity and Provide Tools for Clinical Practice,”American scientist, vol. 89, no. 4, pp. 344–350, 2001
2001
-
[52]
Sixteen Facial Expressions Occur in Similar Contexts Worldwide,
A. S. Cowen, D. Keltner, F. Schroff, B. Jou, H. Adam, and G. Prasad, “Sixteen Facial Expressions Occur in Similar Contexts Worldwide,” Nature, vol. 589, no. 7841, pp. 251–257, 2021
2021
-
[53]
Mapping the Passions: Toward a High-Dimensional Taxonomy of Emotional Experience and Expression,
A. Cowen, D. Sauter, J. L. Tracy, and D. Keltner, “Mapping the Passions: Toward a High-Dimensional Taxonomy of Emotional Experience and Expression,”Psychological Science in the Public Interest, vol. 20, no. 1, pp. 69–90, 2019
2019
-
[54]
Exploring Inter- pretability in Deep Learning for Affective Computing: A Comprehensive Review,
X. Zhang, T. Zhang, L. Sun, J. Zhao, and Q. Jin, “Exploring Inter- pretability in Deep Learning for Affective Computing: A Comprehensive Review,”ACM Transactions on Multimedia Computing, Communica- tions and Applications, 2025
2025
-
[55]
Emotion- Lens: Interactive Visual Exploration of the Circumplex Emotion Space in Literary Works via Affective Word Clouds,
B. Wang, Q. Shi, X. Wang, Y . Zhou, W. Zeng, and Z. Wang, “Emotion- Lens: Interactive Visual Exploration of the Circumplex Emotion Space in Literary Works via Affective Word Clouds,”Visual Informatics, vol. 9, no. 1, pp. 84–98, 2025
2025
-
[56]
This Looks Like That: Deep Learning for Interpretable Image Recognition,
C. Chen, O. Li, D. Tao, A. Barnett, C. Rudin, and J. K. Su, “This Looks Like That: Deep Learning for Interpretable Image Recognition,” Advances in Neural Information Processing Systems, vol. 32, 2019
2019
-
[57]
Interpretable Image Recognition with Hierarchical Prototypes,
P. Hase, C. Chen, O. Li, and C. Rudin, “Interpretable Image Recognition with Hierarchical Prototypes,” inProceedings of the AAAI Conference on Human Computation and Crowdsourcing, vol. 7, 2019, pp. 32–40
2019
-
[58]
Neural Prototype Trees for Interpretable Fine-Grained Image Recognition,
M. Nauta, R. Van Bree, and C. Seifert, “Neural Prototype Trees for Interpretable Fine-Grained Image Recognition,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 14 933–14 943
2021
-
[59]
ProtoDiffusion: Classifier-Free Diffusion Guidance with Prototype Learning,
G. Baykal, H. F. Karagoz, T. Binhuraib, and G. Unal, “ProtoDiffusion: Classifier-Free Diffusion Guidance with Prototype Learning,” inAsian Conference on Machine Learning. PMLR, 2024, pp. 106–120
2024
-
[60]
StyleCLIPDraw: Coupling Content and Style in Text-to-Drawing Translation,
P. Schaldenbrand, Z. Liu, and J. Oh, “StyleCLIPDraw: Coupling Content and Style in Text-to-Drawing Translation,” inProceedings of the Thirty- First International Joint Conference on Artificial Intelligence (IJCAI- 22), 2022, pp. 4966–4972
2022
-
[61]
Semantic space theory: A computational approach to emotion,
A. S. Cowen and D. Keltner, “Semantic space theory: A computational approach to emotion,”Trends in Cognitive Sciences, vol. 25, no. 2, pp. 124–136, 2021
2021
-
[62]
Multimodal physiological analysis of impact of emotion on cognitive control in vr,
M. Li, J. Pan, Y . Li, Y . Gao, H. Qin, and Y . Shen, “Multimodal physiological analysis of impact of emotion on cognitive control in vr,” IEEE Transactions on Visualization and Computer Graphics, vol. 30, no. 5, pp. 2044–2054, 2024
2044
-
[63]
Adding Conditional Control to Text-to-Image Diffusion Models,
L. Zhang, A. Rao, and M. Agrawala, “Adding Conditional Control to Text-to-Image Diffusion Models,” inProceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 3836–3847
2023
-
[64]
Learning Transferable Visual Models from Natural Language Supervision,
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clarket al., “Learning Transferable Visual Models from Natural Language Supervision,” inInternational Conference on Machine Learning. PMLR, 2021, pp. 8748–8763
2021
-
[65]
Gaussian Error Linear Units (GELUs),
D. Hendrycks and K. Gimpel, “Gaussian Error Linear Units (GELUs),” arXiv preprint arXiv:1606.08415, 2016
Pith/arXiv arXiv 2016
-
[66]
3D Gaussian Splatting for Real-Time Radiance Field Rendering,
B. Kerbl, G. Kopanas, T. Leimk ¨uhler, and G. Drettakis, “3D Gaussian Splatting for Real-Time Radiance Field Rendering,”ACM Trans. Graph., vol. 42, no. 4, pp. 139–1, 2023
2023
-
[67]
A Simple Frame- work for Contrastive Learning of Visual Representations,
T. Chen, S. Kornblith, M. Norouzi, and G. Hinton, “A Simple Frame- work for Contrastive Learning of Visual Representations,” inInterna- tional Conference on Machine Learning. PMLR, 2020, pp. 1597–1607
2020
-
[68]
MLP- Mixer: An All-MLP Architecture for Vision,
I. O. Tolstikhin, N. Houlsby, A. Kolesnikov, L. Beyer, X. Zhai, T. Un- terthiner, J. Yung, A. Steiner, D. Keysers, J. Uszkoreitet al., “MLP- Mixer: An All-MLP Architecture for Vision,”Advances in Neural Information Processing Systems, vol. 34, pp. 24 261–24 272, 2021
2021
-
[69]
Diffusion Models Beat GANs on Image Synthesis,
P. Dhariwal and A. Nichol, “Diffusion Models Beat GANs on Image Synthesis,”Advances in Neural Information Processing Systems, vol. 34, pp. 8780–8794, 2021
2021
-
[70]
A. Hurst, A. Lerer, A. P. Goucher, A. Perelman, A. Ramesh, A. Clark, A. Ostrow, A. Welihinda, A. Hayes, A. Radfordet al., “GPT-4o System Card,”arXiv preprint arXiv:2410.21276, 2024
Pith/arXiv arXiv 2024
-
[71]
SDXL Panorama,
J. Bilcke, “SDXL Panorama,” https://huggingface.co/jbilcke-hf/ sdxl-panorama, 2024, hugging Face Model Hub, Accessed: 2025-08- 22
2024
-
[72]
Latent Consistency Mod- els: Synthesizing High-Resolution Images with Few-Step Inference,
S. Luo, Y . Tan, L. Huang, J. Li, and H. Zhao, “Latent Consistency Mod- els: Synthesizing High-Resolution Images with Few-Step Inference,” arXiv preprint arXiv:2310.04378, 2023
Pith/arXiv arXiv 2023
-
[73]
Outpainting with Stable Diffusion,
Hugging Face, “Outpainting with Stable Diffusion,” https://huggingface. co/docs/diffusers/advanced inference/outpaint, 2024, accessed: 2025- 08-22
2024
-
[74]
SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis,
D. Podell, Z. English, K. Lacey, A. Blattmann, T. Dockhorn, J. M ¨uller, J. Penna, and R. Rombach, “SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis,”arXiv preprint arXiv:2307.01952, 2023
Pith/arXiv arXiv 2023
-
[75]
EmoSet: A Large-Scale Visual Emotion Dataset with Rich Attributes,
J. Yang, Q. Huang, T. Ding, D. Lischinski, D. Cohen-Or, and H. Huang, “EmoSet: A Large-Scale Visual Emotion Dataset with Rich Attributes,” inProceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 20 383–20 394
2023
-
[76]
BLIP-2: Bootstrapping Language- Image Pre-Training with Frozen Image Encoders and Large Language Models,
J. Li, D. Li, S. Savarese, and S. Hoi, “BLIP-2: Bootstrapping Language- Image Pre-Training with Frozen Image Encoders and Large Language Models,” inInternational Conference on Machine Learning. PMLR, 2023, pp. 19 730–19 742
2023
-
[77]
LAION-5B: An Open Large-Scale Dataset for Training Next Gener- ation Image-Text Models,
C. Schuhmann, R. Beaumont, R. Vencu, C. Gordon, R. Wightman, M. Cherti, T. Coombes, A. Katta, C. Mullis, M. Wortsmanet al., “LAION-5B: An Open Large-Scale Dataset for Training Next Gener- ation Image-Text Models,”Advances in Neural Information Processing Systems, vol. 35, pp. 25 278–25 294, 2022
2022
-
[78]
Meta Quest 2,
Meta Platforms, Inc., “Meta Quest 2,” 2020, accessed: 2025-08-22. [Online]. Available: https://www.meta.com/quest/products/quest-2/
2020
-
[79]
Unity 2022.3.6f1 Release Notes,
Unity Technologies, “Unity 2022.3.6f1 Release Notes,” 2022, accessed: 2025-08-22. [Online]. Available: https://unity.com/releases/ editor/whats-new/2022.3.6f1
2022
-
[80]
Affective Interac- tions Using Virtual Reality: the Link between Presence and Emotions,
G. Riva, F. Mantovani, C. S. Capideville, A. Preziosa, F. Morganti, D. Villani, A. Gaggioli, C. Botella, and M. Alca ˜niz, “Affective Interac- tions Using Virtual Reality: the Link between Presence and Emotions,” Cyberpsychology & behavior, vol. 10, no. 1, pp. 45–56, 2007
2007
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.