Pith. sign in

REVIEW 4 major objections 4 minor 1 cited by

Fine-grained emotion control in AI image generation is achievable with learnable prototypes

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-03 00:01 UTC pith:POYBWXJZ

load-bearing objection A clever, novel emotion-prototype generation system whose headline numbers are partly measuring its own objective; worth engaging, but the evaluation needs independent emotion assessment before the claims will convince me. the 4 major comments →

arxiv 2602.11658 v2 pith:POYBWXJZ submitted 2026-02-12 cs.CV

EmoSpace: Immersive Affective Image Generation Guided by Fine-Grained Emotion Prototypes

classification cs.CV
keywords emotion generationdiffusion modelsprototype learningaffective computingvirtual realityimage generationtext-to-imageemotional control
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper argues that existing categorical or dimensional emotion models are too coarse to express the nuances needed for immersive VR content. To fix this, it introduces EmoSpace, which learns a dynamic bank of emotion prototypes in a vision-language embedding space, letting users control generation with free-form emotion descriptions. The framework injects emotion into a diffusion model through multi-prototype guidance, temporal blending, prompt refinement, and attention reweighting, and extends to panoramic, stylized, and outpainting tasks. The authors claim EmoSpace outperforms existing emotion-aware generators in quantitative and human evaluations, and that VR viewing changes emotional perception patterns. If true, it would give content creators a flexible, label-free way to steer the emotional tone of generated imagery.

Core claim

The central claim is that emotions in visual content are better represented as a set of learnable prototypes in a shared vision-language embedding space than as fixed categories or points on a valence-arousal plane. EmoSpace learns a bank of such prototypes from an emotional image dataset through a composite loss, letting them merge and split during training to adaptively cover the emotion distribution. At generation time, a user's free-text emotion description is mapped to a weighted combination of the most similar prototypes, which then guide a diffusion model via attention reweighting, iterative prompt refinement that adds sub-emotion descriptors, and temporal blending that shifts from co

What carries the argument

The prototype bank with merge-split dynamics, together with the mapper that transfers vision-language embeddings into the diffusion model's prompt space. The hierarchical emotion representation (basic categories plus a large set of dynamic prototypes) is the central object; it creates a 'structured-yet-continuous' emotion space that is interpretable and controllable. Multi-prototype guidance computes a weighted combination of similar prototypes for conditioning; attention reweighting injects this into the model's cross-attention layers; temporal blending schedules its influence during denoising; and iterative prompt refinement uses a language model to enrich the prompt with sub-emotion words

Load-bearing premise

The paper relies on the assumption that vision-language embedding similarity between a generated image and an emotion-description prompt faithfully measures fine-grained emotional alignment; the prototypes, mapper, and prompt refinement are all optimized to maximize this specific similarity, so if it doesn't track human perception, the quantitative gains become partly self-referential.

What would settle it

A large, pre-registered human study in which independent raters judge whether images generated by EmoSpace versus baseline methods match a set of fine-grained emotion descriptions. If humans show no significant preference for EmoSpace, or if the embedding-similarity scores used in the paper fail to predict human judgments across models, the central claim would be refuted. Alternatively, an ablation that freezes the prototype bank to a fixed set of categorical labels while keeping all other components would reveal whether the dynamic prototypes actually drive the claimed advantage.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Content creators can steer image generation with free-form emotion descriptions rather than fixed labels, making emotional control more intuitive and expressive.
  • The same emotion-conditioning mechanism works across standard text-to-image, panorama, style transfer, and outpainting, suggesting a unified approach to affective content generation.
  • If the reported gains hold, prototype-based emotion spaces could replace categorical and dimensional models for downstream generation tasks.
  • The VR user study suggests that immersive presentation itself shifts emotional perception, so emotion-aware generation for VR may need to account for viewing medium.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The dynamic prototype bank could generalize to other abstract attributes (e.g., cultural tone, aesthetic mood) beyond emotion; the same merge-split mechanism might learn arbitrary semantic spaces from paired image-text data.
  • Replacing the vision-language-similarity evaluation with forced-choice human judgments would test whether the quantitative lead reflects genuine emotional fidelity rather than optimization of the same metric used for conditioning.
  • The VR finding that emotion category selection shifts toward positive emotions suggests that content tuned for desktop may not read identically in VR; a practical extension would be to condition generation on display mode.
  • Because prototypes are interpretable embeddings, they could be probed to reveal which emotional concepts are being activated, enabling cross-cultural adaptation of emotion control.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper proposes EmoSpace, a framework for fine-grained emotion-controlled image generation. It learns a bank of dynamic CLIP-space emotion prototypes via vision-language alignment, then steers a diffusion model using multi-prototype guidance, temporal blending, attention reweighting, and GPT-assisted iterative prompt refinement. The framework is extended to outpainting, stylized generation, and panorama generation. The paper reports quantitative comparisons with EmoGen, EmotiCrafter, and SDXL (Tables II-III), ablations (Table V), a small human ranking study (Sec. V-D), and a VR vs. desktop user study (Sec. VI) with the claim that EmoSpace improves emotional alignment and aesthetic quality while immersive presentation changes emotion perception.

Significance. Fine-grained emotionally controllable generative models for VR are a useful and timely topic, and the paper is ambitious in scope. The strengths are the prototype-based emotion representation, the integration of multiple control pathways, the breadth of downstream tasks, and the inclusion of a VR user study. If the central claim is supported by independent evidence, EmoSpace would be a practical contribution to affective computing and immersive content generation. However, the current evidence is not yet convincing: the automated emotion metrics are computed in the same CLIP space the method explicitly optimizes, baselines differ in base model, and the human evaluation is small and unreleased. The result therefore has moderate-to-high potential but needs additional validation.

major comments (4)
  1. [Sec. V-A / Sec. IV-B / Sec. IV-C2] Table II's headline emotion-accuracy gains are computed with CLIP cosine similarity in CLIP-H-14 space (Emotion-Category and Emotion-Fine-Grained, Sec. V-A). This is the same embedding space where the prototype bank is learned (Sec. IV-B1: d_v=d_t=1024), where multi-prototype guidance retrieves prototypes by cosine similarity, and where iterative prompt refinement explicitly selects candidates maximizing <phi_CLIP(c), p_emo> (Sec. IV-C2). The MLP mapper is also trained to maximize cosine similarity. Thus the reported improvements are partly self-referential: they measure how well the system optimizes its training/selection objective. The only independent evidence is the ranking study in Sec. V-D (28 users, 104 images), but the images and protocol are not released, and the VR user study in Sec. VI is not a head-to-head method comparison. Please add a human emotion-alignment evaluation on
  2. [Tables II and III] The quantitative comparison is confounded by base-model differences. EmoGen uses SD 1.5, while EmoSpace, EmotiCrafter, and SDXL use SDXL (Sec. V-A). The very large CLIP-Prompt gap (17.49 for EmoGen vs. 31.99 for SDXL) and aesthetic gap (6.13 vs. 16.25) are likely inherited from the base model rather than reflecting emotion-control ability. Table III equalizes the prompts but not the base models, so SD 1.5 baselines remain disadvantaged. Please either adapt EmoGen to SDXL (if feasible), provide an SD 1.5 version of EmoSpace, or include a same-base SDXL baseline for categorical/dimensional emotion conditioning. Without this, Table II's cross-method rankings for emotion and aesthetics are not interpretable.
  3. [Sec. IV-B, IV-C and supplementary material] Several mechanisms central to the method are only described by reference to a supplementary file that is not included. Specifically, the merge/split criteria and update rules for the dynamic prototype bank, the exact definitions of L_contrast, L_diversity, and L_dist, the three-phase temporal blending schedule, and the GPT-based iterative prompt refinement details (number of iterations, threshold, prompt template, temperature) are all deferred. In addition, the abstract states 256 prototypes while Sec. IV-B1 sets K=1024. These omissions prevent full assessment and replication of the central representation and steering contributions. Please include the supplementary material or move the detailed formulations into the main text, and resolve the prototype-count inconsistency.
  4. [Sec. VI-B] The VR user study uses independent-samples t-tests on a within-subject design: each participant experienced both VR and desktop conditions, so the observations are paired. The reported values (T1 t=-1.742, p=0.082; T2 t=-0.495, p=0.621; T3 t=-1.161, p=0.246) are also inconsistent with the means (VR > desktop in all cases), suggesting the sign convention or the test may be erroneous. Please use paired-samples tests and report effect sizes. This does not affect the text-to-image comparison in Sec. V, but it is load-bearing for the VR-perception claims (H1 and H2) in Sec. VI.
minor comments (4)
  1. [Sec. V-A] The CLIP-Prompt score for EmoSpace is computed against the original prompt, while the generation actually uses the refined prompt. The paper notes this, but it should also report alignment to the final refined prompt so that the text-image alignment measure is well defined.
  2. [Sec. V-B] The statement that EmoSpace is 'the first framework to support diverse emotional generation tasks' should be qualified. Prior systems support outpainting, style LoRAs, and panorama generation; the novelty is the unified emotion-conditioning mechanism, not the task support per se.
  3. [Fig. 8] The panels in Fig. 8 lack clear axis labels and the caption repeats numeric results from the text. Please make the figure self-contained (e.g., label 'Mean Rating' and 'Selection Distribution' directly on the axes) and avoid duplicating statistics in the caption.
  4. [Sec. VI-C] The 'Limitations and Future Work' subsection addresses only hardware comfort and user-requested features. There is no limitations paragraph for the proposed method or the evaluation (e.g., dependency on CLIP similarity, small human study, lack of release). Adding one would improve scientific transparency.

Circularity Check

1 steps flagged

Head-table emotion-accuracy gains rest on the same CLIP cosine similarity EmoSpace is explicitly trained/selected to maximize; independent human studies provide partial but not complete mitigation.

specific steps
  1. fitted input called prediction [Sec. V-A (Evaluation Metrics); Sec. IV-B2 (Training Objectives); Sec. IV-C2 (Iterative Prompt Refinement)]
    ""Emotion Accuracy: We employ CLIP-based metrics to evaluate emotion alignment at both categorical (Emotion-Category) and fine-grained (Emotion-Fine-Grained) levels using cosine similarity in CLIP feature space (higher is better)." ... "the mapper is trained using a composite loss that combines cosine similarity maximization and Euclidean distance minimization" ... "then selects the dominant emotion c* that maximizes CLIP-space similarity ⟨φ_CLIP(c_i), p_emo⟩""

    The headline emotion-accuracy metrics are CLIP cosine similarities, while the MLP mapper is explicitly trained with cosine-similarity maximization, the iterative prompt refinement selects candidates by CLIP-space similarity, and the prototypes/embeddings live in CLIP-H-14 space (Sec. IV-B1). Thus EmoSpace is directly optimized for the same quantity used as the independent evaluation metric. The Emo-Cat and Emo-FG gains in Tables II and IV are at least partly forced by construction rather than being an external test of human-perceived emotion. The human ranking study in Sec. V-D gives some independent support, but the main quantitative head-to-head superiority claim rests on this self-referential metric.

full rationale

The paper is largely a systems/application paper: prototype learning, multi-prototype guidance, attention reweighting, and VR-specific extensions are evaluated with ablations, qualitative comparisons, and user studies. The central quantitative claim of 'superior performance in emotion accuracy' is where circularity enters. Sec. V-A defines Emotion-Category and Emotion-Fine-Grained as CLIP cosine similarity between generated images and emotion text. The same cosine similarity appears as the training objective for the MLP mapper (Sec. IV-B2) and as the selection criterion in iterative prompt refinement (Sec. IV-C2), with all prototypes and embeddings in CLIP-H-14 space. Consequently, the reported metric gains are partly self-referential: the method explicitly maximizes the measure used to declare it the winner. This is not a fatal flaw because the paper also reports an independent 28-participant ranking study (Sec. V-D) and a VR pairwise preference test (Sec. VI-B) that favor EmoSpace; those are genuine, if small and unreleased, human evaluations. No load-bearing self-citation chain was found: the one overlapping-author reference ([55]) is not used to justify the central derivation, and there is no imported uniqueness theorem or ansatz-by-citation. The score of 6 reflects partial circularity in the quantitative emotion-accuracy evidence, not in the whole framework. The small/unreleased human studies are a validity concern but not themselves a circularity step.

Axiom & Free-Parameter Ledger

8 free parameters · 6 axioms · 1 invented entities

The framework depends on CLIP-space assumptions, hand-set hyperparameters (K, thresholds, loss weights, attention intensity), and an unverified mapping from CLIP to SDXL. The prototype bank is an internal learned entity without independent external validation.

free parameters (8)
  • Number of emotion prototypes K = 1024 (abstract says 256)
    Hand-chosen prototype count; affects granularity and compute (Sec. IV-B).
  • Merge threshold tau_m = 0.3
    Hand-chosen threshold for prototype merging (Sec. IV-B1).
  • Split threshold tau_s = 0.7
    Hand-chosen threshold relative to average usage (Sec. IV-B1).
  • Loss weights alpha_L, beta_L, gamma_L, delta_L = 1.0, 0.5, 0.1, 0.1
    Hand-set weights for classification, contrastive, diversity, and distance losses (Eq. 5).
  • Attention injection intensity alpha_attn = 1.5
    Controls strength of emotional attention reweighting (Eq. 7).
  • Softmax temperature tau_temp = 0.1
    Temperature for prototype weighting in multi-prototype guidance (Eq. 6).
  • Top-k for positive/negative prototypes = not stated
    Number of prototypes combined for guidance is not specified (Sec. IV-C1).
  • Three-phase temporal blending schedule = not specified
    Phase boundaries and blending function are deferred to supplementary (Sec. IV-C1).
axioms (6)
  • domain assumption CLIP-H-14 embedding space is a sufficient representation for fine-grained emotion semantics and for evaluating emotion alignment.
    Used in Sec. IV-B for prototypes/fusion and Sec. V-A for Emotion-Cat/Emotion-FG metrics; if false, the whole learning and evaluation loop is compromised.
  • domain assumption EmoSet-118K's categorical labels plus BLIP-2 captions provide enough emotional supervision to learn generalizable prototypes.
    Training relies solely on this dataset and generated captions (Sec. V-A); no external emotion-recognition validation is provided.
  • domain assumption A lightweight MLP can map CLIP emotion vectors into SDXL text-embedding space without losing emotional content.
    Sec. IV-B2 uses an MLP mapper; no analysis of mapping distortion or failure cases.
  • domain assumption Diffusion denoising exhibits a three-phase hierarchical structure that justifies the temporal blending schedule.
    Sec. IV-C1, citing Dhariwal & Nichol [69]; schedule details are in missing supplementary.
  • domain assumption GPT-4o-mini refinement increases rather than distorts emotion-prompt correspondence.
    Iterative prompt refinement (Sec. IV-C2) uses GPT-4o-mini; no separate human evaluation of the refined prompts alone.
  • domain assumption Prototype learning from object classification transfers to abstract emotion semantics.
    Sec. II-B/III-B motivation; not demonstrated independently of the reported in-system results.
invented entities (1)
  • Dynamic emotion prototype bank with merge/split no independent evidence
    purpose: Provides a continuous, interpretable emotional conditioning signal for diffusion generation
    No external benchmark, physiology measure, or released model confirms that the prototypes correspond to human-perceived emotion clusters; only internal CLIP metrics and internal user tests are provided.

pith-pipeline@v1.3.0-alltime-deepseek · 20542 in / 14084 out tokens · 118353 ms · 2026-08-03T00:01:00.554973+00:00 · methodology

0 comments
read the original abstract

Immersive affective content generation aims to create visually compelling VR imagery with controllable emotional nuance, yet existing methods typically rely on coarse labels or prompt-only control. Although modern diffusion transformers (DiTs) such as FLUX improve visual fidelity, they are not designed to incorporate structured affective representations. We present EmoSpace, an immersive affective content generation framework guided by fine-grained emotion prototypes, transforming free-form text and emotion descriptions into fine-grained affective imagery through three coordinated components. First, to represent sub-emotion variation beyond conventional categorical models, EmoSpace learns a hierarchical bank of 256 prototypes with input-conditioned adaptation through vision-language alignment. Second, Prototype-Conditioned Steering converts these prototypes into DiT-compatible generation signals through multi-pathway injection and temporal blending, while Iterative Prompt Refinement enriches prompts with prototype-aligned sub-emotion descriptors. Third, Affect-Grounded Modulation coordinates emotion conditioning with controllable LoRAs for panoramic, stylized, and multi-conditional generation. Through quantitative and qualitative evaluations, EmoSpace improves fine-grained emotional alignment while maintaining high aesthetic quality. Our user study shows that EmoSpace outputs are perceived as more emotionally aligned than baseline results and more suitable for immersive scene design. Additionally, we find that immersive presentation alters emotional perception and increases emotional engagement. Together, these findings inform the design of emotion-aware generative systems for immersive media, with potential applications including education, immersive storytelling, and artistic creation. We will release our code and models to facilitate future research along this line.

Figures

Figures reproduced from arXiv: 2602.11658 by Bingyuan Wang, Lin-Ping Yuan, Xingbei Chen, Zeyu Wang, Zongyang Qiu.

Figure 1
Figure 1. Figure 1: Example Results Generated by EmoSpace. We demonstrate EmoSpace’s capability for immersive affective content generation in (a) emotional panorama generation from the similar prompt “an emotional panorama” and different fine-grained emotional descriptions (styles for each row: Ghibli, 3D render, toy, ink painting, pixel art), (b) emotional image outpainting from an original image and different emotions in di… view at source ↗
Figure 2
Figure 2. Figure 2: Overview of EmoSpace. Our framework consists of three main components: (a) emotion prototype learning that learns dynamic, interpretable emotion representations through vision-language alignment with rich learnable prototypes, (b) emotion-conditioned generation featuring multi-prototype guidance, iterative prompt refinement ( [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Iterative Prompt Refinement in the Latent Space. We iteratively optimize prompts through GPT association, generating candidate prompts and emotional descriptions, evaluating semantic alignment with target embedding, and selecting optimal solutions until convergence to achieve precise emotion￾prompt correspondence. top-k similar prototypes. Let P + = {i : si ∈ top-k({sj} K j=1)} where si = ⟨eclip, pi⟩. The … view at source ↗
Figure 4
Figure 4. Figure 4: Example Results Generated from Prototypes. We demonstrate EmoSpace’s capability of fine-grained emotion modeling and control from the similar prompt “an emotional face in Studio Ghibli style” (no iterative refinement) and four randomly selected prototypes. Red arrows indicate consistent emotional features of each prototype. to support emotional image outpainting, stylized emotional generation, and emotiona… view at source ↗
Figure 5
Figure 5. Figure 5: Comparative Study Results. Qualitative comparison between different methods for emotional image generation. Among detailed content prompts and fine-grained emotional descriptions, our method demonstrates superior emotion fidelity and visual quality across diverse content and emotional inputs. C. Ablation Study Quantitative Results. We conducted an ablation study on the three main modules of our emotion-con… view at source ↗
Figure 6
Figure 6. Figure 6: Ablation Study. We analyzed the three most important modules of EmoSpace, including multi-prototype guidance, attention reweighting, and iterative prompt refinement. D. Human Evaluation We conducted a user experiment to evaluate the effective￾ness of our method by comparing it to three baseline methods (EmoGen, EmotiCrafter, and SDXL). We generated 104 test images with varying emotional nuances across diff… view at source ↗
Figure 7
Figure 7. Figure 7: Experiment Setting. The participants were equipped with Meta Quest 2 and assessed the panoramas generated by EmoSpace. An example experiment setting is displayed in [PITH_FULL_IMAGE:figures/full_fig_p010_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Quantitative Analysis of Study Results. (a) Human evaluation of EmoSpace against other methods in text-image alignment and emotion accuracy (both basic categories and fine-grained descriptions). (b, c) Differences in emotional category and intensity level perception between desktop and VR. (d, e) Participants’ task performance and subjective perception in desktop and VR. (f) Post-study questionnaire result… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. EmoScene: A Dual-space Dataset for Controllable Affective Image Generation

    cs.CV 2026-04 reject novelty 6.0

    EmoScene contributes 1.2M images annotated with discrete emotions, continuous VAD scores, perceptual attributes, and captions, plus a cross-attention modulation that shifts generated images toward requested affective targets.

Reference graph

Works this paper leans on

80 extracted references · 11 linked inside Pith · cited by 1 Pith paper

  1. [1]

    Feeling Virtually Present Makes Me Happier: The Influence of Immersion, Sense of Presence, and Video Contents on Positive Emotion Induction,

    K. Pavic, L. Chaby, T. Gricourt, and D. Vergilino-Perez, “Feeling Virtually Present Makes Me Happier: The Influence of Immersion, Sense of Presence, and Video Contents on Positive Emotion Induction,” Cyberpsychology, Behavior, and Social Networking, vol. 26, no. 4, pp. 238–245, 2023

  2. [2]

    Virtual Reality for Emotion Elicitation–a Review,

    R. Somarathna, T. Bednarz, and G. Mohammadi, “Virtual Reality for Emotion Elicitation–a Review,”IEEE Transactions on Affective Com- puting, vol. 14, no. 4, pp. 2626–2645, 2022

  3. [3]

    A Survey on Affective and Cognitive VR,

    T. Luong, A. Lecuyer, N. Martin, and F. Argelaguet, “A Survey on Affective and Cognitive VR,”IEEE Transactions on Visualization and Computer Graphics, vol. 28, no. 12, pp. 5154–5171, 2021

  4. [4]

    Emotional qualities of VR space,

    A. Naz, R. Kopper, R. P. McMahan, and M. Nadin, “Emotional qualities of VR space,” in2017 IEEE virtual reality (VR). IEEE, 2017, pp. 3–11

  5. [5]

    Traces in Virtual Environments: A Framework and Exploration to Conceptualize the Design of Social Vir- tual Environments,

    L. Hirsch, C. George, and A. Butz, “Traces in Virtual Environments: A Framework and Exploration to Conceptualize the Design of Social Vir- tual Environments,”IEEE Transactions on Visualization and Computer Graphics, vol. 28, no. 11, pp. 3874–3884, 2022

  6. [6]

    Affective com- puting in virtual reality: emotion recognition from brain and heartbeat dynamics using wearable sensors,

    J. Mar ´ın-Morales, J. L. Higuera-Trujillo, A. Greco, J. Guixeres, C. Llinares, E. P. Scilingo, M. Alca ˜niz, and G. Valenza, “Affective com- puting in virtual reality: emotion recognition from brain and heartbeat dynamics using wearable sensors,”Scientific reports, vol. 8, no. 1, p. 13657, 2018

  7. [7]

    Revolutionizing Virtual Reality With Generative AI: An In-Depth Review,

    M. E. A. Hashim, W. A. W. Mustafa, N. S. Prameswari, M. M. Ghani, and H. F. Hanafi, “Revolutionizing Virtual Reality With Generative AI: An In-Depth Review,”Journal of Advanced Research in Computing and Applications, vol. 30, no. 1, pp. 19–30, 2023. 13

  8. [8]

    Generative AI Meets Virtual Reality: A Comprehensive Survey on Applications, Challenges, and Future Direction,

    F. Rahimi, A. Sadeghi-Niaraki, and S.-M. Choi, “Generative AI Meets Virtual Reality: A Comprehensive Survey on Applications, Challenges, and Future Direction,”IEEE Access, 2025

  9. [9]

    Effects of Virtual Environment Platforms on Emotional Responses,

    K. Kim, M. Z. Rosenthal, D. J. Zielinski, and R. Brady, “Effects of Virtual Environment Platforms on Emotional Responses,”Computer Methods and Programs in Biomedicine, vol. 113, no. 3, pp. 882–893, 2014

  10. [10]

    The Effect of Immersion on Emotional Responses to Film Viewing in a Virtual Environment,

    A. Kim, M. Chang, Y . Choi, S. Jeon, and K. Lee, “The Effect of Immersion on Emotional Responses to Film Viewing in a Virtual Environment,” in2018 IEEE Conference on Virtual Reality and 3D User Interfaces (VR). IEEE, 2018, pp. 601–602

  11. [11]

    Effects of Immersion in a Simulated Natural Environment on Stress Reduction and Emotional Arousal: A Systematic Review and Meta-Analysis,

    H. Li, Y . Ding, B. Zhao, Y . Xu, and W. Wei, “Effects of Immersion in a Simulated Natural Environment on Stress Reduction and Emotional Arousal: A Systematic Review and Meta-Analysis,”Frontiers in psy- chology, vol. 13, p. 1058177, 2023

  12. [12]

    Basic emotions,

    P. Ekman, T. Dalgleish, and M. Power, “Basic emotions,”San Francisco, USA, 1999

  13. [13]

    A Circumplex Model of Affect,

    J. A. Russell, “A Circumplex Model of Affect,”Journal of personality and social psychology, vol. 39, no. 6, p. 1161, 1980

  14. [14]

    Emogen: Emotional Image Content Generation with Text-to-Image Diffusion Models,

    J. Yang, J. Feng, and H. Huang, “Emogen: Emotional Image Content Generation with Text-to-Image Diffusion Models,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 6358–6368

  15. [15]

    EmotiCrafter: Text-to-Emotional-Image Generation Based on Valence-Arousal Model,

    S. Dang, Y . He, L. Ling, Z. Qian, N. Zhao, and N. Cao, “EmotiCrafter: Text-to-Emotional-Image Generation Based on Valence-Arousal Model,” arXiv preprint arXiv:2501.05710, 2025

  16. [16]

    A systematic review on affective computing: Emotion models, databases, and recent advances,

    Y . Wang, W. Song, W. Tao, A. Liotta, D. Yang, X. Li, S. Gao, Y . Sun, W. Ge, W. Zhanget al., “A systematic review on affective computing: Emotion models, databases, and recent advances,”Information Fusion, vol. 83, pp. 19–52, 2022

  17. [17]

    Emotional Category Data on Images from the International Affective Picture System,

    J. A. Mikels, B. L. Fredrickson, G. R. Larkin, C. M. Lindberg, S. J. Maglio, and P. A. Reuter-Lorenz, “Emotional Category Data on Images from the International Affective Picture System,”Behavior research methods, vol. 37, no. 4, pp. 626–630, 2005

  18. [18]

    Pleasure-Arousal-Dominance: A General Framework for Describing and Measuring Individual Differences in Temperament,

    A. Mehrabian, “Pleasure-Arousal-Dominance: A General Framework for Describing and Measuring Individual Differences in Temperament,” Current psychology, vol. 14, no. 4, pp. 261–292, 1996

  19. [19]

    Robust Image Sentiment Analysis Using Progressively Trained and Domain Transferred Deep Networks,

    Q. You, J. Luo, H. Jin, and J. Yang, “Robust Image Sentiment Analysis Using Progressively Trained and Domain Transferred Deep Networks,” inProceedings of the AAAI Conference on Artificial Intelligence, vol. 29, no. 1, 2015

  20. [20]

    MA VEN: Multi-Modal Attention for Valence-Arousal Emotion Network,

    V . Ahire, K. Shah, M. Khan, N. Pakhale, L. Sookha, M. Ganaie, and A. Dhall, “MA VEN: Multi-Modal Attention for Valence-Arousal Emotion Network,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2025, pp. 5789–5799

  21. [21]

    EmoVit: Revolutionizing Emotion Insights with Visual Instruction Tuning,

    H. Xie, C.-J. Peng, Y .-W. Tseng, H.-J. Chen, C.-F. Hsu, H.-H. Shuai, and W.-H. Cheng, “EmoVit: Revolutionizing Emotion Insights with Visual Instruction Tuning,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 26 596–26 605

  22. [22]

    Emotion-LLaMA: Multimodal Emotion Recognition and Reasoning with Instruction Tuning,

    Z. Cheng, Z.-Q. Cheng, J.-Y . He, K. Wang, Y . Lin, Z. Lian, X. Peng, and A. Hauptmann, “Emotion-LLaMA: Multimodal Emotion Recognition and Reasoning with Instruction Tuning,”Advances in Neural Informa- tion Processing Systems, vol. 37, pp. 110 805–110 853, 2024

  23. [23]

    GPT-4V with Emotion: A Zero-Shot Benchmark for Generalized Emotion Recognition,

    Z. Lian, L. Sun, H. Sun, K. Chen, Z. Wen, H. Gu, B. Liu, and J. Tao, “GPT-4V with Emotion: A Zero-Shot Benchmark for Generalized Emotion Recognition,”Information Fusion, vol. 108, p. 102367, 2024

  24. [24]

    MEMO-Bench: A Multiple Benchmark for Text-to-Image and Multimodal Large Language Models on Human Emotion Analysis,

    Y . Zhou, Z. Zhang, J. Cao, J. Jia, Y . Jiang, F. Wen, X. Liu, X. Min, and G. Zhai, “MEMO-Bench: A Multiple Benchmark for Text-to-Image and Multimodal Large Language Models on Human Emotion Analysis,” arXiv preprint arXiv:2411.11235, 2024

  25. [25]

    Make Me Happier: Evoking Emotions Through Image Diffusion Models,

    Q. Lin, J. Zhang, Y . S. Ong, and M. Zhang, “Make Me Happier: Evoking Emotions Through Image Diffusion Models,”arXiv preprint arXiv:2403.08255, 2024

  26. [26]

    Sal-Guide Diffusion: Saliency Maps Guide Emotional Image Generation through Adapter,

    X. Lin, S. Zhong, Y . Liu, and G. Chen, “Sal-Guide Diffusion: Saliency Maps Guide Emotional Image Generation through Adapter,” in2024 IEEE International Conference on Multimedia and Expo (ICME). IEEE, 2024, pp. 1–6

  27. [27]

    TGCA-PVT: Topic-Guided Context-Aware Pyramid Vision Transformer for Sticker Emotion Recognition,

    J. Chen, W. Wang, Y . Hu, J. Chen, H. Liu, and X. Hu, “TGCA-PVT: Topic-Guided Context-Aware Pyramid Vision Transformer for Sticker Emotion Recognition,” inProceedings of the 32nd ACM International Conference on Multimedia, 2024, pp. 9709–9718

  28. [28]

    Knowledge-Aligned Counterfactual-Enhancement Diffusion Perception for Unsupervised Cross-Domain Visual Emotion Recognition,

    W. Yin, Y . Wang, G. Duan, D. Zhang, X. Hu, Y .-F. Li, and T. He, “Knowledge-Aligned Counterfactual-Enhancement Diffusion Perception for Unsupervised Cross-Domain Visual Emotion Recognition,” inPro- ceedings of the Computer Vision and Pattern Recognition Conference, 2025, pp. 3888–3898

  29. [29]

    EmoEdit: Evoking Emotions Through Image Manipulation,

    J. Yang, J. Feng, W. Luo, D. Lischinski, D. Cohen-Or, and H. Huang, “EmoEdit: Evoking Emotions Through Image Manipulation,” inPro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2025, pp. 24 690–24 699

  30. [30]

    Image Re-Emotionalizing,

    M. Xu, B. Ni, J. Tang, and S. Yan, “Image Re-Emotionalizing,” inThe Era of Interactive Media. Springer, 2012, pp. 3–14

  31. [31]

    Affective Image Editing: Shaping Emotional Factors via Text Descriptions,

    P. Zhang, S. Weng, C. Zhu, B. Tang, Z. Jia, S. Li, and B. Shi, “Affective Image Editing: Shaping Emotional Factors via Text Descriptions,”arXiv preprint arXiv:2505.18699, 2025

  32. [32]

    Bridge Then Begin Anew: Generating Target-Relevant Intermediate Model for Source-Free Visual Emotion Adaptation,

    J. Zhu, S. Zhao, J. Jiang, W. Tang, Z. Xu, T. Han, P. Xu, and H. Yao, “Bridge Then Begin Anew: Generating Target-Relevant Intermediate Model for Source-Free Visual Emotion Adaptation,” inProceedings of the AAAI Conference on Artificial Intelligence, vol. 39, no. 2, 2025, pp. 1674–1682

  33. [33]

    It Is Okay to Not Be Okay: Overcoming Emotional Bias in Affective Image Captioning by Contrastive Data Collection,

    Y . Mohamed, F. F. Khan, K. Haydarov, and M. Elhoseiny, “It Is Okay to Not Be Okay: Overcoming Emotional Bias in Affective Image Captioning by Contrastive Data Collection,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 21 263–21 272

  34. [34]

    Improved Emotional Alignment of AI and Humans: Human Ratings of Emotions Expressed by Stable Diffusion v1, DALL-E 2, and DALL-E 3,

    J. D. Lomas, W. van der Maden, S. Bandyopadhyay, G. Lion, N. Patel, G. Jain, Y . Litowsky, H. Xue, and P. Desmet, “Improved Emotional Alignment of AI and Humans: Human Ratings of Emotions Expressed by Stable Diffusion v1, DALL-E 2, and DALL-E 3,”arXiv preprint arXiv:2405.18510, 2024

  35. [35]

    All One Needs to Know About Metaverse: A Complete Survey on Technological Singularity, Virtual Ecosystem, and Research Agenda,

    L.-H. Lee, T. Braud, P. Y . Zhou, L. Wang, D. Xu, Z. Lin, A. Kumar, C. Bermejo, P. Huiet al., “All One Needs to Know About Metaverse: A Complete Survey on Technological Singularity, Virtual Ecosystem, and Research Agenda,”Foundations and Trends in Human-Computer Interaction, vol. 18, no. 2–3, pp. 100–337, 2024

  36. [36]

    PanoWan: Lifting Diffusion Video Generation Models to 360° with Latitude/Longitude-aware Mechanisms,

    Y . Xia, S. Weng, S. Yang, J. Liu, C. Zhu, M. Teng, Z. Jia, H. Jiang, and B. Shi, “PanoWan: Lifting Diffusion Video Generation Models to 360° with Latitude/Longitude-aware Mechanisms,”arXiv preprint arXiv:2505.22016, 2025

  37. [37]

    DreamCube: 3D Panorama Generation via Multi-plane Synchronization,

    Y . Huang, Y . Zhou, J. Wang, K. Huang, and X. Liu, “DreamCube: 3D Panorama Generation via Multi-plane Synchronization,”arXiv preprint arXiv:2506.17206, 2025

  38. [38]

    LayerPano3D: Layered 3D Panorama for Hyper-Immersive Scene Gen- eration,

    S. Yang, J. Tan, M. Zhang, T. Wu, G. Wetzstein, Z. Liu, and D. Lin, “LayerPano3D: Layered 3D Panorama for Hyper-Immersive Scene Gen- eration,” inProceedings of the Special Interest Group on Computer Graphics and Interactive Techniques Conference Conference Papers, 2025, pp. 1–10

  39. [39]

    DynamicScaler: Seamless and Scalable Video Generation for Panoramic Scenes,

    J. Liu, S. Lin, Y . Li, and M.-H. Yang, “DynamicScaler: Seamless and Scalable Video Generation for Panoramic Scenes,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2025, pp. 6144–6153

  40. [40]

    Taming Stable Diffusion for Text to 360 Panorama Image Generation,

    C. Zhang, Q. Wu, C. C. Gambardella, X. Huang, D. Phung, W. Ouyang, and J. Cai, “Taming Stable Diffusion for Text to 360 Panorama Image Generation,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 6347–6357

  41. [41]

    Kiss3DGen: Repurposing Image Diffusion Models for 3D Asset Generation,

    J. Lin, X. Yang, M. Chen, Y . Xu, D. Yan, L. Wu, X. Xu, L. Xu, S. Zhang, and Y .-C. Chen, “Kiss3DGen: Repurposing Image Diffusion Models for 3D Asset Generation,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2025, pp. 5870–5880

  42. [42]

    SceneDreamer: Unbounded 3D Scene Generation from 2D Image Collections,

    Z. Chen, G. Wang, and Z. Liu, “SceneDreamer: Unbounded 3D Scene Generation from 2D Image Collections,”IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 45, no. 12, pp. 15 562–15 576, 2023

  43. [43]

    Matrix-3D: Omnidirectional Explorable 3D World Generation,

    Z. Yang, W. Ge, Y . Li, J. Chen, H. Li, M. An, F. Kang, H. Xue, B. Xu, Y . Yinet al., “Matrix-3D: Omnidirectional Explorable 3D World Generation,”arXiv preprint arXiv:2508.08086, 2025

  44. [44]

    Infinite Photorealistic Worlds Using Procedural Generation,

    A. Raistrick, L. Lipson, Z. Ma, L. Mei, M. Wang, Y . Zuo, K. Kayan, H. Wen, B. Han, Y . Wanget al., “Infinite Photorealistic Worlds Using Procedural Generation,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 12 630–12 641

  45. [45]

    DI-PCG: Diffusion- Based Efficient Inverse Procedural Content Generation for High-Quality 3D Asset Creation,

    W. Zhao, Y .-P. Cao, J. Xu, Y . Dong, and Y . Shan, “DI-PCG: Diffusion- Based Efficient Inverse Procedural Content Generation for High-Quality 3D Asset Creation,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2025, pp. 11 061–11 072

  46. [46]

    CEAP-360VR: A Continuous Physiological and Behavioral Emotion Annotation Dataset for 360 VR Videos,

    T. Xue, A. El Ali, T. Zhang, G. Ding, and P. Cesar, “CEAP-360VR: A Continuous Physiological and Behavioral Emotion Annotation Dataset for 360 VR Videos,”IEEE Transactions on Multimedia, vol. 25, pp. 243–255, 2021

  47. [47]

    An Immer- sive and Interactive VR Dataset to Elicit Emotions,

    W. Jiang, M. Windl, B. Tag, Z. Sarsenbayeva, and S. Mayer, “An Immer- sive and Interactive VR Dataset to Elicit Emotions,”IEEE Transactions on Visualization and Computer Graphics, 2024

  48. [48]

    High-Resolution Image Synthesis with Latent Diffusion Models,

    R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “High-Resolution Image Synthesis with Latent Diffusion Models,” in 14 Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 10 684–10 695

  49. [49]

    An Image Is Worth One Word: Personalizing Text-to-Image Generation Using Textual Inversion,

    R. Gal, Y . Alaluf, Y . Atzmon, O. Patashnik, A. H. Bermano, G. Chechik, and D. Cohen-Or, “An Image Is Worth One Word: Personalizing Text-to-Image Generation Using Textual Inversion,”arXiv preprint arXiv:2208.01618, 2022

  50. [50]

    LoRA: Low-Rank Adaptation of Large Language Models,

    E. J. Hu, Y . Shen, P. Wallis, Z. Allen-Zhu, Y . Li, S. Wang, L. Wang, and W. Chen, “LoRA: Low-Rank Adaptation of Large Language Models,” inICLR. OpenReview.net, 2022. [Online]. Available: http://dblp.uni-trier.de/db/conf/iclr/iclr2022.html#HuSW ALWWC22

  51. [51]

    The Nature of Emotions: Human Emotions Have Deep Evolutionary Roots, a Fact That May Explain Their Complexity and Provide Tools for Clinical Practice,

    R. Plutchik, “The Nature of Emotions: Human Emotions Have Deep Evolutionary Roots, a Fact That May Explain Their Complexity and Provide Tools for Clinical Practice,”American scientist, vol. 89, no. 4, pp. 344–350, 2001

  52. [52]

    Sixteen Facial Expressions Occur in Similar Contexts Worldwide,

    A. S. Cowen, D. Keltner, F. Schroff, B. Jou, H. Adam, and G. Prasad, “Sixteen Facial Expressions Occur in Similar Contexts Worldwide,” Nature, vol. 589, no. 7841, pp. 251–257, 2021

  53. [53]

    Mapping the Passions: Toward a High-Dimensional Taxonomy of Emotional Experience and Expression,

    A. Cowen, D. Sauter, J. L. Tracy, and D. Keltner, “Mapping the Passions: Toward a High-Dimensional Taxonomy of Emotional Experience and Expression,”Psychological Science in the Public Interest, vol. 20, no. 1, pp. 69–90, 2019

  54. [54]

    Exploring Inter- pretability in Deep Learning for Affective Computing: A Comprehensive Review,

    X. Zhang, T. Zhang, L. Sun, J. Zhao, and Q. Jin, “Exploring Inter- pretability in Deep Learning for Affective Computing: A Comprehensive Review,”ACM Transactions on Multimedia Computing, Communica- tions and Applications, 2025

  55. [55]

    Emotion- Lens: Interactive Visual Exploration of the Circumplex Emotion Space in Literary Works via Affective Word Clouds,

    B. Wang, Q. Shi, X. Wang, Y . Zhou, W. Zeng, and Z. Wang, “Emotion- Lens: Interactive Visual Exploration of the Circumplex Emotion Space in Literary Works via Affective Word Clouds,”Visual Informatics, vol. 9, no. 1, pp. 84–98, 2025

  56. [56]

    This Looks Like That: Deep Learning for Interpretable Image Recognition,

    C. Chen, O. Li, D. Tao, A. Barnett, C. Rudin, and J. K. Su, “This Looks Like That: Deep Learning for Interpretable Image Recognition,” Advances in Neural Information Processing Systems, vol. 32, 2019

  57. [57]

    Interpretable Image Recognition with Hierarchical Prototypes,

    P. Hase, C. Chen, O. Li, and C. Rudin, “Interpretable Image Recognition with Hierarchical Prototypes,” inProceedings of the AAAI Conference on Human Computation and Crowdsourcing, vol. 7, 2019, pp. 32–40

  58. [58]

    Neural Prototype Trees for Interpretable Fine-Grained Image Recognition,

    M. Nauta, R. Van Bree, and C. Seifert, “Neural Prototype Trees for Interpretable Fine-Grained Image Recognition,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 14 933–14 943

  59. [59]

    ProtoDiffusion: Classifier-Free Diffusion Guidance with Prototype Learning,

    G. Baykal, H. F. Karagoz, T. Binhuraib, and G. Unal, “ProtoDiffusion: Classifier-Free Diffusion Guidance with Prototype Learning,” inAsian Conference on Machine Learning. PMLR, 2024, pp. 106–120

  60. [60]

    StyleCLIPDraw: Coupling Content and Style in Text-to-Drawing Translation,

    P. Schaldenbrand, Z. Liu, and J. Oh, “StyleCLIPDraw: Coupling Content and Style in Text-to-Drawing Translation,” inProceedings of the Thirty- First International Joint Conference on Artificial Intelligence (IJCAI- 22), 2022, pp. 4966–4972

  61. [61]

    Semantic space theory: A computational approach to emotion,

    A. S. Cowen and D. Keltner, “Semantic space theory: A computational approach to emotion,”Trends in Cognitive Sciences, vol. 25, no. 2, pp. 124–136, 2021

  62. [62]

    Multimodal physiological analysis of impact of emotion on cognitive control in vr,

    M. Li, J. Pan, Y . Li, Y . Gao, H. Qin, and Y . Shen, “Multimodal physiological analysis of impact of emotion on cognitive control in vr,” IEEE Transactions on Visualization and Computer Graphics, vol. 30, no. 5, pp. 2044–2054, 2024

  63. [63]

    Adding Conditional Control to Text-to-Image Diffusion Models,

    L. Zhang, A. Rao, and M. Agrawala, “Adding Conditional Control to Text-to-Image Diffusion Models,” inProceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 3836–3847

  64. [64]

    Learning Transferable Visual Models from Natural Language Supervision,

    A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clarket al., “Learning Transferable Visual Models from Natural Language Supervision,” inInternational Conference on Machine Learning. PMLR, 2021, pp. 8748–8763

  65. [65]

    Gaussian Error Linear Units (GELUs),

    D. Hendrycks and K. Gimpel, “Gaussian Error Linear Units (GELUs),” arXiv preprint arXiv:1606.08415, 2016

  66. [66]

    3D Gaussian Splatting for Real-Time Radiance Field Rendering,

    B. Kerbl, G. Kopanas, T. Leimk ¨uhler, and G. Drettakis, “3D Gaussian Splatting for Real-Time Radiance Field Rendering,”ACM Trans. Graph., vol. 42, no. 4, pp. 139–1, 2023

  67. [67]

    A Simple Frame- work for Contrastive Learning of Visual Representations,

    T. Chen, S. Kornblith, M. Norouzi, and G. Hinton, “A Simple Frame- work for Contrastive Learning of Visual Representations,” inInterna- tional Conference on Machine Learning. PMLR, 2020, pp. 1597–1607

  68. [68]

    MLP- Mixer: An All-MLP Architecture for Vision,

    I. O. Tolstikhin, N. Houlsby, A. Kolesnikov, L. Beyer, X. Zhai, T. Un- terthiner, J. Yung, A. Steiner, D. Keysers, J. Uszkoreitet al., “MLP- Mixer: An All-MLP Architecture for Vision,”Advances in Neural Information Processing Systems, vol. 34, pp. 24 261–24 272, 2021

  69. [69]

    Diffusion Models Beat GANs on Image Synthesis,

    P. Dhariwal and A. Nichol, “Diffusion Models Beat GANs on Image Synthesis,”Advances in Neural Information Processing Systems, vol. 34, pp. 8780–8794, 2021

  70. [70]

    GPT-4o System Card,

    A. Hurst, A. Lerer, A. P. Goucher, A. Perelman, A. Ramesh, A. Clark, A. Ostrow, A. Welihinda, A. Hayes, A. Radfordet al., “GPT-4o System Card,”arXiv preprint arXiv:2410.21276, 2024

  71. [71]

    SDXL Panorama,

    J. Bilcke, “SDXL Panorama,” https://huggingface.co/jbilcke-hf/ sdxl-panorama, 2024, hugging Face Model Hub, Accessed: 2025-08- 22

  72. [72]

    Latent Consistency Mod- els: Synthesizing High-Resolution Images with Few-Step Inference,

    S. Luo, Y . Tan, L. Huang, J. Li, and H. Zhao, “Latent Consistency Mod- els: Synthesizing High-Resolution Images with Few-Step Inference,” arXiv preprint arXiv:2310.04378, 2023

  73. [73]

    Outpainting with Stable Diffusion,

    Hugging Face, “Outpainting with Stable Diffusion,” https://huggingface. co/docs/diffusers/advanced inference/outpaint, 2024, accessed: 2025- 08-22

  74. [74]

    SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis,

    D. Podell, Z. English, K. Lacey, A. Blattmann, T. Dockhorn, J. M ¨uller, J. Penna, and R. Rombach, “SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis,”arXiv preprint arXiv:2307.01952, 2023

  75. [75]

    EmoSet: A Large-Scale Visual Emotion Dataset with Rich Attributes,

    J. Yang, Q. Huang, T. Ding, D. Lischinski, D. Cohen-Or, and H. Huang, “EmoSet: A Large-Scale Visual Emotion Dataset with Rich Attributes,” inProceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 20 383–20 394

  76. [76]

    BLIP-2: Bootstrapping Language- Image Pre-Training with Frozen Image Encoders and Large Language Models,

    J. Li, D. Li, S. Savarese, and S. Hoi, “BLIP-2: Bootstrapping Language- Image Pre-Training with Frozen Image Encoders and Large Language Models,” inInternational Conference on Machine Learning. PMLR, 2023, pp. 19 730–19 742

  77. [77]

    LAION-5B: An Open Large-Scale Dataset for Training Next Gener- ation Image-Text Models,

    C. Schuhmann, R. Beaumont, R. Vencu, C. Gordon, R. Wightman, M. Cherti, T. Coombes, A. Katta, C. Mullis, M. Wortsmanet al., “LAION-5B: An Open Large-Scale Dataset for Training Next Gener- ation Image-Text Models,”Advances in Neural Information Processing Systems, vol. 35, pp. 25 278–25 294, 2022

  78. [78]

    Meta Quest 2,

    Meta Platforms, Inc., “Meta Quest 2,” 2020, accessed: 2025-08-22. [Online]. Available: https://www.meta.com/quest/products/quest-2/

  79. [79]

    Unity 2022.3.6f1 Release Notes,

    Unity Technologies, “Unity 2022.3.6f1 Release Notes,” 2022, accessed: 2025-08-22. [Online]. Available: https://unity.com/releases/ editor/whats-new/2022.3.6f1

  80. [80]

    Affective Interac- tions Using Virtual Reality: the Link between Presence and Emotions,

    G. Riva, F. Mantovani, C. S. Capideville, A. Preziosa, F. Morganti, D. Villani, A. Gaggioli, C. Botella, and M. Alca ˜niz, “Affective Interac- tions Using Virtual Reality: the Link between Presence and Emotions,” Cyberpsychology & behavior, vol. 10, no. 1, pp. 45–56, 2007