REVIEW 4 major objections 5 minor 67 references
Visually grounded emotion regulation via diffusion models and user-driven reappraisal
T0 review · 4 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read The paper claims that coupling spoken cognitive reappraisal with AI-generated images that visually instantiate the reappraisal reduces negative affect more than reappraisal alone, and that the closer the image matches the user's verbal…
desk verdict Genuinely novel application with a real behavioral effect, but the control condition does not isolate visual grounding and the alignment analysis is circular—worth reviewing, not accepting as-is. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery is a fine-tuned IP-Adapter mounted on a frozen Stable Diffusion XL backbone. The IP-Adapter injects a CLIP embedding of the original IAPS image into the U-Net's cross-attention layers through a parallel image-conditioning stream, letting the model redraw the scene according to the user's spoken reappraisal prompt while retaining the original image's structure. A conditioning scale α balances text versus image influence; in the Reappraisal-AI condition α is adapted per image, whereas in the Describe-AI control it is fixed at 0.8 to preserve the original negative content. Semantic alignment between prompt and generated image is then measured as cosine similarity between sentence embeddings of the prompt and a GPT-4V caption of the generated image.
What would settle it
A condition in which participants are shown a positive AI-generated transformation of a negative IAPS image without being asked to reappraise it—matched in depicted positivity to the RAI images rather than locked to α=0.8—would falsify the imagery-alone null if it produced the same affective relief as RAI; alternatively, a preregistered replication with a larger sample that fails to find a significant Neg-RAI versus Neg-R difference would falsify the headline claim.
Extended reading notes
Core claim
In the paper's own terms, the central discovery is that AI-assisted reappraisal (Reappraisal-AI) reduces negative affect for aversive IAPS stimuli relative to traditional reappraisal (mean difference 0.818, t=5.20, p<0.001 Bonferroni-corrected), with a significant three-way emotion × instruction × modality interaction, F(1,19)=8.07, p=0.0104. The authors further report that the semantic alignment between the participant's verbal reappraisal prompt and the generated image correlates with affective relief (Neg-RAI rho=0.56, p=0.020; Neu-RAI rho=0.63, p=0.005), and that sentiment of the spoken reappraisal correlates with outcome only in the AI-visual conditions, suggesting that visual grounding, not verbal content alone, carries the regulatory benefit.
Load-bearing premise
The argument that reappraisal plus imagery, not imagery alone, causes the relief rests on the assumption that the Describe-AI condition—where the conditioning scale was deliberately fixed at 0.8 so generated images kept the original negative content—fairly represents what imagery alone would do if participants never reappraised.
Editorial extensions
If this is right
- Reappraisal training could be delivered with real-time visual feedback instead of relying on abstract verbal reasoning, potentially reducing the executive-function burden that limits the strategy in depression, PTSD, and high-arousal moments.
- Clinical digital tools could use the same pipeline to let users see their own reframes, turning an internal cognitive act into an external, inspectable artifact.
- The alignment–outcome correlation suggests that intervention quality can be monitored automatically by measuring how faithfully the generated image realizes the user's stated reappraisal, providing a target for system tuning.
- The null effect of imagery alone (Describe-AI vs Describe) implies that adding generative visual feedback to an active task will not help unless the user is engaged in intentional reinterpretation.
Reading between the lines
- If the effect replicates, a natural next step is testing whether the visual feedback works via increased emotional engagement, deeper semantic processing, or the externalization of a goal state; the current design cannot separate those routes.
- One could estimate the dose-response by comparing different α values in the RAI condition: if dynamic adaptation matters, a parametric manipulation of conditioning scale against affective outcome would be a direct test.
- The correlation evidence suggests a practical design principle—maximize prompt-image alignment—but alignment itself is measured through AI captioning, so part of the reported effect could be mediated by the captioner rather than by participants' perception; a human-rating validation of alignment would settle this.
- The sample is small and healthy; the strongest translational claim—that this helps people with trauma or depression—remains untested, and a clinical replication would be the key extension.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a system that augments cognitive reappraisal with AI-generated visual feedback. In a within-subject experiment (N=20), participants either described or reappraised negative and neutral IAPS images, with or without real-time diffusion-model images conditioned on the original image and the participant's spoken prompt. The authors report that reappraisal with AI-generated imagery (RAI) reduces negative affect relative to reappraisal alone (R) (mean difference 0.818, t=5.20, p<0.001 Bonferroni-corrected), with a significant three-way emotion by instruction by modality interaction, F(1,19)=8.07, p=0.0104. They also report that the sentiment of the reappraisal prompt correlates with affective outcome in RAI but not R, and that semantic alignment between the prompt and a caption of the generated image predicts affective relief. The paper interprets these findings as evidence that generative visual grounding amplifies cognitive reappraisal, and that the effect is not due to imagery alone because Describe-AI did not differ from Describe.
Significance. If the central effect is robust, the paper makes a useful contribution at the intersection of generative AI and affective computing: it is one of the first controlled demonstrations that user-conditioned diffusion imagery can alter self-reported emotional state in an emotion-regulation task. The study is technically ambitious (real-time speech transcription, fine-tuned IP-Adapter, within-subject design) and the headline behavioral result is supported by standard frequentist statistics. The sentiment and alignment analyses are clearly described and falsifiable. However, the paper's stronger mechanistic claims—that the benefit comes from the interaction of reappraisal with visual grounding rather than from positive imagery alone, and that multimodal coherence predicts relief—rest on design choices and analyses that do not currently rule out plausible confounds, so the significance of the interpretational contribution is not yet established.
major comments (4)
- [Section 2.6 and Section 3.2 (Hypothesis 2)] The Describe-AI condition fixes the conditioning scale at alpha=0.8 specifically to keep the generated image closely aligned with the original stimulus, so the images in DAI preserve the original negative content. Consequently, the null DAI versus D comparison only shows that viewing a negative AI-rendered version of the stimulus has no effect; it does not test whether viewing a positive image without reappraisal would improve mood. In the RAI condition, alpha is dynamically adjusted to produce images that match the participant's positive reappraisal, so the RAI versus R contrast differs simultaneously in the presence of visual feedback and in the valence of that feedback. The claim in the Discussion that the benefits are "not attributable to imagery alone" is therefore not supported by the data; a control condition that presents a positive image without reappraisal (e.g., a generic positive image, or a positive image generated from another participant's reappraisal) is needed to isolate the reappraisal by imagery interaction.
- [Section 2.2] The stated timing match is not complete. In the AI conditions, after the 4-second gray screen the generated image is shown for 3 seconds before the affective rating; in the non-AI conditions, participants rate immediately after the 4-second gray screen. The RAI condition thus includes an extra 3 seconds of visual stimulation (and a longer interval between speech and rating), which could itself affect mood through distraction, positive mood induction, or simple delay. The authors should either present a non-AI condition with a matched 3-second display of a neutral or positive image, or explicitly model and discuss this timing difference as a potential confound.
- [Section 3.4] The alignment metric is computed between the participant's reappraisal prompt and a caption of the image that was generated from that same prompt. This makes the metric partly self-referential: if the generation pipeline is prompt-faithful, alignment will be high by construction. Moreover, the analysis does not control for the sentiment of the prompt, and a more positive prompt will tend to produce a more positive image and a more positive caption, so the reported correlation between alignment and affect could be mediated by prompt sentiment. To support the claim that multimodal coherence drives regulatory efficacy, the authors should report partial correlations controlling for prompt sentiment (and ideally prompt length and complexity), or demonstrate that alignment remains predictive within subsamples matched on prompt sentiment.
- [Section 3.3 and Figure 3] The correlation analysis is under-specified. The text reports rho=0.63, p=0.025 for Neg-RAI and rho=0.64, p=0.018 for Neu-RAI, but the Neu-R entry contains the apparent typo "p = 6.80.49", and the number of comparisons used in the Bonferroni correction is not stated. Please also report the degrees of freedom (or the number of subjects and whether the analysis uses subject-level means or trial-level observations) for each correlation, and ensure the p-values in the text, figures, and supplementary figures are mutually consistent.
minor comments (5)
- [Section 3.2] The paragraph beginning "Specifically, we compared mean emotional responses..." is duplicated verbatim; remove the duplicate.
- [Abstract] The abstract contains the typo "cogitive reappraisal"; it should read "cognitive reappraisal."
- [Figure S3 caption] The caption states that the plot is "within the Reappraisal AI condition," but the figure displays the Describe AI condition; correct the caption.
- [Sections 2.4 and 2.6] The conditioning scale is denoted lambda in Eq. (2) but alpha in Section 2.6; use a single symbol throughout to avoid confusion.
- [Section 3.2] The reported pairwise differences are arithmetically inconsistent: Neg-R versus Neg-D = 0.6 and Neg-RAI versus Neg-R = 0.818 imply Neg-RAI versus Neg-D is approximately 1.42, whereas the reported Neg-RAI versus Neg-DAI difference is 1.15 even though Neg-D and Neg-DAI do not differ (0.05, n.s.). Check the underlying means and report the actual condition means in a table.
Circularity Check
No significant circularity: the behavioral main effect is empirical; the alignment analysis is a confounding/validity concern, not a by-construction reduction.
full rationale
The paper's central result—RAI versus R (mean difference = 0.818, t = 5.20, p < .001)—is a direct within-subject behavioral measurement, not a quantity produced by fitting a model to the outcome; the diffusion pipeline is an intervention generator, and no parameter of that pipeline is tuned to the affective ratings. The secondary alignment analysis (Section 3.4) computes cosine similarity between the participant's prompt embedding and the caption embedding of an image generated from that same prompt, with GPT-4V instructed that 'This image was generated as a positive reinterpretation of a scene,' so alignment partly indexes pipeline self-consistency and prompt positivity rather than an independent multimodal-coherence construct. The paper itself flags this entanglement in Section 4.1: 'future work should more systematically disentangle the unique and interactive effects of alignment, affective tone, and reappraisal intent.' This is a confounding/interpretational limitation, but it is not an equation-level equivalence between the predictor and the outcome, so it does not meet the stated bar for circularity. No load-bearing self-citation appears; the technical citations (SDXL, IP-Adapter, Whisper, GPT-4V, Sentence-BERT) are external. The DAI control is weakened by fixing alpha = 0.8 to preserve negative content, but that is an experimental-design confound (image-valence confound), not a circular derivation.
Assumptions & free parameters
free parameters (3)
- IP-Adapter conditioning scale alpha/lambda =
lambda in [0.3, 0.7]; alpha fixed at 0.8 for Describe-AI, dynamically adapted per image for Reappraisal-AI
- Inference settings for SDXL =
text guidance scale 7.5, DDIM sampler, 40 denoising steps
- IP-Adapter fine-tuning hyperparameters =
200,000 steps, AdamW learning rate 1e-4, batch size 48, four NVIDIA L40 GPUs, about 54 hours
assumptions (5)
- standard math Repeated-measures ANOVA and Pearson correlation assumptions hold for the reported statistics.
- domain assumption Self-reported valence on a visual analog scale is a valid measure of affective state for this task.
- domain assumption Whisper transcription and automatic translation preserve the meaning of spoken reappraisals in all four languages.
- domain assumption GPT-4V captions and sentence embeddings provide a valid semantic alignment measure between prompts and generated images.
- ad hoc to paper The GPT-3.5-generated and GPT-4V-filtered synthetic dataset is a valid training proxy for human reappraisals.
Cite this review
Pith. "Pith review of Visually grounded emotion regulation via diffusion models and user-driven reappraisal." pith.science (2026). https://pith.science/paper/VHGVMKEV
@misc{pith2026250710861,
author = {Pith},
title = {Pith review of: Visually grounded emotion regulation via diffusion models and user-driven reappraisal},
year = {2026},
howpublished = {\url{https://pith.science/paper/VHGVMKEV}},
note = {Machine review of arXiv:2507.10861}
}
read the original abstract
Cognitive reappraisal is a key strategy in emotion regulation, involving reinterpretation of emotionally charged stimuli to alter affective responses. Despite its central role in clinical and cognitive science, real-world reappraisal interventions remain cognitively demanding, abstract, and primarily verbal. This reliance on higher-order cognitive and linguistic processes is often impaired in individuals with trauma or depression, limiting the effectiveness of standard approaches. Here, we propose a novel, visually based augmentation of cognitive reappraisal by integrating large-scale text-to-image diffusion models into the emotional regulation process. Specifically, we introduce a system in which users reinterpret emotionally negative images via spoken reappraisals, which are transformed into supportive, emotionally congruent visualizations using stable diffusion models with a fine-tuned IP-adapter. This generative transformation visually instantiates users' reappraisals while maintaining structural similarity to the original stimuli, externalizing and reinforcing regulatory intent. To test this approach, we conducted a within-subject experiment (N = 20) using a modified cognitive emotion regulation (CER) task. Participants reappraised or described aversive images from the International Affective Picture System (IAPS), with or without AI-generated visual feedback. Results show that AI-assisted reappraisal significantly reduced negative affect compared to both non-AI and control conditions. Further analyses reveal that sentiment alignment between participant reappraisals and generated images correlates with affective relief, suggesting that multimodal coherence enhances regulatory efficacy. These findings demonstrate that generative visual input can support cogitive reappraisal and open new directions at the intersection of generative AI, affective computing, and therapeutic technology.
Figures
Figures from the paper (1 more)
Reference graph
Works this paper leans on
-
[1]
Emotion regulation: Affective, cognitive, and social consequences,
J. J. Gross, “Emotion regulation: Affective, cognitive, and social consequences,” Psychophysiology, vol. 39, no. 3, pp. 281–291, 2002
work page 2002
-
[2]
T. M. Lincoln, L. Schulze, and B. Renneberg, “The role of emotion regulation in the characteriza- tion, development and treatment of psychopathology,” Nature Reviews Psychology, vol. 1, no. 5, pp. 272–286, 2022
work page 2022
-
[3]
Emotion regulation: Current status and future prospects,
J. J. Gross, “Emotion regulation: Current status and future prospects,” Psychological Inquiry, vol. 26, no. 1, pp. 1–26, 2015
work page 2015
-
[4]
Emotion regulation and the electrocortical response to unpleasant pictures,
G. Hajcak, A. MacNamara, and D. M. Olvet, “Emotion regulation and the electrocortical response to unpleasant pictures,” Psychophysiology, vol. 47, no. 3, pp. 432–441, 2010
work page 2010
-
[5]
A. Riepenhausen, C. Wackerhagen, Z. C. Reppmann, H.-C. Deter, R. Kalisch, I. M. Veer, and H. Walter, “Positive cognitive reappraisal in stress resilience, mental health, and well-being: A comprehensive systematic review,” Emotion Review, vol. 14, no. 4, pp. 310–331, 2022
work page 2022
-
[6]
F. Han, R. Duan, B. Huang, and Q. Wang, “Psychological resilience and cognitive reappraisal mediate the effects of coping style on the mental health of children,” Frontiers in Psychology , vol. 14, p. 1110642, 2023
work page 2023
-
[7]
S. In, J.-W. Hur, G. Kim, and J.-H. Lee, “Effects of distraction, cognitive reappraisal, and accep- tance on the urge to self-harm and negative affect in nonsuicidal self-injury,” Current Psychology, pp. 1–8, 2021
work page 2021
-
[8]
M. W. Southward, A. C. Holmes, D. R. Strunk, and J. S. Cheavens, “More and better: reappraisal quality partially explains the effect of reappraisal use on changes in positive and negative affect,” Cognitive Therapy and Research, pp. 1–13, 2022
work page 2022
Show all 67 references
-
[9]
Reappraisal-related neural predictors of treatment response to cognitive behavior therapy for post-traumatic stress disorder,
R. A. Bryant, M. Erlinger, K. Felmingham, A. Klimova, L. M. Williams, G. Malhi, D. Forbes, and M. S. Korgaonkar, “Reappraisal-related neural predictors of treatment response to cognitive behavior therapy for post-traumatic stress disorder,” Psychological Medicine, vol. 51, no....
2021
-
[10]
Cognitive reappraisal of emotion: a meta-analysis of human neuroimaging studies,
J. T. Buhle, J. A. Silvers, T. D. Wager, R. Lopez, C. Onyemekwu, H. Kober, J. Weber, and K. N. Ochsner, “Cognitive reappraisal of emotion: a meta-analysis of human neuroimaging studies,” Cerebral cortex, vol. 24, no. 11, pp. 2981–2990, 2014
2014
-
[11]
Neural correlates of cognitive reappraisal of positive and negative affect in older adults,
K. Halfmann, W. Hedgcock, and N. L. Denburg, “Neural correlates of cognitive reappraisal of positive and negative affect in older adults,” Aging & mental health , vol. 25, no. 1, pp. 126–133, 2021. 14
2021
-
[12]
Unpacking cognitive reappraisal: goals, tactics, and outcomes
K. McRae, B. Ciesielski, and J. J. Gross, “Unpacking cognitive reappraisal: goals, tactics, and outcomes.” Emotion, vol. 12, no. 2, p. 250, 2012
2012
-
[13]
The role of cognitive reappraisal strategies in modulating traumatic memories,
J. Montfort and L. Sondheim, “The role of cognitive reappraisal strategies in modulating traumatic memories,” Studies in Psychological Science , vol. 2, no. 2, pp. 1–8, 2024
2024
-
[14]
How to regulate emotion? neural networks for reappraisal and distraction,
P. Kanske, J. Heissler, S. Sch¨ onfelder, A. Bongers, and M. Wessa, “How to regulate emotion? neural networks for reappraisal and distraction,” Cerebral cortex, vol. 21, no. 6, pp. 1379–1388, 2011
2011
-
[15]
International affective picture system (iaps): Technical manual and affective ratings,
P. J. Lang, M. M. Bradley, B. N. Cuthbert et al. , “International affective picture system (iaps): Technical manual and affective ratings,” NIMH Center for the Study of Emotion and Attention , vol. 1, no. 39-58, p. 3, 1997
1997
-
[16]
Acute and sustained effects of cognitive emotion regulation in major depression,
S. Erk, A. Mikschl, S. Stier, A. Ciaramidaro, V. Gapp, B. Weber, and H. Walter, “Acute and sustained effects of cognitive emotion regulation in major depression,” Journal of Neuroscience , vol. 30, no. 47, pp. 15 726–15 734, 2010
2010
-
[17]
Cognitive emotion regulation withstands the stress test: An fmri study on the effect of acute stress on distraction and reappraisal,
M. Sandner, P. Zeier, G. Lois, and M. Wessa, “Cognitive emotion regulation withstands the stress test: An fmri study on the effect of acute stress on distraction and reappraisal,” Neuropsychologia, vol. 157, p. 107876, 2021
2021
-
[18]
Amygdala–frontal con- nectivity during emotion regulation,
S. J. Banks, K. T. Eddy, M. Angstadt, P. J. Nathan, and K. L. Phan, “Amygdala–frontal con- nectivity during emotion regulation,” Social cognitive and affective neuroscience, vol. 2, no. 4, pp. 303–312, 2007
2007
-
[19]
Predicting cognitive behavioral therapy response in social anxiety disorder with anterior cingulate cortex and amygdala during emotion regulation,
H. Klumpp, J. M. Fitzgerald, K. L. Kinney, A. E. Kennedy, S. A. Shankman, S. A. Langenecker, and K. L. Phan, “Predicting cognitive behavioral therapy response in social anxiety disorder with anterior cingulate cortex and amygdala during emotion regulation,”NeuroImage: Clinical...
2017
-
[20]
Anatomical insights into the interaction of emotion and cognition in the prefrontal cortex,
R. D. Ray and D. H. Zald, “Anatomical insights into the interaction of emotion and cognition in the prefrontal cortex,” Neuroscience & Biobehavioral Reviews, vol. 36, no. 1, pp. 479–501, 2012
2012
-
[21]
Changes in effective connectivity between dorsal and ventral prefrontal regions moderate emotion regulation,
C. Morawetz, S. Bode, J. Baudewig, E. Kirilina, and H. R. Heekeren, “Changes in effective connectivity between dorsal and ventral prefrontal regions moderate emotion regulation,”Cerebral Cortex, vol. 26, no. 5, pp. 1923–1937, 2016
1923
-
[22]
Relationship between emotion comprehension, vocabulary, and verbal working memory in intellectual developmental disorders: involvement of verbal reasoning skills,
M. Vy, S. Ferrara, N. Dollion, and C. Declercq, “Relationship between emotion comprehension, vocabulary, and verbal working memory in intellectual developmental disorders: involvement of verbal reasoning skills,” Cognition and Emotion , pp. 1–9, 2024
2024
-
[23]
Emotional verbal fluency: A new task on emotion and executive function interaction,
K. Sass, K. Fetz, S. Oetken, U. Habel, and S. Heim, “Emotional verbal fluency: A new task on emotion and executive function interaction,” Behavioral sciences, vol. 3, no. 3, pp. 372–387, 2013
2013
-
[24]
Cognitive emotion regulation: A review of theory and scientific findings,
K. McRae, “Cognitive emotion regulation: A review of theory and scientific findings,” Current Opinion in Behavioral Sciences , vol. 10, pp. 119–124, 2016
2016
-
[25]
Emotion regulation and mental health
J. J. Gross and R. F. Mu˜ noz, “Emotion regulation and mental health.” Clinical psychology: Science and practice, vol. 2, no. 2, p. 151, 1995
1995
-
[26]
The behavioral emotion regulation questionnaire: development, psychometric properties and relationships with emotional problems and the cognitive emotion regulation questionnaire,
V. Kraaij and N. Garnefski, “The behavioral emotion regulation questionnaire: development, psychometric properties and relationships with emotional problems and the cognitive emotion regulation questionnaire,” Personality and Individual Differences , vol. 137, pp. 56–61, 2019
2019
-
[27]
A study on cognitive emotion regulation and anxiety and depression in adults,
S. Jacob and M. M. Anto, “A study on cognitive emotion regulation and anxiety and depression in adults,” The International Journal of Indian Psychology , vol. 3, no. 2, pp. 118–24, 2016
2016
-
[28]
The digital revolution and its impact on mental health care,
S. Bucci, M. Schwannauer, and N. Berry, “The digital revolution and its impact on mental health care,” Psychology and Psychotherapy: Theory, Research and Practice , vol. 92, no. 2, pp. 277–297, 2019. 15
2019
-
[29]
Digitally assisted mindfulness in training self-regulation skills for sustainable mental health: a systematic review,
E. Mitsea, A. Drigas, and C. Skianis, “Digitally assisted mindfulness in training self-regulation skills for sustainable mental health: a systematic review,” Behavioral Sciences, vol. 13, no. 12, p. 1008, 2023
2023
-
[30]
The application of artificial intelligence in the field of mental health: a systematic review,
R. Dehbozorgi, S. Zangeneh, E. Khooshab, D. H. Nia, H. R. Hanif, P. Samian, M. Yousefi, F. H. Hashemi, M. Vakili, N. Jamalimoghadam et al. , “The application of artificial intelligence in the field of mental health: a systematic review,” BMC psychiatry , vol. 25, p. 132, 2025
2025
-
[31]
Talking to machines about personal mental health problems,
A. S. Miner, A. Milstein, and J. T. Hancock, “Talking to machines about personal mental health problems,” Jama, vol. 318, no. 13, pp. 1217–1218, 2017
2017
-
[32]
Comparing the value of perceived human versus ai-generated empathy,
M. Rubin, J. Z. Li, F. Zimmerman, D. C. Ong, A. Goldenberg, and A. Perry, “Comparing the value of perceived human versus ai-generated empathy,” Nature Human Behaviour , pp. 1–15, 2025
2025
-
[33]
Influencing human–ai interaction by prim- ing beliefs about ai can increase perceived trustworthiness, empathy and effectiveness,
P. Pataranutaporn, R. Liu, E. Finn, and P. Maes, “Influencing human–ai interaction by prim- ing beliefs about ai can increase perceived trustworthiness, empathy and effectiveness,” Nature Machine Intelligence , vol. 5, no. 10, pp. 1076–1086, 2023
2023
-
[34]
How human–ai feedback loops alter human perceptual, emotional and social judgements,
M. Glickman and T. Sharot, “How human–ai feedback loops alter human perceptual, emotional and social judgements,” Nature Human Behaviour , vol. 9, no. 2, pp. 345–359, 2025
2025
-
[35]
Human–ai collaboration enables more empathic conversations in text-based peer-to-peer mental health support,
A. Sharma, I. W. Lin, A. S. Miner, D. C. Atkins, and T. Althoff, “Human–ai collaboration enables more empathic conversations in text-based peer-to-peer mental health support,” Nature Machine Intelligence, vol. 5, no. 1, pp. 46–57, 2023
2023
-
[36]
Guiding large language models to perform cognitive reappraisal,
Z. Zhan, Z. Lin, Y. Shen, Y. Zhang, and K. Peng, “Guiding large language models to perform cognitive reappraisal,” arXiv preprint arXiv:2403.09798 , 2024
2024 arXiv
-
[37]
A computational framework for behavioral assessment of llm therapists,
Y. Y. Chiu, A. Sharma, I. W. Lin, and T. Althoff, “A computational framework for behavioral assessment of llm therapists,” arXiv preprint arXiv:2401.00820 , 2024
2024 arXiv
-
[38]
A therapeutic relational agent for reducing problematic substance use (woebot): development and usability study,
J. J. Prochaska, E. A. Vogel, A. Chieng, M. Kendra, M. Baiocchi, S. Pajarito, and A. Robinson, “A therapeutic relational agent for reducing problematic substance use (woebot): development and usability study,” Journal of medical Internet research , vol. 23, no. 3, p. e24850, 2021
2021
-
[39]
Chatbots and conversational agents in mental health: a review of the psychiatric landscape,
A. N. Vaidyam, H. Wisniewski, J. D. Halamka, M. S. Kashavan, and J. B. Torous, “Chatbots and conversational agents in mental health: a review of the psychiatric landscape,” The Canadian Journal of Psychiatry , vol. 64, no. 7, pp. 456–464, 2019
2019
-
[40]
Stable video diffusion: Scaling latent video diffusion models to large datasets,
A. Blattmann, T. Dockhorn, S. Kulal, D. Mendelevitch, M. Kilian, D. Lorenz, Y. Levi, Z. English, V. Voleti, A. Letts et al. , “Stable video diffusion: Scaling latent video diffusion models to large datasets,” arXiv preprint arXiv:2311.15127 , 2023
2023 arXiv
-
[41]
Adding conditional control to text-to-image diffusion mod- els,
L. Zhang, A. Rao, and M. Agrawala, “Adding conditional control to text-to-image diffusion mod- els,” in Proceedings of the IEEE/CVF international conference on computer vision , 2023, pp. 3836–3847
2023
-
[42]
Sdxl: Improving latent diffusion models for high-resolution image synthesis,
D. Podell, Z. English, K. Lacey, A. Blattmann, T. Dockhorn, J. M¨ uller, J. Penna, and R. Rombach, “Sdxl: Improving latent diffusion models for high-resolution image synthesis,”
-
[43]
Ip-adapter: Text-to-image diffusion models are zero-shot image-to-image trans- lators,
X. Xu and et al., “Ip-adapter: Text-to-image diffusion models are zero-shot image-to-image trans- lators,” arXiv preprint arXiv:2308.06700 , 2023
2023 arXiv
-
[44]
Emogen: Emotional image content generation with text-to- image diffusion models,
J. Yang, J. Feng, and H. Huang, “Emogen: Emotional image content generation with text-to- image diffusion models,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 6358–6368
2024
-
[45]
Improved emotional alignment of ai and humans: Human ratings of emotions expressed by stable diffusion v1, dall-e 2, and dall-e 3,
J. D. Lomas, W. van der Maden, S. Bandyopadhyay, G. Lion, N. Patel, G. Jain, Y. Litowsky, H. Xue, and P. Desmet, “Improved emotional alignment of ai and humans: Human ratings of emotions expressed by stable diffusion v1, dall-e 2, and dall-e 3,”arXiv preprint arXiv:2405.18510,...
2024 arXiv
-
[46]
Stable bias: Evaluating societal represen- tations in diffusion models,
S. Luccioni, C. Akiki, M. Mitchell, and Y. Jernite, “Stable bias: Evaluating societal represen- tations in diffusion models,” Advances in Neural Information Processing Systems , vol. 36, pp. 56 338–56 351, 2023
2023
-
[47]
Ai-assisted character design in medical storytelling with stable diffusion,
S. Mittenentzwei, L. A. Garrison, B. Budich, K. Lawonn, A. Dockhorn, B. Preim, and M. Meuschke, “Ai-assisted character design in medical storytelling with stable diffusion,”Available at SSRN 4772811 , 2024
2024
-
[48]
Positive visual reframing: A randomised controlled trial using drawn visual imagery to defuse the intensity of negative experiences and regulate emotions in healthy adults,
J. C. Ruppert, F. J. Eiroa-Orosa et al., “Positive visual reframing: A randomised controlled trial using drawn visual imagery to defuse the intensity of negative experiences and regulate emotions in healthy adults,” Anales de Psicolog ´ ıa/Annals of Psychology, vol. 34, no. 2,...
2018
-
[49]
Exploratory study of imagery rescripting without focusing on early traumatic memories for major depressive disorder,
F. Yamada, Y. Hiramatsu, T. Murata, Y. Seki, M. Yokoo, R. Noguchi, T. Shibuya, M. Tanaka, R. Takanashi, and E. Shimizu, “Exploratory study of imagery rescripting without focusing on early traumatic memories for major depressive disorder,” Psychology and Psychotherapy: Theory, ...
2018
-
[50]
Mental imagery: functional mecha- nisms and clinical applications,
J. Pearson, T. Naselaris, E. A. Holmes, and S. M. Kosslyn, “Mental imagery: functional mecha- nisms and clinical applications,” Trends in cognitive sciences, vol. 19, no. 10, pp. 590–602, 2015
2015
-
[51]
Advances in the use of virtual reality to treat mental health conditions,
I. H. Bell, R. Pot-Kolder, A. Rizzo, M. Rus-Calafell, V. Cardi, M. Cella, T. Ward, S. Riches, M. Reinoso, A. Thompson et al. , “Advances in the use of virtual reality to treat mental health conditions,” Nature Reviews Psychology, vol. 3, no. 8, pp. 552–567, 2024
2024
-
[52]
Mental imagery in the science and practice of cognitive behaviour therapy: Past, present, and future perspectives,
S. E. Blackwell, “Mental imagery in the science and practice of cognitive behaviour therapy: Past, present, and future perspectives,” International Journal of Cognitive Therapy , vol. 14, no. 1, pp. 160–181, 2021
2021
-
[53]
Robust speech recognition via large-scale weak supervision,
A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust speech recognition via large-scale weak supervision,” arXiv preprint arXiv:2212.03322 , 2022. [Online]. Available: https://arxiv.org/abs/2212.03322
2022 arXiv
-
[54]
Enhancing multimodal understanding with clip- based image-to-text transformation,
C. Che, Q. Lin, X. Zhao, J. Huang, and L. Yu, “Enhancing multimodal understanding with clip- based image-to-text transformation,” in Proceedings of the 2023 6th International Conference on Big Data Technologies, 2023, pp. 414–418
2023
-
[55]
Learning transferable visual models from natural language supervi- sion,
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al. , “Learning transferable visual models from natural language supervi- sion,” in International conference on machine learning . PmLR, 2021, pp. 8748–8763
2021
-
[56]
Gpt-3.5 technical overview,
OpenAI, “Gpt-3.5 technical overview,” 2023. [Online]. Available: https://platform.openai.com/ docs/models/gpt-3-5
2023
-
[57]
Gpt-4 with vision: System card,
——, “Gpt-4 with vision: System card,” 2023. [Online]. Available: https://openai.com/index/ gpt-4-system-card
2023
-
[58]
Classifier-free diffusion guidance,
J. Ho and T. Salimans, “Classifier-free diffusion guidance,” arXiv preprint arXiv:2207.12598 , 2022
2022 arXiv
-
[59]
Denoising diffusion implicit models,
J. Song, C. Meng, and S. Ermon, “Denoising diffusion implicit models,” arXiv preprint arXiv:2010.02502, 2020
2010 arXiv
-
[60]
Tweeteval: Unified benchmark and comparative evaluation for tweet classification,
F. Barbieri, J. Camacho-Collados, L. Neves, and L. Espinosa-Anke, “Tweeteval: Unified benchmark and comparative evaluation for tweet classification,” in Proceedings of the 12th Language Resources and Evaluation Conference (LREC 2020) . Santa Monica, CA, USA: Snap Inc., Cardiff...
2020
-
[61]
textstat: Python package for readability and text statistics,
K. P. R. and textstat contributors, “textstat: Python package for readability and text statistics,” 2020, accessed: 2025-07-08. [Online]. Available: https://github.com/shivam5992/textstat
2020
-
[62]
Sentence-bert: Sentence embeddings using siamese bert-networks,
N. Reimers and I. Gurevych, “Sentence-bert: Sentence embeddings using siamese bert-networks,” arXiv preprint arXiv:1908.10084 , 2019. 17
1908 arXiv
-
[63]
Physiological and self-report measures in emotion studies: Method- ological considerations,
P. Korpal and K. Jankowiak, “Physiological and self-report measures in emotion studies: Method- ological considerations,” 2018
2018
-
[64]
Regulating emotion through distancing: A taxonomy, neurocog- nitive model, and supporting meta-analysis,
J. P. Powers and K. S. LaBar, “Regulating emotion through distancing: A taxonomy, neurocog- nitive model, and supporting meta-analysis,” Neuroscience & Biobehavioral Reviews, vol. 96, pp. 155–173, 2019
2019
-
[65]
Linguistic measures of psychological distance track symptom levels and treatment outcomes in a large set of psychotherapy transcripts,
E. C. Nook, T. D. Hull, M. K. Nock, and L. H. Somerville, “Linguistic measures of psychological distance track symptom levels and treatment outcomes in a large set of psychotherapy transcripts,” Proceedings of the National Academy of Sciences , vol. 119, no. 13, p. e2114737119, 2022
2022
-
[66]
Putting feelings into words: affect labeling disrupts amygdala activity in response to affective stimuli. psycholsci18 (5): 421–428,
M. Lieberman, N. Eisenberger, M. Crockett et al. , “Putting feelings into words: affect labeling disrupts amygdala activity in response to affective stimuli. psycholsci18 (5): 421–428,” 2007. 18 6 Supplementary figures Figure S1: Visual Comparison of pre-trained and fine-tuned...
2007
-
[2023]
Available: https://arxiv.org/abs/2307.01952
[Online]. Available: https://arxiv.org/abs/2307.01952
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.