REVIEW 4 major objections 6 minor 1 cited by
EmoStyle closes the test-time control gap in emotional art generation by turning compact prompts into structured affective plans that modulate style-specialist generators.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-14 13:47 UTC pith:HPSH6ATN
load-bearing objection Solid challenge-systems win with a clean style/affect split; the ablation does not cleanly prove residual affective modulation is the causal bridge. the 4 major comments →
EmoStyle: Affective Conditioning of Style-Specialist Experts for Emotional Image Generation
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper claims that an LLM-inferred affective plan—valence-arousal scores, dominant emotion, therapeutic-effect labels, and aspect ratio—encoded as a condition vector and injected into Z-Image denoising blocks through residual AdaLN-style scale-gate offsets, together with deterministic bucket-specific LoRA style experts and VLM-guided candidate selection, supplies the missing fine-grained controls that pure test prompts lack and thereby enables emotion-aware artistic generation that ranked first on Track 1 of the AffectiveArt Challenge 2026.
What carries the argument
Affective condition modulation: a sinusoidal valence-arousal coordinate embedding is fused with learnable dominant-emotion and therapeutic-label embeddings into a shared affective vector; residual heads then offset the frozen backbone’s native scale and gate parameters inside each block, while a dedicated LoRA expert per style bucket supplies style-specific color, texture, brushwork, and composition priors. Style selection and emotion expression stay separate.
Load-bearing premise
The method assumes that LLM-predicted valence, arousal, emotion, and therapeutic labels, when injected as residual modulation, are a faithful substitute for the missing training annotations rather than mostly soft enrichment whose gains partly come from style LoRAs and multi-sample selection alone.
What would settle it
Run a controlled ablation that freezes or randomizes the affective condition vector while keeping style LoRAs and VLM candidate ranking fixed; if FID and attribute alignment barely change, residual affective modulation is not causally carrying the claimed control.
If this is right
- Test-time emotion-aware art generation can be controlled without requiring the full EmoArt annotation suite at inference.
- Style and affect can be factorized: bucket LoRAs own style priors while residual modulation owns emotional expression.
- VLM multi-candidate ranking improves joint content-style-affect-quality satisfaction without extra training.
- Model-agnostic pieces (planning plus refinement) transfer to closed APIs; the full modulation stack transfers to open diffusion backbones.
- First-place Track 1 scores (overall 0.80, FID 66.12, AAS 0.99) show structured conditioning beats prompt-only baselines on this benchmark.
Where Pith is reading between the lines
- The same plan-then-modulate-then-rank pattern may transfer to other control-gap settings such as emotional video or style-plus-mood music generation.
- Zero-initialized residual modulation is a general recipe for adding continuous control axes to frozen diffusion transformers without immediately disturbing pretrained behavior.
- Aspect-ratio planning’s large FID effect in the ablations suggests layout is an under-appreciated affective channel that future art benchmarks should score explicitly.
- If LLM affective labels are noisy, the VLM judge may absorb much of the gain; isolating judge credit from modulation credit would clarify the causal chain.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. EmoStyle addresses emotion-aware artistic image generation under a test-time control gap: training data (EmoArt) provide valence–arousal, dominant emotion, therapeutic labels, visual attributes, and style buckets, but test inputs supply only a compact caption and a style bucket. The method (i) uses an LLM reasoner to predict compact affective fields and aspect ratio (§3.2), (ii) trains one LoRA expert per style bucket on frozen Z-Image (§3.3), (iii) encodes affective fields into a condition vector and injects residual scale/gate offsets into native Z-Image modulation (Eqs. 9–15, §3.4), and (iv) ranks multi-seed candidates with a VLM judge (§3.5). On AffectiveArt Challenge 2026 Track 1 the system ranks first (overall 0.80; FID 66.12; AAS 0.99). Local validation (Tables 2–3, Fig. 2) reports FID/AAS gains over base Z-Image and ablations of layout, planning, LoRA, and refinement.
Significance. If the causal story holds, the paper offers a practical and modular recipe for challenge-style affective artistic generation: separate style specialization (bucket LoRAs) from affect control (residual modulation), keep the backbone frozen, and use structured LLM planning plus test-time VLM selection. First place on Track 1 is a concrete, falsifiable outcome. The residual zero-init design and deterministic bucket-to-expert mapping are clean engineering choices that other controllable diffusion systems can reuse. The contribution is primarily systems/challenge-oriented rather than a new theoretical principle of affect in vision; its lasting value depends on whether residual affective modulation is shown to be more than soft enrichment plus multi-sample ranking.
major comments (4)
- Table 3 undercuts the central causal claim that LLM-inferred affective fields injected via residual modulation (§3.2–3.4, Eqs. 9–15) are the bridge for the missing EmoArt annotations. Removing prompt-to-affect planning improves AAS (0.7957 vs 0.7864), content, and style, while only hurting FID and Attribute. That pattern is consistent with style LoRAs, aspect-ratio layout, and VLM multi-sample selection carrying most of the win, with residual affect acting as a secondary regularizer (or mild distractor for automatic alignment). The paper should either (a) isolate residual modulation from the LLM plan (e.g., plan-as-text only vs residual injection with fixed plan; freeze LoRAs and ablate only M_y / g_ϕ), or (b) revise the claim to match the evidence that planning mainly helps FID/attribute realization rather than AAS/content/style.
- §3.4 never isolates the residual AdaLN-style path from the rest of the pipeline. Residual heads are zero-initialized and trained with frozen backbone and style LoRAs, so weak gradients leave the path near-identity. There is no ablation that keeps the LLM plan and style LoRAs fixed while removing only residual offsets (or replacing them with prompt-text serialization of the same fields). Without that control, the claim that injection into intermediate features (vs prompt enrichment) is load-bearing remains untested, especially given the closed-source backbone results in Table 2 where only plan serialization + VLM ranking are applied.
- §4.1 and Table 2: local AAS uses original MiniCPM-V-2.6 without the official EmoArt fine-tune used for leaderboard AAS. Content/Style/Attribute in Tables 2–3 are therefore not the same metric as the challenge AAS that underwrites first place. The paper should either re-evaluate local ablations with a closer proxy to the official evaluator, report correlation between local and official scores on a held-out set, or clearly demote local AAS components to development proxies and base causal claims primarily on FID plus official Track-1 metrics.
- No multi-seed variance, error bars, or fixed-seed protocol is reported for Tables 2–3 or the leaderboard numbers. Given that §3.5 explicitly relies on multi-candidate generation and VLM ranking, and FID is known to be seed-sensitive, the magnitude of gains (e.g., FID 75.84→58.64) cannot be assessed for stability. At minimum, report mean±std over several seeds for the full model and the two most contested ablations (w/o prompt-to-affect; w/o residual modulation if added).
minor comments (6)
- §3.4 / Eq. (6): the frequency set B and band count K are not specified numerically; free parameters (LoRA rank is given as 128, but K, λ_p/λ_s/λ_a/λ_q, candidate count, regeneration policy) should be listed in §4.2 for reproducibility.
- Figure 1 is described but the manuscript text does not specify how many candidates are generated per prompt or the regeneration threshold when the VLM judge rejects all samples.
- §4.1 Gongbi substitute protocol (China Image samples) should state whether those substitutes appear only in train/val or could leak style priors into the LoRA expert used at test time.
- Related work §2.2–2.3 is thorough; a short explicit comparison to EmoGen / CoEmoGen on what is unique about residual VA+label modulation versus emotion-token or prompt-only baselines would help non-challenge readers.
- Typographical consistency: “AdaLN-style” is used in the abstract/intro while §3.4 states residual offsets to native scale-gate modulation rather than an inserted AdaLN layer—align terminology to avoid overstating architectural novelty.
- Table 1 leaderboard: report whether all teams used the same resolution/aspect policy; otherwise FID comparisons across ranks are partly confounded by layout (your own ablation shows aspect-ratio removal hurts FID most).
Circularity Check
No circular derivation: empirical systems pipeline with external challenge metrics, not a self-defined or fitted-as-prediction result.
full rationale
EmoStyle is an engineering method paper, not a first-principles derivation. Its load-bearing claims are empirical: (i) style-bucket LoRAs trained with standard flow-matching loss on EmoArt subsets (Eqs. 3–5), (ii) residual affective modulation trained with the same loss while the backbone is frozen (Eqs. 9–15), and (iii) first place on the external AffectiveArt Challenge 2026 Track 1 leaderboard (Table 1). Training uses EmoArt’s provided style buckets and affective fields as supervised conditions; test-time LLM plans and VLM ranking are external model calls, not quantities fitted to the reported metrics and then re-labeled as predictions. EmoArt and the challenge evaluator are from other authors; self-citations (EmoVerse, CreatiParser, etc.) appear only as related-work context and do not force uniqueness or the central result. Ablation Table 3 may weaken causal attribution of the residual path, but that is an effectiveness/correctness concern, not circularity by construction. No step reduces a claimed prediction to its own inputs.
Axiom & Free-Parameter Ledger
free parameters (4)
- LoRA rank and training schedule =
rank 128; 3 epochs; 1e-4
- VA sinusoidal frequency set B / band count K =
K frequency bands (value not numerically fixed in text)
- VLM judge weights λ_p, λ_s, λ_a, λ_q
- Number of candidates / seeds and regeneration policy
axioms (5)
- domain assumption EmoArt style-bucket labels are a reliable specialization signal so one LoRA expert per bucket with deterministic mapping is sufficient (no learned router).
- domain assumption Compact LLM-inferred valence-arousal, dominant emotion, and therapeutic labels (without free-form visual attributes) are adequate control variables for emotional artistic expression at test time.
- ad hoc to paper Residual offsets to native Z-Image scale/gate modulation can inject affect while preserving frozen backbone and style LoRA behavior (zero-init residual heads).
- domain assumption A VLM judge scoring prompt/style/affect/quality is a valid surrogate for selecting emotionally and stylistically correct artworks.
- standard math Standard text-conditioned flow-matching loss on Z-Image is an appropriate training objective for both style LoRAs and affective residual heads.
invented entities (3)
-
Affective condition vector c_i,aff and shared encoder g_ϕ
no independent evidence
-
Bucket-specific residual modulation modules M_y
no independent evidence
-
Style-bucket LoRA expert set {E_y}
no independent evidence
read the original abstract
Emotion-aware artistic image generation requires an image to match the input prompt, follow the specified artistic style, and convey the target emotion. In this challenge, the main difficulty is that the visual and affective attributes available in the training data are not explicitly provided at test time. Without these attributes, the generator has to decide not only what to depict, but also how the target emotion should be expressed through color, lighting, brushwork, composition, line, and layout. This creates a control gap between the available test prompt and the fine-grained conditions needed for emotion-aware artistic generation. To bridge this gap, we propose EmoStyle, a Z-Image-based framework that converts the input prompt into a structured generation state. An LLM reasoner first predicts affective cues (valence-arousal, dominant emotion, and therapeutic-effect labels) and an aspect-ratio decision. Instead of using these predictions only as additional prompt text, we encode the affective fields into an affective condition vector and inject it into the denoising blocks through AdaLN-style modulation. This allows the inferred control variables to directly guide the generation of intermediate features. Since emotional expression is also style-dependent, we further train a dedicated LoRA adapter for each artistic style bucket and select the corresponding expert during inference, enabling the same affective cues to be rendered with bucket-specific priors for color, texture, brushwork, and composition. Finally, a lightweight VLM-guided candidate selection step ranks the generated images based on prompt alignment, style consistency, emotional expression, and visual quality. In Track 1 of the AffectiveArt Challenge 2026, our USTC\_PI\_LAB\_TEAM submission achieved first place.
Figures
Forward citations
Cited by 1 Pith paper
-
HTT-Net: Hierarchical Text-guided Transition Modeling for Surgical Video Phase Recognition
A hierarchical text-guided network that constructs and calibrates phase segments improves surgical phase recognition, setting a high Jaccard on Cholec80 and reporting large gains on a private LCRS-100 benchmark.
Reference graph
Works this paper leans on
-
[1]
Panos Achlioptas, Maks Ovsjanikov, Kilichbek Haydarov, Mohamed Elhoseiny, and Leonidas Guibas. 2021. ArtEmis: Affective Language for Visual Art. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. IEEE, Virtual Conference, 11569–11579. doi:10.1109/CVPR46437.2021.01140
-
[2]
2026.AffectiveArt Challenge 2026: Emotion- A ware Artistic Image Generation
AffectiveArt Challenge Organizers. 2026.AffectiveArt Challenge 2026: Emotion- A ware Artistic Image Generation. ACM Multimedia Grand Challenge. https: //affectiveart-challenge.github.io/
2026
-
[3]
Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. 2023. Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.arXiv preprint arXiv:2308.12966(2023). https://arxiv.org/abs/2308.12966
Pith/arXiv arXiv 2023
-
[4]
Le, Christo- pher Ré, and Azalia Mirhoseini
Bradley Brown, Jordan Juravsky, Ryan Ehrlich, Ronald Clark, Quoc V. Le, Christo- pher Ré, and Azalia Mirhoseini. 2024. Large Language Monkeys: Scaling Infer- ence Compute with Repeated Sampling.arXiv preprint arXiv:2407.21787(2024). https://arxiv.org/abs/2407.21787
Pith/arXiv arXiv 2024
-
[5]
Tingfeng Cao, Chengyu Wang, Bingyan Liu, Ziheng Wu, Jinhui Zhu, and Jun Huang. 2023. BeautifulPrompt: Towards Automatic Prompt Engineering for Text- to-Image Synthesis. InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing: Industry Track. Association for Computational Linguistics, Singapore, 1–11
2023
-
[6]
Dar-Yen Chen, Hamish Tennent, and Ching-Wen Hsu. 2023. ArtAdapter: Text-to- Image Style Transfer using Multi-Level Style Encoder and Explicit Adaptation. arXiv preprint arXiv:2312.02109(2023). https://arxiv.org/abs/2312.02109
Pith/arXiv arXiv 2023
-
[7]
Weidong Chen, Dexiang Hong, Zhendong Mao, Yutao Cheng, Xinyan Liu, Lei Zhang, and Yongdong Zhang. 2026. CreatiParser: Generative Image Parsing of Raster Graphic Designs into Editable Layers.arXiv preprint arXiv:2604.19632 (2026). https://arxiv.org/abs/2604.19632
Pith/arXiv arXiv 2026
-
[8]
Weidong Chen, Dexiang Hong, Yuankai Qi, Zhenjun Han, Shuhui Wang, Laiyun Qing, Qingming Huang, and Guorong Li. 2022. Multi-Attention Network for Compressed Video Referring Object Segmentation. InProceedings of the 30th ACM International Conference on Multimedia. Association for Computing Machinery, 4416–4425. doi:10.1145/3503161.3547761
-
[9]
Weidong Chen, Guorong Li, Xinfeng Zhang, Shuhui Wang, Liang Li, and Qing- ming Huang. 2023. Weakly Supervised Text-based Actor-Action Video Segmen- tation by Clip-level Multi-instance Learning.ACM Transactions on Multimedia Computing, Communications, and Applications19, 1, Article 12 (Jan. 2023), 12:1– 12:22 pages. doi:10.1145/3514250
-
[10]
Weidong Chen, Guorong Li, Xinfeng Zhang, Hongyang Yu, Shuhui Wang, and Qingming Huang. 2021. Cascade Cross-modal Attention Network for Video Actor and Action Segmentation from a Sentence. InProceedings of the 29th ACM International Conference on Multimedia. Association for Computing Machinery, 4053–4062. doi:10.1145/3474085.3475534
-
[11]
Weidong Chen, Cheng Ye, Zhendong Mao, Peipei Song, Xinyan Liu, Lei Zhang, Xiaojun Chang, and Yongdong Zhang. 2026. FACE-net: Factual Calibration and Emotion Augmentation for Retrieval-enhanced Emotional Video Captioning. arXiv preprint arXiv:2603.17455(2026). https://arxiv.org/abs/2603.17455
arXiv 2026
-
[12]
Weidong Chen, Cheng Ye, Zhendong Mao, Liping Wang, Xinyan Liu, and Yong- dong Zhang. 2026. Towards Accurate Emotion-Attributed Video Captioning via Fine-grained Emotion-Cause Pair Extraction.arXiv preprint arXiv:2606.08566 (2026). https://arxiv.org/abs/2606.08566
Pith/arXiv arXiv 2026
-
[13]
Weidong Chen, Cheng Ye, Peipei Song, Lei Zhang, Yongdong Zhang, and Zhen- dong Mao. 2026. Subjective-Objective Emotion-Correlated Generation Network for Subjective Video Captioning.IEEE Transactions on Image Processing35 (2026), 540–555. doi:10.1109/TIP.2025.3649363
-
[14]
Siddhartha Datta, Alexander Ku, Deepak Ramachandran, and Peter Anderson
-
[15]
https://arxiv.org/abs/2312.16720
Prompt Expansion for Adaptive Text-to-Image Generation.arXiv preprint arXiv:2312.16720(2023). https://arxiv.org/abs/2312.16720
Pith/arXiv arXiv 2023
-
[16]
Prafulla Dhariwal and Alexander Nichol. 2021. Diffusion Models Beat GANs on Image Synthesis.Advances in Neural Information Processing Systems34 (2021), 8780–8794
2021
-
[17]
Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas Müller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, Dustin Podell, Tim Dockhorn, Zion English, Kyle Lacey, Alex Goodwin, Yannik Marek, and Robin Rombach. 2024. Scaling Rectified Flow Transformers for High-Resolution Image Synthesis.arXiv preprint arXiv:2403.03206...
Pith/arXiv arXiv 2024
-
[18]
William Fedus, Jeff Dean, and Barret Zoph. 2022. A Review of Sparse Expert Models in Deep Learning.arXiv preprint arXiv:2209.01667(2022). https://arxiv. org/abs/2209.01667
Pith/arXiv arXiv 2022
-
[19]
William Fedus, Barret Zoph, and Noam Shazeer. 2022. Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity.Journal of Machine Learning Research23, 120 (2022), 1–39. https://jmlr.org/papers/v23/21- 0998.html
2022
-
[20]
Fengyi Fu, Shancheng Fang, Weidong Chen, and Zhendong Mao. 2024. Sentiment- Oriented Transformer-Based Variational Autoencoder Network for Live Video Commenting.ACM Transactions on Multimedia Computing, Communications, and Applications20, 4, Article 104 (2024), 104:1–104:24 pages. doi:10.1145/3633334
-
[21]
Bermano, Gal Chechik, and Daniel Cohen-Or
Rinon Gal, Yuval Alaluf, Yuval Atzmon, Or Patashnik, Amit H. Bermano, Gal Chechik, and Daniel Cohen-Or. 2023. An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion. InInternational Conference MM ’26, November 10–14, 2026, Rio de Janeiro, Brazil Hong et al. on Learning Representations. Kigali, Rwanda. https://openreview...
2023
-
[22]
Google. 2026. Gemini 3.5 Flash. https://ai.google.dev/gemini-api/docs/models/ gemini-3.5-flash. Gemini API documentation. Model ID: gemini-3.5-flash; last updated June 24, 2026; accessed June 25, 2026
2026
-
[23]
Xin Gu, Congcong Li, Xinyao Wang, Dexiang Hong, Libo Zhang, Tiejian Luo, Longyin Wen, and Heng Fan. 2026. Structured Context Learning for Generic Event Boundary Detection. InProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision. 4808–4817
2026
-
[24]
Yuchao Gu, Xintao Wang, Jay Zhangjie Wu, Yujun Shi, Yunpeng Chen, Zihan Fan, Wuyou Xiao, Rui Zhao, Shuning Chang, Weijia Wu, Yixiao Ge, Ying Shan, and Mike Zheng Shou. 2023. Mix-of-Show: Decentralized Low-Rank Adaptation for Multi-Concept Customization of Diffusion Models. InAdvances in Neural Infor- mation Processing Systems. Neural Information Processin...
2023
-
[25]
Yijie Guo, Dexiang Hong, Weidong Chen, Zihan She, Cheng Ye, Xiaojun Chang, and Zhendong Mao. 2025. EmoVerse: A MLLMs-Driven Emotion Representation Dataset for Interpretable Visual Emotion Analysis.arXiv preprint arXiv:2511.12554 (2025). https://arxiv.org/abs/2511.12554
Pith/arXiv arXiv 2025
-
[26]
Jingxuan Han, Quan Wang, Licheng Zhang, Weidong Chen, Yan Song, and Zhen- dong Mao. 2023. Text Style Transfer with Contrastive Transfer Pattern Mining. InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, Toronto, Canada, 7914–7927. doi:10.18653/v1/202...
-
[27]
Yaru Hao, Zewen Chi, Li Dong, and Furu Wei. 2023. Optimizing Prompts for Text-to-Image Generation. InAdvances in Neural Information Processing Systems. Neural Information Processing Systems Foundation, New Orleans, LA, USA
2023
-
[28]
Amir Hertz, Andrey Voynov, Shlomi Fruchter, and Daniel Cohen-Or. 2024. Style Aligned Image Generation via Shared Attention. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. IEEE, Seattle, WA, USA, 4775–4785
2024
-
[29]
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. 2017. GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium. InAdvances in Neural Information Processing Systems, Vol. 30. Curran Associates, Inc., Long Beach, California, USA, 6626–6637
2017
-
[30]
Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020. Denoising Diffusion Probabilistic Models. InAdvances in Neural Information Processing Systems. Curran Associates, Inc., Virtual Conference
2020
-
[31]
Dexiang Hong, Congcong Li, Longyin Wen, Xinyao Wang, and Libo Zhang
-
[32]
Generic Event Boundary Detection Challenge at CVPR 2021 Techni- cal Report: Cascaded Temporal Attention Network (CastaNet).arXiv preprint arXiv:2107.00239(2021)
Pith/arXiv arXiv 2021
-
[33]
Dexiang Hong, Guorong Li, Kai Xu, Li Su, and Qingming Huang. 2021. Siamese Dynamic Mask Estimation Network for Fast Video Object Segmentation. In2020 25th International Conference on Pattern Recognition (ICPR). IEEE, 9476–9482
2021
-
[34]
Dexiang Hong, Guorong Li, Bineng Zhong, Zhenjun Han, Li Su, and Qingming Huang. 2022. CRNet: Collaborative Refinement Network for Self-Supervised Video Object Segmentation. In2022 IEEE 5th International Conference on Multi- media Information Processing and Retrieval (MIPR). IEEE, 172–177
2022
-
[35]
Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen
Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022. LoRA: Low-Rank Adaptation of Large Language Models. InInternational Conference on Learning Representations. Virtual Conference. https://openreview.net/forum?id=nZeVKeeFYf9
2022
-
[36]
Xiwei Hu, Haokun Chen, Zhongqi Qi, Hui Zhang, Dexiang Hong, Jie Shao, and Xinglong Wu. 2025. DreamPoster: A Unified Framework for Image-Conditioned Generative Poster Design.arXiv preprint arXiv:2507.04218(2025)
Pith/arXiv arXiv 2025
-
[37]
Chengsong Huang, Qian Liu, Bill Yuchen Lin, Tianyu Pang, Chao Du, and Min Lin. 2024. LoraHub: Efficient Cross-Task Generalization via Dynamic LoRA Composition. InFirst Conference on Language Modeling. https://openreview.net/ forum?id=TrloAXEJ2B
2024
-
[38]
Xiaoyu Huang, Weidong Chen, Bo Hu, and Zhendong Mao. 2025. Graph Mixture of Experts and Memory-Augmented Routers for Multivariate Time Series Anom- aly Detection. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 39. 17476–17484. doi:10.1609/AAAI.V39I16.33921
-
[39]
Yuda Jin, Weidong Chen, Yuanhe Tian, Yan Song, and Chenggang Yan. 2024. Improving Radiology Report Generation with Multi-Grained Abnormality Pre- diction.Neurocomputing600 (2024), 128122. doi:10.1016/j.neucom.2024.128122
-
[40]
Yuda Jin, Weidong Chen, Yuanhe Tian, Yan Song, Chenggang Yan, and Zhendong Mao. 2024. Improving Radiology Report Generation with𝐷 2-Net: When Diffusion Meets Discriminator. InICASSP 2024 – 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2215–2219. https:// ieeexplore.ieee.org/document/10448326
arXiv 2024
-
[41]
Alvarez, Adria Recasens, and Agata Lapedriza
Ronak Kosti, Jose M. Alvarez, Adria Recasens, and Agata Lapedriza. 2020. Context Based Emotion Recognition using EMOTIC Dataset.IEEE Transactions on Pattern Analysis and Machine Intelligence42, 11 (2020), 2755–2766. doi:10.1109/TPAMI. 2019.2916866
doi:10.1109/tpami 2020
-
[42]
Nupur Kumari, Bingliang Zhang, Richard Zhang, Eli Shechtman, and Jun-Yan Zhu
-
[43]
InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
Multi-Concept Customization of Text-to-Image Diffusion. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. IEEE, Vancouver, BC, Canada, 1931–1941
1931
-
[44]
Congcong Li, Xinyao Wang, Dexiang Hong, Yufei Wang, Libo Zhang, Tiejian Luo, and Longyin Wen. 2022. Structured Context Transformer for Generic Event Boundary Detection.arXiv preprint arXiv:2206.02985(2022)
Pith/arXiv arXiv 2022
-
[45]
Congcong Li, Xinyao Wang, Longyin Wen, Dexiang Hong, Tiejian Luo, and Libo Zhang. 2022. End-to-end Compressed Video Representation Learning for Generic Event Boundary Detection. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 13967–13976
2022
-
[46]
Dengchun Li, Yingzi Ma, Naizheng Wang, Zhengmao Ye, Zhiyuan Cheng, Yinghao Tang, Yan Zhang, Lei Duan, Jie Zuo, Cal Yang, and Mingjie Tang. 2024. MixLoRA: Enhancing Large Language Models Fine-Tuning with LoRA-based Mixture of Experts.arXiv preprint arXiv:2404.15159(2024). https://arxiv.org/abs/2404.15159
Pith/arXiv arXiv 2024
-
[47]
Guorong Li, Dexiang Hong, Kai Xu, Bineng Zhong, Li Su, Zhenjun Han, and Qingming Huang. 2022. Self Supervised Progressive Network for High Perfor- mance Video Object Segmentation.IEEE Transactions on Neural Networks and Learning Systems35, 6 (2022), 7671–7684
2022
-
[48]
Jingyu Li, Zhendong Mao, Hao Li, Weidong Chen, and Yongdong Zhang. 2024. Exploring Visual Relationships via Transformer-Based Graphs for Enhanced Image Captioning.ACM Transactions on Multimedia Computing, Communications, and Applications20, 5, Article 133 (2024), 133:1–133:23 pages. doi:10.1145/3638558
-
[49]
Zhe Li, Lei Zhang, Kun Zhang, Weidong Chen, Yongdong Zhang, and Zhendong Mao. 2025. Rethinking Pseudo Word Learning in Zero-Shot Composed Image Retrieval: From an Object-Aware Perspective. InProceedings of the 48th Interna- tional ACM SIGIR Conference on Research and Development in Information Retrieval. Association for Computing Machinery, 833–843. doi:1...
-
[50]
Zefeng Lin, Weidong Chen, Yan Song, and Yongdong Zhang. 2024. Prompting Few- Shot Multi-Hop Question Generation via Comprehending Type-Aware Semantics. InFindings of the Association for Computational Linguistics: NAACL 2024. Associ- ation for Computational Linguistics, 3730–3740. doi:10.18653/v1/2024.findings- naacl.236
-
[51]
Chang Liu, Yuanhe Tian, Weidong Chen, Yan Song, and Yongdong Zhang. 2024. Bootstrapping Large Language Models for Radiology Report Generation. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 38. 18635–18643. doi:10.1609/AAAI.V38I17.29826
-
[52]
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023. Visual In- struction Tuning. InAdvances in Neural Information Processing Systems. Curran Associates, Inc., New Orleans, LA, USA, 34892–34916. doi:10.52202/075280-1516
-
[53]
Xinyan Liu, Weidong Chen, Zhaobo Qi, Beichen Zhang, and Weigang Zhang
-
[54]
Matching Street View and Satellite Images via Drone Imagery and Semantic Descriptions. InUA VM 2025 – Proceedings of the 3rd International Workshop on UA Vs in Multimedia: Capturing the World from a New Perspective, Co-located with MM 2025. Association for Computing Machinery, 4–9. doi:10.1145/3728482. 3757391
-
[55]
Zihan Meng, Dexiang Hong, Weidong Chen, Ziyu Zhou, Bo Hu, and Zhendong Mao. 2026. Audio-Visual Exchange-Aware Token Pruning for Efficient Audio- Visual Captioning.arXiv preprint arXiv:2606.10533(2026). https://arxiv.org/abs/ 2606.10533
Pith/arXiv arXiv 2026
-
[56]
William Peebles and Saining Xie. 2023. Scalable Diffusion Models with Trans- formers. InProceedings of the IEEE/CVF International Conference on Computer Vision. IEEE, Paris, France, 4195–4205
2023
-
[57]
Xilin Qin, Dexiang Hong, Weidong Chen, Cheng Ye, Xinyan Liu, Peipei Song, and Lei Zhang. 2025. Query-Based Collaborative Multimodal Token Pruning for Audio-Visual Question Answering. In2025 4th International Conference on Artificial Intelligence, Human-Computer Interaction and Robotics (AIHCIR). IEEE. doi:10.1109/AIHCIR67580.2025.11405267
-
[58]
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021. Learning Transferable Visual Models From Natural Language Supervision. InProceedings of the 38th International Conference on Machine Learning. PMLR, Virtual Conferen...
2021
-
[59]
Carlos Riquelme, Joan Puigcerver, Basil Mustafa, Maxim Neumann, Rodolphe Jenatton, André Susano Pinto, Daniel Keysers, and Neil Houlsby. 2021. Scaling Vision with Sparse Mixture of Experts.Advances in Neural Information Processing Systems34 (2021), 8583–8595
2021
-
[60]
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. 2022. High-Resolution Image Synthesis with Latent Diffusion Models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. IEEE, New Orleans, LA, USA, 10684–10695. doi:10.1109/CVPR52688.2022.01042
-
[61]
Nataniel Ruiz, Yuanzhen Li, Varun Jampani, Yael Pritch, Michael Rubinstein, and Kfir Aberman. 2023. DreamBooth: Fine Tuning Text-to-Image Diffusion Models for Subject-Driven Generation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. IEEE, Vancouver, BC, Canada, 22500– 22510. EmoStyle: Affective Conditioning of Style-S...
2023
-
[62]
Sara Mah- davi, Raphael Gontijo-Lopes, Tim Salimans, Jonathan Ho, David J
Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, S. Sara Mah- davi, Raphael Gontijo-Lopes, Tim Salimans, Jonathan Ho, David J. Fleet, and Mohammad Norouzi. 2022. Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding. InAdvances in Neural Informatio...
2022
-
[63]
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean. 2017. Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer. InInternational Conference on Learning Representations. https://openreview.net/forum?id=B1ckMDqlg
2017
-
[64]
Charlie Snell, Jaehoon Lee, Kelvin Xu, and Aviral Kumar. 2024. Scaling LLM Test- Time Compute Optimally can be More Effective than Scaling Model Parameters. arXiv preprint arXiv:2408.03314(2024). https://arxiv.org/abs/2408.03314
Pith/arXiv arXiv 2024
-
[65]
Peipei Song, Zhiyan Zhang, Weidong Chen, Jinpeng Hu, Xun Yang, and Xiaojun Chang. 2026. Bridging Subjectivity in Affective Explanation Captioning via Consensus-Prompted Emotion Reasoning.IEEE Transactions on Image Processing PP (2026). Online ahead of print; IEEE Xplore document 11574588. doi:10.1109/ TIP.2026.3703790
arXiv 2026
-
[66]
Yuanhe Tian, Weidong Chen, Bo Hu, Yan Song, and Fei Xia. 2023. End-to-end Aspect-based Sentiment Analysis with Combinatory Categorial Grammar. In Findings of the Association for Computational Linguistics: ACL 2023. Association for Computational Linguistics, Toronto, Canada, 13597–13609. doi:10.18653/v1/ 2023.findings-acl.859
doi:10.18653/v1/ 2023
-
[67]
Chuang Wang, Weidong Chen, Xu Cui, Yiming Zhao, Zhaobo Qi, Pengqi Huang, Xinyan Liu, and Weigang Zhang. 2025. Combatting Data Imbalance and Noise in Micro-Action Recognition. InMM 2025 – Proceedings of the 33rd ACM International Conference on Multimedia, Co-Located with MM 2025. Association for Computing Machinery, 14229–14235. doi:10.1145/3746027.3762088
-
[68]
Liping Wang, Cheng Ye, Weidong Chen, Peipei Song, Bo Hu, and Zhendong Mao. 2026. A Multi-Agent Framework with Structured Reasoning and Reflective Refinement for Multimodal Empathetic Response Generation.arXiv preprint arXiv:2604.18988(2026). https://arxiv.org/abs/2604.18988
Pith/arXiv arXiv 2026
-
[69]
Ting Wang, Weidong Chen, Yuanhe Tian, Yan Song, and Zhendong Mao. 2023. Improving Image Captioning via Predicting Structured Concepts. InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, 360–370. doi:10.18653/v1/2023.emnlp- main.25
-
[70]
Xinyao Wang, Longyin Wen, Congcong Li, and Dexiang Hong. 2025. Feature Extraction Method and Apparatus for Video, Slicing Method and Apparatus for Video, and Electronic Device and Storage Medium. US Patent App. 18/837,577
2025
-
[71]
Yunlong Wang, Shuyuan Shen, and Brian Y. Lim. 2023. RePrompt: Automatic Prompt Editing to Refine AI-Generative Art Towards Precise Expressions. In Proceedings of the CHI Conference on Human Factors in Computing Systems. Asso- ciation for Computing Machinery, Hamburg, Germany, 1–29
2023
-
[72]
Longyin Wen, Xinyao Wang, Dexiang Hong, and Congcong Li. 2025. Informa- tion Segmentation Methods, Apparatuses, and Electronic Devices. US Patent 12,481,631
2025
-
[73]
Xun Wu, Shaohan Huang, and Furu Wei. 2024. Mixture of LoRA Experts.arXiv preprint arXiv:2404.13628(2024). https://arxiv.org/abs/2404.13628
Pith/arXiv arXiv 2024
-
[74]
Jingyuan Yang, Jiawei Feng, and Hui Huang. 2024. EmoGen: Emotional Im- age Content Generation with Text-to-Image Diffusion Models.arXiv preprint arXiv:2401.04608(2024). https://arxiv.org/abs/2401.04608
Pith/arXiv arXiv 2024
-
[75]
Jingyuan Yang, Qirui Huang, Tingting Ding, Dani Lischinski, Daniel Cohen-Or, and Hui Huang. 2023. EmoSet: A Large-scale Visual Emotion Dataset with Rich Attributes. InProceedings of the IEEE/CVF International Conference on Computer Vision. 20326–20337. doi:10.1109/ICCV51070.2023.01864
-
[76]
Yang Yang, Wen Wang, Liang Peng, Chaotian Song, Yao Chen, Hengjia Li, Xiao- long Yang, Qinglin Lu, Deng Cai, Boxi Wu, and Wei Liu. 2025. LoRA-Composer: Leveraging Low-Rank Adaptation for Multi-Concept Customization in Training- Free Diffusion Models.IEEE Transactions on Image Processing34 (2025), 8145–8158. doi:10.1109/TIP.2025.3633153
-
[77]
Yuan Yao, Tianyu Yu, Ao Zhang, Chongyi Wang, Junbo Cui, Hongji Zhu, Tianchi Cai, Haoyu Li, Weilin Zhao, Zhihui He, et al . 2024. MiniCPM-V: A GPT-4V Level MLLM on Your Phone.arXiv preprint arXiv:2408.01800(2024). https: //arxiv.org/abs/2408.01800
Pith/arXiv arXiv 2024
-
[78]
Cheng Ye, Weidong Chen, Bo Hu, Lei Zhang, Yongdong Zhang, and Zhendong Mao. 2025. Improving Video Summarization by Exploring the Coherence Between Corresponding Captions.IEEE Transactions on Image Processing34 (2025), 5369–
2025
-
[79]
doi:10.1109/TIP.2025.3598709
-
[80]
Cheng Ye, Weidong Chen, Jingyu Li, Lei Zhang, and Zhendong Mao. 2024. Dual- path Collaborative Generation Network for Emotional Video Captioning. In Proceedings of the 32nd ACM International Conference on Multimedia. Association for Computing Machinery, 496–505. doi:10.1145/3664647.3681603
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.