REVIEW 3 major objections 5 minor 61 references
PRAC improves personalized aesthetic rating prediction by mining preference-rich images and merging similar users' models.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 22:25 UTC pith:H46E2NR7
load-bearing objection Solid PIAA method with consistent gains, but the profile-dependent PDM is undefined for three of the four benchmarks, so the reported results are not yet reproducible. the 3 major comments →
Personalized Image Aesthetic Assessment via Preference-rich Sample Mining and Cohort Merging
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
PRAC, a multimodal large language model personalized aesthetic assessment framework, reports state-of-the-art Spearman rank correlation on all four benchmarks. On PARA it reaches 0.707 (10-shot) and 0.733 (100-shot), beating previous bests of 0.702 and 0.716; on FLICKR-AES 0.692/0.778 versus 0.668/0.748; on AADB 0.597/0.671; and on the cross-database REAL-CUR test 0.585/0.631. The claim is that these gains come from two targeted operations: selecting images whose ratings best expose a user's taste (high collective controversy plus high personalized deviation from public consensus), and merging the fine-tuned models of a small cohort of users whose preference embeddings are most similar, whil
What carries the argument
The engine is a two-metric selection score, Pscore, combined with a cohort-merging rule. Pscore combines CCM, the standard deviation of the model's predicted public aesthetic rating distribution, with PDM, the KL divergence between the generic distribution and the distribution the model gives when the prompt is conditioned on the user's textual profile (age, gender, expertise, Big-Five traits). The selected high-scoring images form the user's query set. Each user is then represented by a Fisher Information Matrix computed from the gradients of a LoRA fine-tuned on that query set; cosine similarity between these matrices measures cross-user taste similarity. PreferMerge selects K cohort membe
Load-bearing premise
The Personalized Deviation Metric assumes that prompting the generic predictor with a user's demographic and personality profile yields a rating distribution close enough to that user's real judgments to identify which images are actually informative.
What would settle it
Swap user profiles between users during PreferSelect: if SRCC does not drop meaningfully, PDM is not using profile information and the reported gains must come from somewhere else. A cleaner direct check is to have users rate a large pool of images and test whether PDM-ranked images match images where users actually disagree.
If this is right
- With as few as 10 user-rated images, PRAC beats previous personalized aesthetic assessment methods on PARA, FLICKR-AES, and AADB in the paper's experiments.
- At 100 shots, the margin over prior best grows on all four benchmarks, indicating that both sample selection and cohort merging scale with data.
- PreferSelect alone improves over randomly sampled fine-tuning; adding PreferMerge adds a further gain on every dataset, so both mechanisms are load-bearing.
- The cross-database results show that a model trained on one dataset's users transfers to unseen users on other datasets, suggesting the pipeline captures transferable preference structure.
- Because the MLLM is prompted with a user profile, PRAC can provide a written rationale for a rating, making the prediction interpretable.
Where Pith is reading between the lines
- If profile-conditioned distributions approximate true user taste, the same PDM-style metric could rank content for other subjective tasks such as recommendation or moderation, where disagreement is informative.
- A natural extension is an online version: update a user's Fisher-embedding as new ratings arrive, recompute the cohort, and re-merge LoRAs, turning the few-shot model into a continuously adapting one.
- Because Fisher-information similarity does not depend on the concrete rating scale, the cohort-merging step could in principle operate across datasets with different scales, enabling cross-platform personalization.
- A direct test would compare cohort selection by Fisher-embedding similarity against selection by observed agreement on a held-out set of divisive images; if the two disagree often, the embedding similarity may be capturing model artifacts rather than taste.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes PRAC, an MLLM-based framework for personalized image aesthetic assessment (PIAA). PRAC has two main components: PreferSelect, which ranks a user's available annotated images by a preference-richness score combining Collective Controversy (CCM) and Personalized Deviation (PDM), and PreferMerge, which finds 'aesthetically-resonant' users via Fisher-information-derived preference embeddings and merges their LoRA adapters to improve the target user's model. The method is evaluated on PARA, FLICKR-AES, AADB, and REAL-CUR in 10-shot and 100-shot settings, reporting SRCC gains over prior PIAA methods, with ablations and computational-cost comparisons. The central claim is that PRAC consistently achieves the best performance across all four benchmarks.
Significance. If the method is fully specified and the reported gains are reproducible, PRAC would be a meaningful advance in few-shot PIAA: it reframes sample selection as a dual-metric preference-richness problem and introduces cohort-based LoRA merging, which is a timely and interpretable use of MLLMs for aesthetic personalization. The paper provides broad comparisons on four datasets, ablations of both components, qualitative rationales for divergent judgments, and a computational-cost analysis. The idea that certain images are more informative for personalization, and that similar users can be found from gradient-based preference embeddings, is well motivated by cognitive-aesthetics literature. However, the main results currently rest on a reproducibility gap: PDM is defined only when rich user profiles exist, yet three of the four benchmark datasets lack such profiles, and the manuscript does not explain how PreferSelect operates there. This must be resolved before the cross-dataset claims can be accepted.
major comments (3)
- [Sec. 3.2, Eqs. (3)–(4); Tables 3–6] PDM is defined in Eq. (3) as the KL divergence between the generic aesthetic distribution and a distribution obtained from the profile-conditioned prompt 'You are a <User Profile>'. Section 4.1 states that only PARA provides rich user profiles; FLICKR-AES, AADB, and REAL-CUR do not. Nevertheless, Tables 3–5 report PRAC on all three datasets, and Table 6 attributes substantial gains to PreferSelect on those datasets. The manuscript does not specify how PDM is computed in the absence of profiles. If PDM is omitted, Pscore in Eq. (4) reduces to (1−α)·CCM, making the tuned α=0.3 from Figure 5 meaningless for those datasets; if an implicit proxy is used, it is not described. As written, the method cannot be applied to three of the four evaluation benchmarks, so the abstract's claim of 'consistent' state-of-the-art performance and the attribution of gains to the dual-metric PreferSelect are no
- [Sec. 4.3, Figures 5–6] The hyperparameters α, K, and β are selected via ablations on PARA, but no held-out validation split is described. Since PARA has fixed testing users (Table 2), tuning on the same testing distribution risks selection on the evaluation set. In addition, no error bars, standard deviations, or significance tests are reported on any main table. The baseline in Table 6 is said to be averaged over 10 random runs, but the variance is not given. This matters for small margins such as the 10-shot PARA comparison (PRAC 0.707 vs PIAA-MIR 0.702) and the ablation increment 0.640 vs 0.634 in Table 6. Please report mean±std over repeated trials and define a validation protocol for hyperparameter selection.
- [Sec. 3.3, Eqs. (5)–(8)] PreferMerge requires solving a combinatorial subset-selection problem over the User Pool in Eq. (7), but the manuscript does not state how this optimization is performed in practice. It also does not specify whether the Fisher Information Matrix F_u is approximated by a diagonal or empirical Fisher, how the preference embedding is reduced to a comparable vector, or how the merging weights w_i in Eq. (8) are computed from the similarities in Eq. (6). These implementation details are necessary to reproduce the cohort-merging step, particularly on PARA with 398 training users, and to understand the computational cost reported in Figure 9.
minor comments (5)
- [Table 2 caption] The rows 'PARA (unconditional)', 'PARA (artistic)', 'PARA (photographic)', and 'PARA (personality)' are listed among methods but appear to be variants or protocol settings, not prior PIAA methods. Please clarify in the caption or re-organize the table.
- [Eq. (7)] The subset notation should use S ⊆ P rather than S ⊂ P, since S can equal the full pool if K equals |P|.
- [References] Reference [16] gives the arXiv ID as '2307.117606'; the correct EmotionPrompt ID appears to be '2307.11760'. Please verify.
- [Sec. 4.1, REAL-CUR] For REAL-CUR, the table shows no training images/users, so the evaluation is inherently cross-database. The description in Section 4.1 could state this more explicitly, since the later cross-database experiment relies on it.
- [Sec. 5] The conclusion calls PRAC 'the first MLLM-based PIAA model'; this is a strong novelty claim that is not formally supported by the related-work survey. Consider softening it.
Circularity Check
No circularity identified: PRAC's reported SRCC values are external held-out evaluations, not quantities defined from the selection or merging metrics by construction.
full rationale
The paper's derivation chain is: (1) train a Generic Aesthetic Predictor with KL loss on distribution labels (Eq. 1); (2) PreferSelect computes CCM as the standard deviation of the predicted distribution (Eq. 2) and PDM as the KL divergence between the generic distribution and the profile-conditioned distribution (Eq. ̃3), combined into Pscore via Eq. 4; (3) selected samples are used to fine-tune user-specific LoRAs and compute FIM preference embeddings (Eq. 5) and cosine similarities (Eq. 6); (4) a cohort is selected by Eq. 7 and merged via Eq. 8. The final evaluated quantities are SRCC on held-out test images (Tables 2–5), not the Pscore, CCM, PDM, FIM, or cosine-similarity values themselves. No equation defines the predicted rating as equal to a fitted selection score, so the reported gains do not reduce to the method's inputs by construction. The ablations in Table 6 also show that removing PreferSelect or PreferMerge decreases SRCC, which would not happen if the final metric were definitionally identical to a fixed component. The close self-references—PDM uses the same generic predictor that is later personalized, and several related-work citations are from the same group—are not load-bearing: PDM is a sample-selection heuristic, and the cited datasets/methods (PARA, MTCL, etc.) are external benchmarks. The concern that PDM is unspecified for FLICKR-AES, AADB, and REAL-CUR is a reproducibility/correctness gap, not a circularity, because it does not make the reported predictions equal to the inputs by definition. No uniqueness theorem or author-imported ansatz is invoked to force the architecture. Therefore no significant circularity is present.
Axiom & Free-Parameter Ledger
free parameters (3)
- alpha (alpha) =
0.3
- cohort size K =
6
- beta (beta) =
0.5
axioms (5)
- domain assumption MLLM in-context prompting with a textual user profile produces an aesthetic distribution approximating that user's taste.
- domain assumption FIM of LoRA gradients on a few annotated images is a faithful preference embedding for cross-user similarity.
- domain assumption Linearly merging LoRA weights of aesthetically similar users transfers preference patterns to the target user.
- ad hoc to paper Images with high collective controversy or high personalized deviation are more informative for fine-tuning a user model.
- domain assumption The generic aesthetic predictor's distribution predictions are well-calibrated enough that STD and KL on them are meaningful.
read the original abstract
Personalized Image Aesthetic Assessment (PIAA) aims to predict aesthetic ratings of images that vary across individuals. The aesthetic preferences manifest to different extents across distinct visual stimuli and exhibit cohort-specific patterns. Motivated by the above fact, this paper presents a Multimodal Large Language Model (MLLM)-based approach, which models individual aesthetic preferences by Preference-Rich sample mining and Aesthetically-resonant Cohort merging (PRAC). Specifically, PRAC first identifies preference-rich samples by analyzing both Collective Controversy and Personalized Deviation of images, maximizing the utility of limited user data. Based upon the preference-rich samples, cross-user preference similarities are measured by comparing preference embeddings. Then, a cohort-based model merging strategy, is proposed by aggregating preference patterns from aesthetically-resonant users, which further enhances the personalization for the target individual. Extensive experiments and comparisons on four benchmark PIAA databases demonstrate the superiority of the proposed PRAC model over the state-of-the-arts. The code and model will be public at https://github.com/yzc-ippl/PRAC.
Figures
Reference graph
Works this paper leans on
-
[1]
Shun-Ichi Amari. 1998. Natural gradient works efficiently in learning.Neural computation10, 2 (1998), 251–276
1998
-
[2]
Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, et al. 2023. Qwen technical report.arXiv preprint arXiv:2309.16609(2023)
Pith/arXiv arXiv 2023
-
[3]
Elif Celikors and David J Field. 2025. Beauty is in the eye of your cohort: Structured individual differences allow predictions of individualized aesthetic ratings of images.Cognition256 (2025), 106036
2025
-
[4]
Li-Wei Chen, Ombretta Strafforello, Anne-Sofie Maerten, Tinne Tuytelaars, and Johan Wagemans. 2025. On the Role of Individual Differences in Current Ap- proaches to Computational Image Aesthetics.arXiv preprint arXiv:2502.20518 (2025)
arXiv 2025
-
[5]
Zijie Chen, Lichao Zhang, Fangsheng Weng, Lili Pan, and Zhenzhong Lan. 2024. Tailored visions: Enhancing text-to-image generation with personalized prompt rewriting. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 7727–7736
2024
-
[6]
Yubin Deng, Chen Change Loy, and Xiaoou Tang. 2017. Image aesthetic assess- ment: An experimental survey.IEEE Signal Processing Magazine34, 4 (2017), 80–106
2017
-
[7]
Qingxiu Dong, Lei Li, Damai Dai, Ce Zheng, Jingyuan Ma, Rui Li, Heming Xia, Jingjing Xu, Zhiyong Wu, Baobao Chang, et al. 2024. A Survey on In-context Learning. InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 1107–1128
2024
-
[8]
Maszalida Hamzah. 2012. Objectifying subjectivity in images through aesthetics: A Bergsonian approach. In2012 International Conference on Innovation Manage- ment and Technology Research. IEEE, 247–252
2012
-
[9]
Jingwen Hou, Weisi Lin, Guanghui Yue, Weide Liu, and Baoquan Zhao. 2022. Interaction-matrix based personalized image aesthetics assessment.IEEE Trans- actions on Multimedia25 (2022), 5263–5278
2022
-
[10]
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen, et al. 2022. Lora: Low-rank adaptation of large language models.ICLR1, 2 (2022), 3
2022
-
[11]
Yipo Huang, Xiangfei Sheng, Zhichao Yang, Quan Yuan, Zhichao Duan, Pengfei Chen, Leida Li, Weisi Lin, and Guangming Shi. 2024. Aesexpert: Towards multi- modality foundation model for image aesthetics perception. InProceedings of the 32nd ACM International Conference on Multimedia. 5911–5920
2024
-
[12]
Yipo Huang, Quan Yuan, Xiangfei Sheng, Zhichao Yang, Haoning Wu, Pengfei Chen, Yuzhe Yang, Leida Li, and Weisi Lin. 2024. Aesbench: An expert benchmark for multimodal large language models on image aesthetics perception.arXiv preprint arXiv:2401.08276(2024)
Pith/arXiv arXiv 2024
-
[13]
Guangyuan Jiang, Manjie Xu, Song-Chun Zhu, Wenjuan Han, Chi Zhang, and Yixin Zhu. 2023. Evaluating and Inducing Personality in Pre-trained Language Models. InAdvances in Neural Information Processing Systems, A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine (Eds.), Vol. 36. Curran Associates, Inc., 10622–10643
2023
-
[14]
Won-Hee Kim, Jun-Ho Choi, and Jong-Seok Lee. 2018. Objectivity and subjectivity in aesthetic quality assessment of digital photographs.IEEE Transactions on Affective Computing11, 3 (2018), 493–506
2018
-
[15]
Shu Kong, Xiaohui Shen, Zhe Lin, Radomir Mech, and Charless Fowlkes. 2016. Photo aesthetics ranking network with attributes and content adaptation. In European conference on computer vision. Springer, 662–679
2016
-
[16]
Cheng Li, Jindong Wang, Kaijie Zhu, Yixuan Zhang, Wenxin Hou, Jianxun Lian, and Xing Xie. 2023. Emotionprompt: Leveraging psychology for large language models enhancement via emotional stimulus.arXiv preprint arXiv:2307.117606 (2023)
Pith/arXiv arXiv 2023
-
[17]
Leida Li, Xiangfei Sheng, Pengfei Chen, Jinjian Wu, and Weisheng Dong. 2024. Towards explainable image aesthetics assessment with attribute-oriented cri- tiques generation.IEEE Transactions on Circuits and Systems for Video Technology 35, 2 (2024), 1464–1477
2024
-
[18]
Leida Li, Hancheng Zhu, Sicheng Zhao, Guiguang Ding, and Weisi Lin. 2020. Personality-Assisted Multi-Task Learning for Generic and Personalized Image Aesthetics Assessment.IEEE Transactions on Image Processing29 (2020), 3898–
2020
-
[19]
Weishi Li, Yong Peng, Miao Zhang, Liang Ding, Han Hu, and Li Shen. 2023. Deep model fusion: A survey.arXiv preprint arXiv:2309.15698(2023)
Pith/arXiv arXiv 2023
-
[20]
Yaohui Li, Yuzhe Yang, Huaxiong Li, Haoxing Chen, Liwu Xu, Leida Li, Yaqian Li, and Yandong Guo. 2022. Transductive aesthetic preference propagation for personalized image aesthetics assessment. InProceedings of the 30th ACM International Conference on Multimedia. 896–904
2022
-
[21]
Jiahong Liu, Zexuan Qiu, Zhongyang Li, Quanyu Dai, Jieming Zhu, Minda Hu, Menglin Yang, and Irwin King. 2025. A survey of personalized large language models: Progress and future directions.arXiv preprint arXiv:2502.11528(2025)
arXiv 2025
-
[22]
Qijiong Liu, Jieming Zhu, Yanting Yang, Quanyu Dai, Zhaocheng Du, Xiao-Ming Wu, Zhou Zhao, Rui Zhang, and Zhenhua Dong. 2024. Multimodal pretraining, adaptation, and generation for recommendation: A survey. InProceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. 6566–6576
2024
-
[23]
Zhenyi Lu, Chenghao Fan, Wei Wei, Xiaoye Qu, Dangyang Chen, and Yu Cheng
-
[24]
Pei Lv, Jianqi Fan, Xixi Nie, Weiming Dong, Xiaoheng Jiang, Bing Zhou, Mingliang Xu, and Changsheng Xu. 2021. User-guided personalized image aesthetic assess- ment based on deep reinforcement learning.IEEE Transactions on Multimedia25 (2021), 736–749
2021
-
[25]
Pei Lv, Meng Wang, Yongbo Xu, Ze Peng, Junyi Sun, Shimei Su, Bing Zhou, and Mingliang Xu. 2018. USAR: An interactive user-specific aesthetic ranking framework for images. InProceedings of the 26th ACM international conference on Multimedia. 1328–1336
2018
-
[26]
Anne-Sofie Maerten, Li-Wei Chen, Stefanie De Winter, Christophe Bossens, and Johan Wagemans. 2025. LAPIS: A novel dataset for personalized image aes- thetic assessment. InProceedings of the Computer Vision and Pattern Recognition Conference. 6302–6311
2025
-
[27]
Naila Murray, Luca Marchesotti, and Florent Perronnin. 2012. AVA: A large-scale database for aesthetic visual analysis. In2012 IEEE conference on computer vision and pattern recognition. IEEE, 2408–2415
2012
-
[28]
Shijia Ni, Feng Shao, Xiongli Chai, Hangwei Chen, and Yo-Sung Ho. 2022. Composition-guided neural network for image cropping aesthetic assessment. IEEE Transactions on Multimedia25 (2022), 6836–6851
2022
-
[29]
Jian Ren, Xiaohui Shen, Zhe Lin, Radomir Mech, and David J Foran. 2017. Per- sonalized image aesthetics. InProceedings of the IEEE international conference on computer vision. 638–647
2017
-
[30]
Evan F Risko, Nicola C Anderson, Sophie Lanthier, and Alan Kingstone. 2012. Cu- rious eyes: Individual differences in personality predict eye movement behavior in scene-viewing.Cognition122, 1 (2012), 86–90
2012
-
[31]
Xiangfei Sheng, Zhichao Duan, Xiaofeng Pan, Yipo Huang, Zhichao Yang, Pengfei Chen, and Leida Li. 2026. TuningIQA: Fine-grained blind image quality assess- ment for livestreaming camera tuning. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 40. 17679–17687
2026
-
[32]
Xiangfei Sheng, Xiaofeng Pan, Zhichao Yang, Pengfei Chen, and Leida Li. 2026. Fine-grained image quality assessment for perceptual image restoration. InPro- ceedings of the AAAI Conference on Artificial Intelligence, Vol. 40. 8914–8922
2026
-
[33]
Xiangfei Sheng, Pangu Xie, Weidong Zou, Pengfei Chen, Tong Zhu, and Leida Li. 2025. InstructCrop: Teaching Multimodal Large Language Models to Crop Aesthetic Images. InProceedings of the 33rd ACM International Conference on Multimedia. 6830–6839
2025
-
[34]
Zhaoxuan Tan, Qingkai Zeng, Yijun Tian, Zheyuan Liu, Bing Yin, and Meng Jiang
-
[35]
Anke Tang, Li Shen, Yong Luo, Yibing Zhan, Han Hu, Bo Du, Yixin Chen, and Dacheng Tao. 2023. Parameter efficient multi-task model fusion with partial linearization.arXiv preprint arXiv:2310.04742(2023)
Pith/arXiv arXiv 2023
-
[36]
Guolong Wang, Junchi Yan, and Zheng Qin. 2018. Collaborative and Attentive Learning for Personalized Image Aesthetic Assessment.. InIJCAI. 957–963
2018
-
[37]
InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing
Democratizing large language models via personalized parameter-efficient fine-tuning. InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 6476–6491
2024
-
[38]
Shoya Washizu, Yoshia Abe, Tatsuya Daikoku, and Yasuo Kuniyoshi. 2025. Bod- ily sensations, emotions, and personality traits in the aesthetic experience of everyday photographs.Scientific Reports(2025)
2025
-
[39]
Yuxiang Wei, Yiheng Zheng, Yabo Zhang, Ming Liu, Zhilong Ji, Lei Zhang, and Wangmeng Zuo. 2025. Personalized image generation with deep generative models: A decade survey.arXiv preprint arXiv:2502.13081(2025)
Pith/arXiv arXiv 2025
-
[40]
Weiyun Wang, Zhangwei Gao, Lixin Gu, Hengjun Pu, Long Cui, Xingguang Wei, Zhaoyang Liu, Linglin Jing, Shenglong Ye, Jie Shao, et al. 2025. Internvl3. 5: Ad- vancing open-source multimodal models in versatility, reasoning, and efficiency. arXiv preprint arXiv:2508.18265(2025)
Pith/arXiv arXiv 2025
-
[41]
Haoning Wu, Zicheng Zhang, Weixia Zhang, Chaofeng Chen, Liang Liao, Chunyi Li, Yixuan Gao, Annan Wang, Erli Zhang, Wenxiu Sun, et al. 2024. Q-Align: Teach- ing LMMs for Visual Scoring via Discrete Text-Defined Levels. InInternational Conference on Machine Learning. PMLR, 54015–54029. MM ’26, November 10–November 14, 2026, Rio de Janeiro, Brazil. Zhichao Y...
2024
-
[42]
Yiyan Xu, Wenjie Wang, Yang Zhang, Biao Tang, Peng Yan, Fuli Feng, and Xiangnan He. 2025. Personalized image generation with large multimodal models. InProceedings of the ACM on Web Conference 2025. 264–274
2025
-
[43]
Chengyue Wu, Teng Wang, Yixiao Ge, Zeyu Lu, Ruisong Zhou, Ying Shan, and Ping Luo. 2023. pi-Tuning: Transferring Multimodal Foundation Models with Optimal Multi-task Interpolation. InInternational Conference on Machine Learning. PMLR, 37713–37727
2023
-
[44]
Yuzhe Yang, Liwu Xu, Leida Li, Nan Qie, Yaqian Li, Peng Zhang, and Yandong Guo. 2022. Personalized image aesthetics assessment with rich attributes. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 19861–19869
2022
-
[45]
Zhichao Yang, Tianjiao Gu, Jianjie Wang, Feiyu Lin, Xiangfei Sheng, Pengfei Chen, and Leida Li. 2026. Longt2ibench: A benchmark for evaluating long text- to-image generation with graph-structured annotations. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 40. 11820–11828
2026
-
[46]
Enneng Yang, Li Shen, Guibing Guo, Xingwei Wang, Xiaochun Cao, Jie Zhang, and Dacheng Tao. 2024. Model merging in llms, mllms, and beyond: Methods, theories, applications and opportunities.arXiv preprint arXiv:2408.07666(2024)
Pith/arXiv arXiv 2024
-
[47]
Zhichao Yang, Leida Li, Pengfei Chen, Jinjian Wu, and Giuseppe Valenzise. 2025. Language-Guided Visual Perception Disentanglement for Image Quality Assess- ment and Conditional Image Generation.arXiv preprint arXiv:2503.02206(2025)
Pith/arXiv arXiv 2025
-
[48]
Zhichao Yang, Leida Li, Yuzhe Yang, Yaqian Li, and Weisi Lin. 2024. Multi-Level Transitional Contrast Learning for Personalized Image Aesthetics Assessment. IEEE Transactions on Multimedia26 (2024), 1944–1956
2024
-
[49]
Zhichao Yang, Leida Li, Pengfei Chen, Jinjian Wu, and Weisheng Dong. 2024. Semantics-aware image aesthetics assessment using tag matching and contrastive ranking. InProceedings of the 32nd ACM International Conference on Multimedia. 2632–2641
2024
-
[50]
Jiabo Ye, Haiyang Xu, Haowei Liu, Anwen Hu, Ming Yan, Qi Qian, Ji Zhang, Fei Huang, and Jingren Zhou. 2024. mplug-owl3: Towards long image- sequence understanding in multi-modal large language models.arXiv preprint arXiv:2408.04840(2024)
Pith/arXiv arXiv 2024
-
[51]
Zhiyuan You, Xin Cai, Jinjin Gu, Tianfan Xue, and Chao Dong. 2025. Teaching large language models to regress accurate image quality scores using score distri- bution. InProceedings of the Computer Vision and Pattern Recognition Conference. 14483–14494
2025
-
[52]
Zhichao Yang, Jianjie Wang, Zhixianhe Zhang, Pangu Xie, Xiangfei Sheng, Pengfei Chen, and Leida Li. 2026. Fine-grained Image Aesthetic Assessment: Learning Dis- criminative Scores from Relative Ranks. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 145–155
2026
-
[53]
Kai Zhang, Yejin Kim, and Xiaozhong Liu. 2024. Personalized llm response generation with parameterized memory injection.arXiv preprint arXiv:2404.03565 (2024)
Pith/arXiv arXiv 2024
-
[54]
Mingrui Zhang, Mading Li, Jiahao Yu, and Li Chen. 2022. Aesthetic photo collage with deep reinforcement learning.IEEE Transactions on Multimedia25 (2022), 4653–4664
2022
-
[55]
Jooyeol Yun and Jaegul Choo. 2024. Scaling Up Personalized Image Aesthetic Assessment via Task Vector Customization. InEuropean Conference on Computer Vision. Springer, 323–339
2024
-
[56]
Hancheng Zhu, Leida Li, Jinjian Wu, Sicheng Zhao, Guiguang Ding, and Guang- ming Shi. 2020. Personalized image aesthetics assessment via meta-learning with bilevel gradient optimization.IEEE Transactions on Cybernetics52, 3 (2020), 1798–1811
2020
-
[57]
Hancheng Zhu, Yong Zhou, Leida Li, Yaqian Li, and Yandong Guo. 2021. Learning personalized image aesthetics from subjective and objective attributes.IEEE Transactions on Multimedia25 (2021), 179–190
2021
-
[58]
Zhehao Zhang, Ryan A Rossi, Branislav Kveton, Yijia Shao, Diyi Yang, Hamed Zamani, Franck Dernoncourt, Joe Barrow, Tong Yu, Sungchul Kim, et al. 2024. Per- sonalization of large language models: A survey.arXiv preprint arXiv:2411.00027 (2024)
Pith/arXiv arXiv 2024
-
[61]
Hancheng Zhu, Yong Zhou, Zhiwen Shao, Wenliang Du, Guangcheng Wang, and Qiaoyue Li. 2022. Personalized image aesthetics assessment via multi-attribute interactive reasoning.Mathematics10, 22 (2022), 4181
2022
-
[2024]
Advances in Neural Information Processing Systems37 (2024), 78905–78935
Twin-merging: Dynamic integration of modular expertise in model merging. Advances in Neural Information Processing Systems37 (2024), 78905–78935
2024
-
[3910]
doi:10.1109/TIP.2020.2968285
arXiv 2020
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.