Pith. sign in

REVIEW 4 major objections 6 minor 95 references

A Contrastive Pretrain Model with Prompt Tuning for Multi-center Medication Recommendation

T0 review · 4 major / 6 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read A shared pretrained medical encoder plus a small learnable prompt per hospital can serve many hospitals at once, with the largest accuracy gains where data is scarce.

desk verdict Solid empirical paper on a new multi-center medication recommendation setting, but the record-level split risks patient leakage and needs to be fixed before the SOTA claim is credible. read the letter →

arxiv 2412.20040 v1 pith:HRCWXEFZ submitted 2024-12-28 cs.IR

classification cs.IR
keywords medicationrecommendationmulti-centerlearningelectronichealthrecordscontrastivepretrainingprompttuningtransformerencodereICUdatasetcatastrophicforgetting
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper aims to show that medication recommendation can work across many hospitals at once, even when most hospitals have few records. It proposes TEMPT, a two-stage model: first pretrain a transformer encoder on diagnosis and procedure codes from all hospitals using masked-code prediction and a contrastive task that aligns diagnosis and procedure representations; then, for each hospital, freeze the pretrained weights and learn only a small hospital-specific prompt vector plus a prediction head. The authors claim this beats standard medication-recommendation models, multi-domain recommendation models, LLM baselines, and full finetuning on the eICU dataset, and that the gain is largest for small hospitals. A sympathetic reader should care because most hospitals lack enough data to train their own recommender, and the paper offers a way to share medical knowledge across hospitals without overwriting it.

What carries the argument

The load-bearing mechanism is a two-stage training scheme built on a shared transformer medical encoder. In pretraining, a mask prediction task randomly masks diagnosis and procedure codes and reconstructs them, capturing intra-set co-occurrence, while a contrastive task aligns the diagnosis representation and procedure representation of the same record against negatives from other records, capturing inter-set relationships. In the second stage, each hospital gets a small set of learnable prompt embeddings inserted at the start of the diagnosis and procedure sequences; the pretrained embedding matrices and encoder are frozen, and only the prompts and a medication-prediction MLP are updated. The prompt vector is what carries the hospital-specific adaptation without modifying the shared medical knowledge.

What would settle it

Re-run the comparison with the original baseline components intact, giving GAMENet its DDI graph and dynamic memory, COGNet its copy module, and G-Bert its historical medications, and check whether each still trails TEMPT on eICU. If any of these full baselines matches or exceeds TEMPT's PRAUC of 0.5468, the reported superiority would be an artifact of weakened competitors rather than a property of TEMPT's design.

Watch

Extended reading notes

Core claim

On the paper's own terms, the discovery is that hospital heterogeneity in prescribing can be handled by separating general medical knowledge from hospital-specific adaptation: pretrain on all hospitals with two self-supervised tasks, then adapt per hospital with continuous prompts rather than finetuning all parameters. The empirical claim is that TEMPT reaches PRAUC 0.5468, Jaccard 0.3318, and F1-score 0.4790 on eICU, with statistically significant improvements over the best baseline, and improves small-hospital PRAUC by 9.20% over Single-Train. The authors also claim the prompt-tuning stage is cheaper in both training time and storage than full finetuning and that it relieves catastrophic forgetting.

Load-bearing premise

The comparison against prior methods assumes the re-implemented baselines are faithful to their original designs, but several were run in stripped-down form: GAMENet without its drug-drug interaction graph and memory module, COGNet without its copy module, and G-Bert with procedures substituted for medications.

Editorial extensions

If this is right

  • If TEMPT's result holds, hospitals with fewer than 1,000 records gain the most from multi-center pretraining, with a 9.20% PRAUC improvement over training on their own data.
  • Freezing the shared medical encoder and learning only small prompts means the per-hospital adaptation cost is a small fraction of full finetuning, making deployment across dozens of hospitals feasible.
  • The two self-supervised pretraining tasks each contribute to the result; removing mask prediction or the contrastive task lowers performance, so learning relationships among diagnosis and procedure codes is part of the recipe.
  • The result suggests that a unified pretrained medication model can serve many hospitals, challenging the assumption that each hospital needs its own recommender trained from scratch.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The same pretrain-and-prompt recipe could carry over to other multi-site clinical prediction tasks, such as predicting in-hospital mortality or length of stay, where site differences are known to hurt a single global model.
  • Because the prompt embeddings are indexed by hospital, TEMPT opens a natural path to federated learning: hospitals could update only their own prompts locally and share the pretrained encoder without transferring raw patient data.
  • The paper's own stated limitation, ignoring drug-drug interactions, suggests a testable extension: add a DDI-constrained loss at the recommendation head and measure whether the multi-center gains survive when safety is enforced.
  • A direct testable implication is that the contrastive task's benefit should grow with distribution shift; comparing TEMPT's margin over pretraining without the contrastive task on hospital pairs with high versus low Jensen-Shannon divergence would isolate that effect.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. TEMPT is a two-stage model for multi-center (multi-hospital) medication recommendation. In stage 1, a shared transformer encoder is pretrained on diagnosis and procedure sets from all hospitals with a masked-token prediction loss and a contrastive alignment loss between diagnosis and procedure representations; the pretraining does not use medication labels. In stage 2, the encoder and code embeddings are frozen and, for each hospital, only a small set of prompt vectors and a medication-prediction MLP are trained. Experiments on eICU (102,363 records from 80 hospitals) compare TEMPT against medication recommendation baselines (Leap, GAMENet, G-Bert, COGNet), LLM baselines (GPT-4, TALLRec), multi-domain recommenders (MMOE, PLE, STAR), and TEMPT variants, using PRAUC, Jaccard, and F1. The paper reports statistically significant improvements over all baselines, with particularly large gains for small hospitals, plus ablations, PEFT comparisons, efficiency analysis, and hyperparameter sensitivity. Code is released.

Significance. If the reported results are valid, TEMPT makes a useful contribution: it identifies a realistic low-resource multi-center medication recommendation setting, shows that self-supervised pretraining on diagnosis and procedure co-occurrence transfers across hospitals, and demonstrates that per-hospital prompt tuning is more parameter-efficient than full finetuning. The paper is strong on reproducibility: code is public, five random splits and seeds are used, standard deviations and t-tests are reported, and the ablation and hyperparameter studies are reasonably comprehensive. The main caveat is that the headline SOTA claim rests on an evaluation protocol whose leakage risk is not addressed, a nonstandard PRAUC definition, and several heavily modified baselines; these issues affect the interpretation of nearly every table.

major comments (4)
  1. [Section 4.1.1, Table 2] The train/validation/test split is described only at the level of records (we divide the data into train/validation/test by the ratio of 8:1:1 for each hospital), and Table 2 reports 102,363 records but no patient count. eICU contains multiple ICU stays per patient, so a record-level split can place the same patient's records in both training and test sets. Because medication regimens are strongly patient-specific and repeat across stays, this can inflate all metrics in Tables 3-7 and may exaggerate TEMPT's advantage, especially given its pretraining and hospital-specific prompts. Please re-run the experiments with a patient-level split (or otherwise demonstrate that patients do not cross splits) and report the patient counts per split.
  2. [Section 4.1.4, Eq. (17)] The PRAUC definition is not the standard area under the precision-recall curve. Eq. (17) sums Precision_i times Recall_i over k=1 to |M| without defining a threshold index or a change in recall, whereas PR-AUC normally integrates Precision as a function of Recall (for example, average precision with (Recall_k - Recall_{k-1}) times Precision_k). Since PRAUC is the headline metric in Tables 3, 4, 5, and 7, please clarify the exact computation, including how thresholds are generated, and if the intention is a different metric, rename it and justify the choice.
  3. [Section 4.1.2, Table 3] The claim that TEMPT outperforms state-of-the-art medication recommendation models is weakened by the baseline adaptations: GAMENet is run without its DDI graph and dynamic memory module, COGNet without its copy module, and G-Bert with procedures substituted for medications. These are central components of the original methods, so the comparisons are against substantially altered versions. TEMPT's advantage over MMOE, PLE, and STAR is less affected, but the state-of-the-art medication recommendation wording should be revised, or faithful implementations should be provided.
  4. [Section 3.3.2, Eq. (9)] In the contrastive loss, the denominator sums over j not equal to i only and omits the positive pair exp(sim(u_i^d, u_i^p)/tau). Standard InfoNCE-style losses include the positive in the denominator. Please clarify whether this omission is intentional and how it affects the optimization; if it is a typo, correct the equation and re-run the pretraining experiments.
minor comments (6)
  1. [Algorithm 1, line 6] Algorithm 1 updates only E_d, E_p, and Theta_encoder, but the pretraining objective in Eq. (11) also contains the mask-prediction MLPs and the contrastive projectors; please list all updated parameters.
  2. [Section 3.5] The inference paragraph says the recommending probability is given as Equation (12) shows, but Eq. (12) is the non-prompted finetuning formulation; inference should use the prompted representation in Eq. (14).
  3. [Table 1] Table 1 lists E_p^h as prompt embedding matrices of diagnosis and medication; medication should be procedure.
  4. [Table 7, Section 4.7] Table 7 introduces TEMPT (finetune-freeze) without defining it in Section 4.7 or in the implementation details; please clarify what this variant is.
  5. [Figure 7] The Figure 7 caption says randomly sampled 8 hospitals, but the figure contains nine panels labeled (a) through (i).
  6. [Section 4.1.4] The text says the reported metrics are averaged over all of the records and hospitals; please specify whether the averaging is per record then per hospital or per record pooled across hospitals, since the two procedures weight hospitals differently.

Circularity Check

0 steps flagged · score 0.0 of 10

No circular reasoning found: TEMPT's pretraining tasks are label-free, prompt tuning is evaluated on held-out hospital records, and no claimed prediction reduces to a fitted input or self-citation chain.

full rationale

The paper's derivation chain is empirical rather than analytic, and none of its load-bearing steps reduces to its own inputs. In the pretraining stage, the mask prediction task reconstructs masked diagnosis and procedure codes (Eqs. 5-7) and the contrastive task aligns diagnosis/procedure representations within the same record (Eqs. 8-10); neither objective uses the medication labels y that are the target of the downstream recommendation task. In the prompt tuning stage, hospital-specific prompt embeddings O_h and MLP parameters Theta_MLP are trained on each hospital's training split and evaluated on that hospital's held-out test split, so the reported PRAUC, Jaccard, and F1 numbers in Tables 3-7 are not fit values renamed as predictions. Hyper-parameters (gamma, tau, b) are selected on validation performance, and results are averaged over five random split/seed runs. The self-citations in the related-work section (e.g., the authors' own LLM-distillation and MoE papers) are contextual and do not carry any uniqueness claim or theorem that forces the proposed design. The reviewers' concerns about a record-level rather than patient-level 8:1:1 split, and about modified baselines such as GAMENet without its DDI graph, are validity or fairness concerns, not circularity: they question whether the comparison is favorable, not whether the model's output is definitionally identical to a fitted input. No equation in the paper defines the prediction in terms of the target, and no fitted parameter is relabeled as a prediction. Therefore the appropriate finding is no significant circularity.

Assumptions & free parameters 5 free parameters · 4 assumptions · 0 invented entities

The model introduces no new physical entities; prompt embeddings are learned parameter matrices. The central claim depends on validation-selected hyper-parameters and on the eICU dataset as a proxy for multi-center practice, plus the fairness of the modified baselines.

free parameters (5)
  • Temperature of contrastive loss (tau) = 0.8
    Selected by validation performance; Figure 6b shows a peak at tau=0.8.
  • Weight of contrastive loss (gamma) = 1e-2
    Selected by validation performance; Figure 6a shows best PRAUC at 1e-2.
  • Number of prompt vectors (b) = 2
    Selected by validation performance; Figure 6c shows best PRAUC with b=2.
  • Recommendation threshold (t) = 0.3
    Chosen because the authors state 'this value performs better in total'; it is not tuned per model or per hospital.
  • Mask proportion in mask prediction task = not reported
    The paper randomly masks a proportion of diagnosis and procedure tokens but does not state the proportion, which is a free choice that can affect pretraining.
assumptions (4)
  • domain assumption eICU is representative of real-world multi-center hospital medication practices
    All conclusions rest on this single ICU database from 2014-2015; results may not generalize to other coding systems or outpatient settings.
  • domain assumption Diagnosis and procedure codes are sufficient inputs for medication recommendation
    The model ignores demographics, allergies, labs, and drug-drug interactions; the authors acknowledge this limitation in Section 6.
  • domain assumption Modified baselines retain the core capabilities of the original models
    GAMENet, COGNet, and G-Bert are altered to fit the instance-based setting; if these modifications weaken them, the SOTA comparison is invalid. Section 4.1.2.
  • domain assumption Random per-hospital 8:1:1 split produces independent training and test distributions
    No temporal ordering is applied, so patients with multiple visits could straddle the split and leak information. Section 4.1.1.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A Contrastive Pretrain Model with Prompt Tuning for Multi-center Medication Recommendation." pith.science (2026). https://pith.science/paper/HRCWXEFZ

@misc{pith2026241220040,
  author       = {Pith},
  title        = {Pith review of: A Contrastive Pretrain Model with Prompt Tuning for Multi-center Medication Recommendation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/HRCWXEFZ}},
  note         = {Machine review of arXiv:2412.20040}
}
read the original abstract

Medication recommendation is one of the most critical health-related applications, which has attracted extensive research interest recently. Most existing works focus on a single hospital with abundant medical data. However, many small hospitals only have a few records, which hinders applying existing medication recommendation works to the real world. Thus, we seek to explore a more practical setting, i.e., multi-center medication recommendation. In this setting, most hospitals have few records, but the total number of records is large. Though small hospitals may benefit from total affluent records, it is also faced with the challenge that the data distributions between various hospitals are much different. In this work, we introduce a novel conTrastive prEtrain Model with Prompt Tuning (TEMPT) for multi-center medication recommendation, which includes two stages of pretraining and finetuning. We first design two self-supervised tasks for the pretraining stage to learn general medical knowledge. They are mask prediction and contrastive tasks, which extract the intra- and inter-relationships of input diagnosis and procedures. Furthermore, we devise a novel prompt tuning method to capture the specific information of each hospital rather than adopting the common finetuning. On the one hand, the proposed prompt tuning can better learn the heterogeneity of each hospital to fit various distributions. On the other hand, it can also relieve the catastrophic forgetting problem of finetuning. To validate the proposed model, we conduct extensive experiments on the public eICU, a multi-center medical dataset. The experimental results illustrate the effectiveness of our model. The implementation code is available to ease the reproducibility https://github.com/Applied-Machine-Learning-Lab/TEMPT.

Figures

Figures reproduced from arXiv: 2412.20040 by the authors.

Figure 1
Figure 1. The histogram of medical record counts of each hospital and the histogram of the average number of [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. The heatmap to visualize Jensen–Shannon divergence of prescription distribution for all pairs of [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. The pretraining stage for TEMPT. 3.3 Pretraining During the pretraining stage, we train the model on the whole multi-center records V to learn the general medical knowledge from all hospitals. However, it is challenging to capture such knowledge because of the data sparsity [94]. Besides, the correlation between diagnosis and procedures has rarely been considered before. Therefore, we design two self-supervised task… view at source ↗
Figures from the paper (4 more)
Figure 4
Figure 4. Figure 4: The prompt tuning stage for TEMPT. In practice, we adopt the gradient descent method to train the whole model. The learnt general medical knowledge are embedded into the parameters E𝑑 , E𝑝 and Θ𝑒𝑛𝑐𝑜𝑑𝑒𝑟. 3.4 Prompt Tuning For recommending a proper medication set, we use…
Figure 5
Figure 5. Figure 5: The computation and storage cost of each model. [PITH_FULL_IMAGE:figures/full_fig_p017_5.png]
Figure 6
Figure 6. Figure 6: The performance of the model with various values of three hyper-parameters, [PITH_FULL_IMAGE:figures/full_fig_p018_6.png]
Figure 7
Figure 7. Figure 7: The performance on randomly sampled 8 hospitals. “rec” represents the record number of the hospital [PITH_FULL_IMAGE:figures/full_fig_p021_7.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

95 extracted references · 59 canonical work pages

  1. [1]

    Zafar Ali, Yi Huang, Irfan Ullah, Junlan Feng, Chao Deng, Nimbeshaho Thierry, Asad Khan, Asim Ullah Jan, Xiaoli Shen, Wu Rui, et al. 2023. Deep learning for medication recommendation: a systematic survey. Data Intelligence 5, 2 (2023), 303–354

  2. [2]

    Parham Abed Azad and Hamid Beigy. 2024. Multi-BERT: Leveraging Adapters and Prompt Tuning for Low-Resource Multi-Domain Adaptation. arXiv preprint arXiv:2404.02335 (2024)

  3. [3]

    Keqin Bao, Jizhi Zhang, Yang Zhang, Wenjie Wang, Fuli Feng, and Xiangnan He. 2023. Tallrec: An effective and efficient tuning framework to align large language model with recommendation. In Proceedings of the 17th ACM Conference on Recommender Systems. 1007–1014

  4. [4]

    Nick Barber, M Rawlins, and B Dean Franklin. 2003. Reducing prescribing error: competence, control, and culture. BMJ Quality & Safety 12, suppl 1 (2003), i29–i32

  5. [5]

    Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems 33 (2020), 1877–1901

  6. [6]

    Jiangxia Cao, Xixun Lin, Xin Cong, Jing Ya, Tingwen Liu, and Bin Wang. 2022. DisenCDR: Learning Disentangled Representations for Cross-Domain Recommendation. In Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval . 267–277

  7. [7]

    Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. 2020. A simple framework for contrastive learning of visual representations. In International conference on machine learning . PMLR, 1597–1607. 24 A Contrastive Pretrain Model with Prompt Tuning for Multi-center Medication Recommendation XX, June 03–05,2018, Woodstock, NY

  8. [8]

    Xinlei Chen, Haoqi Fan, Ross Girshick, and Kaiming He. 2020. Improved baselines with momentum contrastive learning. arXiv preprint arXiv:2003.04297 (2020)

Show all 95 references
  1. [9]

    Zeyu Cui, Jianxin Ma, Chang Zhou, Jingren Zhou, and Hongxia Yang. 2022. M6-Rec: Generative Pretrained Language Models are Open-Ended Recommender Systems. arXiv preprint arXiv:2205.08084 (2022)

  2. [10]

    Sunhao Dai, Ninglu Shao, Haiyuan Zhao, Weijie Yu, Zihua Si, Chen Xu, Zhongxiang Sun, Xiao Zhang, and Jun Xu. 2023. Uncovering chatgpt’s capabilities in recommender systems. In Proceedings of the 17th ACM Conference on Recommender Systems. 1126–1132

  3. [11]

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of naacL-HLT, Vol. 1. Minneapolis, Minnesota, 2

  4. [12]

    Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. 2020. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. In Internation...

  5. [13]

    Wenqi Fan, Tyler Derr, Xiangyu Zhao, Yao Ma, Hui Liu, Jianping Wang, Jiliang Tang, and Qing Li. 2021. Attacking black-box recommendations via copying cross-domain user profiles. In 2021 IEEE 37th international conference on data engineering (ICDE). IEEE, 1583–1594

  6. [14]

    Wenqi Fan, Xiangyu Zhao, Qing Li, Tyler Derr, Yao Ma, Hui Liu, Jianping Wang, and Jiliang Tang. 2023. Adversarial attacks for black-box recommender systems via copying transferable cross-domain user profiles. IEEE Transactions on Knowledge and Data Engineering 35, 12 (2023), 1...

  7. [15]

    Zichuan Fu, Xiangyang Li, Chuhan Wu, Yichao Wang, Kuicai Dong, Xiangyu Zhao, Mengchen Zhao, Huifeng Guo, and Ruiming Tang. 2023. A unified framework for multi-domain ctr prediction via large language models. ACM Transactions on Information Systems (2023)

  8. [16]

    Jingtong Gao, Bo Chen, Menghui Zhu, Xiangyu Zhao, Xiaopeng Li, Yuhao Wang, Yichao Wang, Huifeng Guo, and Ruiming Tang. 2024. HierRec: Scenario-Aware Hierarchical Modeling for Multi-scenario Recommendations. In Proceedings of the 33rd ACM International Conference on Information...

  9. [17]

    Jingtong Gao, Xiangyu Zhao, Bo Chen, Fan Yan, Huifeng Guo, and Ruiming Tang. 2023. AutoTransfer: Instance Transfer for Cross-Domain Recommendations. In Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval . 1478–1487

  10. [18]

    Shijie Geng, Shuchang Liu, Zuohui Fu, Yingqiang Ge, and Yongfeng Zhang. 2022. Recommendation as language processing (rlp): A unified pretrain, personalized prompt & predict paradigm (p5). In Proceedings of the 16th ACM Conference on Recommender Systems . 299–315

  11. [19]

    Fan Gong, Meng Wang, Haofen Wang, Sen Wang, and Mengyue Liu. 2021. SMR: medical knowledge graph embedding for safe medicine recommendation. Big Data Research 23 (2021), 100174

  12. [20]

    Bowen Hao, Jing Zhang, Hongzhi Yin, Cuiping Li, and Hong Chen. 2021. Pre-training graph neural networks for cold-start users and items representation. In Proceedings of the 14th ACM International Conference on Web Search and Data Mining. 265–273

  13. [21]

    Yupeng Hou, Shanlei Mu, Wayne Xin Zhao, Yaliang Li, Bolin Ding, and Ji-Rong Wen. 2022. Towards Universal Sequence Representation Learning for Recommender Systems. In Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining . 585–593

  14. [22]

    Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. 2019. Parameter-efficient transfer learning for NLP. In International conference on machine learning. PMLR, 2790–2799

  15. [23]

    Edward J Hu, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen, et al. 2021. LoRA: Low-Rank Adaptation of Large Language Models. In International Conference on Learning Representations

  16. [24]

    Menglin Jia, Luming Tang, Bor-Chun Chen, Claire Cardie, Serge Belongie, Bharath Hariharan, and Ser-Nam Lim. 2022. Visual prompt tuning. In European Conference on Computer Vision . Springer, 709–727

  17. [25]

    Pengyue Jia, Yichao Wang, Shanru Lin, Xiaopeng Li, Xiangyu Zhao, Huifeng Guo, and Ruiming Tang. 2024. D3: A Methodological Exploration of Domain Division, Modeling, and Balance in Multi-Domain Recommendations. In Proceedings of the AAAI Conference on Artificial Intelligence , ...

  18. [26]

    Yuchen Jiang, Qi Li, Han Zhu, Jinbei Yu, Jin Li, Ziru Xu, Huihui Dong, and Bo Zheng. 2022. Adaptive Domain Interest Network for Multi-domain Recommendation. In Proceedings of the 31st ACM International Conference on Information & Knowledge Management. 3212–3221

  19. [27]

    Alistair EW Johnson, Tom J Pollard, Lu Shen, Li-wei H Lehman, Mengling Feng, Mohammad Ghassemi, Benjamin Moody, Peter Szolovits, Leo Anthony Celi, and Roger G Mark. 2016. MIMIC-III, a freely accessible critical care database. Scientific data 3, 1 (2016), 1–9

  20. [28]

    Hung Le, Truyen Tran, and Svetha Venkatesh. 2018. Dual memory neural computer for asynchronous two-view sequential learning. In Proceedings of the 24th ACM SIGKDD international conference on knowledge discovery & data mining. 1637–1645. 25 XX, June 03–05,2018, Woodstock, NY Qi...

  21. [29]

    Penny J Lewis and Mary P Tully. 2009. Uncomfortable prescribing decisions in hospitals: the impact of teamwork. Journal of the Royal Society of Medicine 102, 11 (2009), 481–488

  22. [30]

    Lei Li, Yongfeng Zhang, and Li Chen. 2023. Personalized prompt learning for explainable recommendation. ACM Transactions on Information Systems 41, 4 (2023), 1–26

  23. [31]

    Muyang Li, Zijian Zhang, Xiangyu Zhao, Wanyu Wang, Minghao Zhao, Runze Wu, and Ruocheng Guo. 2023. Automlp: Automated mlp for sequential recommendations. In Proceedings of the ACM Web Conference 2023 . 1190–1198

  24. [32]

    Xinhang Li, Zhaopeng Qiu, Xiangyu Zhao, Zihao Wang, Yong Zhang, Chunxiao Xing, and Xian Wu. 2022. Gromov- wasserstein guided representation learning for cross-domain recommendation. In Proceedings of the 31st ACM Interna- tional Conference on Information & Knowledge Management...

  25. [33]

    Xinhang Li, Zhaopeng Qiu, Xiangyu Zhao, Yong Zhang, Chunxiao Xing, and Xian Wu. 2023. REST: Drug-Drug Inter- action Prediction via Reinforced Student-Teacher Curriculum Learning. In Proceedings of the 32nd ACM International Conference on Information and Knowledge Management . ...

  26. [34]

    Xiaopeng Li, Fan Yan, Xiangyu Zhao, Yichao Wang, Bo Chen, Huifeng Guo, and Ruiming Tang. 2023. HAMUR: Hyper Adapter for Multi-Domain Recommendation. In Proceedings of the 32nd ACM International Conference on Information and Knowledge Management. 1268–1277

  27. [35]

    Xinhang Li, Xiangyu Zhao, Yong Zhang, and Chunxiao Xing. 2023. Towards Automatic ICD Coding via Knowledge Enhanced Multi-Task Learning. In Proceedings of the 32nd ACM International Conference on Information and Knowledge Management. 1238–1248

  28. [36]

    Xiang Lisa Li and Percy Liang. 2021. Prefix-Tuning: Optimizing Continuous Prompts for Generation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Lo...

  29. [37]

    Vladislav Lialin, Vijeta Deshpande, and Anna Rumshisky. 2023. Scaling down to scale up: A guide to parameter-efficient fine-tuning. arXiv preprint arXiv:2303.15647 (2023)

  30. [38]

    Jiahao Liang, Xiangyu Zhao, Muyang Li, Zijian Zhang, Wanyu Wang, Haochen Liu, and Zitao Liu. 2023. Mmmlp: Multi-modal multilayer perceptron for sequential recommendations. In Proceedings of the ACM Web Conference 2023 . 1109–1117

  31. [39]

    Jianhua Lin. 1991. Divergence measures based on the Shannon entropy. IEEE Transactions on Information theory 37, 1 (1991), 145–151

  32. [40]

    Weilin Lin, Xiangyu Zhao, Yejing Wang, Yuanshao Zhu, and Wanyu Wang. 2023. Autodenoise: Automatic data instance denoising for recommendations. In Proceedings of the ACM Web Conference 2023 . 1003–1011

  33. [41]

    Langming Liu, Liu Cai, Chi Zhang, Xiangyu Zhao, Jingtong Gao, Wanyu Wang, Yifu Lv, Wenqi Fan, Yiqi Wang, Ming He, et al. 2023. LinRec: Linear Attention Mechanism for Long-term Sequential Recommender Systems. In Proceedings of the 46th International ACM SIGIR Conference on Rese...

  34. [42]

    Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. 2023. Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing. Comput. Surveys 55, 9 (2023), 1–35

  35. [43]

    Qidong Liu, Jiaxi Hu, Yutian Xiao, Xiangyu Zhao, Jingtong Gao, Wanyu Wang, Qing Li, and Jiliang Tang. 2024. Multimodal recommender systems: A survey. Comput. Surveys 57, 2 (2024), 1–17

  36. [44]

    Qidong Liu, Feng Tian, Qinghua Zheng, and Qianying Wang. 2023. Disentangling interest and conformity for eliminating popularity bias in session-based recommendation. Knowledge and Information Systems 65, 6 (2023), 2645–2664

  37. [45]

    Qidong Liu, Xian Wu, Xiangyu Zhao, Yuanshao Zhu, Derong Xu, Feng Tian, and Yefeng Zheng. 2024. When MOE Meets LLMs: Parameter Efficient Fine-tuning for Multi-task Medical Applications. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in...

  38. [46]

    Qidong Liu, Xian Wu, Xiangyu Zhao, Yuanshao Zhu, Zijian Zhang, Feng Tian, and Yefeng Zheng. 2024. Large Language Model Distilling Medication Recommendation Model. arXiv preprint arXiv:2402.02803 (2024)

  39. [47]

    Qidong Liu, Fan Yan, Xiangyu Zhao, Zhaocheng Du, Huifeng Guo, Ruiming Tang, and Feng Tian. 2023. Diffusion augmentation for sequential recommendation. In Proceedings of the 32nd ACM International Conference on Information and Knowledge Management. 1576–1586

  40. [48]

    Xiao Liu, Yanan Zheng, Zhengxiao Du, Ming Ding, Yujie Qian, Zhilin Yang, and Jie Tang. 2023. GPT understands, too. AI Open (2023)

  41. [49]

    Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. 2021. Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF International Conference on Computer Vision. 10012–10022

  42. [50]

    Ziru Liu, Jiejie Tian, Qingpeng Cai, Xiangyu Zhao, Jingtong Gao, Shuchang Liu, Dayou Chen, Tonghao He, Dong Zheng, Peng Jiang, et al. 2023. Multi-task recommendations with reinforcement learning. In Proceedings of the ACM Web Conference 2023. 1273–1282. 26 A Contrastive Pretra...

  43. [51]

    Jiaqi Ma, Zhe Zhao, Xinyang Yi, Jilin Chen, Lichan Hong, and Ed H Chi. 2018. Modeling task relationships in multi-task learning with multi-gate mixture-of-experts. In Proceedings of the 24th ACM SIGKDD international conference on knowledge discovery & data mining . 1930–1939

  44. [52]

    Arthur Mai, Karen Voigt, Jeannine Schübel, and Felix Gräßer. 2023. A drug recommender system for the treatment of hypertension. BMC Medical Informatics and Decision Making 23, 1 (2023), 89

  45. [53]

    Tom J Pollard, Alistair EW Johnson, Jesse D Raffa, Leo A Celi, Roger G Mark, and Omar Badawi. 2018. The eICU Collaborative Research Database, a freely available multi-center database for critical care research. Scientific data 5, 1 (2018), 1–13

  46. [54]

    Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019. Language models are unsupervised multitask learners. Technical Report (2019)

  47. [55]

    Atefeh Rajabalizadeh, Javad Norouzi Nia, Nima Safaei, Mojtaba Talafidaryani, Reyhaneh Bijari, Atousa Zarindast, Fateme Fotouhi, Masud Salehi, and Mahdi Moqri. 2020. An exploratory analysis of electronic intensive care unit (EICU) collaborative research database. (2020)

  48. [56]

    Vinay Venkatesh Ramasesh, Aitor Lewkowycz, and Ethan Dyer. 2021. Effect of scale on catastrophic forgetting in neural networks. In International Conference on Learning Representations

  49. [57]

    Matthew Reynolds, Seetal Jheeta, Jonathan Benn, Inderjit Sanghera, Ann Jacklin, Digby Ingle, and Bryony Dean Franklin. 2017. Improving feedback on junior doctors’ prescribing errors: mixed-methods evaluation of a quality improvement project. BMJ Quality & Safety 26, 3 (2017), 240–247

  50. [58]

    Nikunj Saunshi, Orestis Plevrakis, Sanjeev Arora, Mikhail Khodak, and Hrishikesh Khandeparkar. 2019. A theoretical analysis of contrastive unsupervised representation learning. In International Conference on Machine Learning . PMLR, 5628–5637

  51. [59]

    Tala B Shahin, Baran Balkan, Jarrod Mosier, Vignesh Subbian, et al. 2019. The connected intensive care unit patient: exploratory analyses and cohort discovery from a critical care telemedicine database. JMIR medical informatics 7, 1 (2019), e13006

  52. [60]

    Junyuan Shang, Tengfei Ma, Cao Xiao, and Jimeng Sun. 2019. Pre-training of graph augmented transformers for medication recommendation. In 28th International Joint Conference on Artificial Intelligence, IJCAI 2019 . International Joint Conferences on Artificial Intelligence, 5953–5959

  53. [61]

    Junyuan Shang, Cao Xiao, Tengfei Ma, Hongyan Li, and Jimeng Sun. 2019. Gamenet: Graph augmented memory networks for recommending medication combination. In proceedings of the AAAI Conference on Artificial Intelligence , Vol. 33. 1126–1133

  54. [62]

    Qijie Shen, Wanjie Tao, Jing Zhang, Hong Wen, Zulong Chen, and Quan Lu. 2021. SAR-Net: A scenario-aware ranking network for personalized fair recommendation in hundreds of travel scenarios. In Proceedings of the 30th ACM International Conference on Information & Knowledge Mana...

  55. [63]

    Xiang-Rong Sheng, Liqin Zhao, Guorui Zhou, Xinyao Ding, Binding Dai, Qiang Luo, Siran Yang, Jingshan Lv, Chi Zhang, Hongbo Deng, et al. 2021. One model to serve all: Star topology adaptive recommender for multi-domain ctr prediction. In Proceedings of the 30th ACM Internationa...

  56. [64]

    Yu Sun, Shuohuan Wang, Yukun Li, Shikun Feng, Hao Tian, Hua Wu, and Haifeng Wang. 2020. Ernie 2.0: A continual pre-training framework for language understanding. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 34. 8968–8975

  57. [65]

    Yanchao Tan, Chengjun Kong, Leisheng Yu, Pan Li, Chaochao Chen, Xiaolin Zheng, Vicki S Hertzberg, and Carl Yang

  58. [66]

    Yanchao Tan, Carl Yang, Xiangyu Wei, Chaochao Chen, Weiming Liu, Longfei Li, Jun Zhou, and Xiaolin Zheng. 2022. Metacare++: Meta-learning with hierarchical subtyping for cold-start diagnosis prediction in healthcare data. In Proceedings of the 45th International ACM SIGIR Conf...

  59. [67]

    Hongyan Tang, Junning Liu, Ming Zhao, and Xudong Gong. 2020. Progressive layered extraction (ple): A novel multi-task learning (mtl) model for personalized recommendations. In Fourteenth ACM Conference on Recommender Systems. 269–278

  60. [68]

    Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971 (2023)

  61. [69]

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in neural information processing systems 30 (2017)

  62. [70]

    Feng Wang and Huaping Liu. 2021. Understanding the behaviour of contrastive loss. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 2495–2504. 27 XX, June 03–05,2018, Woodstock, NY Qidong Liu et al

  63. [71]

    Yanda Wang, Weitong Chen, Dechang PI, Lin Yue, Sen Wang, and Miao Xu. 2021. Self-supervised adversarial distribution regularization for medication recommendation. In IJCAI International Joint Conference on Artificial Intelligence. International Joint Conferences on Artificial ...

  64. [72]

    Yejing Wang, Zhaocheng Du, Xiangyu Zhao, Bo Chen, Huifeng Guo, Ruiming Tang, and Zhenhua Dong. 2023. Single- shot feature selection for multi-task recommendations. In Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieva...

  65. [73]

    Yejing Wang, Shen Ge, Xiangyu Zhao, Xian Wu, Tong Xu, Chen Ma, and Zhi Zheng. 2023. Doctor specific tag recommendation for online medical record management. In Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining . 5150–5161

  66. [74]

    Yuhao Wang, Ziru Liu, Yichao Wang, Xiangyu Zhao, Bo Chen, Huifeng Guo, and Ruiming Tang. 2024. Diff-MSR: A Diffusion Model Enhanced Paradigm for Cold-Start Multi-Scenario Recommendation. In Proceedings of the 17th ACM International Conference on Web Search and Data Mining . 779–787

  67. [75]

    Yuhao Wang, Yichao Wang, Zichuan Fu, Xiangyang Li, Wanyu Wang, Yuyang Ye, Xiangyu Zhao, Huifeng Guo, and Ruiming Tang. 2024. Llm4msr: An llm-enhanced paradigm for multi-scenario recommendation. In Proceedings of the 33rd ACM International Conference on Information and Knowledg...

  68. [76]

    Yuhao Wang, Xiangyu Zhao, Bo Chen, Qidong Liu, Huifeng Guo, Huanshuo Liu, Yichao Wang, Rui Zhang, and Ruiming Tang. 2023. PLATE: A Prompt-Enhanced Paradigm for Multi-Scenario Recommendations. In Proceedings of the 46th International ACM SIGIR Conference on Research and Develop...

  69. [77]

    Max T Wayne, Sarah Seelye, Daniel Molling, Cainnear K Hogan, Thomas S Valley, Douglas A Arenberg, Jose De Carde- nas, and Hallie C Prescott. 2022. Variation in US hospital practices for bronchoscopy in the intensive care unit. Annals of the American Thoracic Society 19, 6 (202...

  70. [78]

    Jin Wen, Yongzhong Cheng, Xiuying Hu, Ping Yuan, Tianyou Hao, and Yingkang Shi. 2016. Workload, burnout, and medical mistakes among physicians in China: a cross-sectional study. Bioscience trends 10, 1 (2016), 27–33

  71. [79]

    Michael Wornow, Alejandro Lozano, Dev Dash, Jenelle Jindal, Kenneth W Mahaffey, and Nigam H Shah. 2024. Zero-shot clinical trial patient matching with llms. arXiv preprint arXiv:2402.05125 (2024)

  72. [80]

    Jiancan Wu, Xiang Wang, Fuli Feng, Xiangnan He, Liang Chen, Jianxun Lian, and Xing Xie. 2021. Self-supervised graph learning for recommendation. In Proceedings of the 44th international ACM SIGIR conference on research and development in information retrieval. 726–735

  73. [81]

    Likang Wu, Zhi Zheng, Zhaopeng Qiu, Hao Wang, Hongchao Gu, Tingjia Shen, Chuan Qin, Chen Zhu, Hengshu Zhu, Qi Liu, et al. 2024. A survey on large language models for recommendation. World Wide Web 27, 5 (2024), 60

  74. [82]

    Rui Wu, Zhaopeng Qiu, Jiacheng Jiang, Guilin Qi, and Xian Wu. 2022. Conditional Generation Net for Medication Recommendation. In Proceedings of the ACM Web Conference 2022 . 935–945

  75. [83]

    Derong Xu, Ziheng Zhang, Zhihong Zhu, Zhenxi Lin, Qidong Liu, Xian Wu, Tong Xu, Wanyu Wang, Yuyang Ye, Xiangyu Zhao, et al. 2024. Editing factual knowledge and explanatory ability of medical large language models. In Proceedings of the 33rd ACM International Conference on Info...

  76. [84]

    Derong Xu, Ziheng Zhang, Zhihong Zhu, Zhenxi Lin, Qidong Liu, Xian Wu, Tong Xu, Xiangyu Zhao, Yefeng Zheng, and Enhong Chen. 2024. Mitigating Hallucinations of Large Language Models in Medical Information Extraction via Contrastive Decoding. In Findings of the Association for ...

  77. [85]

    Chaoqi Yang, Cao Xiao, Lucas Glass, and Jimeng Sun. 2021. Change Matters: Medication Change Prediction with Recurrent Residual Networks. In 30th International Joint Conference on Artificial Intelligence, IJCAI 2021 . International Joint Conferences on Artificial Intelligence, ...

  78. [86]

    Chaoqi Yang, Cao Xiao, Fenglong Ma, Lucas Glass, and Jimeng Sun. 2021. SafeDrug: Dual Molecular Graph Encoders for Recommending Effective and Safe Drug Combinations. In 30th International Joint Conference on Artificial Intelligence, IJCAI 2021. International Joint Conferences ...

  79. [87]

    Yutao Zhang, Robert Chen, Jie Tang, Walter F Stewart, and Jimeng Sun. 2017. LEAP: learning to prescribe effective and safe treatment combinations for multimorbidity. In proceedings of the 23rd ACM SIGKDD international conference on knowledge Discovery and data Mining . 1315–1324

  80. [88]

    Yingying Zhang, Xian Wu, Quan Fang, Shengsheng Qian, and Chengsheng Xu. 2022. Knowledge-enhanced Attributed Multi-Task Learning for Medicine Recommendation. ACM Transactions on Information Systems (TOIS) (2022)

  81. [89]

    Zijian Zhang, Shuchang Liu, Jiaao Yu, Qingpeng Cai, Xiangyu Zhao, Chunxu Zhang, Ziru Liu, Qidong Liu, Hongwei Zhao, Lantao Hu, et al. 2024. M3oE: Multi-Domain Multi-Task Mixture-of Experts Recommendation Framework. In Proceedings of the 47th International ACM SIGIR Conference ...

  82. [90]

    Zijian Zhang, Xiangyu Zhao, Qidong Liu, Chunxu Zhang, Qian Ma, Wanyu Wang, Hongwei Zhao, Yiqi Wang, and Zitao Liu. 2023. PromptST: Prompt-Enhanced Spatio-Temporal Multi-Attribute Prediction. In Proceedings of the 32nd ACM International Conference on Information and Knowledge M...

  83. [91]

    Zhi Zheng, Zhaopeng Qiu, Hui Xiong, Xian Wu, Tong Xu, Enhong Chen, and Xiangyu Zhao. 2022. Ddr: Dialogue based doctor recommendation for online medical service. In Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining . 4592–4600

  84. [92]

    Zhi Zheng, Zhaopeng Qiu, Tong Xu, Xian Wu, Xiangyu Zhao, Enhong Chen, and Hui Xiong. 2022. CBR: context bias aware recommendation for debiasing user modeling and click prediction. In Proceedings of the ACM Web Conference

  85. [93]

    Zhi Zheng, Chao Wang, Tong Xu, Dazhong Shen, Penggang Qin, Xiangyu Zhao, Baoxing Huai, Xian Wu, and Enhong Chen. 2022. Interaction-aware Drug Package Recommendation via Policy Gradient. ACM Transactions on Information Systems (TOIS) (2022)

  86. [94]

    Kun Zhou, Hui Wang, Wayne Xin Zhao, Yutao Zhu, Sirui Wang, Fuzheng Zhang, Zhongyuan Wang, and Ji-Rong Wen. 2020. S3-rec: Self-supervised learning for sequential recommendation with mutual information maximization. In Proceedings of the 29th ACM International Conference on Info...

  87. [2022]

    In Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining

    4SDrug: Symptom-based Set-to-set Small and Safe Drug Recommendation. In Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining . 3970–3980

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.