REVIEW 2 cited by
HistLLM: A Unified Framework for LLM-Based Multimodal Recommendation with User History Encoding and Compression
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
HistLLM: A Unified Framework for LLM-Based Multimodal Recommendation with User History Encoding and Compression
read the original abstract
While large language models (LLMs) have proven effective in leveraging textual data for recommendations, their application to multimodal recommendation tasks remains relatively underexplored. Although LLMs can process multimodal information through projection functions that map visual features into their semantic space, recommendation tasks often require representing users' history interactions through lengthy prompts combining text and visual elements, which not only hampers training and inference efficiency but also makes it difficult for the model to accurately capture user preferences from complex and extended prompts, leading to reduced recommendation performance. To address this challenge, we introduce HistLLM, an innovative multimodal recommendation framework that integrates textual and visual features through a User History Encoding Module (UHEM), compressing multimodal user history interactions into a single token representation, effectively facilitating LLMs in processing user preferences. Extensive experiments demonstrate the effectiveness and efficiency of our proposed mechanism.
Forward citations
Cited by 2 Pith papers
-
HTT-Net: Hierarchical Text-guided Transition Modeling for Surgical Video Phase Recognition
A hierarchical text-guided network that constructs and calibrates phase segments improves surgical phase recognition, setting a high Jaccard on Cholec80 and reporting large gains on a private LCRS-100 benchmark.
-
VENOMREC: Cross-Modal Interactive Poisoning for Targeted Promotion in Multimodal LLM Recommender Systems
Synchronized text+image poisoning steers multimodal LLM recommenders to promote target items, reaching 0.73 mean exposure@20.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.