REVIEW 16 cited by
A Content-Driven Micro-Video Recommendation Dataset at Scale
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Micro-videos have recently gained immense popularity, sparking critical research in micro-video recommendation with significant implications for the entertainment, advertising, and e-commerce industries. However, the lack of large-scale public micro-video datasets poses a major challenge for developing effective recommender systems. To address this challenge, we introduce a very large micro-video recommendation dataset, named "MicroLens", consisting of one billion user-item interaction behaviors, 34 million users, and one million micro-videos. This dataset also contains various raw modality information about videos, including titles, cover images, audio, and full-length videos. MicroLens serves as a benchmark for content-driven micro-video recommendation, enabling researchers to utilize various modalities of video information for recommendation, rather than relying solely on item IDs or off-the-shelf video features extracted from a pre-trained network. Our benchmarking of multiple recommender models and video encoders on MicroLens has yielded valuable insights into the performance of micro-video recommendation. We believe that this dataset will not only benefit the recommender system community but also promote the development of the video understanding field. Our datasets and code are available at https://github.com/westlake-repl/MicroLens.
Forward citations
Cited by 16 Pith papers
-
Bridging Language and Items for Retrieval and Recommendation: Benchmarking LLMs as Semantic Encoders
BLaIR is a new benchmark and 570M-review dataset showing that LLM performance rankings on recommendation tasks have little correlation with rankings on general embedding benchmarks like MTEB.
-
Will It Go Viral? Grounding Micro-Video Popularity Prediction on the Open Web
WEBSHORTS dataset and SHORTS-CAST framework ground micro-video popularity prediction in structured open-web context collected at upload time and enable selective online adaptation using delayed labels.
-
Sparse Contrastive Learning for Content-Based Cold Item Recommendation
SEMCo uses sparse entmax contrastive learning for purely content-based cold-start item recommendation, outperforming standard methods in ranking accuracy.
-
UniRank: Benchmarking Ranking Models for Unified Sequential Modeling and Feature Interaction
UniRank is an open benchmark that standardizes chronological autoregressive supervision, multi-task evaluation, and capacity controls for 15 unified ranking models on five large datasets.
-
Learning Sparse Representations of Multimodal Content for Enhanced Cold Item Recommendation
Sparse content embeddings with a pre-sparsification alpha-entmax activation outperform dense embeddings for cold-start item recommendation at lower storage cost, especially for users with multiple interests.
-
Prediction Is Not Memory: Dual-Timescale Gated Profile Writing for Persistent User Modeling
A lightweight write-risk gate reduces harmful persistent-profile updates from 22.45% to about 14.5% on MicroLens-100K, and next-item ranking confidence is a poor substitute for write-risk scoring.
-
Popcorn: A Configurable Benchmark for Visual Evidence in Multimodal Movie Recommendation
Popcorn is a new benchmark standardizing modality assembly, fusion, and evaluation of thumbnails, trailers, and full movies encoded by VLMs for multimodal movie recommendation.
-
OmniTrend: Content-Context Modeling for Scalable Social Popularity Prediction
OmniTrend predicts popularity by combining separate content attractiveness and contextual exposure predictors using cross-modal and exogenous signals.
-
RecBase: Generative Foundation Model Pretraining for Zero-Shot Recommendation
A from-scratch model that tokenizes items into hierarchical codes and predicts next-item codes reaches higher average zero-shot AUC on 8 datasets than LLM recommenders up to 7B parameters.
-
EGRA:Toward Enhanced Behavior Graphs and Representation Alignment for Multimodal Recommendation
EGRA improves multimodal recommendation by using pretrained-model item embeddings to build the item-item graph and by dynamically weighting modality-behavior alignment per entity and per epoch.
-
CaIRec: Calibrated Modality Imputation for Incomplete Multimodal Recommendation
Structurally calibrating imputed item modalities and then adapting them with pseudo-missing alignment and completion-aware graphs improves incomplete multimodal recommendation on Amazon benchmarks.
-
CaIRec: Calibrated Modality Imputation for Incomplete Multimodal Recommendation
CaIRec consistently beats eleven baselines on three Amazon datasets (about 2-7% relative gains) by combining latent imputation, spectral cross-modal calibration, and recommendation-space alignment of recovered item features.
-
Multimodal Large Language Models with Adaptive Preference Optimization for Sequential Recommendation
HaNoRec dynamically weights harder preference samples and applies Gaussian perturbations to output distributions to improve multimodal LLM performance on sequential recommendation tasks.
-
Describe What You See with Multimodal Large Language Models to Enhance Video Recommendations
Replacing raw video and audio features with MLLM-generated natural-language captions improves hit rate and nDCG for two-tower and SASRec recommenders on MicroLens-100K.
-
Beyond Noisy Signals: Dual-Level Denoising for Multi-modal Sequential Recommendation
A dual-level denoising framework that combines graph Laplacian smoothing and learnable FFT filtering improves multi-modal sequential recommendation on four benchmarks.
-
ViLLA-MMBench: A Unified Benchmark Suite for LLM-Augmented Multimodal Movie Recommendation
ViLLA-MMBench is an open, YAML-configured benchmark combining audio, visual, and LLM-enriched text embeddings for movie recommendation, reporting cold-start and coverage gains from LLM augmentation.
Discussion (0). Sign in to comment.