REVIEW 13 cited by
A Content-Driven Micro-Video Recommendation Dataset at Scale
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Micro-videos have recently gained immense popularity, sparking critical research in micro-video recommendation with significant implications for the entertainment, advertising, and e-commerce industries. However, the lack of large-scale public micro-video datasets poses a major challenge for developing effective recommender systems. To address this challenge, we introduce a very large micro-video recommendation dataset, named "MicroLens", consisting of one billion user-item interaction behaviors, 34 million users, and one million micro-videos. This dataset also contains various raw modality information about videos, including titles, cover images, audio, and full-length videos. MicroLens serves as a benchmark for content-driven micro-video recommendation, enabling researchers to utilize various modalities of video information for recommendation, rather than relying solely on item IDs or off-the-shelf video features extracted from a pre-trained network. Our benchmarking of multiple recommender models and video encoders on MicroLens has yielded valuable insights into the performance of micro-video recommendation. We believe that this dataset will not only benefit the recommender system community but also promote the development of the video understanding field. Our datasets and code are available at https://github.com/westlake-repl/MicroLens.
Forward citations
Cited by 13 Pith papers
-
UniRank: Benchmarking Ranking Models for Unified Sequential Modeling and Feature Interaction
UniRank is an open benchmark that standardizes chronological autoregressive supervision, multi-task evaluation, and capacity controls for 15 unified ranking models on five large datasets.
-
Learning Sparse Representations of Multimodal Content for Enhanced Cold Item Recommendation
Sparse content embeddings with a pre-sparsification alpha-entmax activation outperform dense embeddings for cold-start item recommendation at lower storage cost, especially for users with multiple interests.
-
Prediction Is Not Memory: Dual-Timescale Gated Profile Writing for Persistent User Modeling
A lightweight write-risk gate reduces harmful persistent-profile updates from 22.45% to about 14.5% on MicroLens-100K, and next-item ranking confidence is a poor substitute for write-risk scoring.
-
RecBase: Generative Foundation Model Pretraining for Zero-Shot Recommendation
A from-scratch model that tokenizes items into hierarchical codes and predicts next-item codes reaches higher average zero-shot AUC on 8 datasets than LLM recommenders up to 7B parameters.
-
EGRA:Toward Enhanced Behavior Graphs and Representation Alignment for Multimodal Recommendation
EGRA improves multimodal recommendation by using pretrained-model item embeddings to build the item-item graph and by dynamically weighting modality-behavior alignment per entity and per epoch.
-
Graph World Model
The Graph World Model uses action nodes and graph message passing to unify multimodal and graph-structured tasks, but its 'outperforms or matches' claim is contradicted by results on Goodreads.
-
VRAgent-R1: Boosting Video Recommendation with MLLM-based Agents via Reinforcement Learning
VRAgent-R1 uses an MLLM agent to summarize videos and a reinforcement-learned agent to simulate user choices, improving video recommendation and user-decision simulation on MicroLens-100K.
-
CaIRec: Calibrated Modality Imputation for Incomplete Multimodal Recommendation
Structurally calibrating imputed item modalities and then adapting them with pseudo-missing alignment and completion-aware graphs improves incomplete multimodal recommendation on Amazon benchmarks.
-
Describe What You See with Multimodal Large Language Models to Enhance Video Recommendations
Replacing raw video and audio features with MLLM-generated natural-language captions improves hit rate and nDCG for two-tower and SASRec recommenders on MicroLens-100K.
-
RAG-VisualRec: An Open Resource for Vision- and Text-Enhanced Retrieval-Augmented Generation in Recommendation
RAG-VisualRec is an open, auditable multimodal benchmark and pipeline for movie recommendation that fuses LLM-generated text with trailer embeddings and reports accuracy and beyond-accuracy metrics.
-
Beyond Noisy Signals: Dual-Level Denoising for Multi-modal Sequential Recommendation
A dual-level denoising framework that combines graph Laplacian smoothing and learnable FFT filtering improves multi-modal sequential recommendation on four benchmarks.
-
ViLLA-MMBench: A Unified Benchmark Suite for LLM-Augmented Multimodal Movie Recommendation
ViLLA-MMBench is an open, YAML-configured benchmark combining audio, visual, and LLM-enriched text embeddings for movie recommendation, reporting cold-start and coverage gains from LLM augmentation.
-
FindRec: Stein-Guided Entropic Flow for Multi-Modal Sequential Recommendation
FindRec combines Mamba temporal encoding, RBF-kernel cross-modal alignment, and expert routing to improve multimodal sequential recommendation, reporting 1.0 to 3.3 percent relative gains over baselines, with no proof...
Discussion (0). Sign in to comment.