Pith. sign in

REVIEW 13 cited by

A Content-Driven Micro-Video Recommendation Dataset at Scale

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.15379 v1 pith:5R6UIZRG submitted 2023-09-27 cs.IR

classification cs.IR
keywords micro-videorecommendationdatasetmicrolensvideorecommenderchallengecontent-driven
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Micro-videos have recently gained immense popularity, sparking critical research in micro-video recommendation with significant implications for the entertainment, advertising, and e-commerce industries. However, the lack of large-scale public micro-video datasets poses a major challenge for developing effective recommender systems. To address this challenge, we introduce a very large micro-video recommendation dataset, named "MicroLens", consisting of one billion user-item interaction behaviors, 34 million users, and one million micro-videos. This dataset also contains various raw modality information about videos, including titles, cover images, audio, and full-length videos. MicroLens serves as a benchmark for content-driven micro-video recommendation, enabling researchers to utilize various modalities of video information for recommendation, rather than relying solely on item IDs or off-the-shelf video features extracted from a pre-trained network. Our benchmarking of multiple recommender models and video encoders on MicroLens has yielded valuable insights into the performance of micro-video recommendation. We believe that this dataset will not only benefit the recommender system community but also promote the development of the video understanding field. Our datasets and code are available at https://github.com/westlake-repl/MicroLens.

Discussion (0). Sign in to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. UniRank: Benchmarking Ranking Models for Unified Sequential Modeling and Feature Interaction

    cs.IR 2026-07 conditional novelty 6.0 of 10

    UniRank is an open benchmark that standardizes chronological autoregressive supervision, multi-task evaluation, and capacity controls for 15 unified ranking models on five large datasets.

  2. Learning Sparse Representations of Multimodal Content for Enhanced Cold Item Recommendation

    cs.IR 2026-07 conditional novelty 6.0 of 10

    Sparse content embeddings with a pre-sparsification alpha-entmax activation outperform dense embeddings for cold-start item recommendation at lower storage cost, especially for users with multiple interests.

  3. Prediction Is Not Memory: Dual-Timescale Gated Profile Writing for Persistent User Modeling

    cs.IR 2026-07 conditional novelty 6.0 of 10

    A lightweight write-risk gate reduces harmful persistent-profile updates from 22.45% to about 14.5% on MicroLens-100K, and next-item ranking confidence is a poor substitute for write-risk scoring.

  4. RecBase: Generative Foundation Model Pretraining for Zero-Shot Recommendation

    cs.IR 2025-09 conditional novelty 6.0 of 10

    A from-scratch model that tokenizes items into hierarchical codes and predicts next-item codes reaches higher average zero-shot AUC on 8 datasets than LLM recommenders up to 7B parameters.

  5. EGRA:Toward Enhanced Behavior Graphs and Representation Alignment for Multimodal Recommendation

    cs.IR 2025-08 conditional novelty 6.0 of 10

    EGRA improves multimodal recommendation by using pretrained-model item embeddings to build the item-item graph and by dynamically weighting modality-behavior alignment per entity and per epoch.

  6. Graph World Model

    cs.LG 2025-07 reject novelty 6.0 of 10

    The Graph World Model uses action nodes and graph message passing to unify multimodal and graph-structured tasks, but its 'outperforms or matches' claim is contradicted by results on Goodreads.

  7. VRAgent-R1: Boosting Video Recommendation with MLLM-based Agents via Reinforcement Learning

    cs.MM 2025-07 conditional novelty 6.0 of 10

    VRAgent-R1 uses an MLLM agent to summarize videos and a reinforcement-learned agent to simulate user choices, improving video recommendation and user-decision simulation on MicroLens-100K.

  8. CaIRec: Calibrated Modality Imputation for Incomplete Multimodal Recommendation

    cs.IR 2026-07 conditional novelty 5.5 of 10

    Structurally calibrating imputed item modalities and then adapting them with pseudo-missing alignment and completion-aware graphs improves incomplete multimodal recommendation on Amazon benchmarks.

  9. Describe What You See with Multimodal Large Language Models to Enhance Video Recommendations

    cs.IR 2025-08 conditional novelty 5.0 of 10

    Replacing raw video and audio features with MLLM-generated natural-language captions improves hit rate and nDCG for two-tower and SASRec recommenders on MicroLens-100K.

  10. RAG-VisualRec: An Open Resource for Vision- and Text-Enhanced Retrieval-Augmented Generation in Recommendation

    cs.IR 2025-06 conditional novelty 5.0 of 10

    RAG-VisualRec is an open, auditable multimodal benchmark and pipeline for movie recommendation that fuses LLM-generated text with trailer embeddings and reports accuracy and beyond-accuracy metrics.

  11. Beyond Noisy Signals: Dual-Level Denoising for Multi-modal Sequential Recommendation

    cs.IR 2026-07 conditional novelty 4.0 of 10

    A dual-level denoising framework that combines graph Laplacian smoothing and learnable FFT filtering improves multi-modal sequential recommendation on four benchmarks.

  12. ViLLA-MMBench: A Unified Benchmark Suite for LLM-Augmented Multimodal Movie Recommendation

    cs.IR 2025-08 unverdicted novelty 4.0 of 10

    ViLLA-MMBench is an open, YAML-configured benchmark combining audio, visual, and LLM-enriched text embeddings for movie recommendation, reporting cold-start and coverage gains from LLM augmentation.

  13. FindRec: Stein-Guided Entropic Flow for Multi-Modal Sequential Recommendation

    cs.IR 2025-07 conditional novelty 4.0 of 10

    FindRec combines Mamba temporal encoding, RBF-kernel cross-modal alignment, and expert routing to improve multimodal sequential recommendation, reporting 1.0 to 3.3 percent relative gains over baselines, with no proof...

Pith tools