REVIEW 6 cited by
When Video Coding Meets Multimodal Large Language Models: A Unified Paradigm for Video Coding
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
When Video Coding Meets Multimodal Large Language Models: A Unified Paradigm for Video Coding
read the original abstract
Existing codecs are designed to eliminate intrinsic redundancies to create a compact representation for compression. However, strong external priors from Multimodal Large Language Models (MLLMs) have not been explicitly explored in video compression. Herein, we introduce a unified paradigm for Cross-Modality Video Coding (CMVC), which is a pioneering approach to explore multimodality representation and video generative models in video coding. Specifically, on the encoder side, we disentangle a video into spatial content and motion components, which are subsequently transformed into distinct modalities to achieve very compact representation by leveraging MLLMs. During decoding, previously encoded components and video generation models are leveraged to create multiple encoding-decoding modes that optimize video reconstruction quality for specific decoding requirements, including Text-Text-to-Video (TT2V) mode to ensure high-quality semantic information and Image-Text-to-Video (IT2V) mode to achieve superb perceptual consistency. In addition, we propose an efficient frame interpolation model for IT2V mode via Low-Rank Adaption (LoRA) tuning to guarantee perceptual quality, which allows the generated motion cues to behave smoothly. Experiments on benchmarks indicate that TT2V achieves effective semantic reconstruction, while IT2V exhibits competitive perceptual consistency. These results highlight potential directions for future research in video coding.
Forward citations
Cited by 6 Pith papers
-
NeuralLVC: Neural Lossless Video Compression via Masked Diffusion with Temporal Conditioning
NeuralLVC achieves better lossless compression than H.264 and H.265 on video sequences by combining masked diffusion with temporal conditioning on frame differences.
-
UniDomain: Pretraining a Unified PDDL Domain from Real-World Demonstrations for Generalizable Robot Task Planning
UniDomain extracts atomic PDDL domains from 12,393 robot videos to create a unified domain of 3137 operators and 2875 predicates, then retrieves and fuses relevant parts to enable zero-shot planning on unseen real-wor...
-
VesselRW: Weakly Supervised Subcutaneous Vessel Segmentation via Learned Random Walk Propagation
VesselRW expands sparse vessel annotations into dense probabilistic supervision via a jointly trained differentiable random walk model with uncertainty weighting and topology regularization for CNN-based subcutaneous ...
-
DualResolution Residual Architecture with Artifact Suppression for Melanocytic Lesion Segmentation
Dual-resolution residual architecture with boundary-aware connections, channel attention, artifact suppression, and combined Dice-Tversky plus boundary and contrastive losses improves lesion boundary precision over st...
-
Edge Detection for Organ Boundaries via Top Down Refinement and SubPixel Upsampling
A top-down backward refinement network with subpixel upsampling generates crisp high-resolution organ boundaries in medical images and improves downstream segmentation and registration performance.
-
Deeply Dual Supervised learning for melanoma recognition
A dual-pathway deep learning model with attention mechanisms and multi-scale feature aggregation claims superior accuracy and fewer false positives for melanoma detection on benchmark datasets.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.