REVIEW 2 cited by
Recent Standard Development Activities on Video Coding for Machines
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
In recent years, video data has dominated internet traffic and becomes one of the major data formats. With the emerging 5G and internet of things (IoT) technologies, more and more videos are generated by edge devices, sent across networks, and consumed by machines. The volume of video consumed by machine is exceeding the volume of video consumed by humans. Machine vision tasks include object detection, segmentation, tracking, and other machine-based applications, which are quite different from those for human consumption. On the other hand, due to large volumes of video data, it is essential to compress video before transmission. Thus, efficient video coding for machines (VCM) has become an important topic in academia and industry. In July 2019, the international standardization organization, i.e., MPEG, created an Ad-Hoc group named VCM to study the requirements for potential standardization work. In this paper, we will address the recent development activities in the MPEG VCM group. Specifically, we will first provide an overview of the MPEG VCM group including use cases, requirements, processing pipelines, plan for potential VCM standards, followed by the evaluation framework including machine-vision tasks, dataset, evaluation metrics, and anchor generation. We then introduce technology solutions proposed so far and discuss the recent responses to the Call for Evidence issued by MPEG VCM group.
Forward citations
Cited by 2 Pith papers
-
Multi-task Just Recognizable Difference for Video Coding for Machines: Database, Model, and Coding Application
AMT-JRD model predicts multi-task JRDs with MAE 3.781, outperforming single-task baselines by 6.7% and delivering 3.861% BD-mAP gain over VVC in VCM coding.
-
BLUE: Semantics-Preserving Video Compression for Efficient Vision-Language Surveillance Analytics
BLUE-compressed H.265 video yields VLM semantic scores within 0.01-0.06 points of raw H.265 on surveillance clips while enabling an estimated 53% reduction in VLM calls via small P-frame skipping.
Discussion (0). Sign in to comment.