{"id":"5b615294-6f21-47b9-a85c-88ad03cf01db","arxiv_id":"2411.15801","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A comprehensive review of ML-based QoE modeling and adaptive streaming for conventional and 360 degree video, including datasets and open challenges.","lead":"This paper is a survey of machine learning methods for predicting video Quality of Experience (QoE) and for adaptive streaming of 2D and 360 degree video. It reviews QoE models, streaming techniques, datasets, and open challenges in the field.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The survey's comprehensiveness claim hinges on an unverified literature selection and untraceable quantitative summaries; a systematic spot-check of Tables 5/7/8 against cited sources would settle whether the synthesis is reliable.","rationale":"The reader's weakest assumption identifies exactly the same point: the survey's comprehensiveness and accuracy depend on a literature-selection and summarization process that is neither documented nor verified. My read agrees with that judgment and does not move the verdict. The article is a review, so its central claim is about coverage and synthesis rather than a new scientific result. That claim stands or falls on whether the included papers and the numbers attributed to them are representative and correct. I focused on the quantitative tables because they are the most consequential and checkable manifestation of this assumption: if a researcher uses Table 5 or Table 8 to benchmark a new model, a single misreported accuracy value can propagate into the literature. The small empirical analysis in Section 3.2 (Figure 7) is also under-specified, but it is illustrative and does not carry the same weight as the comparative tables. The proposed test is deliberately narrow: a 10-row spot-check against cited sources is sufficient to establish whether the reporting is trustworthy. If the check passes, the conditional acceptance stands; if it fails, the tables would need to be corrected and re-verified before the survey can be used as a reference. I therefore recommend no change to the reader's conditional verdict, with the condition being the verifiability of the survey's synthesis.","tokens_in":49309,"tokens_out":5687,"duration_ms":49531,"concrete_test":"Randomly select 10 rows spanning Tables 5, 7, and 8 (covering overall QoE, continuous QoE, 2D ABR, and 360-degree streaming). For each row, retrieve the cited paper and independently locate the quantitative result claimed in the table (accuracy percentage, relative gain, or influence factors). Record whether the table entry exactly matches the source, overstates it, or cannot be found. If 2 or more of the 10 entries are misreported or unsupported, the survey's synthesis is unreliable and the comprehensiveness claim fails; if all 10 match, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central contribution is a comprehensive, accurate synthesis of ML-based QoE modeling and streaming techniques for conventional and 360-degree video. This rests on two unverified premises: (i) the selected literature is representative, and (ii) the reported methods and performance numbers are faithfully extracted from the sources. The manuscript provides no systematic search strategy, no inclusion/exclusion criteria, and no data-extraction protocol. The comparative Tables 5, 7, and 8 list concrete quantitative claims (e.g., '4.8-20% relative PLCC gain', '8.52%, 48.18% lowered bit-rate', '33-71% relative LCC increase') without any indication of how each number was obtained or where it appears in the cited paper. In addition, the paper's own works (MO-QoE [14], DeSVQ [51], M-3R [133], MAIVS [37], I2MB [23]) appear prominently across the key tables, raising the possibility that the 'state-of-the-art' comparison is drawn from a convenience sample rather than a balanced field-level survey. If a spot-check reveals even a few misreported figures or a systematic omission of major works, the survey's value as a reference is materially weakened, and the claimed comprehensiveness would no longer be supported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript is a survey of machine-learning-based QoE modeling and user-centric adaptive streaming for conventional and 360-degree video. It covers subjective and objective quality assessment, overall and continuous time-varying QoE prediction, adaptive bit-rate streaming, tile-based 360-degree streaming, public datasets, and open research challenges. The paper positions itself as a comprehensive overview, and it includes comparative tables (Tables 5, 7, and 8) with reported performance gains, a short head-orientation analysis in Section 3.2, and an inventory of datasets in Tables 9-11.","tokens_in":49559,"tokens_out":5876,"duration_ms":52169,"significance":"If the synthesis is reliable, this survey would be a useful entry point for researchers, especially because it combines QoE modeling and ML-based streaming for both conventional and 360-degree video in one place. The dataset inventory in Tables 9-11 and the list of open challenges are practically valuable. The paper also makes a clear organizational effort, separating overall from continuous QoE prediction and conventional from immersive streaming. However, the reference value of the survey depends on the accuracy and representativeness of the literature synthesis; the absence of a documented selection methodology and the untraceable quantitative entries in the comparison tables currently weaken that value.","major_comments":[{"comment":"The paper claims to provide 'a comprehensive overview' but never describes a systematic search strategy, inclusion/exclusion criteria, or a data extraction protocol. For a survey whose central claim is coverage and synthesis, this omission is load-bearing. The authors should either add a methodology subsection (e.g., databases searched, search strings, time window, screening criteria) or explicitly temper the comprehensiveness claim.","section":"Section 1.1 (general methodology)"},{"comment":"The Accuracy columns report concrete numerical gains without traceable provenance. For example, Table 5 lists DeSVQ as '4.8-20% relative PLCC gain', Table 7 lists MAIVS as '8.52%, 48.18% lowered bit-rate', and Table 8 lists the [214] row as '33-71% relative LCC increase'; none of these entries states the baseline, dataset split, or metric definition used to compute them. Please add a mapping from each table entry to the exact figure, table, or section of the cited paper, or remove the overly precise values. This traceability is necessary for the survey to serve as a reliable reference.","section":"Tables 5, 7, and 8"},{"comment":"The small empirical study of 55 viewings is not reproducible: the text does not specify which video content from datasets [64, 65, 23] was used, the tile grid geometry, the projection format, the QP encoding configuration, or the algorithm that maps head orientation to tile numbers. In addition, Figure 7(a) shows 61 tiles on the x-axis but the text states that tiles 57-64 are least viewed. Please provide the processing details or remove the analysis; as written, it does not substantiate the claim that FoV-based bit-rate adaptation is motivated by these statistics.","section":"Section 3.2, Figure 7"},{"comment":"The authors' own prior work appears frequently in the central comparison tables and is consistently presented as best-performing (e.g., MO-QoE in Table 5 and Figures 11-12, M-3R and DeSVQ in Table 5, MAIVS in Tables 7-8, and I2MB in Section 2.2). Self-citation is not improper, but without inclusion criteria the prominence of these works raises a selection-bias concern that should be addressed explicitly. The authors should state how all entries were chosen for the tables and whether any independent re-evaluation or replication of the numbers was performed.","section":"Tables 5-8 and Section 5.1"}],"minor_comments":[{"comment":"There is a typo in the text: 'WCPPPSNR' should be 'WCP-PSNR'. Additionally, several equations in Table 2 contain garbled subscripts and notation (e.g., NC-PSNR and WS-SSIM entries) that are hard to parse; please clean up the LaTeX.","section":"Section 2.1, Table 2"},{"comment":"The SROCC values in Table 3 (e.g., PSNR 0.3527, SSIM 0.5119) are presented without a citation or a description of the computation. Please verify these numbers against the cited source and specify the exact database version and metric configuration.","section":"Section 2.3, Table 3"},{"comment":"There is a typo: 'It is uded to measure' should be 'It is used to measure'.","section":"Section 4.2"},{"comment":"Figure 4 appears to contain original analysis, but the caption cites [48, 49, 50] ambiguously. Please clarify whether the correlation curves are reproduced from those papers or computed by the authors, and describe the video samples and preprocessing.","section":"Figure 4"},{"comment":"The rating type 'overall+cont.' in Table 9 is not explained in the caption; please expand the abbreviation to 'overall and continuous' or define it in the table notes.","section":"Table 9"}],"recommendation":"major_revision","confidential_remarks":"This is a survey paper, so novelty is not the issue. The main question is whether the synthesis can be trusted. I would be willing to accept after the authors add a reproducible literature-selection methodology, make the quantitative entries in Tables 5, 7, and 8 traceable to their sources, and either document or remove the ad hoc analysis in Section 3.2. The self-citation pattern is worth monitoring, but it is not by itself disqualifying if the selection criteria are made explicit."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThis is a survey of ML-based QoE modeling and adaptive streaming for both conventional and 360-degree video. If you need a single place to find the main models, ABR schemes, and datasets, it's a useful starting point. The organization is sensible: QoE definitions, overall and continuous prediction models, adaptive streaming techniques, datasets, and open challenges. Tables 5, 7, and 8 give a quick comparative picture of methods, and Tables 9 and 11 list the common datasets with descriptions. The paper also includes a small empirical analysis of tile popularity (Figure 7), which is a nice addition even if not fully described.\n\nThe main soft spot is that the authors don't describe any systematic literature selection process. There's no search strategy, no inclusion/exclusion criteria, and no data-extraction protocol. That makes the claim of 'comprehensive overview' hard to verify. I'm also concerned about the quantitative entries in the comparison tables. Numbers like '4.8-20% relative PLCC gain' or '8.52%, 48.18% lowered bit-rate' appear without any indication of where in the cited paper they come from. A reader can't check whether they're correctly extracted. The stress-test note is right that a spot-check of a few of these figures against the cited sources would settle the survey's reliability.\n\nRelated to that, the authors' own works (MO-QoE, DeSVQ, M-3R, MAIVS, I2MB) show up frequently and are presented favorably. That's not inherently wrong, but combined with the lack of a transparent selection process it does raise a question about whether the 'state of the art' is a convenience sample. The paper would be stronger with a more balanced comparative analysis and with the authors' works presented as part of the field rather than as the centerpiece.\n\nNone of this is fatal. The survey is coherent and covers a lot of ground. It's just that the central value—being a reliable reference—rests on assumptions that aren't documented. If the numbers check out and the selection bias is addressed, this could be a solid reference for graduate students and engineers entering the area. As is, I'd recommend peer review with major revisions, specifically asking for a methodology section, sourcing for the table values, and a rebalance of self-citations.\n\nI wouldn't cite this in my own work until I can verify the numbers, but I'd consider bringing it to a reading group if we want to discuss survey methodology.\n\nBest,\n[Your name]","headline":"A broad but uneven survey of ML-based QoE and streaming; the tables are handy but the lack of methodology and self-citation bias keep it from being the go-to reference.","tokens_in":50068,"tokens_out":3765,"would_cite":false,"duration_ms":32224,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A survey that maps machine-learning QoE prediction and adaptive streaming for 2D and 360-degree video, and argues that continuous, time-varying QoE modeling should drive user-centric streaming decisions.","keywords":["Quality of Experience","QoE prediction","adaptive video streaming","machine learning","360-degree video","viewport prediction","reinforcement learning","video quality assessment"],"falsifier":"A systematic re-review of the same literature using a transparent search and selection protocol (e.g., strict inclusion criteria, PRISMA-style flow diagram, multiple independent coders) that a year later produces a materially different landscape of what counts as the leading ML-based QoE/streaming techniques would show whether the survey's coverage is representative. Alternatively, reproducing the Figure 7 analysis with a different head-orientation dataset and showing that the tile popularity distribution is substantially different would cast doubt on the implied motivating observation for viewport-adaptive streaming.","tokens_in":49121,"feed_emoji":"📺","tokens_out":3911,"duration_ms":29496,"temperature":0.7,"pith_summary":"This survey paper sets out to organize the fast-growing body of research on machine-learning-based Quality-of-Experience (QoE) prediction and adaptive streaming for both conventional and 360-degree video. Its central claim is that QoE should be treated not as a one-time overall score but as a continuous, time-varying quantity that can be learned from objective video metrics, impairment events such as rebuffering, and user behavior, and then used to drive bit-rate and tile-selection decisions in real time. The paper argues that a user-centric paradigm, where the client actively adapts based on predicted perceptual quality rather than just network throughput, is the natural direction for future multimedia streaming systems. It matters because streaming video dominates internet traffic, and the proposed ML-based approaches promise to improve viewer experience while making more efficient use of bandwidth. The survey's value lies in its synthesis: it brings together QoE definitions, assessment methodologies, ML-based QoE prediction models, adaptive streaming algorithms, and publicly available datasets under one coherent framework.","feed_headline":"Machine learning remaps how video QoE is measured and streamed","feed_subtitle":"A survey of ML-based QoE prediction and adaptive streaming for 2D and 360-degree video.","key_machinery":"The central organizing object is the ML-based QoE prediction model that maps input features—objective VQA metrics (PSNR, SSIM, MS-SSIM, VMAF, etc.), impairment factors (stalling, rebuffering, quality switches), and user-centric signals (viewport, head movement, gaze)—to a predicted QoE score, either overall or frame-by-frame. For continuous QoE, the key mechanism is temporal modeling with recurrent or convolutional architectures (LSTM, bidirectional LSTM, temporal convolutional networks) that capture the hysteresis effect and long-term dependencies in perceived quality. For adaptive streaming, the central mechanism is the bit-rate adaptation (ABR) algorithm, increasingly formulated as a reinforcement learning problem where an agent selects bit-rates or tile qualities to maximize QoE. For 360-degree video, the additional key mechanism is viewport prediction (using head movement traces, saliency maps, and eye tracking) combined with tile-based streaming, where only the tiles in the predicted viewport are streamed at high quality.","core_discovery":"The paper claims that the state of the art in multimedia streaming is converging on a user-centric, machine-learning-driven approach in which QoE is modeled as a continuous, time-varying signal rather than a single overall score. It argues that ML models—especially recurrent architectures like LSTM and attention-based models—can capture the temporal dependencies, hysteresis effects, and non-linear interactions between video quality, rebuffering, and bit-rate switches that traditional parametric QoE models miss. For 360-degree video, the paper contends that viewport-adaptive, tile-based streaming driven by ML-based viewport prediction is the key to reducing bandwidth requirements while maintaining immersion. The survey catalogs a wide range of ML-based QoE prediction models (e.g., ATLAS, DEMI, DeSVQ, MO-QoE, NARX, LSTM-based models, bidirectional LSTM, temporal convolutional networks) and adaptive streaming techniques (e.g., Pensieve, Comyco, SAC-ABR, ABRaider, D-DASH, FReD-ViQ, and 360-degree-specific systems like DRL360, PARSEC, NOVA, RoSal360), showing that the field has moved from heuristic buffer- and throughput-based adaptation to learning-based policies that optimize perceptual quality directly.","pith_inferences":["One implicit consequence of the survey's framing is that QoE prediction should be evaluated not only by correlation with subjective scores, but by its downstream impact on adaptation decisions; a model that predicts QoE well on offline datasets may still fail when used in a closed-loop streaming system, since the adaptation changes what the user sees.","The survey's emphasis on datasets suggests a testable extension: a standardized, continuously updated benchmark that combines QoE prediction accuracy with streaming performance metrics (rebuffering, quality switches, bandwidth efficiency) across diverse network traces and content types would be a natural next step for the field.","The paper's treatment of 360-degree streaming implies that viewport prediction accuracy is the single most load-bearing component for bandwidth savings; yet most systems still rely on historical head movement data, and the open challenge of predicting viewport under volatile head motion suggests that robust uncertainty-aware prediction could be a more fruitful direction than chasing marginal accur"],"forward_implications":["If ML-based continuous QoE prediction becomes reliable enough, streaming clients can make bit-rate decisions that optimize perceived quality over time, rather than just minimizing rebuffering or maximizing throughput.","For 360-degree video, accurate ML-based viewport prediction combined with tile-based adaptive streaming could substantially reduce bandwidth requirements while maintaining or improving immersive quality.","The availability of public datasets (subjective scores, network traces, head movement, eye tracking) could enable standardized benchmarking and reproducible comparison of QoE prediction models and streaming policies.","Integration of ML-based QoE prediction with ABR algorithms could shift the field from network-centric to truly user-centric streaming, where each user's experience is individually optimized.","The trend toward deep learning models for QoE prediction may continue, with hybrid approaches (e.g., CNN+LSTM, fuzzy logic+RL) addressing both accuracy and computational efficiency."],"supporting_citations":[{"why":"The LIVE NFLX II dataset with continuous subjective quality scores, used as a key testbed for the ML-based QoE models the survey reviews.","marker":"[48]"},{"why":"The LIVE Netflix dataset, one of the primary sources of continuous-time subjective QoE scores used to train and evaluate the models discussed in Section 5.","marker":"[49]"},{"why":"The Mobile stall II dataset, providing continuous and overall QoE scores with stalling patterns, used to evaluate time-varying QoE models.","marker":"[50]"},{"why":"Eswara et al.'s LSTM-based QoE prediction model, a central example of the ML-based continuous QoE prediction approach the survey highlights.","marker":"[60]"},{"why":"Pensieve, the reinforcement learning-based ABR algorithm that serves as the canonical example of ML-driven adaptive streaming.","marker":"[154]"},{"why":"The LFOVIA dataset and associated continuous QoE evaluation framework, used as a standard benchmark for time-varying QoE models.","marker":"[123]"},{"why":"The MO-QoE framework, an optimized multi-feature fusion QoE prediction model, cited as a state-of-the-art example in the overall QoE prediction category.","marker":"[14]"},{"why":"Bampis et al.'s NARX continuous QoE prediction model, which the survey discusses as a key autoregressive approach to time-varying QoE.","marker":"[124]"},{"why":"DeSVQ, a two-stage deep learning framework combining CNN and LSTM for continuous streaming QoE estimation, used as a representative of the hybrid approach.","marker":"[51]"},{"why":"Barman and Martini's survey of QoE modeling for HTTP adaptive streaming, which the paper positions as a complementary prior review it extends.","marker":"[44]"}],"fun_headline_variants":["ML survey: continuous QoE for 2D and 360 video","Machine learning makes QoE a live signal for streaming","ML-driven adaptive streaming puts user experience first"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The survey's accuracy and usefulness depend on the assumption that the authors' selection and summaries of the cited literature accurately reflect the state of the art, and that the small empirical analysis of tile viewing patterns generalizes beyond the specific dataset and 360-degree video used.","fun_headline_variants_meta":{"raw":{"variants":["ML survey: continuous QoE for 2D and 360 video","Machine learning makes QoE a live signal for streaming","ML-driven adaptive streaming puts user experience first"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000717,"raw_usage":{"total_tokens":3302,"prompt_tokens":1108,"completion_tokens":2194,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":724,"completion_tokens_details":{"reasoning_tokens":2140}},"tokens_in":724,"tokens_out":2194,"duration_ms":16577,"temperature":1.0,"reasoning_tokens":2140,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T13:52:11.895367+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A systematic re-review of the same literature using a transparent search and selection protocol (e.g., strict inclusion criteria, PRISMA-style flow diagram, multiple independent coders) that a year later produces a materially different landscape of what counts as the leading ML-based QoE/streaming techniques would show whether the survey's coverage is representative. Alternatively, reproducing the Figure 7 analysis with a different head-orientation dataset and showing that the tile popularity distribution is substantially different would cast doubt on the implied motivating observation for viewport-adaptive streaming.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The LIVE Netflix dataset, one of the primary sources of continuous-time subjective QoE scores used to train and evaluate the models discussed in Section 5."}],"review_version":1}