{"id":"098dc021-c5b9-4e78-9471-02a9f9b6c16f","arxiv_id":"2506.23852","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A 2,100-video database with human opinions shows that current video quality models underperform on robot-generated content, motivating a new VQA subfield.","lead":"The paper builds a new database of 2,100 videos recorded by robots, with human quality ratings for each clip, and shows that existing video quality assessment models perform poorly on this content. It matters because robot-generated video is growing fast across delivery, surveillance, and teleoperation, yet no dedicated quality benchmark existed before.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Wheeled-robot category appears to include human-driven dashcam footage, not robot-generated content; if so, RGCD's central claim and RGC benchmark are diluted.","rationale":"The reader's CONDITIONAL verdict is appropriate. I considered three candidate concerns: the MOS formula in Eq. (3), the humanoid-robot label, and the source composition of the wheel category. Eq. (3) is internally inconsistent as written, but since every subject rates every video, subject means and standard deviations are per-subject constants, and the formula reduces to an affine transform of the mean raw rating. This does not change Spearman or Pearson correlations, so it is a processing discrepancy rather than a threat to the benchmark conclusions. The humanoid-robot concern is real but the central claim and Table 1 label the category as 'Robot'/robot arms, and the authors explicitly acknowledge the absence of authentic humanoid locomotion data; this is an overclaim of representativeness but not a failure of the database's RGC definition. The wheel contamination is more serious because Section 3.1 cites JAAD, BDD100K, and a traffic-accident dataset as sources for 'wheeled robots', and these are widely used dashcam datasets from human-driven vehicles. If a significant fraction of the wheel category is not robot-generated, then RGCD is not purely RGC, and the empirical evidence for 'existing models are unsuited to RGC' is confounded by ordinary driving-scene content. This is verifiable from released metadata and can be fixed by re-benchmarking on a filtered subset. I therefore keep the reader's CONDITIONAL verdict but shift the primary burden of proof to source-platform verification.","tokens_in":14899,"tokens_out":8289,"duration_ms":92010,"concrete_test":"Obtain RGCD metadata from the provided GitHub repository and map each wheel-category video to its source dataset and clip ID. For each source among JAAD, BDD100K, the traffic-accident dataset, KITTI, and Ford Campus, classify the capture platform as a robot (autonomous or remote-controlled) vs. a human-driven dashcam. Compute the fraction of wheel videos from non-robot platforms. If that fraction is substantial (e.g., >10%), re-run the Table 1 zero-shot and fine-tuned benchmarks on the robot-only subset; if the model rankings or the overall conclusion that VQA models are unsuited to RGC change materially, the central claim is weakened.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing concern is the provenance of the 'wheel' category. Section 3.1 states that wheeled-robot videos are sourced from 'autonomous driving datasets', citing refs [1, 11, 30, 31, 48]. Among these are JAAD, BDD100K, and a traffic-accident anticipation dataset, which are commonly captured with dashboard cameras in human-driven vehicles, not by robots operating autonomously or under remote control. Since the paper defines RGC as egocentric video captured by robots operating autonomously or under remote control, including ordinary dashcam clips would mean a substantial part of the database does not satisfy its own inclusion criterion. This goes to the core of the central claim: if the 700-wheel-video category is largely human driving footage, then RGCD is not purely robotic-generated, and the benchmark conclusion that existing VQA models are unsuited to RGC is confounded by driving-domain content rather than by robotic generation. This is more central than the 'robot' category mislabeling because it attacks the definition of the database itself. The concern is verifiable from released source metadata and can be remedied by re-benchmarking on a robot-only subset.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces Robotic-Generated Content (RGC) as a new content taxonomy and presents RGCD, a database of 2,100 egocentric videos from drones, wheeled robots, and robot arms, together with subjective quality ratings from 15 lab subjects. The authors describe the collection process, provide statistical content and MOS analyses, and benchmark 11 existing VQA models plus five zero-shot evaluations, concluding that current VQA models are not well suited to robotic-generated video quality assessment.","tokens_in":15099,"tokens_out":5184,"duration_ms":55084,"significance":"If the provenance and analysis issues are resolved, RGCD would be a genuinely useful community resource: it is, to my knowledge, the first public database explicitly targeting robot-generated video quality, it covers diverse sources and platform types, and the benchmark experiments give a first indication of a domain gap for existing VQA models. The paper is empirical and does not rely on circular derivation; the database release and the concrete benchmark numbers are the substantive contributions. The main risks are whether the data actually satisfy the paper's own definition of RGC and whether the reported correlations are stable.","major_comments":[{"comment":"The MOS formula multiplies each rescaled z-score by the subject standard deviation sigma_i, which contradicts the described averaging of z-scores and the stated [0, 100] scale. As written, MOS_j is not the average of the rescaled z-scores z'_ij and can fall outside the plotted MOS range used in Figure 4 and Table 1. Please correct Eq. (3) to a plain average over z'_ij, or explicitly justify the sigma_i factor, and then recompute any affected MOS-based analyses and benchmark correlations.","section":"Section 3.3, Eq. (3)"},{"comment":"The 'wheeled robot' category is sourced from autonomous-driving datasets [1, 11, 30, 31, 48], but several of these, including JAAD [31], BDD100K [48], and the traffic-accident anticipation dataset [1], are typically recorded by dashboard cameras in human-driven vehicles. Such footage does not satisfy the paper's own definition of RGC as egocentric video captured by robots operating autonomously or under remote control. Since the wheel category constitutes one third of RGCD, the central claim that RGCD is a robotic-generated-content database is at stake. Please provide per-source evidence that each included clip was captured by a robot, or re-run the database construction and benchmark on a subset that excludes human-driven footage.","section":"Section 3.1"},{"comment":"The benchmark results are reported from a single 4:1 train/test split with no confidence intervals, significance tests, or repeated-split statistics. The paper's main conclusion is that existing VQA models are unsuited to RGC; that conclusion needs to be robust to the choice of split. Please report results over multiple random splits with mean and standard deviation, or otherwise demonstrate that the numerical differences in Table 1 are not artifacts of a single data partition.","section":"Section 4.1.3, Table 1"},{"comment":"The category labeled 'robot' (and described as 'humanoid robot' in the text) is acknowledged to consist of wrist-camera and other egocentric views from robotic-arm manipulation datasets such as X-Embodiment and ARIO, with the authors explicitly stating that authentic locomotion data for humanoid or quadruped robots could not be found. The abstract and conclusion nevertheless describe the database as covering three robot types including humanoid robots. This overstates the content taxonomy; the category name and the associated claims about humanoid-robot video should be corrected.","section":"Section 3.1, Figure 2"}],"minor_comments":[{"comment":"There are several typographical and grammatical errors, e.g., 'A extensive benchmark experiment' (Section 4.1.1), 'out database' and 'posse' (Section 3.4.1), and 'in lab subjective experiment' (Section 3.2). The manuscript would benefit from a careful proofreading pass.","section":"Throughout"},{"comment":"The text in Section 2.1 and related work refers to 'Ego4D' when discussing first-person human perspective, but reference [13] is the Ego-Exo4D paper. Please correct the citation or the reference entry so that the cited work matches the text.","section":"References"},{"comment":"The subjective experiment uses 15 subjects, all described as college students, with no report of inter-rater agreement, per-video confidence intervals, or screen/calibration details beyond the viewing distance and resolution. Since the MOS values are the ground truth for the whole database, adding such details and confidence information would substantially improve the reproducibility and credibility of the database.","section":"Section 3.2"},{"comment":"The source-proportion pie charts in Figure 2 use a 'bilibili+others' label without a legend explaining what 'others' includes. Given that Section 3.1 states that 'nearly 200' videos came from Bilibili, clarifying the exact contribution of Bilibili and the composition of 'others' would remove an apparent ambiguity.","section":"Figure 2"},{"comment":"The distinction between the zero-shot evaluation of the five deep models and the fine-tuned rows marked with '*' is not fully explained in the text. Please state explicitly which rows in Table 1 correspond to zero-shot inference with official pretrained weights and which rows were trained or fine-tuned on RGCD, since this affects how readers interpret the generalization claim.","section":"Section 4.3, Table 1"}],"recommendation":"major_revision","confidential_remarks":"The provenance issue in Section 3.1 is the most serious risk. If the released per-source metadata confirm that the 'wheel' category is predominantly human-driven dashcam footage, the database's core identity as robotic-generated content would be compromised, and the paper would need substantial reframing or removal of that category. I would request the source metadata and a re-analysis on a verified robot-only subset as part of the revision. The MOS formula issue in Eq. (3) is likely a typographical error but must be corrected before the MOS values can be trusted."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing you should know: this is a genuinely useful dataset contribution. RGCD is the first MOS-labeled VQA database aimed at robot-generated video — 2,100 clips across drones, wheeled robots, and robotic arms, with a real in-lab subjective experiment and 31.5K ratings. The benchmark showing that existing VQA models transfer poorly to this content is a find worth reporting. But the central provenance claim is shaky. The 'wheel' category is sourced from autonomous-driving datasets [1,11,30,31,48]; at least JAAD, BDD100K, KITTI, and Ford Campus are ordinary dashcam recordings from human-driven cars, not video captured by robots operating autonomously or under remote control — which is the paper's own definition of RGC. If that is correct, a third of the database is not RGC, and the benchmark conclusion is confounded by driving-domain content rather than robotic generation. This is the most serious issue, and it is verifiable from the source metadata. The paper is honest elsewhere: in Sec. 3.1 it admits it could not find authentic egocentric locomotion data for humanoid or quadruped robots, so the 'robot' category is largely wrist-camera footage from robotic-arm datasets. That honesty is good, but it narrows the claim. What the paper does well: the gap is real, no prior VQA dataset covers this content, and the subjective experiment follows ITU-R BT.500 with outlier rejection and a sensible rating protocol. The feature and MOS distribution analyses are reasonable. Benchmarking 11 methods with official implementations is standard practice. The soft spots beyond provenance: Eq. (3) is internally inconsistent — it multiplies the rescaled z-scores by the subject's standard deviation, which contradicts the described z-score averaging; if literally used, the MOS is not a simple mean and the standard analysis is off. The benchmark uses a single 4:1 split with no error bars or significance tests, so small SRCC differences (around 0.02) are not meaningful. The dataset URL and code cannot be verified from the paper alone. These are all addressable. The wheel category can be re-curated or re-benchmarked on a robot-only subset; the MOS formula can be corrected; multi-split with confidence intervals is routine. If the authors fix these, this is a solid MM-level dataset paper. As is, I would want major revision, not desk rejection. For peer review: yes, send it — the database fills a real niche and the issues are testable, not fatal. But the referee should demand provenance verification for all 700 wheel videos and a corrected MOS pipeline.","headline":"A useful first dataset for robot-generated video quality, but the wheeled-robot category looks like ordinary dashcam footage from human-driven cars, and the MOS formula has a bug; fixable but requires major revision.","tokens_in":15662,"tokens_out":2421,"would_cite":false,"duration_ms":28588,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Robotic-generated video is a distinct content class; the paper builds the first 2,100-video quality benchmark for it and finds all 11 existing VQA models underperform.","keywords":["Video quality assessment","Robotic-generated content","RGCD database","Mean opinion score","Egocentric video","No-reference VQA benchmark","Human-robot interaction"],"falsifier":"Collect authentic egocentric video from humanoid or quadruped robots in locomotion, rate it under the same SSACR protocol, and rerun the benchmark; if these clips produce MOS distributions and model rankings that match the current robot category, the category assumption holds, while a divergence would show that RGCD's robot arm results do not generalize to mobile humanoid robots.","tokens_in":14696,"feed_emoji":"🤖","tokens_out":6434,"duration_ms":65270,"temperature":0.7,"pith_summary":"The paper proposes Robotic-Generated Content (RGC) as a category of video distinct from professionally generated and user generated content: footage captured from the egocentric perspective of robots. To make that category measurable, it builds RGCD, the first quality-assessment database for such video, with 2,100 clips evenly split among drones, wheeled robots, and robot arms, each scored by human viewers in a controlled lab study yielding mean opinion scores. Benchmarking 11 state-of-the-art video quality models on RGCD, it finds that none predicts human quality judgments reliably across all categories, with the largest gaps on drone and wheeled-robot footage. The paper's case is that this failure is systematic rather than incidental, so RGC needs models that account for machine-specific motion, egocentric viewpoints, and device-induced artifacts.","feed_headline":"2100 robot-made videos expose blind spot in video quality AI","feed_subtitle":"First MOS-rated benchmark for drone, wheeled-robot, and robot-arm footage shows current VQA models underperform.","key_machinery":"The central object is RGCD itself, a corpus built to be the measuring instrument: 700 drone clips, 700 wheeled-robot clips, and 700 robotic-arm clips, trimmed to 4-12 seconds and drawn from SLAM, driving, aerial, and manipulation datasets plus real streaming-platform footage. Quality is captured through a single-stimulus absolute category rating (SSACR) experiment in which each clip is shown alone and rated on a continuous 0-5 scale, then z-score normalized and averaged into mean opinion scores. The benchmark protocol standardizes a 4:1 train/test split and reports SRCC, KRCC, and PLCC for 11 models, which is what exposes the performance gap.","core_discovery":"On the paper's own terms, the central claim is that robot-view video forms a coherent but unstudied content regime whose perceptual quality cannot be captured by models built for professional or user content. The evidence is RGCD: 2,100 in-the-wild clips from three robot types, 31,500 lab ratings from 15 subjects, and a benchmark showing that the best existing no-reference models reach only moderate correlation with human scores, with performance varying sharply by robot category. From this the authors conclude that no existing VQA model is extensively reliable for RGC and that RGC-tailored evaluation is a necessary next step.","pith_inferences":["An immediate testable extension is to fine-tune the strongest baselines (SimpleVQA, FAST-VQA) on half of RGCD and evaluate on the held-out half; if a generic architecture then reaches near-human correlation, the bottleneck is domain training data rather than model architecture.","Because the robot category is assembled from wrist-camera manipulation footage, the benchmark says little about humanoid or quadruped locomotion; authentic egocentric locomotion data, when it appears, might change the category rankings.","The RGC concept plausibly extends beyond the three device types to autonomous vehicle camera feeds, surgical robots, and inspection bots, so the database format could be reused rather than rebuilt."],"forward_implications":["RGCD gives the VQA community a public testbed on which claims of general video-quality ability can be checked against robot-view content.","Existing no-reference VQA models cannot be assumed ready for teleoperation, surveillance, delivery, or search-and-rescue footage without retraining or adaptation.","Models that perform well on robot-arm clips, such as TLVQM and SimpleVQA, may benefit from the strong repetitive motion cues in that category; drone and wheeled footage remains the harder test.","DOVER's weaker showing hints that aesthetic-oriented quality factors can hurt on RGC, where task-relevant clarity matters more than framing."],"supporting_citations":[{"why":"Curated SLAM dataset list from which drone and wheeled-robot RGB sequences were selected.","marker":"[5]"},{"why":"Open X-Embodiment supplies cross-robot manipulation data used for the humanoid/robot-arm egocentric view collection.","marker":"[6]"},{"why":"ARIO provides unified embodied-agent data, another source of wrist-camera and egocentric robot views.","marker":"[42]"},{"why":"ITU-R BT.500 recommendation governs the lab subjective experiment design and outlier rejection.","marker":"[32]"},{"why":"TLVQM is the strongest traditional VQA baseline whose category-level correlations anchor the benchmark comparison.","marker":"[18]"},{"why":"SimpleVQA is a top deep baseline, used both zero-shot and fine-tuned, whose results mark the state of the art being challenged.","marker":"[35]"},{"why":"FAST-VQA is the other leading deep baseline demonstrating the generalization gap on drone and wheeled content.","marker":"[43]"}],"fun_headline_variants":["2,100 robot videos reveal quality-AI gap","RGCD: robot videos expose VQA limits","Robot-made videos elude quality-check AI","VQA models underperform on robot video"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The 'robot' category is assumed to represent humanoid-robot video, but it is built almost entirely from robotic-arm wrist cameras because authentic humanoid locomotion footage could not be found.","fun_headline_variants_meta":{"raw":{"variants":["2,100 robot videos reveal quality-AI gap","RGCD: robot videos expose VQA limits","Robot-made videos elude quality-check AI","VQA models underperform on robot video"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001171,"raw_usage":{"total_tokens":4825,"prompt_tokens":912,"completion_tokens":3913,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":528,"completion_tokens_details":{"reasoning_tokens":3855}},"tokens_in":528,"tokens_out":3913,"duration_ms":31460,"temperature":1.0,"reasoning_tokens":3855,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T21:30:15.854545+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Collect authentic egocentric video from humanoid or quadruped robots in locomotion, rate it under the same SSACR protocol, and rerun the benchmark; if these clips produce MOS distributions and model rankings that match the current robot category, the category assumption holds, while a divergence would show that RGCD's robot arm results do not generalize to mobile humanoid robots.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Curated SLAM dataset list from which drone and wheeled-robot RGB sequences were selected."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"ITU-R BT.500 recommendation governs the lab subjective experiment design and outlier rejection."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"TLVQM is the strongest traditional VQA baseline whose category-level correlations anchor the benchmark comparison."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"SimpleVQA is a top deep baseline, used both zero-shot and fine-tuned, whose results mark the state of the art being challenged."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"FAST-VQA is the other leading deep baseline demonstrating the generalization gap on drone and wheeled content."}],"review_version":1}