{"id":"1126a5d6-f320-4d0b-89e6-74c255f4215e","arxiv_id":"2502.07327","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Text-video retrieval models systematically rank AI-generated videos above semantically matched real videos, driven by both visual and temporal cues and amplified by AI content in training data.","lead":"This paper builds a video-search benchmark pairing 13,000 AI-generated clips with real videos and finds that three standard retrieval models rank the AI clips higher for the same text query, a bias that grows when AI clips enter the training data. It adds new bias metrics and a contrastive fine-tuning fix, but the fix appears to invert the bias rather than remove it.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Equal-relevance premise for the headline bias metric is contradicted by the paper's own human evaluation (Table 2), so 'same relevance level' and all reported magnitudes are not established; a human-labeled equal-subset test is needed.","rationale":"The reader's weakest assumption is exactly the load-bearing one. The central assertion in Section 3.2 and the abstract explicitly invokes 'same relevance level', and the only basis for assigning equal relevance is Benchmark Requirement (1). Table 2 shows that requirement is not met for a substantial share of pairs: real videos are judged more relevant far more often than AI videos. Since every reported metric is computed under the equal-relevance label assignment, the headline is not established as stated. I considered whether this is fatal: it is not, because if human raters say the real video is more relevant, a model ranking the AI video above it is even more clearly making a source-biased error, so the direction may survive. However, the magnitude and the precise framing need repair. NormalizedDelta was designed to fix the semantic gap, but its random-interleaving null model is unvalidated, so it cannot rescue the measurement. The debiasing overcorrection in Table 7 and the causal p-vector claim in Section 5.2 are real secondary issues, but they do not threaten the existence claim as directly as the equal-relevance premise. The reader's CONDITIONAL verdict is therefore appropriate, and no adjustment is needed.","tokens_in":20867,"tokens_out":5380,"duration_ms":52817,"concrete_test":"Use the existing Table 2 annotation protocol to build an equal-relevance subset: for each dataset, retain only pairs where at least two of three raters judged the real and AI videos equally relevant to the query. Re-run the mixed retrieval experiments for all three models on this subset and recompute RelativeDelta, NormalizedDelta, and MixR. If the AI-over-real preference persists (e.g., NormalizedDelta MixR remains negative) on the equal-relevance subset, the central claim survives; if the effect attenuates or flips, the headline claim and reported bias magnitudes are artifacts of the equal-relevance label assumption.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim depends on Benchmark Requirement (1), 'Identical Semantics': each AI-generated video is assumed to share the query relevance of its real counterpart, so ranking the AI video above the real video is counted as source bias. The paper's own human evaluation (Table 2) contradicts this premise: raters judged the real video more relevant in 32%–47% of sampled pairs and the AI video in only 13%–18%, with 'Equal' in 40%–53%. Because RelativeDelta, NormalizedDelta, and MixR (Eqs. 1–5) assign equal relevance to every real/AI pair, the abstract and Section 3.2 claim 'even when both have the same relevance level' is not established by the benchmark, and the reported magnitudes conflate source bias with label misspecification. NormalizedDelta was designed to correct for semantic discrepancy, but its correction relies on the random-interleaving null model in Eqs. (2)–(3), whose validity is not demonstrated; the signed NormalizedDelta values in Table 3 even flip across models and datasets. The direction of the effect may survive on an equal-relevance subset, but the quantitative claims and the 'same relevance level' framing inherit an unmet assumption.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper introduces a benchmark of 13,000 AI-generated videos derived from MSR-VTT using CogVideoX and OpenSora, with four test conditions (text-conditioned, image-conditioned, video-extended, plus a 9,000-video training split), and proposes RelativeΔ, NormalizedΔ, and MixR metrics to quantify whether text-video retrieval models rank AI-generated videos above real videos under an 'identical semantics' assumption. Using Alpro, Frozen in Time, and InternVideo, the paper reports predominantly negative metric values, interprets them as Visual-Temporal Induced Source Bias, reports that training with increasing AI-video proportions intensifies the bias, attributes the bias to extra visual and temporal information, and proposes a contrastive-learning debiasing method plus a debiasing vector p.","tokens_in":20944,"tokens_out":5644,"duration_ms":50933,"significance":"If the central claim held, this would be a timely and useful contribution: video retrieval on mixed real-and-AI libraries is an emerging problem, the benchmark is large and publicly released, and the use of three retrieval models, two generators, and a PIKA spot-check provides valuable external grounding. The paper is also honest in its limitation appendices. The directional finding—AI-generated videos tend to be ranked above real counterparts in mixed retrieval—is likely robust. However, the headline 'same relevance level' claim and the quantitative bias magnitudes rest on an unverified equal-relevance premise that the paper's own human evaluation contradicts. At present the contribution is a promising measurement framework plus a directional finding, rather than an established quantitative bias result.","major_comments":[{"comment":"The benchmark's first requirement, 'Identical Semantics,' is load-bearing: the claim that retrieval models rank AI-generated videos higher 'even when both have the same relevance level' depends on every real/AI pair sharing the query's relevance label. Table 2 shows that human evaluators judged the real video more relevant in 32%–47% of sampled pairs and the AI video in 13%–18%, with equal judgments in 40%–53%. Because RelativeΔ, NormalizedΔ, and MixR (Eqs. 1–5) assign equal relevance to every pair, the headline comparison is made under a label assignment that the authors' own data contradict. The paper needs a human-labeled equal-relevance subset (or per-pair relevance stratification) and must report the bias metrics on that subset; without this, the reported magnitudes conflate source bias with relevance misspecification. The direction of the effect may survive, but the quantitative claims and the 'same relevance level' framing do not.","section":"§2 Requirement (1) and Table 2; abstract and §3.2"},{"comment":"NormalizedΔ is introduced as a correction for semantic discrepancy, but it relies on the random-interleaving null model in Eqs. (2)–(3), in which mixed rankings are simulated by doubling independent ranks and subtracting a random offset c. No justification is given that this interleaving matches the candidate-pool size and score competition of the actual mixed list, and the distribution of c is unspecified (the notation c ∈ 0,1 is ambiguous between {0,1} and the interval [0,1]). Without validation of this null model, the signed NormalizedΔ values—which flip across models and datasets (e.g., InternVideo R@1 +34.67 on OpenSora TextCond versus Frozen's –64.89)—cannot be interpreted as calibrated bias measurements.","section":"§2.4, Eqs. (2)–(3)"},{"comment":"The training-loop corollary, that the model's preference for AI-generated content strengthens as the AI share grows from 0% to 80%, is supported by a figure plus a detailed 20% example rather than by a full quantitative comparison. The text's summary numbers are also hard to parse: it says NormalizedΔ R@1 'increases by 49.29 points' when fine-tuning on real videos, and then reports a 'decrease of 89.52' when 20% AI is added. The paper should report the complete 0/20/40/60/80 table with significance and variability, and state explicitly whether the trend is monotonic for every metric. As written, the central training-loop claim is not fully established.","section":"§3.3 and Figure 3"},{"comment":"The debiasing loss in Eq. (10) is constructed from Δr = score(AI) – score(real) and is applied only when Δr ≥ 0, so it directly penalizes the model for ranking AI videos above real videos. Table 7 then reports mixed-AI R@1 = 0 and RelativeΔ = 200, a saturated outcome; this demonstrates that the objective was optimized, not that an independently measured source bias was reduced. Relatedly, the p vector in Eq. (11) is the difference between the original encoder and this same debiased encoder, so the claim that p captures 'additional information embedded by generation encoders' risks circularity: adding p_avg to real representations is equivalent to adding the learned debiasing direction. Independent evidence, such as probing generation encoders directly or controlling for relevance, is needed before p can be interpreted as the cause of the bias.","section":"§5.1, Eq. (10); §5.2, Eq. (11)"}],"minor_comments":[{"comment":"Figure 1 contains the typo 'addition imformation,' and the caption states the causal mechanism ('extra visual and temporal information embedded by the generation model') more strongly than the experiments independently support.","section":"Figure 1"},{"comment":"Please specify whether c is drawn from {0,1} or from the interval [0,1], and report sensitivity of LocationΔ and NormalizedΔ to the choice of c.","section":"§2.4, Eqs. (2)–(3)"},{"comment":"Table 1 reports CLIP similarity between generated and real videos (0.72–0.87), while Table 2 shows that human raters often prefer the real video's query relevance; the paper should clarify that these measure different constructs and discuss the apparent tension.","section":"§2.3, Tables 1 and 2"},{"comment":"The statistical significance appendix reports 'paired T-tests' but gives no sample size, no correction for multiple comparisons, and no effect sizes; please add these details.","section":"Appendix A"},{"comment":"The sentence 'when Δr < 0, we do not apply this loss function, ensuring that during contrastive learning, the model still favors generated videos' appears to contradict the stated goal of reducing preference for AI-generated videos and should be rephrased.","section":"§5.1"},{"comment":"The PIKA check uses only 100 videos and a single retrieval model; it is useful as a spot check, but the statement that 'Visual-Temporal Induced Source Bias might be further amplified' in commercial models goes beyond what this check can support.","section":"Appendix B"}],"recommendation":"major_revision","confidential_remarks":"The equal-relevance premise is the main barrier to acceptance. The reported direction of the effect is likely robust, and the manuscript already contains the human-evaluation data needed to diagnose the problem. If the authors re-run the analysis on a human-labeled equal-relevance subset and temper the 'same relevance level' wording, the paper could become publishable. I do not see a load-bearing error that is unfixable within the manuscript's scope. The relation to prior image source-bias work is acknowledged, and the video-specific benchmark and temporal analysis are sufficiently novel for a multimedia venue if the central measurement is re-grounded."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a useful empirical paper, likely right about the direction of the effect, but the headline magnitudes rest on an assumption the paper's own data contradict, and the strongest model shows mixed results.\n\nThe real contributions are the benchmark and the measurement toolkit. Thirteen thousand generated videos from two open-source generators across four conditions, tested on three retrieval models, plus a PIKA robustness check: that is a solid base. The NormalizedΔ and MixR metrics are a reasonable attempt to separate source bias from semantic mismatch, and the training-loop experiment (Figure 3) is a nice demonstration that AI-generated training data amplifies the effect. The visual/temporal decomposition via frame shuffling and single-frame retrieval is genuinely informative. Public code and data are a plus.\n\nThe soft spots are real. First, the equal-relevance premise (Benchmark Requirement 1) is not met. Table 2 shows human raters judge the real video more relevant in 32–47% of pairs, versus 13–18% for the AI video. The headline RelativeΔ and NormalizedΔ are computed under the assumption that each pair shares the query relevance. The authors acknowledge the difficulty in §2.3 but still present 'even when both have the same relevance level' as a finding. The direction of the bias may survive on an equal-relevance subset—if anything, ranking AI videos above semantically superior real videos makes the bias more striking—but the magnitudes, and the 'same relevance' framing, are not established. Second, NormalizedΔ's correction depends on a random-interleaving null model that is asserted rather than validated. Third, InternVideo shows positive RelativeΔ for R@1 on three of four datasets, so the 'clear preference for AI' is not uniform across models; the paper overstates. Fourth, the debiasing method in Table 7 pushes AI videos to R@1=0 with median ranks over 200—an inverted anti-AI bias that is not acknowledged. Fifth, the ImageCond keyframe position was selected on the test data, and the statistical appendix is a one-paragraph assertion of p<0.05 with no effect sizes or test details. Finally, §5.2 calls the clustering of the p-vector a 'direct cause'; the shuffle and t-SNE evidence support association, not causation, and Appendix D itself concedes the mechanisms need more work.\n\nWho should read this: anyone working on AIGC bias in retrieval or building mixed real/AI video collections. The benchmark is a reusable asset and the phenomenon is worth knowing about. But the quantitative claims should be treated as provisional until the equal-relevance issue is addressed with a human-labeled equal-relevance subset.\n\nRecommendation: yes, send to serious peer review. The core measurement is externally grounded and the benchmark is a contribution, but the authors need to fix the framing, add error bars, and either demonstrate the equal-relevance subset or soften the claim.","headline":"Useful benchmark, likely real effect, but the headline bias magnitudes rest on an unmet equal-relevance assumption and the strongest model shows mixed results.","tokens_in":21694,"tokens_out":4275,"would_cite":true,"duration_ms":38023,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that text-video retrieval models rank AI-generated videos above equally relevant real videos, calls the effect Visual-Temporal Induced Source Bias, and shows that mixing AI videos into training data strengthens it while…","keywords":["text-video retrieval","AI-generated content","source bias","Visual-Temporal Induced Source Bias","retrieval fairness","contrastive debiasing","benchmark construction","temporal information"],"falsifier":"Filter the benchmark to pairs that human annotators judge equally relevant (or replace equal-relevance labels with human relevance scores), recompute Normalized $\\Delta$ on that subset for the same three retrieval models, and check whether AI-generated videos still rank above real videos; if the preference disappears, the paper's central claim fails on its own benchmark.","tokens_in":20477,"feed_emoji":"🎬","tokens_out":6664,"duration_ms":54875,"temperature":0.7,"pith_summary":"The paper argues that text-video retrieval models systematically rank AI-generated videos above real videos that carry the same relevance label, and that this preference comes from extra visual and temporal signals that generative models embed in their output. The authors construct a benchmark pairing real videos from MSR-VTT with AI videos made by two open-source generators, and they measure retrieval with a new metric, Normalized $\\Delta$, designed to separate source bias from quality and semantic gaps. Across three retrieval models and four test sets, the AI videos rank higher, and the preference strengthens as the share of AI-generated videos in the training set grows from 0% to 80%. A contrastive fine-tuning step mostly reverses the preference. If the claim holds, any retrieval system indexing a mixed real-and-AI library will silently over-expose synthetic content and may feed that bias back into future training data.","feed_headline":"AI-generated videos outrank real videos in retrieval tests","feed_subtitle":"Ranking gap grows as AI clips enter training data; contrastive fine-tuning can close it.","key_machinery":"The quantitative core is the Normalized $\\Delta$ metric, defined as Relative $\\Delta$ minus Location $\\Delta$. Relative $\\Delta$ compares the mixed-retrieval rankings of real and AI videos, while Location $\\Delta$ simulates a no-interference baseline by interleaving independently computed rankings with a random offset, so the subtraction is meant to isolate genuine source bias from semantic or quality gaps. The supporting machinery includes frame-shuffle and single-frame ablations that separate the temporal and visual components of the bias, and a contrastive objective $\\Delta r = E_V(V_G) - E_V(V_R)$ that is added to the retrieval loss to push real videos above AI-generated ones.","core_discovery":"On the paper's own terms, text-video retrieval models are biased toward AI-generated videos: for a query, an AI video that shares the relevance label of its real counterpart is ranked above that counterpart across essentially all model-dataset combinations. The authors name this effect Visual-Temporal Induced Source Bias and locate its causes in both visual content and temporal structure, showing that randomizing frame order reduces the temporal contribution and that single-frame retrieval leaves a residual visual preference. The bias grows monotonically as the proportion of AI-generated videos in the training set increases, with a visible effect already at 20% AI content. The authors attribute the bias to additional, highly consistent information that video generators encode into their outputs, and they demonstrate a contrastive debiasing objective that substantially reduces the preference by shifting scores toward real videos.","pith_inferences":["Because the paper's human evaluation contradicts its equal-relevance premise, the headline bias magnitudes are likely inflated; the direction of the effect would probably survive a relevance-matched re-test, but that is an editorial reading, not the paper's claim.","The clustering of the debiasing vector $p$ suggests AI video generators imprint a common statistical signature on their output, one that could serve as a provenance detector or watermark outside retrieval tasks.","The same visual-temporal preference likely extends to video recommendation and autoplay ranking, not just text-video retrieval; a logged-interaction study on a mixed catalogue would test that.","If better commercial generators embed richer temporal information, the bias may grow rather than shrink as generation quality improves, since the paper's mechanism treats temporal richness as high-relevance signal."],"forward_implications":["Video search engines that index mixed libraries will systematically place AI-generated clips above equally relevant real footage, shaping which content users see first.","As AI videos accumulate online and enter training sets, the preference self-amplifies: even a 20% training share changes ranking behavior, so the bias compounds across model generations.","Retrieval evaluation on mixed real-and-AI corpora needs relevance-matched human judgment or metrics like Normalized $\\Delta$; otherwise measured bias is conflated with quality differences.","Contrastive debiasing works as a post-hoc fine-tuning step, and the extracted debiasing vector can be transferred to other videos, offering a practical mitigation without architectural changes."],"supporting_citations":[{"why":"Establishes source bias in text-image retrieval, which this paper extends to the video modality.","marker":"[33]"},{"why":"Shows neural retrievers favor LLM-generated text, motivating the investigation of retrieval bias toward AI-generated content.","marker":"[10]"},{"why":"Supplies the Frozen in Time retrieval model and the MSR-VTT split used in the benchmark.","marker":"[3]"},{"why":"Supplies the ALPRO retrieval model used in all main experiments.","marker":"[17]"},{"why":"Supplies the InternVideo retrieval model, also used for the training-loop and debiasing experiments.","marker":"[27]"},{"why":"One of the two video generators used to produce the benchmark's AI videos.","marker":"[35]"},{"why":"The other video generator, used for most of the AI-generated datasets including the training set.","marker":"[20]"},{"why":"Provides the real videos and captions from MSR-VTT that anchor the benchmark.","marker":"[31]"},{"why":"Provides the contrastive learning objective used in the debiasing fine-tuning.","marker":"[6]"}],"fun_headline_variants":["AI videos beat real ones in retrieval rankings","Retrieval models prefer AI-generated videos","Ranking bias favors AI videos; training data worsens it","Contrastive fine-tuning counters AI video retrieval bias","AI video bias: visual and temporal cues boost rankings"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The benchmark labels every AI-generated video as equally relevant to its real counterpart, but the paper's own human evaluation judged the real video more relevant in 32% to 47% of pairs (versus 13% to 18% for the AI video), so the equal-relevance premise underlies all reported bias magnitudes.","fun_headline_variants_meta":{"raw":{"variants":["AI videos beat real ones in retrieval rankings","Retrieval models prefer AI-generated videos","Ranking bias favors AI videos; training data worsens it","Contrastive fine-tuning counters AI video retrieval bias","AI video bias: visual and temporal cues boost rankings"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000156,"raw_usage":{"total_tokens":1246,"prompt_tokens":998,"completion_tokens":248,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":614,"completion_tokens_details":{"reasoning_tokens":174}},"tokens_in":614,"tokens_out":248,"duration_ms":2754,"temperature":1.0,"reasoning_tokens":174,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T13:07:03.426766+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Filter the benchmark to pairs that human annotators judge equally relevant (or replace equal-relevance labels with human relevance scores), recompute Normalized $\\Delta$ on that subset for the same three retrieval models, and check whether AI-generated videos still rank above real videos; if the preference disappears, the paper's central claim fails on its own benchmark.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes source bias in text-image retrieval, which this paper extends to the video modality."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the Frozen in Time retrieval model and the MSR-VTT split used in the benchmark."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the ALPRO retrieval model used in all main experiments."},{"cited_title":"Open-Sora: Democratizing Efficient Video Production for All","cited_arxiv_id":null,"evidence_quote":"The other video generator, used for most of the AI-generated datasets including the training set."}],"review_version":1}