{"id":"486c683a-4274-460f-b9b0-579d0de7a8b6","arxiv_id":"2507.20518","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"T2VParser uses shared learnable decomposition tokens to parse video and text into multiview embeddings and align matching views, improving text-to-video retrieval on four benchmarks.","lead":"This paper introduces T2VParser, a method that splits video and text into several shared semantic views using learnable tokens, then aligns only the matching views for retrieval. It reports higher retrieval accuracy than several CLIP-based baselines on four video-text benchmarks and on noisy text settings.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Claimed SOTA and consistent gains rely on mismatched backbones: Table 1's 58.4 for T2VParser+CLIP-VIP/Mug-STAN matches B/16+DSL rows, while baselines are B/32; Table 2 B/32 variants underperform published baselines.","rationale":"The paper's primary contribution is an empirically demonstrated retrieval improvement: the abstract and introduction claim state-of-the-art results, and Table 1 is presented as evidence of consistent gains across baselines. However, a careful cross-check of Table 1 against Table 2 reveals that the largest reported gains (e.g., 58.4 R@1 for T2VParser+CLIP-VIP and T2VParser+Mug-STAN) correspond to the CLIP-ViT-B/16 plus DSL configurations, while the baselines are the published B/32 numbers. This is a benchmarking mismatch, not a matter of missing error bars or unreleased documents. The paper itself contains the controlled B/32 numbers showing T2VParser below the published baselines (55.2 and 55.7 vs. 57.7 and 57.3 on MSR-VTT), so the overclaim is internally visible. The reader's identified weak point (unverified cross-modal token correspondence) is a legitimate theoretical concern, but it is secondary: if the empirical comparison is unfair, the central claim fails regardless of whether the decomposition mechanism is sound. The manuscript could potentially be revised with corrected comparisons and tempered claims, but the current version's central empirical assertion is not supported. Hence the verdict should move from CONDITIONAL to REJECT.","tokens_in":15448,"tokens_out":17325,"duration_ms":173748,"concrete_test":"Retrain the relevant Table 1 rows under a single matched protocol for both baseline and T2VParser: same CLIP ViT-B/32 backbone, same DSL usage (either both with or both without), same frame sampling, and same training epochs. Compare T2VParser+CLIP-VIP and T2VParser+Mug-STAN directly against the published baselines (57.7 and 57.3 R@1 on MSR-VTT). If T2VParser does not exceed these published numbers under matched conditions, the central claim of consistent state-of-the-art improvements is unsupported and the tables and abstract must be corrected.","verdict_should_be":"REJECT","load_bearing_attack":"The paper's central empirical claim is that T2VParser 'achieves state-of-the-art performance' and yields consistent improvements across baselines (Secs. 4.1-4.2). This claim is not supported under matched experimental conditions. In Table 1, T2VParser+CLIP-VIP and T2VParser+Mug-STAN report 58.4 R@1 on MSR-VTT, but these exact numbers appear only in Table 2's CLIP-ViT-B/16 rows with the DSL trick (T2VParser+CLIP-VIP†* and T2VParser+Mug-STAN†*), whereas the baseline scores in Table 1 (57.7 and 57.3) are the published B/32 numbers. For the B/32 setting explicitly reported in Table 2, T2VParser+CLIP-VIP†* scores 55.2 and T2VParser+Mug-STAN†* scores 55.7, both below the published baselines. Thus the headline improvements are an artifact of comparing a stronger backbone plus DSL against B/32 published numbers, and the paper does not disclose the backbone/DSL used for the T2VParser entries in Table 1. A controlled B/32 comparison against their own re-implemented baselines shows smaller gains (e.g., CLIP-VIP from 50.1 to 51.3 without DSL), but those re-implemented baselines are far below published scores, so the absolute SOTA claim still fails. The reader's concern about unverified cross-modal token correspondence is real but secondary: even if the mechanism were sound, the current experiments do not establish the claimed retrieval advantage.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes T2VParser, a model-agnostic framework for text-to-video retrieval that decomposes video and text into multiview embeddings using a set of modality-shared learnable Adaptive Decomposition Tokens (ADTs). A Dual Communication Mechanism exchanges and filters information between the resulting view sets, and training combines contrastive alignment with a diversity loss and with synthetic 'video documents' generated by key-frame captioning and LLM summarization. The authors report consistent gains over several baselines on MSR-VTT, MSVD, DiDeMo, and ActivityNet, as well as on self-created long-text versions of MSR-VTT and MSVD, and they claim state-of-the-art performance.","tokens_in":15808,"tokens_out":6200,"duration_ms":59959,"significance":"If the reported gains hold under matched experimental conditions, the ADT decomposition and Dual Communication Mechanism constitute a reusable approach to partial text-video alignment that preserves the knowledge of pretrained image-text encoders. The paper includes ablations for the main components, releases code, and evaluates robustness to noisy and partially relevant content, which is a practically relevant scenario. However, the headline state-of-the-art claim currently rests on comparisons that mix backbones and the DSL trick, and the central premise that same-index token views correspond semantically across modalities is not validated. The experiments also lack error bars or significance testing, which matters because several controlled gains are only about one R@1 point.","major_comments":[{"comment":"The claimed state-of-the-art and consistent improvement results are not established under matched experimental conditions. In Table 1, T2VParser+Mug-STAN and T2VParser+CLIP-VIP report 58.4 R@1 on MSR-VTT-1k against baselines of 57.3 and 57.7, but those T2VParser numbers appear in Table 2 only in the CLIP-ViT-B/16 rows with the DSL trick (T2VParser+Mug-STAN†* and T2VParser+CLIP-VIP†*), while the Table 1 baselines are not identified as B/16 or B/32 or as with/without DSL. In the controlled B/32 rows of Table 2, T2VParser+Mug-STAN† improves only from 50.9 to 51.6 and T2VParser+CLIP-VIP† from 50.1 to 51.3, which are far below published baseline numbers. The absolute state-of-the-art claim is also contradicted by ActivityNet, where CLIP-VIP* (61.4) exceeds T2VParser+Mug-STAN†* (60.5). Please report every comparison with matched backbone, input resolution, and DSL usage, and avoid claiming state-of-the-art on the basis of the mixed-condition numbers in Table 1.","section":"Section 4.2, Tables 1 and 2"},{"comment":"The method assumes that the i-th adaptive decomposition token in the video parser and the i-th token in the text parser capture corresponding semantic views, because the final matching score uses the diagonal of the similarity matrix in Eq. (9). Nothing in the formulation enforces or measures this cross-modal correspondence: the diversity loss in Eq. (17) separates views only within each modality, and the parsers are trained independently with supervision applied only to the final aggregated vectors e_fv and e_ft. Figure 4 demonstrates within-modality diversity but does not show cross-modal view correspondence. This is a load-bearing assumption for the paper's 'partial alignment' interpretation. Please provide quantitative evidence of cross-modal token correspondence (for example, per-view retrieval analysis, attention-overlap statistics, or canonical correlation), or reformulate the claim so that it does not depend on same-index semantic matching.","section":"Section 3.1, Eqs. (1)-(5) and Section 3.2, Eqs. (6)-(10)"},{"comment":"The central empirical demonstration relies on self-created long-text datasets (MSR-VTT-Doc and MSVD-Concat), but these datasets are not released and the code repository link does not guarantee access to them. Since the paper argues that the method's advantage grows with text complexity, the long-text benchmarks are necessary for reproduction. In addition, all reported numbers appear to be single runs with no error bars or significance testing; given that the matched-condition gains in Table 2 are around 0.7 to 1.9 R@1 points, the paper should report multiple seeds and variance, or a significance test. Please release the generated long-text datasets (or the full generation pipeline with seeds) and add variance estimates.","section":"Section 4.1 and Section 4.2"}],"minor_comments":[{"comment":"The text in Section 5 says T2VParser introduces a 15% longer per-query retrieval time, but Table 7 shows 283 ms versus 192 ms for CLIP-VIP (47% longer) and 834 ms versus 584 ms (43% longer). Please reconcile the text with the table.","section":"Section 5 and Table 7"},{"comment":"References [21] and [22] are duplicates: both entries list the same title, authors, venue, and page range for 'Revisiting Temporal Modeling for CLIP-Based Image-to-Video Knowledge Transferring', yet Table 2 cites [22] as X-CLIP (CVPR'23). The X-CLIP citation should point to the correct paper.","section":"References"},{"comment":"There is a discrepancy about when the generated documents are used: Section 4 says 'the extra captions only use in training stage', but Section 4.1 constructs MSR-VTT-Doc by generating documents for both training and testing splits and evaluates on them. Please clarify whether the long-text test sets are used at test time and how this is consistent with the statement that extra captions are training-only.","section":"Implementation Details and Section 4.1"},{"comment":"Several table captions and headers contain fragments rather than complete sentences, such as Table 1's caption ending with 'MSVD-Concat(288.3)' and Table 2's stray 'Caption' column label. Please clean up the table formatting.","section":"Tables"}],"recommendation":"major_revision","confidential_remarks":"The backbone/DSL mismatch in Tables 1 and 2 is the main obstacle: the current write-up claims state-of-the-art on the basis of numbers that are only obtained with a stronger backbone plus the DSL trick, while the controlled B/32 gains are much smaller. The cross-modal correspondence assumption is a genuine scientific concern that should be addressed with evidence, not just asserted. If the authors fix the experimental presentation and release the long-text datasets, the paper could become publishable, but the claims as written are not supportable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing to know: the paper's central selling point—consistent SOTA on MSR-VTT, MSVD, DiDeMo, ActivityNet—does not survive a matched comparison. In Table 1 the T2VParser rows for Mug-STAN and CLIP-VIP report 58.4 R@1 on MSR-VTT-1k, but those exact numbers come from Table 2's ViT-B/16 rows with the DSL trick. The baselines they are compared against are the published B/32 numbers (57.3 and 57.7). When you look at the B/32 rows in Table 2, T2VParser+Mug-STAN†* scores 55.7 and T2VParser+CLIP-VIP†* 55.2, both below the published baselines. The paper never says which backbone/DSL was used in Table 1. That is a load-bearing omission, not a nitpick.\n\nWhat is genuinely new: a shared set of learnable decomposition tokens that parse video and text into multiview embeddings, a dual communication mechanism to filter alignment-relevant views, and training with LLM-generated video documents. The idea of partial alignment through semantic decomposition is reasonable, and the ablations show each piece contributes something. The code is promised, which helps. The efficiency section and the feature-collapse analysis are nice touches.\n\nSoft spots beyond the backbone mismatch: the re-implemented baselines in Table 2 are far below published numbers (e.g., CLIP-VIP† 50.1 vs published 57.7). That suggests their training recipe is not reproducing the baselines, so even the relative gains (e.g., +1.2 for CLIP-VIP B/32) are hard to interpret. There are no error bars or significance tests anywhere. The self-created long-text datasets (MSR-VTT-Doc, MSVD-Concat) are not released, so Section 4.1 is not independently checkable. And the central assumption—that the same token index corresponds across modalities—is never enforced or tested; the diversity loss only separates views within each modality. If the token indices don't correspond, the dual communication mechanism's softmax weighting loses its semantic foundation. That concern is real but secondary to the experimental mismatch.\n\nThis is a paper with a useful idea and a broken headline. The right path is peer review with a clear request: redo the comparisons with matched backbones and DSL settings, report variances, release the generated documents and code, and either verify or drop the cross-modal correspondence claim. The method itself is worth engaging with once the numbers are trustworthy.\n\nI'd bring it to reading group as a case study in how evaluation choices can overturn a SOTA claim. I would not cite the numbers until they're fixed.","headline":"The core idea is sound, but the headline SOTA numbers rest on a backbone mismatch that needs fixing before the paper's empirical claims can be trusted.","tokens_in":16340,"tokens_out":3799,"would_cite":false,"duration_ms":34950,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Decomposing text and video into shared semantic views lets retrieval align only the parts that match.","keywords":["text-to-video retrieval","partial alignment","multiview embeddings","Adaptive Decomposition Tokens","Dual Communication Mechanism","CLIP knowledge transfer","video-text pretraining","representation diversity"],"falsifier":"If, after permuting the order of the decomposition tokens on only the text side (or only the video side) at inference, retrieval recall is unchanged, then same-index correspondence between views carries no signal and the performance would be attributable to the attention-weighted aggregation rather than to the claimed semantic decomposition.","tokens_in":15257,"feed_emoji":"🎬","tokens_out":8861,"duration_ms":74429,"temperature":0.7,"pith_summary":"The paper argues that standard text-to-video retrieval trains on wrong supervision because a caption usually describes only part of a video, yet the whole text embedding is pushed toward the whole video embedding. It proposes T2VParser, which uses a small set of learnable tokens shared by both modalities to decompose text and video into multiview embeddings and then aligns only the corresponding views. A Dual Communication Mechanism weights the views before matching, and a diversity loss keeps the views from collapsing into a single representation. The paper reports consistent top-1 recall gains over its own baselines on MSR-VTT, MSVD, DiDeMo, and ActivityNet, with larger gains on longer texts and on noisy or partially relevant queries.","feed_headline":"Aligning only matching views lifts video-text retrieval","feed_subtitle":"Decomposing video and text into shared semantic views improves recall on MSR-VTT, DiDeMo, and noisy partial matches.","key_machinery":"The load-bearing object is the Adaptive Decomposition Token (ADT): a set of $k$ learnable tokens shared across modalities, used as queries in DETR-style cross-attention parsers that pull distinct semantic views out of the video and text representations. The same token index on the two sides is meant to correspond to the same view, which is what makes comparing the resulting embeddings meaningful. The Dual Communication Mechanism then exchanges information between the two modalities, extracts the diagonal of a video–text similarity matrix as a soft weighting over views, and produces the aggregated representations $\\tilde f_v$ and $\\tilde f_t$ used for retrieval. A representation diversity loss keeps the $k$ views mutually dissimilar, preventing the parser from collapsing all queries into one generic embedding.","core_discovery":"The central claim is that partial alignment—matching only the semantically overlapping components of a video and its caption—is both necessary and sufficient for better text-to-video retrieval, and that a shared set of Adaptive Decomposition Tokens can produce those components. On each side, the same $k$ learnable query tokens attend over the pretrained encoder's video or text features, yielding $k$ perspective-specific embeddings that are concatenated with the encoder's local features. The Dual Communication Mechanism computes a similarity matrix between the two sets of views, turns it into attention weights, and aggregates each side so irrelevant views are down-weighted instead of being forced to match. Contrastive training on the aggregated representations, plus a diversity loss on the views, gives the final retrieval score at inference. In the paper's experiments the gains grow as the text becomes richer, from roughly one recall point on the short-caption MSR-VTT-1k setting to over three points on ActivityNet, and the same model attached to CLIP4Clip, CLIP-VIP, or Mug-STAN improves each baseline.","pith_inferences":["A direct test of the token-correspondence assumption is to permute the ADT order on only one modality at inference: if same-index views are truly aligned, recall should drop noticeably under permutation; if not, the gains likely come from attention-weighted pooling rather than semantic decomposition.","The same shared-token decomposition idea could transfer to other retrieval settings where one side contains strictly more content than the other, such as image–text retrieval with long captions, document–query retrieval, or moment localization inside long videos.","Because no explicit cross-modal correspondence loss is applied to the tokens, the paper's framing is stronger than what the training objective enforces; probing the learned views with caption fragments or frame subsets would show which views actually specialize.","The document-generation step depends on an LLM's summary style, and the paper's ablation shows LLM choice matters little; an interesting extension is whether the structural prompt constraints, rather than the LLM itself, are what the parser benefits from."],"forward_implications":["T2VParser can be attached to existing CLIP-based video–text encoders without retraining the pretrained backbone, and it improves top-1 recall for each of the four encoder/baseline pairs tested.","The benefit is largest when the input content is information-rich: gains on DiDeMo, ActivityNet, and long-text versions of MSR-VTT and MSVD exceed gains on the short-caption MSR-VTT-1k split.","On partially relevant data, such as ActivityNet queries with 25 percent of captions replaced by captions from other videos and the PRVR benchmark, T2VParser reduces the retrieval drop relative to the same baselines.","Inference still uses a single similarity score between the two aggregated view representations, so the framework changes how representations are built rather than how retrieval is performed.","Without the representation diversity loss, the decomposition tokens collapse toward redundant embeddings and recall drops, indicating that view separation is needed for the alignment mechanism to help."],"supporting_citations":[{"why":"Supplies the STAN temporal encoder whose video and text features the parsers decompose, and serves as a global-matching baseline.","marker":"[21]"},{"why":"CLIP4Clip is the CLIP frame-averaging baseline that T2VParser is attached to and compared with.","marker":"[24]"},{"why":"CLIP-VIP provides a token-wise matching baseline and an encoder whose temporal features T2VParser wraps.","marker":"[33]"},{"why":"Mug-STAN is the token-wise matching baseline used as the main encoder in ablations and parameter studies.","marker":"[20]"},{"why":"DETR-style learnable query tokens are the stated inspiration for the Adaptive Decomposition Tokens.","marker":"[3]"},{"why":"BLIP-2 captions the eight key frames that feed the video document generator.","marker":"[16]"},{"why":"DeepSeek-V2 generates the training video documents from captions and key-frame descriptions; LLM choice is ablated in the paper.","marker":"[7]"},{"why":"Provides the PRVR partially relevant video retrieval benchmark used in the noise and partial-alignment evaluation.","marker":"[8]"},{"why":"TSDPC selects the eight key frames per video that are captioned and turned into documents.","marker":"[28]"}],"fun_headline_variants":["Shared tokens align only matching views in video-text retrieval","Partial alignment via adaptive decomposition boosts video-text recall","Decomposing text and video into views improves partial matching","Adaptive tokens down-weight mismatched views for retrieval","Same tokens decompose both modalities for precise video-text alignment"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The method assumes that the same learnable token, after separately attending to text and to video, comes to represent the same semantic view in both modalities, so that comparing same-index views is meaningful.","fun_headline_variants_meta":{"raw":{"variants":["Shared tokens align only matching views in video-text retrieval","Partial alignment via adaptive decomposition boosts video-text recall","Decomposing text and video into views improves partial matching","Adaptive tokens down-weight mismatched views for retrieval","Same tokens decompose both modalities for precise video-text alignment"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000162,"raw_usage":{"total_tokens":1254,"prompt_tokens":977,"completion_tokens":277,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":593,"completion_tokens_details":{"reasoning_tokens":201}},"tokens_in":593,"tokens_out":277,"duration_ms":2917,"temperature":1.0,"reasoning_tokens":201,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T17:42:05.292688+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"If, after permuting the order of the decomposition tokens on only the text side (or only the video side) at inference, retrieval recall is unchanged, then same-index correspondence between views carries no signal and the performance would be attributable to the attention-weighted aggregation rather than to the claimed semantic decomposition.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"CLIP4Clip is the CLIP frame-averaging baseline that T2VParser is attached to and compared with."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"CLIP-VIP provides a token-wise matching baseline and an encoder whose temporal features T2VParser wraps."},{"cited_title":"Mug-STAN: Adapting Image-Language Pretrained Models for General Video Understanding","cited_arxiv_id":"2311.15075","evidence_quote":"Mug-STAN is the token-wise matching baseline used as the main encoder in ablations and parameter studies."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"DETR-style learnable query tokens are the stated inspiration for the Adaptive Decomposition Tokens."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"BLIP-2 captions the eight key frames that feed the video document generator."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the PRVR partially relevant video retrieval benchmark used in the noise and partial-alignment evaluation."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"TSDPC selects the eight key frames per video that are captioned and turned into documents."}],"review_version":1}