{"id":"a85647dd-2cd9-4ee6-8216-d23c359b57b2","arxiv_id":"2604.23321","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"On 190 tasks and a 12-direction cross-modal diagnostic, seven embedding models frequently fail to honor explicit target-modality instructions: retrieval is biased toward the query modality and instruction-induced shifts are misaligned.","lead":"MMEB-V3 is a new 190-task benchmark for multimodal embedding models — text, image, video, audio, and agent tasks — plus OmniSET, a diagnostic made of semantically identical content across modalities. Evaluating seven models, it finds they often ignore explicit modality instructions, retrieve the wrong modality, and behave asymmetrically across directions.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"OmniSET's synthetic video/audio generation (Veo from image, TTS from caption) makes the headline I→V/A→T successes and T→A/V→T failures potentially measure generation fidelity rather than modality-instruction following; the central claim depends on untested semantic equivalence.","rationale":"The reader's weakest_assumption is correct and I agree with it. This concern is load-bearing because the three headline findings are all measured on OmniSET, and the only directions with non-trivial success coincide with the synthetic generation edges. The paper's own appendix caveat is good faith but insufficient; the main-text conclusion overstates generality. However, the central negative finding—that instruction-conditioned retrieval often fails even when the target is generated from the query—has some support in unconfounded directions, so the paper should not be rejected. The appropriate disposition remains CONDITIONAL: the synthetic-equivalence premise and the undisclosed instruction templates must be addressed before the diagnostic is taken at full strength. My concrete human-equivalence audit would settle whether the asymmetry is a model property or a dataset artifact. If the audit and a real-data control preserve the pattern, the claim stands; if not, the abstract's broad conclusion needs scaling back.","tokens_in":24826,"tokens_out":5935,"duration_ms":71474,"concrete_test":"Human-equivalence audit: sample 30 OmniSET tuples; for each, show 3 annotators the source image/caption alongside the generated video/audio and ask whether the generated instance preserves the same entities, actions, and setting with no material additions or omissions. Require ≥90% agreement and ≥95% equivalence; then re-run Table 4 on the subset of tuples that pass. If I→V/A→T drop materially or T→A/V→T rise, the generation graph is responsible for the headline pattern.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Table 4's most extreme asymmetries align exactly with the generation graph described in §A.2.1: videos are Veo-generated from the query image, audio is TTS-generated from the caption. I→V and A→T are the two directions where the target is a deterministic function of the source; they are also the only near-perfect directions (Hit@1 = 100 for Omni-Embed-Nemotron; I→V = 100 for Qwen3-VL). If the generated video is a near-copy of its source image and the TTS audio is a spoken version of its caption, these scores can be obtained by low-level modality matching (frame similarity, speech-to-text) without the model ever using the instruction to select the target modality. Conversely, the headline failures (T→A = 0.0, V→T = 0.0) could be inflated if TTS/Veo outputs are poor semantic renderings of the captions: the model may be retrieving a faithful target that isn't in the pool. §A.2.2 concedes this: the construction 'may introduce a form of modality preference' and the near-perfect scores 'may partially reflect dataset construction effects'. Yet the abstract and §5.1 present the asymmetry and instruction-failure as general model limitations. The diagnostic can only separate semantics from modality if the equivalence tuples are actually semantically equivalent; the paper provides no human or automatic verification of that premise.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces MMEB-V3, an extension of the MMEB-V2 benchmark to 190 tasks spanning text, image, video, audio, visual documents, and agent-centric retrieval. It also constructs OmniSET, a diagnostic dataset of approximately 100 'semantic equivalence tuples' in which the same content is rendered as text, image, video, and audio; videos are synthesized from images with Veo-3.1 and audio from captions with Gemini TTS. Using OmniSET, the paper reports three headline findings: (1) explicit modality instructions often fail (Hit@1 near zero in most cross-modal directions), (2) cross-modal retrieval is asymmetric and dominated by query-modality bias, and (3) instruction-induced embedding shifts do not consistently move queries toward the target modality. The conclusion is that current omni-modality embeddings cannot reliably enforce modality constraints. The main benchmark evaluation uses standard public datasets and task-appropriate metrics; the diagnostic analysis uses the newly constructed OmniSET.","tokens_in":25131,"tokens_out":8819,"duration_ms":92864,"significance":"The benchmark resource itself is a genuine contribution: it assembles a large, heterogeneous collection of tasks with clear protocols, reports detailed per-task results for seven models, and extends coverage to audio and agent scenarios that previous MMEB versions lack. The OmniSET design is also valuable in principle: controlled cross-modal tuples can expose modality bias in a way that standard retrieval benchmarks cannot. If the three findings survive the confounding concerns below, they would be an important signal for the multimodal embedding community. However, the abstract and conclusion generalize from a diagnostic set whose construction is confounded with the very modality asymmetries the paper reports; until that is addressed, the headline claims are not yet established. I also credit the authors for explicitly discussing the synthetic-data limitation in §A.2.2; the problem is that the main text does not carry that caveat into the abstract and conclusions.","major_comments":[{"comment":"The near-perfect directions in Table 4 are exactly the generation edges: I→V (video is Veo-generated from the query image) and A→T (audio is TTS-generated from the caption). Conversely, the headline failures (T→A, V→T) are directions where the target is a generated artifact whose semantic fidelity to the query is unverified. The paper itself concedes in §A.2.2 that the construction 'may introduce a form of modality preference' and that I→V/A→T scores 'may partially reflect dataset construction effects.' Without human or automatic verification that the generated video/audio are semantically equivalent to their source and to the other tuple members, Hit@1 in these directions measures generation fidelity plus model behavior, not modality-instruction following. The abstract and §5.1 nevertheless present asymmetry and instruction failure as general model limitations. Please provide equivalenc","section":"§A.2.1, §A.2.2, Table 4"},{"comment":"OmniSET is built from roughly 100 hand-curated queries (Table 1 reports 1.2K directed query instances, §A.2 says 100 base queries). Table 4 reports Hit@1/MRR with no confidence intervals, significance tests, or per-query variability. With 100 base queries, the difference between Hit@1=0.0 and Hit@1=3.0 is three successes; a single query changes scores by 1 percentage point. This is particularly problematic for the small differences in §5.3 ('all below 0.09') and for the WAVE T2V/A2V values. Please report bootstrap confidence intervals or per-query breakdowns and state the effective sample size for each directional score. The current presentation makes it impossible to know which asymmetries are robust.","section":"§A.2, Table 4"},{"comment":"§5.3 describes T→V as 'a small improvement (+0.041)' and V→T as 'degradation (−0.158)', implying positive = better. The caption of Figure 10 states 'Negative values indicate that the instruction-augmented query moves closer to the target modality, while positive values indicate increased distance', i.e., negative = better. If Figure 10's convention is correct, the two examples in §5.3 are reversed and the qualitative claim about which directions improve is inverted. If §5.3's convention is correct, Figure 4a and Figure 10 captions are wrong. Please make the sign convention consistent across the text and figures and recompute the affected statements.","section":"§5.3 vs. Figure 10"},{"comment":"The mitigation says the key phenomena are 'consistent across multiple models and modality directions, including those not directly affected by synthetic generation.' This is not supported by the data. Every direction involving V or A is affected because both V and A are synthetic; only T→I and I→T use the original MSCOCO image/text pair exclusively. Table 4 shows T→I Hit@1=0.0 and I→T=0.0 for all models, but those are the only unaffected directions, and they too come from the same 100-query instrument. The claimed consistency across 'unaffected' directions therefore cannot be checked from the reported results. Either list which directions are considered unaffected and report them separately, or drop this mitigation.","section":"§A.2.2"}],"minor_comments":[{"comment":"The 'All' column treats missing audio as 0 for Qwen/VLM2Vec/GME, which conflates lack of modality support with poor performance. All* (average over available tasks) is more interpretable; recommend reporting it as primary or adding a footnote with per-modality coverage.","section":"§3.1, Table 3"},{"comment":"The number of OmniSET queries is not consistent: 'approximately 100' high-quality samples vs. 1.2K query count in Table 1. Clarify that 1.2K = 100 base tuples × 12 directed tasks (or state the actual base count).","section":"§A.2, Table 1"},{"comment":"Hit@1 values with two decimals (e.g., 68.32) are confusing for count-based metrics; report as fractions (e.g., 820/1200) or state the sample size so readers can interpret.","section":"Table 4"},{"comment":"'Sensitivity' measured as cosine distance is a magnitude, not necessarily 'responsiveness' in terms of instruction following; consider renaming or clarifying to avoid implying effectiveness.","section":"§5.2, Figure 3"},{"comment":"Typos and formatting: 'T uples' should be 'Tuples' in the abstract and §3.1; Figure 4a's heatmap has no colorbar or units. A final proofread would help.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The benchmark contribution is solid and the OmniSET concept is worth publishing, but the three headline findings are currently over-stated relative to the evidence. The confounding from synthetic generation is the central issue; it is fixable with additional validation or by sharply narrowing the claims. If the authors can add equivalence verification and control directions, I would be willing to accept a revised version."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's the short version. MMEB-V3 is a competent extension of an established benchmark line, and OmniSET is a genuinely useful diagnostic construct. The main claim — that current omni-modal embedding models often retrieve the wrong modality even when explicitly instructed which one to return — is supported. But the paper's most eye-catching asymmetry numbers are partly built into the dataset construction, and the authors know it; they just don't carry that caveat into the abstract.\n\nWhat's actually new: 111 new tasks over MMEB-V2, with real audio coverage, agent tasks, and complex text retrieval, run on public datasets with task-appropriate metrics. The per-task appendix tables look honest. Credit where due: §A.2.2 explicitly concedes the synthetic-data confound, and the design choices — shared mixed-modality candidate pool, excluding same-modality retrieval, including the source instance as a distractor — are thoughtful. OmniSET is framed as a diagnostic rather than a leaderboard, which is the right call.\n\nSoft spots, in order of importance.\n\nFirst, the striking I→V and A→T near-perfect scores are exactly the directions where video is generated from the query image and audio from the caption. Those numbers can be explained partly by low-level modality matching, since the target is a deterministic function of the source. The stress-test note is right about that. But it doesn't sink the paper, because the failure findings also hold on directions untouched by the generation graph: T→I and I→T are 0.0 across models, and those use original MSCOCO captions and images. The central claim survives; the asymmetry claim should be toned down.\n\nSecond, OmniSET is 100 hand-curated queries with no error bars. With 100 queries, 0 versus 100 is a big gap, but the §5.2–5.3 shift numbers (0.08–0.4 cosine) come without uncertainty. The instruction templates are never specified, so those measurements are conditional on an undisclosed prompt design. That is a reproducibility problem.\n\nThird, no artifacts are released. For a benchmark paper, that is the thing I'd most want before accepting.\n\nWho this is for: anyone working on multimodal embeddings or retrieval for RAG and agentic systems. It deserves serious peer review — the benchmark is useful even though the diagnostic needs a revision pass. I would send it out, and I'd ask the authors to release the harness, disclose the templates, add uncertainty estimates, and move the §A.2.2 caveat into the main text.","headline":"Useful benchmark extension with a genuinely good diagnostic idea; the headline asymmetry numbers are partly construction artifacts, but the central failure finding survives on unconfounded directions.","tokens_in":25769,"tokens_out":2614,"would_cite":true,"duration_ms":32117,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Current multimodal embeddings cannot reliably follow modality instructions, according to a new 190-task benchmark and its controlled diagnostic set.","keywords":["multimodal embeddings","benchmark","modality instruction following","cross-modal retrieval","embedding evaluation","agent retrieval","semantic equivalence tuples","modality bias"],"falsifier":"Take a set of human-verified quadruples where the video is a real recording of the scene in the image and the audio is a genuine spoken version of the caption, so no generation artifacts tie any modality pair. If, on these, all 12 directed retrieval directions show Hit@1 well above chance—or if the query-modality bias disappears—then the paper's central claim of systematic instruction-failure is falsified; if the same near-zero and asymmetric pattern persists, the claim is confirmed.","tokens_in":24609,"feed_emoji":"🎯","tokens_out":6562,"duration_ms":62760,"temperature":0.7,"pith_summary":"MMEB-V3 is a 190-task benchmark that evaluates embedding models across text, image, video, audio, and agent-centric settings, and it carries a diagnostic component, OmniSET, that pairs semantically identical content across all four modalities. By controlling semantics while varying modality, OmniSET lets the authors ask a precise question: when an instruction names a target modality, does the retrieved instance actually come from that modality? Their answer is no. On 12 directed cross-modal retrieval tasks, three representative models score near zero in most directions, and the modality that dominates the top-10 results tracks the query's own modality far more often than the instructed target. Instruction-augmented queries do shift in embedding space—sometimes substantially—but the shifts rarely reduce distance to the target and often increase it. The paper concludes that today's omni-embedding models treat modality as a byproduct of semantics rather than as an enforceable instruction constraint, which matters because agents and retrieval-augmented systems increasingly depend on retrieving the right sensory format, not just the right content.","feed_headline":"Ignore instructed modality, retrievers fail 190-task benchmark","feed_subtitle":"Embedding models return the query's own format instead of the target, breaking agent retrieval pipelines.","key_machinery":"OmniSET (Omni-modality Semantic Equivalence Tuples) — aligned quadruples {text, image, video, audio} with the same semantics, each query expanded into 12 directed cross-modal retrieval tasks sharing one mixed-modality candidate pool; source instance kept as a distractor, same-modality pairs excluded, 15–20 human-verified hard negatives per query. It isolates the modality dimension by holding semantics fixed, and the 12-direction design exposes directionality and query-modality bias. Its acknowledged limitation is that video and audio are synthesized from images and captions, which may tie certain pairs together and partly drive the I→V and A→T near-perfect scores.","core_discovery":"On OmniSET, ~100 queries with hard negatives exist in all four modalities, giving 12 directed retrieval tasks against a shared mixed-modality pool that keeps the source as a distractor. Across three models, Hit@1 is 0.0 in most directions (text→image, text→audio, video→text); only generation-linked directions (image→video, audio→text) succeed. Top-10 dominant modality follows the query, not the target: Nemotron retrieves 82.7% text for text queries even when images are requested; WAVE retrieves video 99.9% of the time. Instruction augmentation shifts queries up to 0.4 cosine distance, but distance to target improves in only a few directions (all <0.09). The claim: current omni-modality embed","pith_inferences":["Because OmniSET's video and audio are generated (Veo-3.1 from images, Gemini-2.5-Flash-TTS from captions), the paper's headline asymmetry may partly reflect generation fidelity: I→V and A→T look easy precisely because the generated targets are near-copies of their sources. A follow-up with naturally occurring quadruples—real video of the same event as the image, real speech of the caption—would se","The dominant-modality statistic (e.g., 99.9% video for WAVE) suggests query-modality bias is not soft preference but a broken retrieval policy: for a text→video query, a model that retrieves only video would have Hit@1 = 0; the embedding geometry appears clustered by modality rather than aligned by content, which would explain why semantic equivalence across modalities is invisible to these models","A testable extension: train or fine-tune an omni-embedding model with an explicit loss that penalizes retrieving an instance whose modality differs from the instruction, then re-run the OmniSET directions. If modality-instruction following improves sharply, the gap is a training objective problem; if it does not, it is an architectural or capacity constraint.","For agent applications, the practical implication the paper leaves implicit is that modality must be enforced downstream—e.g., by reranking candidates with a modality classifier or by constraining the candidate set—since the embedding alone cannot be trusted to honor the instruction."],"forward_implications":["Retrieval pipelines relying on omni-modality embeddings for modality-specified queries will silently return wrong-format results: agents asking for an audio clip or a video will often get text or the query's own modality instead.","Because the failure is consistent across very different models, it points to a shared training-objective gap—contrastive alignment over semantic similarity does not teach modality-conditioning—rather than to a defect of one architecture.","The instruction-shift analysis implies that naive prompt augmentation (appending 'retrieve a video' to a query) is not sufficient: embedding shifts must be oriented toward the target modality, not merely increased in magnitude.","OmniSET provides a reusable diagnostic: future embedding models can be tested on the same 12 directed tasks with a shared candidate pool, making modality-instruction following directly measurable.","Benchmark scores on standard cross-modal retrieval may overstate capability, because retrieval of the correct semantic content is not the same as retrieval of content in the instructed modality—a distinction MMEB-V3 makes visible."],"fun_headline_variants":["Retrievers follow query, not target, in omni-modality benchmark","MMEB-V3: instruction shifts fail to fix cross-modal retrieval","Query bias dominates: text queries pull text even when images asked","Omni-embedding models ignore instructed target; 190-task test"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The findings rest on the assumption that the pipeline that generates videos from images and speech from captions yields faithful semantic equivalents, so that cross-modal scores measure instruction-following rather than the fidelity of the generated candidates—an assumption the paper itself flags as a possible source of modality bias.","fun_headline_variants_meta":{"raw":{"variants":["Retrievers follow query, not target, in omni-modality benchmark","MMEB-V3: instruction shifts fail to fix cross-modal retrieval","Query bias dominates: text queries pull text even when images asked","Omni-embedding models ignore instructed target; 190-task test"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000353,"raw_usage":{"total_tokens":1793,"prompt_tokens":812,"completion_tokens":981,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":556,"completion_tokens_details":{"reasoning_tokens":904}},"tokens_in":556,"tokens_out":981,"duration_ms":9653,"temperature":1.0,"reasoning_tokens":904,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T15:32:36.260289+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a set of human-verified quadruples where the video is a real recording of the scene in the image and the audio is a genuine spoken version of the caption, so no generation artifacts tie any modality pair. If, on these, all 12 directed retrieval directions show Hit@1 well above chance—or if the query-modality bias disappears—then the paper's central claim of systematic instruction-failure is falsified; if the same near-zero and asymmetric pattern persists, the claim is confirmed.","supporting_citations":[],"review_version":2}