{"id":"494a1827-17fc-4362-b825-6c5f3187dde6","arxiv_id":"2505.04620","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A new five-level evaluation framework and a 702-task multimodal benchmark show that current multimodal LLMs rarely surpass task-specific specialists, so none reach the highest 'total synergy' level.","lead":"This paper introduces General-Level, a five-level scale that ranks multimodal AI models by 'synergy', meaning whether they beat specialized models and transfer skills across tasks and modalities. It also presents General-Bench, a 702-task benchmark across image, video, audio, 3D, and language, and finds that most current models, including GPT-4o, lack this synergy.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'no synergy' conclusion is defined by whether zero-shot generalists beat hand-picked open-source SoTA specialists; that equivalence is unvalidated, so Level-3+ rankings may reflect baseline selection rather than cross-task transfer.","rationale":"The reader's verdict is CONDITIONAL and I agree with the identified weakest assumption. The benchmark artifact itself is substantial: 702 tasks, 145 skills, original-format evaluation, and 100+ models is a real contribution, and the directional observation that zero-shot generalists rarely beat fine-tuned specialists is plausible. But the paper's headline inference from that observation to 'no synergy' rests on the unvalidated §3.2.2 equivalence. The direct definition of synergy (joint modeling of A and B exceeding solo modeling) is acknowledged infeasible, and the relaxation to beating specialists is neither calibrated nor given a positive control. This makes the Level-3+ leaderboard and the Level-5 non-finding sensitive to specialist selection and to the zero-shot/fine-tuned asymmetry. The recommended CONDITIONAL verdict stands: the framework and benchmark are publishable as infrastructure, but the central empirical claim about synergy absence should be presented as conditional on the relaxation, with sensitivity analysis and/or a controlled transfer experiment. Secondary proof issues reinforce the need for revision but do not change the verdict.","tokens_in":65188,"tokens_out":6298,"duration_ms":64897,"concrete_test":"Run a controlled calibration on a subset of General-Bench image-comprehension tasks (e.g., 10 tasks from Table 6 with specialists in Tables 21–37). Starting from one small open MLLM, fine-tune variant A on task i alone, variant B on tasks i and j jointly, and variant C on task j alone. Define true transfer on task i as B(i) − A(i). Then compute the §3.2.2 Level-3 mask for B on task i, i.e., whether B(i) ≥ SoTA_i. If the mask and the sign of true transfer disagree on more than a small fraction of tasks, the relaxation is invalid and Level-3+ rankings cannot support the synergy-absence conclusion. As a cheaper robustness check, recompute Tables 14 and 16 after replacing each specialist with the next-best open model; if the Level-4 trio or Level-3 top-10 changes materially, the result is baseline-sensitive.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—'most MLLMs lack the cross-task/cross-modal synergy ability' and 'no model has demonstrated the ability to enhance language intelligence through non-language modalities'—is not measured directly. Under §3.2.2, a model is credited with synergy on a task exactly when its zero-shot score reaches or exceeds the SoTA specialist score the authors selected for that task. The authors concede in §6 that this 'avoids a direct measurement of the synergy effect.' The load-bearing problem is the asymmetry: generalists are run zero-shot (§5.3), while specialists are task-fine-tuned and chosen for top public performance. A zero-shot generalist failing to beat a fine-tuned specialist cannot distinguish 'no learned cross-task transfer' from 'the specialist bar is high by construction.' Conversely, a win over a selected specialist is not positive evidence of transfer, because the specialist set excludes closed-source models (§5.1), and a weak or contaminated baseline can produce a 'win.' Every Level-3, Level-4, and Level-5 conclusion inherits this equivalence, so the leaderboard numbers and the 'no reverse modality-to-language synergy' finding are not yet supported as evidence about synergy. Secondary but related: the S4≤S3 proof in §3.2.3 contains an algebraic error, and Stotal in S5 is undefined, so the formal monotonicity scaffolding around these scores is also unreliable.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes General-Level, a five-level taxonomy for ranking multimodal large language models (MLLMs) by a construct called 'synergy,' and General-Bench, a benchmark of 702 tasks and 325,876 instances spanning image, video, audio, 3D, and language modalities in native task formats. At Level-2 and above, scores are computed from task performance relative to selected SoTA specialists: a model is credited with synergy on a task when its zero-shot score reaches or exceeds the specialist's score. The authors evaluate over 100 LLM/MLLM systems, report leaderboards at Levels 2-4, and conclude that most MLLMs lack cross-task and cross-modal synergy, that even GPT-4V and GPT-4o do not rank at the top, and that no model has demonstrated enhancement of language intelligence through non-language modalities.","tokens_in":65370,"tokens_out":7060,"duration_ms":69945,"significance":"If the central measurement were valid, the paper would offer a substantively new way to evaluate generalists beyond raw accuracy, and its negative result about current MLLMs would be a notable challenge to prevailing benchmark narratives. The benchmark itself is a large and potentially useful resource: 702 tasks in original free-form formats, 172 specialist references, 102 evaluated generalists, and broad modality and domain coverage. The observation that closed models such as GPT-4V rank below several open models under a specialist-relative scoring rule is also interesting and falsifiable. However, the headline synergy claims rest entirely on an unvalidated equivalence between beating a selected specialist and exhibiting cross-task transfer, and the formal monotonicity scaffolding contains algebraic errors. The benchmark and leaderboards can survive a reformulation, but the current framing overstates what the data establish.","major_comments":[{"comment":"The central measurement assumption is unvalidated. The paper defines a task-level 'win' over a SoTA specialist as evidence of a synergy effect, and all Level-3, Level-4, and Level-5 scores, together with the headline conclusions ('most MLLMs lack synergy'; 'no model has demonstrated the ability to enhance language intelligence through non-language modalities'), are built on that equivalence. The comparison is asymmetric: generalists are evaluated zero-shot (§5.3), while specialists are fine-tuned and selected for top public performance (§5.1), and the specialist set excludes closed-source models. A zero-shot generalist failing to beat a fine-tuned specialist cannot distinguish 'no learned cross-task transfer' from 'a high specialist bar by construction,' and a win over a weak or contaminated selected baseline cannot by itself establish transfer. The authors acknowledge in §6 that the relaxation 'avoids a direct measurement of the synergy effect.' As written, the Level-3+ rankings and the 'no reverse synergy' result are not measurements of synergy as defined in §3.1.2. I ask for either a validation study (for example, controlled single-task versus multi-task training where ground-truth transfer is known, checking whether beating a specialist predicts transfer) or a re-framing of all Level-3+ claims as 'zero-shot specialist-relative performance' with the term synergy removed.","section":"§3.2.2, §5.4, §6"},{"comment":"The monotonicity proofs contain algebraic errors and do not establish the claimed strict decline. In the proof of S4 ≤ S3, multiplying the displayed inequality by 4(SC + SG) yields (SC + SG)^3 ≥ 8 SC SG, not ≥ 8 SC SG (SC + SG); the subsequent factorization is not a consequence of the displayed inequality. The correct AM-HM argument does show S4 ≤ S3 for positive scores, but only as a non-strict inequality. Similarly, the proof of S3 ≤ S2 establishes ≤, and equality occurs whenever a model exceeds every specialist threshold. The paper's Property-2 claim that 'S_{k-1} > S_k' is therefore not proven. Additionally, Table 1 defines w_L = S_L / S_total without defining S_total, so the Level-5 score is not computable as specified. These issues do not necessarily invalidate the leaderboard numbers, but the formal support claimed in §3.2.3 should be corrected or downgraded to non-strict claims.","section":"§3.2.3, Table 1"},{"comment":"The claimed minimum of 500 samples per task is inconsistent with the reported totals. Section 4.1.2 states 'We ensure that each task includes (at least) 500 data samples,' while Section 4.2 says 'For most of the tasks, we maintain around 500 testing instances.' Table 2 reports 702 tasks and 325,876 instances, an average of about 464 instances per task; the largest cell (271 image-comprehension tasks with 124,880 instances) averages about 461. These numbers are incompatible with a 500-sample minimum unless many tasks exceed 500 and others fall below. The paper should report the actual per-task distribution and reconcile the text. This matters because per-task specialist comparisons with small samples have high variance, which directly affects the stability of the Level-3 single-task 'win' decisions.","section":"§4.1.2, §4.2, Table 2"},{"comment":"The leaderboard rankings are sensitive to the choice of the specialist baseline set, and no sensitivity analysis is reported. The paper excludes closed-source models from the specialist pool (§5.1) and selects specialists by public benchmark recognition, so the 'win-over-specialist' counts (Observation-2 in §5.3) and all Level-3+ scores are relative to this particular set. For example, GPT-4V and GPT-4o are evaluated as generalists but cannot serve as specialist references; had they been included for the tasks they support, several Level-3 wins could disappear. The paper should report how rankings change under alternative specialist choices (for example, best open-source model per task versus best available model, or including closed-source API baselines where feasible) before claiming that the rankings are stable characterizations of generalist capability.","section":"§5.1, §5.3, Tables 12-16"}],"minor_comments":[{"comment":"The section heading 'Receipt to Leveling Upper in General-Level' appears to be a typo; it should read 'Recipe for Leveling Up in General-Level.'","section":"§3.3"},{"comment":"The backbone name 'Qwev-7B' appears twice (models 10 and 22); this is presumably 'Qwen-7B' and should be corrected.","section":"Table 4"},{"comment":"Observation-2 says 'few models capable of surpassing the SoTA generalist,' but the comparison is against specialists; the terminology should be fixed to 'SoTA specialist.'","section":"§5.3, Observation-2"},{"comment":"The model is named 'SEED-LLaMA-13B' in Table 4 but 'SEED-LLaMA-14B' in Tables 6 and 7; the naming should be consistent across the paper.","section":"Tables 6-7 vs. Table 4"},{"comment":"The text says 'General-Bench comprises 130 multimodal skills, containing 702 tasks,' while Table 2 and Figure 7 report 145 skills; please clarify whether language skills are excluded from the 130 figure.","section":"§4.3"},{"comment":"The 'More Task, The Better' argument is not generally true as stated: since S2 is an average over all benchmark tasks, adding a task where the model scores zero cannot increase its score, and a model supporting more tasks can still have a lower average if the added task scores are low. The claim should be reworded to describe an incentive under specific support/score conditions.","section":"§3.2.3, Property-3"}],"recommendation":"major_revision","confidential_remarks":"For the editor: the General-Bench resource is substantial and likely citable even if the synergy framing is rejected. I would encourage the authors to consider reframing the paper as a zero-shot generalist-versus-specialist benchmark with an explicit 'relative to selected open-source specialists' caveat, rather than claiming direct measurement of synergy. The current abstract and Section 5.4 conclusions overstate what the data can establish, but the underlying benchmark and the raw specialist-relative leaderboards are salvageable within the manuscript's scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is the rare benchmark paper that is actually useful despite its central claim being shaky. The resource itself is the contribution: 702 tasks, 145 skills, free-form formats, modalities beyond images (audio, 3D, video, language), and 100+ models evaluated, with code, data, and a leaderboard. If you work on multimodal evaluation, you will want this on your desk. The 5-level General-Level taxonomy is a reasonable framing for talking about generality, even if the levels themselves are not empirically validated.\n\nThe soft spot is exactly where the stress-test places it. The paper defines synergy as beating a SoTA specialist on a task, runs generalists zero-shot, and uses fine-tuned open-source specialists as the bar. As the authors concede in §6, this is not a direct measurement of synergy. A zero-shot generalist losing to a fine-tuned specialist tells you something about the difficulty of the bar, not about whether cross-task transfer happened. The headline findings—'most MLLMs lack synergy' and 'no reverse modality-to-language synergy'—are therefore partly built into the definition. The empirical result that survives is narrower but still real: current MLLMs rarely beat fine-tuned open-source specialists zero-shot, especially outside image comprehension.\n\nThe formal scaffolding has real errors. The S4≤S3 proof in §3.2.3 contains wrong algebra (the factorized expression does not follow from the previous line), the claimed strict monotonicity only shows ≤, and Stotal in Level-5 is undefined. Also, the paper says each task has at least 500 samples, but 702 tasks × 500 = 351K exceeds the stated 325.8K total. These are fixable, but they mean the specific leaderboard numbers should not be taken as reliable without revision.\n\nNone of this kills the paper. The benchmark itself is a substantial, reusable asset, and the observation that GPT-4V/4o rank low due to task support is a valuable counterpoint to raw-accuracy leaderboards. What needs to change: fix the proofs, define Stotal, reconcile the sample counts, and add sensitivity analysis around specialist selection. In short: send it to peer review, but expect major revision.","headline":"A large, genuinely useful multimodal benchmark whose 'synergy' leaderboard rests on an admitted and unvalidated equivalence; worth refereeing, but not as-is.","tokens_in":66133,"tokens_out":2147,"would_cite":true,"duration_ms":24977,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A five-level synergy test finds most AI generalists fall short","keywords":["multimodal large language models","multimodal generalist","synergy","evaluation benchmark","five-level taxonomy","comprehension and generation","AGI","cross-modal transfer"],"falsifier":"Train two copies of a top Level-3 model—one jointly on a set of tasks, one on each task independently—and compare their performance against the same specialists; if the separately trained copy matches or beats the joint model, then the measured \"synergy\" is not transfer and the Level-3 and higher hierarchy loses its evidential basis.","tokens_in":64861,"feed_emoji":"🧩","tokens_out":6852,"duration_ms":64138,"temperature":0.7,"pith_summary":"The paper argues that conventional MLLM benchmarks mislead: high average accuracy across tasks does not tell you whether a model is a true multimodal generalist or merely a collection of competences. To fix this, it proposes General-Level, a five-level taxonomy whose criterion is \"synergy\"—the ability of knowledge learned in one task or modality to lift performance in another, operationalized as beating the state-of-the-art specialist on a given task. To drive the taxonomy, it builds General-Bench, a benchmark of 702 tasks and 325,800 instances spanning image, video, audio, 3D, and language, in original output formats rather than forced multiple choice. Evaluating over 100 models, the paper finds that most MLLMs support few tasks and rarely surpass specialists, that GPT-4V and GPT-4o do not lead the generalist ranking, and that no current model reaches Level-5: none improves language intelligence through non-language modalities. If the framework is right, the path to AGI is not measured by raw task accuracy but by demonstrable cross-task and cross-modal transfer.","feed_headline":"A five-level synergy test finds most AI generalists fall short","feed_subtitle":"A 702-task benchmark ranks multimodal models by whether they beat specialists; GPT-4-class models don't lead.","key_machinery":"The load-bearing object is the General-Level scoring ladder. Level-2 averages normalized scores over all supported comprehension and generation tasks; Level-3 re-scores each task as zero unless the generalist beats the task's SoTA specialist, so the score counts only \"winning\" tasks; Level-4 takes the harmonic mean of comprehension and generation scores, rewarding balance; Level-5 multiplies the Level-4 score by a normalized weight equal to the model's rate of beating NLP SoTA specialists. Because higher levels are built from masked or combined lower-level scores, the framework mathematically guarantees scores decrease monotonically as levels rise. The synergy concept is what carries the argument: the paper treats beating a specialist as observable evidence of transfer, and General-Bench supplies the task surface—702 tasks in native formats, grouped by modality and by comprehension and generation—on which that evidence is collected.","core_discovery":"On the paper's own terms, the central discovery is that \"synergy\"—defined as a generalist outperforming the state-of-the-art (SoTA) specialist on a task, taken as evidence of transfer from other tasks or modalities—is rare and shallow in current MLLMs. At Level-2 (basic unified comprehension and generation), models like Unified-io-2 and AnyGPT outrank GPT-4V and GPT-4o because breadth of task and modality support outweighs single-task strength. At Level-3 (cross-task synergy), top ranks go to Sa2VA-26B, LLaVA-One-Vision-72B, and Qwen2-VL-72B, while GPT-4V and GPT-4o place lower. Only Mini-Gemini, Vitron-V1, and Emu2-37B reach Level-4, meaning synergy across comprehension and generation. No model earns a non-zero Level-5 score: no tested system outperforms NLP SoTA specialists on language tasks, so there is no evidence that non-language modalities enhance language intelligence.","pith_inferences":["Because \"synergy\" is measured against a chosen pool of open-source SoTA specialists, the same model could land at different levels if that pool is swapped; a natural test is to re-run the Level-3 and higher scoring against a stronger or weaker specialist pool and watch the rankings move.","The framework's own monotonicity proofs imply that higher levels are definitionally harder to reach, meaning a model's level is only partly an empirical fact about the model and partly a design choice about task coverage and specialist baselines.","The paper's observed image-video synergy clustering suggests visual modalities share transferable features; a testable implication is that joint image-video training should produce the largest specialist-beating gains, while audio-language joint training should show the smallest, guiding where to invest in architecture.","The empty Level-5 predicts that simply adding more multimodal pretraining to an LLM will not improve its core NLP performance; this could be tested by measuring an MLLM's NLP scores before and after multimodal training under controlled data budgets."],"forward_implications":["Rankings of MLLMs change once synergy is the criterion: models that support many modalities and beat specialists on many tasks rise, while high-scoring but narrow models such as GPT-4V fall.","Future training of multimodal generalists should explicitly target cross-task and cross-modal transfer, because Level-3 and Level-4 cannot be reached by adding parameters or data within a single task.","The absence of any Level-5 model implies that current language-centric MLLM architectures are not yet producing bidirectional modality-to-language transfer, so improving that direction becomes a named research goal.","Benchmarks should preserve native task formats rather than coercing everything into multiple choice, since forced QA hides generation and fine-grained-output failures.","Because specialist baselines update over time, a generalist's level is not permanent: as SoTA specialists improve, models must keep improving to hold their Level-3 and higher status."],"supporting_citations":[{"why":"Unified-io-2-XXL is the top-ranked Level-2 generalist; its evaluation establishes that breadth beats GPT-4V at Level-2.","marker":"(Lu et al., 2024a)"},{"why":"AnyGPT ranks second at Level-2 and supplies the any-to-any discrete-sequence baseline that the framework's breadth scoring rewards.","marker":"(Zhan et al., 2024)"},{"why":"NExT-GPT-V1.5 is a core any-to-any generalist evaluated across image, video, and audio, anchoring cross-modal support comparisons.","marker":"(Wu et al., 2024a)"},{"why":"Vitron-V1 reaches Level-4 and supplies the strongest pixel-level comprehension plus generation evidence for cross-comprehension-generation synergy.","marker":"(Fei et al., 2024a)"},{"why":"Mini-Gemini tops Level-4; its comprehension and generation balance carries the harmonic-mean scoring demonstration.","marker":"(Li et al., 2024c)"},{"why":"Emu2-37B is one of only three models reaching Level-4, providing evidence that generation-capable generalists are rare.","marker":"(Sun et al., 2024)"},{"why":"Sa2VA-26B is the top Level-3 model, supporting the claim that open-source dense-grounded models beat GPT-4V on cross-task synergy.","marker":"(Yuan et al., 2025)"},{"why":"InternVL2.5-8B is the open-source model with the highest image-comprehension task support rate, used in the task-support observations.","marker":"(Chen et al., 2024c)"},{"why":"Supplies the automotive five-level taxonomy that inspires General-Level's leveled structure.","marker":"(Yurtsever et al., 2020)"}],"fun_headline_variants":["No AI model passes level 5 in 700-task generality test","Synergy is rare: generalists fail to beat specialists in new benchmark","5-level framework ranks AI generalists: only 3 reach level 4","General-Bench: 700 tasks show no model achieves top synergy","Most multimodal generalists lack cross-modal synergy, study finds"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole measurement of synergy assumes that a generalist beating a chosen specialist on a task proves it transferred knowledge from other tasks or modalities, rather than simply being larger, better trained, or matched against weaker baselines.","fun_headline_variants_meta":{"raw":{"variants":["No AI model passes level 5 in 700-task generality test","Synergy is rare: generalists fail to beat specialists in new benchmark","5-level framework ranks AI generalists: only 3 reach level 4","General-Bench: 700 tasks show no model achieves top synergy","Most multimodal generalists lack cross-modal synergy, study finds"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00066,"raw_usage":{"total_tokens":3079,"prompt_tokens":1069,"completion_tokens":2010,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":685,"completion_tokens_details":{"reasoning_tokens":1918}},"tokens_in":685,"tokens_out":2010,"duration_ms":15218,"temperature":1.0,"reasoning_tokens":1918,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T23:23:47.506868+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train two copies of a top Level-3 model—one jointly on a set of tasks, one on each task independently—and compare their performance against the same specialists; if the separately trained copy matches or beats the joint model, then the measured \"synergy\" is not transfer and the Level-3 and higher hierarchy loses its evidential basis.","supporting_citations":[],"review_version":1}