Pith. sign in

REVIEW 16 cited by

SEED-Bench-2: Benchmarking Multimodal Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.17092 v1 pith:BSGFVEPI submitted 2023-11-28 cs.CV

classification cs.CV
keywords mllmsmodelsseed-bench-2capabilitiesevaluationhumanlanguagelarge
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
abstract

Multimodal large language models (MLLMs), building upon the foundation of powerful large language models (LLMs), have recently demonstrated exceptional capabilities in generating not only texts but also images given interleaved multimodal inputs (acting like a combination of GPT-4V and DALL-E 3). However, existing MLLM benchmarks remain limited to assessing only models' comprehension ability of single image-text inputs, failing to keep up with the strides made in MLLMs. A comprehensive benchmark is imperative for investigating the progress and uncovering the limitations of current MLLMs. In this work, we categorize the capabilities of MLLMs into hierarchical levels from $L_0$ to $L_4$ based on the modalities they can accept and generate, and propose SEED-Bench-2, a comprehensive benchmark that evaluates the \textbf{hierarchical} capabilities of MLLMs. Specifically, SEED-Bench-2 comprises 24K multiple-choice questions with accurate human annotations, which spans 27 dimensions, including the evaluation of both text and image generation. Multiple-choice questions with groundtruth options derived from human annotation enables an objective and efficient assessment of model performance, eliminating the need for human or GPT intervention during evaluation. We further evaluate the performance of 23 prominent open-source MLLMs and summarize valuable observations. By revealing the limitations of existing MLLMs through extensive evaluations, we aim for SEED-Bench-2 to provide insights that will motivate future research towards the goal of General Artificial Intelligence. Dataset and evaluation code are available at \href{https://github.com/AILab-CVC/SEED-Bench}

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 16 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SocialOmni: Benchmarking Audio-Visual Social Interactivity in Omni Models

    cs.AI 2026-03 conditional novelty 6.5 of 10

    SocialOmni jointly evaluates speaker ID, turn-entry timing, and interruption phrasing on 12 OLMs and finds perception accuracy decouples from socially appropriate generation.

  2. TSHA: A Benchmark for Visual Language Models in Trustworthy Safety Hazard Assessment Scenarios

    cs.CV 2026-03 unverdicted novelty 6.0 of 10

    TSHA is a new 80,000-pair benchmark for indoor safety hazard assessment; current vision-language models score roughly 45-85, and fine-tuning on TSHA raised Qwen2.5-VL-3B by 18.3 points on TSHA's test set.

  3. CARES: Context-Aware Resolution Selector for VLMs

    cs.CV 2025-10 conditional novelty 6.0 of 10

    A lightweight selector trained on target-VLM rollouts routes each image-query pair to the minimal sufficient resolution, cutting prefill compute by 63-80% across five benchmarks at roughly matched accuracy.

  4. GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling

    cs.CL 2025-04 conditional novelty 6.0 of 10

    A document-intelligence benchmark decouples visual and reasoning complexity, and a parameter-freezing fine-tuning method improves an 8B model without catastrophic forgetting.

  5. Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward

    cs.CV 2025-04 conditional novelty 6.0 of 10

    V2R-Bench shows that 21 large vision-language models are markedly less accurate on simple object and direction tasks when object position, scale, orientation, or context is varied, and attributes the failure to multim...

  6. MM-RLHF: The Next Step Forward in Multimodal LLM Alignment

    cs.CL 2025-02 conditional novelty 6.0 of 10

    A human-annotated multimodal preference dataset plus critique-based reward modeling and reward-margin-weighted DPO improves MLLM performance across many benchmarks.

  7. Detailed Object Description with Controllable Dimensions

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A training-free post-processing pipeline improves how well multimodal LLMs stick to user-selected object dimensions such as color, texture, and pose.

  8. NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects?

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A new benchmark shows that multimodal LLMs, including GPT-4o, consistently fail to recognize objects when their colors are modified, and that larger language models can degrade the vision encoder's performance during ...

  9. SURDS: Benchmarking Spatial Understanding and Reasoning in Driving Scenarios with Vision Language Models

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A new driving-scene benchmark of 41,080 training and 9,250 evaluation spatial questions, plus a GRPO alignment method that lifts a 3B VLM's overall score from 26.94 to 40.80.

  10. End-to-End LLM Flight Planning with RAG-based Memory and Multi-modal Coach Agent

    cs.RO 2026-07 conditional novelty 5.0 of 10

    FRAMe combines an LLM planner with RAG-based memory and a multi-modal coach agent to generate valid, preference-aligned eVTOL flight plans, achieving up to 93.8% validity across four LLMs.

  11. SEFE: Superficial and Essential Forgetting Eliminator for Multimodal Continual Instruction Tuning

    cs.LG 2025-05 conditional novelty 5.0 of 10

    SEFE reduces both answer-format drift and knowledge forgetting in multimodal continual instruction tuning by mixing question styles across tasks and penalizing changes at high-importance LoRA weight positions.

  12. Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey

    cs.CL 2024-12 conditional novelty 4.0 of 10

    A comprehensive survey that frames multimodal understanding and generation as next token prediction and proposes a five-part taxonomy.

  13. Explainable and Interpretable Multimodal Large Language Models: A Comprehensive Survey

    cs.CL 2024-12 conditional novelty 4.0 of 10

    A survey maps the field of MLLM explainability and interpretability into data, model, and training and inference perspectives.

  14. MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs

    cs.CV 2024-11 conditional novelty 4.0 of 10

    A broad survey that organizes MLLM evaluation benchmarks into capability categories, explains benchmark construction and scoring methods, and identifies gaps in current evaluation practice.

  15. PSA-VLM: Enhancing Vision-Language Model Safety through Progressive Concept-Bottleneck-Driven Alignment

    cs.CV 2024-11 conditional novelty 4.0 of 10

    PSA-VLM adds a safety classifier and prompt-rewriting gate to LLaVA-style vision-language models, reporting state-of-the-art average safety scores on RTVLM and large gains on politics, porn, and cyberbullying datasets.

  16. Natural Language Understanding and Inference with MLLM in Visual Question Answering: A Survey

    cs.CL 2024-11 conditional novelty 1.0 of 10

    A survey that organizes VQA methods from feature extraction through MLLM reasoning, datasets, and metrics, without introducing new experimental results.

Pith tools