Pith. sign in

REVIEW 4 cited by

Lumina-OmniLV: A Unified Multimodal Framework for General Low-Level Vision

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2504.04903 v2 pith:ZRUTVQOY submitted 2025-04-07 cs.CV cs.AI

classification cs.CVcs.AI
keywords low-levelvisionframeworkgenerativeimagemulti-taskmultimodalomnilv
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present Lunima-OmniLV (abbreviated as OmniLV), a universal multimodal multi-task framework for low-level vision that addresses over 100 sub-tasks across four major categories: image restoration, image enhancement, weak-semantic dense prediction, and stylization. OmniLV leverages both textual and visual prompts to offer flexible and user-friendly interactions. Built on Diffusion Transformer (DiT)-based generative priors, our framework supports arbitrary resolutions -- achieving optimal performance at 1K resolution -- while preserving fine-grained details and high fidelity. Through extensive experiments, we demonstrate that separately encoding text and visual instructions, combined with co-training using shallow feature control, is essential to mitigate task ambiguity and enhance multi-task generalization. Our findings also reveal that integrating high-level generative tasks into low-level vision models can compromise detail-sensitive restoration. These insights pave the way for more robust and generalizable low-level vision systems.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Universal Image Restoration via Internalized Chain-of-Thought Reasoning

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    CoTIR fine-tunes a pre-trained image editing model using a differentiable CoT-style objective inspired by Lagrangian optimization to enable single-pass universal image restoration, supported by a new 5.2M-sample bench...

  2. Rethinking Where to Edit: Task-Aware Localization for Instruction-Based Image Editing

    cs.CV 2026-04 unverdicted novelty 6.0 of 10

    Task-aware localization via attention cues and feature centroids from source/target streams in IIE models improves non-edit consistency while preserving instruction following.

  3. TIR-Agent: Training an Explorative and Efficient Agent for Image Restoration

    cs.CV 2026-03 conditional novelty 6.0 of 10

    A vision-language agent trained with SFT plus RL, exploration-driven trajectory perturbation, and adaptive multi-metric rewards learns direct tool selection for composite image restoration, beating training-free agent...

  4. Step1X-Edit: A Practical Framework for General Image Editing

    cs.CV 2025-04 unverdicted novelty 4.0 of 10

    Step1X-Edit integrates a multimodal LLM with a diffusion decoder, trained on a custom high-quality dataset, to deliver image editing performance that surpasses open-source baselines and approaches proprietary models o...

Pith tools