A single unified multimodal model matches leading task-specialized vision systems across detection, segmentation, dense geometry, and multi-view 3D by casting all outputs as native text or image generation.
hub
Semantic image synthesis with spatially-adaptive normalization
6 Pith papers cite this work, alongside 2,832 external citations. Polarity classification is still indexing.
hub tools
years
2026 6representative citing papers
Introduces OCR-Robust benchmark and evaluates 18 VLMs showing clean accuracy does not guarantee robustness with charts and tables more fragile than documents under selected perturbations.
A model-agnostic Geometric Risk Controller reduces extreme errors in VLM-based OCR by requiring cross-view consensus before accepting outputs.
Introduces the CIFAR Synthetic Evidence Corpus, a multi-family dataset of AI-manipulated documents with source-separated train/test splits for evaluating detectors of AI-generated legal evidence.
Classifies all simple Whittaker bar S_2-modules in each block Omega and establishes two category equivalences, one to finite-dimensional modules over the parabolic subalgebra bar S_2^{>=0} and one to H_1-fmod.
SPADE-LDM conditional synthesis from composite semantic masks produces realistic 3D LGE MRI that raises LA cavity Dice from 0.908 to 0.936.
citing papers explorer
-
Vision as Unified Multimodal Generation
A single unified multimodal model matches leading task-specialized vision systems across detection, segmentation, dense geometry, and multi-view 3D by casting all outputs as native text or image generation.
-
How Robust is OCR-Reasoning? Evaluating OCR-Reasoning Robustness of Vision-Language Models under Visual Perturbations
Introduces OCR-Robust benchmark and evaluates 18 VLMs showing clean accuracy does not guarantee robustness with charts and tables more fragile than documents under selected perturbations.
-
From Plausibility to Verifiability: Risk-Controlled Generative OCR with Vision-Language Models
A model-agnostic Geometric Risk Controller reduces extreme errors in VLM-based OCR by requiring cross-view consensus before accepting outputs.
-
The CIFAR Synthetic Evidence Corpus for Detecting AI-Generated Evidence
Introduces the CIFAR Synthetic Evidence Corpus, a multi-family dataset of AI-manipulated documents with source-separated train/test splits for evaluating detectors of AI-generated legal evidence.
-
The category of Whittaker modules over the Cartan Type Lie algebra $\bar{S}_2$
Classifies all simple Whittaker bar S_2-modules in each block Omega and establishes two category equivalences, one to finite-dimensional modules over the parabolic subalgebra bar S_2^{>=0} and one to H_1-fmod.
-
3D Conditional Image Synthesis of Left Atrial LGE MRI from Composite Semantic Masks
SPADE-LDM conditional synthesis from composite semantic masks produces realistic 3D LGE MRI that raises LA cavity Dice from 0.908 to 0.936.