Pith. sign in

REVIEW 26 cited by

Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2211.01335 v3 pith:V4JKKJPJ submitted 2022-11-02 cs.CV cs.CL

Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese

classification cs.CV cs.CL
keywords chineseclipachievemodelsperformancepretrainingcontrastivedataset
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

The tremendous success of CLIP (Radford et al., 2021) has promoted the research and application of contrastive learning for vision-language pretraining. In this work, we construct a large-scale dataset of image-text pairs in Chinese, where most data are retrieved from publicly available datasets, and we pretrain Chinese CLIP models on the new dataset. We develop 5 Chinese CLIP models of multiple sizes, spanning from 77 to 958 million parameters. Furthermore, we propose a two-stage pretraining method, where the model is first trained with the image encoder frozen and then trained with all parameters being optimized, to achieve enhanced model performance. Our comprehensive experiments demonstrate that Chinese CLIP can achieve the state-of-the-art performance on MUGE, Flickr30K-CN, and COCO-CN in the setups of zero-shot learning and finetuning, and it is able to achieve competitive performance in zero-shot image classification based on the evaluation on the ELEVATER benchmark (Li et al., 2022). We have released our codes, models, and demos in https://github.com/OFA-Sys/Chinese-CLIP

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 26 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Disentangling Fact from Sentiment: A Dynamic Conflict-Consensus Framework for Multimodal Fake News Detection

    cs.LG 2025-12 unverdicted novelty 7.0

    DCCF disentangles fact and sentiment in multimodal data, applies dynamic polarization to extract conflicts, and uses a conflict-consensus mechanism to improve fake news detection accuracy by 3.52% on average over baselines.

  2. Illuminating Visual Identity in Universal Multimodal Embeddings

    cs.CV 2026-08 conditional novelty 6.0

    By adding identity-aware sampling and a contrastive loss on a new 28-dataset benchmark, the authors build multimodal embeddings that are far better at visual identity matching without losing general retrieval accuracy.

  3. GALA: Generative Aligned Learning for Adaptive Multimodal Representation in the Taobao Shangou Recommender System

    cs.IR 2026-07 conditional novelty 6.0

    A three-stage multimodal recommender pipeline with GRPO-based behavior alignment and adaptive ID-content fusion claims a 0.55% online order-volume increase and small offline AUC gains at Taobao Shangou.

  4. RecGPT-V3 Technical Report

    cs.IR 2026-07 conditional novelty 6.0

    A stateful LLM recommender with memory, text-plus-Semantic-ID grounding, and latent reasoning reports higher Taobao engagement and sales at ~52% lower serving compute than its predecessor.

  5. Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation

    cs.CV 2026-06 unverdicted novelty 6.0

    Qwen-RobotWorld is a language-conditioned video world model using Double-Stream MMDiT, an 8.6M-frame embodied corpus, and progressive curriculum training that ranks first on EWMBench and DreamGen Bench.

  6. Fine-grained Fragment Retrieval in Multi-modal Long-form Dialogues

    cs.CL 2026-06 unverdicted novelty 6.0

    Introduces FFR task, F2RVLM and FFRS models, and MLDR dataset for retrieving coherent multi-modal dialogue fragments, reporting superior performance on single-dialogue and corpus benchmarks.

  7. MyoSem: Aligning Electromyography to Natural-Language Action Semantics for Hand Action Understanding

    cs.CV 2026-05 unverdicted novelty 6.0

    MyoSem is a multimodal alignment framework that maps EMG signals to text-based action semantics for bidirectional retrieval and improved generalization in hand action understanding.

  8. MindAlign: Bridging EEG, Vision, and Language for Zero-Shot Visual Decoding

    cs.LG 2026-05 unverdicted novelty 6.0

    A tri-modal contrastive learning method for EEG-based zero-shot visual decoding reports 54.1% top-1 accuracy on the Things-EEG2 200-way benchmark, outperforming prior baselines of 32.4%.

  9. TIGER-FG: Text-Guided Implicit Fine-Grained Grounding for E-commerce Retrieval

    cs.IR 2026-05 unverdicted novelty 6.0

    TIGER-FG proposes text-guided implicit fine-grained grounding with dual distillation to address modality and granularity asymmetries in image-to-multimodal e-commerce retrieval, reporting Recall@1 gains of 6.1 and 34....

  10. Text-Guided Visual Representation Learning for Robust Multimodal E-Commerce Recommendation

    cs.IR 2026-05 unverdicted novelty 6.0

    TGQ-Former uses metadata-guided hybrid queries and dual-gated modulation to improve visual token selection in multimodal e-commerce retrieval, raising average Hit Rate@100 by 6.04% over baselines.

  11. DRG-Font: Dynamic Reference-Guided Few-shot Font Generation via Contrastive Style-Content Disentanglement

    cs.CV 2026-04 unverdicted novelty 6.0

    DRG-Font generates stylistically consistent glyphs from few references by decomposing style and content via contrastive disentanglement, dynamic reference selection, and multi-scale fusion blocks.

  12. Maximizing Mutual Information Between Prompt and Response Improves LLM Performance With No Additional Data

    cs.LG 2026-03 unverdicted novelty 6.0

    MIPO constructs contrastive preference pairs from correct versus random prompts and uses DPO to maximize mutual information between prompts and responses, producing 3-40% gains on personalization and 1-18% on math tas...

  13. Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

    cs.CV 2025-11 conditional novelty 6.0

    A 6B single-stream diffusion transformer trained with heavily curated data reaches top open-source image-generation quality in 314K H800 GPU hours, releasing Turbo and Edit variants.

  14. FG-CLIP 2: A Bilingual Fine-grained Vision-Language Alignment Model

    cs.CV 2025-10 conditional novelty 6.0

    A two-stage bilingual CLIP-style model with region-text supervision and a new text-side contrastive loss outperforms prior open models on fine-grained vision-language tasks in English and Chinese.

  15. Qwen-Audio-VAE Technical Report

    eess.AS 2026-07 conditional novelty 5.0

    A 12.5 Hz continuous audio VAE reconstructs speech, music, and sound well while encoding 64×30s clips in 541 ms after latency-aware encoder pruning.

  16. BamiBERT: A New BERT-based Language Model for Vietnamese

    cs.CL 2026-07 unverdicted novelty 5.0

    BamiBERT is a new base-sized Vietnamese BERT model trained on raw text that outperforms PhoBERT on 11 of 15 metrics across 8 benchmarks.

  17. JuZhou 1.0 Technical Report: The First Edge-Native Text-to-Image Foundation Model Trained Entirely on China-Developed AI Accelerators

    cs.CV 2026-06 conditional novelty 5.0

    A compact Chinese-native T2I U-Net (~0.387B) trained on Sugon K100, distilled to 4 steps, reports GenEval 0.69 and ~1.6–4.5s offline mobile generation.

  18. CHaystack: Benchmarking Chinese Document Retrieval and VQA

    cs.IR 2026-05 conditional novelty 5.0

    CHaystack, a Chinese document retrieval-and-VQA benchmark, shows Qwen3-VL reaching 71.91 Recall@1 versus 14.40 for the best non-Qwen model, and its VLM filter improves recall.

  19. Maximizing Mutual Information Between Prompt and Response Improves LLM Performance With No Additional Data

    cs.LG 2026-03 unverdicted novelty 5.0

    Contrastive preference pairs from correct vs random prompts, optimized with DPO, maximize base-model PMI and improve personalization, math, and QA without extra data or verifiers.

  20. JARVIS: An Evidence-Grounded Retrieval System for Interpretable Deceptive Reviews Adjudication

    cs.IR 2026-02 unverdicted novelty 5.0

    JARVIS combines hybrid retrieval and evidence graphs with LLMs to raise deceptive-review detection precision from 0.953 to 0.988 and recall from 0.830 to 0.901 on a custom dataset while cutting manual inspection time ...

  21. Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

    cs.CV 2025-11 unverdicted novelty 5.0

    Z-Image is an efficient 6B-parameter foundation model for image generation that rivals larger commercial systems in photorealism and bilingual text rendering through a new single-stream diffusion transformer and strea...

  22. FORGE: Forming Semantic Identifiers for Generative Retrieval in Industrial Datasets

    cs.IR 2025-09 conditional novelty 5.0

    FORGE shows that balancing codebook usage and adding multimodal side information improves semantic identifiers for generative retrieval, validated offline and on Taobao.

  23. InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

    cs.CV 2023-12 unverdicted novelty 5.0

    InternVL scales a vision model to 6B parameters and aligns it with LLMs using web data to achieve state-of-the-art results on 32 visual-linguistic benchmarks.

  24. JuZhou 1.0 Technical Report: The First Edge-Native Text-to-Image Foundation Model Trained Entirely on China-Developed AI Accelerators

    cs.CV 2026-06 unverdicted novelty 4.0

    JuZhou 1.0 is a 0.387B-parameter T2I diffusion model with 4-step inference achieving 0.69 GenEval, trained on 9M Chinese pairs using Sugon K100 accelerators and deployable on Android/iOS devices.

  25. UniNote: A Unified Embedding Model for Multimodal Representation and Ranking

    cs.IR 2026-05 unverdicted novelty 4.0

    UniNote proposes a two-stage trained unified embedding model (contrastive SFT then RL) for multimodal I2I retrieval that claims SOTA results and was deployed at Xiaohongshu with MRL for improved quality and efficiency.

  26. Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

    cs.CV 2025-02 unverdicted novelty 4.0

    Step-Video-T2V describes a 30B-parameter text-to-video model with custom Video-VAE, 3D DiT, flow matching, and Video-DPO that claims state-of-the-art results on a new internal benchmark.