Pith. sign in

REVIEW 10 cited by

PanGu-{\Sigma}: Towards Trillion Parameter Language Model with Sparse Heterogeneous Computing

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.10845 v1 pith:3GWPVC6O submitted 2023-03-20 cs.CL

classification cs.CL
keywords languagemodelpangu-sigmacomputinggenerationheterogeneousparameter
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The scaling of large language models has greatly improved natural language understanding, generation, and reasoning. In this work, we develop a system that trained a trillion-parameter language model on a cluster of Ascend 910 AI processors and MindSpore framework, and present the language model with 1.085T parameters named PanGu-{\Sigma}. With parameter inherent from PanGu-{\alpha}, we extend the dense Transformer model to sparse one with Random Routed Experts (RRE), and efficiently train the model over 329B tokens by using Expert Computation and Storage Separation(ECSS). This resulted in a 6.3x increase in training throughput through heterogeneous computing. Our experimental findings show that PanGu-{\Sigma} provides state-of-the-art performance in zero-shot learning of various Chinese NLP downstream tasks. Moreover, it demonstrates strong abilities when fine-tuned in application data of open-domain dialogue, question answering, machine translation and code generation.

Discussion (0). Sign in to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. HetCCL: Enabling Collective Communication For Mixed-Vendor Heterogeneous Clusters

    cs.NI 2026-05 unverdicted novelty 7.0 of 10

    HetCCL enables efficient collective communication across mixed-vendor GPU clusters via P2P transport and a border-communicator mechanism, delivering 17-19x higher bandwidth than Gloo and up to 16.9% faster end-to-end ...

  2. SAMoRA: Semantic-Aware Mixture of LoRA Experts for Task-Adaptive Learning

    cs.CL 2026-04 unverdicted novelty 6.0 of 10

    SAMoRA is a parameter-efficient fine-tuning framework that uses semantic-aware routing and task-adaptive scaling within a Mixture of LoRA Experts to improve multi-task performance and generalization over prior methods.

  3. Fine-grained Approaches for Confidence Calibration of LLMs in Automated Code Revision

    cs.SE 2026-04 unverdicted novelty 6.0 of 10

    Local Platt scaling on three fine-grained confidence scores reduces calibration error for LLM-based automated code revision across tasks and models compared to global scaling alone.

  4. Universal Pansharpening Model

    cs.CV 2026-03 conditional novelty 6.0 of 10

    A single pansharpening model works across 4-, 7-, 8-, and 10-band satellite images by projecting arbitrary-band MS data into a fixed latent space and fusing with PAN via a latent diffusion bridge.

  5. DEER: Disentangled Mixture of Experts with Instance-Adaptive Routing for Generalizable Machine-Generated Text Detection

    cs.CL 2025-11 conditional novelty 6.0 of 10

    DEER, a disentangled mixture-of-experts detector with RL-based instance routing, reports F1 gains of about 1.4 in-domain and 5.3 points out-of-domain over prior MGT detectors.

  6. DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models

    cs.CL 2024-01 unverdicted novelty 5.0 of 10

    DeepSeekMoE 2B matches GShard 2.9B performance and approaches a dense 2B model; the 16B version matches LLaMA2-7B at 40% compute by using fine-grained expert segmentation plus shared experts.

  7. From Chatbot to Digital Colleague: The Paradigm Shift Toward Persistent Autonomous AI

    cs.AI 2026-06 conditional novelty 4.0 of 10

    Autonomous AI becomes dependable when tool use is embedded in persistent workspaces with reusable skills, shifting evaluation from answers to task closure.

  8. StackingNet: Collective Inference Across Independent AI Foundation Models

    cs.AI 2026-02 conditional novelty 4.0 of 10

    A lightweight weighted-average 'meta-model' over black-box LLM/VLM outputs improves accuracy, reduces bias, and ranks/prunes unreliable models across regression and classification tasks.

  9. A Survey of Large Language Models

    cs.CL 2023-03 accept novelty 3.0 of 10

    This survey reviews the background, key techniques, and evaluation methods for large language models, emphasizing emergent abilities that appear at large scales.

  10. A Comprehensive Overview of Large Language Models

    cs.CL 2023-07 unverdicted novelty 2.0 of 10

    A survey paper providing an overview of Large Language Models, their background, and recent advances in the field.

Pith tools