GLM-130B
GLM-130B: An open 130B-parameter bilingual transformer for high-performance research and real-world NLP
Quick facts
- Best for
- GLM-130B: An open 130B-parameter bilingual transformer for high-performance research and real-world NLP
- Pricing
- Freemium
- Editor rating
- 4.5 / 5
- Community saves
- 0
About GLM-130B
GLM-130B is a 130-billion-parameter bilingual (Chinese–English) transformer-based General Language Model from THUDM/THUKEG, released as an open model for academic research and certain commercial use. Trained on 400B tokens with the GLM library, it delivers strong results on a range of NLP benchmarks and ships with downloadable checkpoints and inference/deployment code for efficient multi-GPU and mixed-precision serving.
Pros
- Open bilingual pre-trained model supporting Chinese and English
- 130B-parameter transformer-based large language model
- Open repository with model/code licenses; permitted research and certain commercial uses per repo
- Trained on 400B text tokens with the GLM training framework
- Optimized large-scale training using the GLM library (parallelism and efficiency)
- Strong reported performance across multiple NLP benchmarks
- Inference and deployment scripts, including multi-GPU and mixed-precision
- Downloadable checkpoints/weights for immediate use
- Instruction-style and few-shot prompting capabilities
- Associated ICLR 2023 paper describing model design and training
