Complete catalog of 64 models: text chat, reasoning, vision, image/video generation, voice, AI search, code, embedding.
Data updated August 2026
Latest flagship model and the most capable open-weights model for coding: +50% over GLM-5.2 on Z.ai Code Bench, open-source SOTA on Terminal-Bench 3.0 and Agents' Last Exam, state of the art on CyberGym (84.5) for vulnerability discovery
First natively multimodal GLM-5 model: 320B total / 18B active parameters, first open frontier model with hybrid sparse-linear attention, vision native to the coding loop, beats GLM-5.2 at one-tenth the price
Previous-generation flagship: 1M lossless context, coding capability open-source SOTA of its generation, supports stable execution of complex long-term tasks — the base that GLM-5.3 post-training builds on
Coding capability aligned with Claude Opus 4.6, significant improvement in long-term tasks, can work independently for up to 8 hours
Programming capability aligned with Claude Opus 4.5, proficient in Agentic long-term planning and execution
Special optimization for long tasks, good continuity in executing complex long-term tasks
Comprehensive upgrade of general dialogue, reasoning, and agent capabilities, programming stronger and more stable, better aesthetics
Lightweight and high-speed, small size with strong capability, suitable for general scenarios such as Chinese writing, translation, role-playing, etc.
Free model, providing inclusive experience, extending the general capabilities of the GLM-4.7 base
Context enhanced to 200K, proficient in advanced coding, complex reasoning, and tool invocation
A balanced performance and cost general model, suitable for dialogue, analysis, and document processing
High-performance and cost-effective lightweight model, with stable performance in inference, coding, and agent tasks
High-performance and cost-effective ultra-fast version, suitable for business scenarios with low latency and high responsiveness
Free model, supports deep thinking mode, supports 128K context processing
Supports 1M context length, designed for processing long texts and memory-intensive tasks
Fourth-generation flagship model, strong comprehensive capabilities (old version)
Lightweight model, fast speed and low cost (old version)
Lightweight Fast Version, Faster Response Speed (Old Version)
Agent-Specific Model, Optimized for Agent Scenarios (Old Version)
Flash Enhanced High-Speed Version, Fast Inference Speed, Suitable for High-Concurrency Call Scenarios
Free Model, Supports Long Context Processing, Suitable for Multilingual Understanding and Tool Invocation Scenarios
Old Version Flagship Text Model, Discontinued on December 30, 2025
Deep Inference Model, High-Performance Version, retired on November 15, 2025
Deep Inference Model, Ultra-Fast Version, retired on November 15, 2025
Deep Inference Model, Fast and Cost-Effective Version, retired on November 15, 2025
First Multimodal Agent Base Model, balancing visual understanding, reasoning, and code generation, supporting multimodal inputs of images, videos, and text
Enhanced visual reasoning capability, natively supporting tool calls and long contexts, with more stable frontend code cloning effects
High-performance visual reasoning model, supporting tool invocation and long context
Free visual reasoning model, supporting tool invocation and long context, with flexible switching of thinking mode
Visual understanding model, supporting input of image and video files
Lightweight visual reasoning model, proficient in understanding complex scenes and multi-step analysis
Free visual reasoning model, proficient in understanding complex scenes and multi-step analysis
Old version of image-text video recognition model (old version)
Free image understanding model with basic multimodal question-answering capabilities
First-generation multimodal visual model (old version)
Lightweight image-text parsing model, balancing high precision and efficiency in document understanding, supporting complex layout parsing
Mobile intelligent assistant framework, supporting natural language to complete App operation tasks, covering the complete set of mobile operation commands
Flagship image generation model, stronger in complex instruction following and knowledge-intensive generation, outstanding in text rendering, especially for Chinese characters
General image generation model, high-quality generation, rich and diverse style expression, more complete image details
Enhanced image generation model, improved image quality, supports more artistic styles
General image generation model, supports multiple resolutions and styles (old version)
Free image generation model, flexible creative generation, fast generation speed
High-intelligence flagship video model, with clear image quality, stronger command compliance and physical simulation, supporting generation of first and last frames
High-quality video generation model, with clear image quality, smooth transitions, richer style expression, and can reduce image distortion
High-speed and low-cost video model, with fast generation speed, natural connection of first and last frames, and stronger consistency with multiple reference images
Standard video generation model, supporting multi-resolution output
Free video generation model, supporting AI sound effects, 4K image quality, and 60fps, with a maximum video length of 10 seconds
Voice synthesis model, supporting ultra-human voice generation and emotional expression, providing streaming and non-streaming interfaces
Voice cloning model, capable of quickly generating similar voice in 3 seconds, supporting delicate emotional expression
High-precision speech recognition model, with low character error rate, supporting custom vocabulary and various dialects
Real-time speech dialogue model, supporting Chinese and English speech understanding and generation, adjustable for emotion, tone, and speed
Real-time audio-video model, supporting video calls and long-term dialogue memory, cross-text, audio, and video real-time inference
Enhanced AI search tool, deeply retrieves web information, supports complex queries and multi-round searches
Standard AI search tool, quickly retrieves web information, suitable for general search scenarios
Sogou Engine search enhancement, combined with Sogou search results to provide more accurate information retrieval
Quark Engine search enhancement, combined with Quark search results to provide more comprehensive content retrieval
Code completion model, suitable for code auto-completion and development assistance, improving coding efficiency
Third-generation vector model, used for semantic retrieval enhancement, clustering, topic modeling, and classification, significantly improving the effect
Second-generation vector model, used for semantic retrieval enhancement, clustering, and classification (old version)
Text reordering model, calculates text relevance score, optimizes recall result sorting and matching effect
Animate conversation model, suitable for emotional companionship and virtual character interaction, supporting more natural character expression
Psychological and emotional support model, with professional counseling and emotional guidance capabilities, helping users understand emotions and cope with problems
Open-source 9B parameter model, supports 128K context, can be locally deployed, MIT license
Third-generation open-source 6B dialogue model, the first generation ChatGLM series, MIT license