← Back to Home

All Models

Complete catalog of 64 models: text chat, reasoning, vision, image/video generation, voice, AI search, code, embedding.

64
models
12
categories
8
Free
⚠ Some legacy models are being deprecated. Migrate to newer versions.
Go to Playground

Data updated August 2026

Text & Chat

22

GLM-5.3

FlagshipNew

Latest flagship model and the most capable open-weights model for coding: +50% over GLM-5.2 on Z.ai Code Bench, open-source SOTA on Terminal-Bench 3.0 and Agents' Last Exam, state of the art on CyberGym (84.5) for vulnerability discovery

Context: 1M Max Output: 128K
¥8 / 28 per million tokens

GLM-5.3-Flash

NewMultimodal

First natively multimodal GLM-5 model: 320B total / 18B active parameters, first open frontier model with hybrid sparse-linear attention, vision native to the coding loop, beats GLM-5.2 at one-tenth the price

Context: 1M Max Output: 128K
¥0.8 / 2.8 per million tokens

GLM-5.2

Reasoning

Previous-generation flagship: 1M lossless context, coding capability open-source SOTA of its generation, supports stable execution of complex long-term tasks — the base that GLM-5.3 post-training builds on

Context: 1M Max Output: 128K
¥8 / 28 per million tokens

GLM-5.1

Reasoning

Coding capability aligned with Claude Opus 4.6, significant improvement in long-term tasks, can work independently for up to 8 hours

Context: 200K Max Output: 128K
¥6 / 24 per million tokens

GLM-5

Reasoning

Programming capability aligned with Claude Opus 4.5, proficient in Agentic long-term planning and execution

Context: 200K Max Output: 128K
¥4 / 18 per million tokens

GLM-5-Turbo

Reasoning

Special optimization for long tasks, good continuity in executing complex long-term tasks

Context: 200K Max Output: 128K
¥5 / 22 per million tokens

GLM-4.7

Reasoning

Comprehensive upgrade of general dialogue, reasoning, and agent capabilities, programming stronger and more stable, better aesthetics

Context: 200K Max Output: 128K
¥2 / 8 per million tokens

GLM-4.7-FlashX

Light

Lightweight and high-speed, small size with strong capability, suitable for general scenarios such as Chinese writing, translation, role-playing, etc.

Context: 200K Max Output: 128K
¥0.5 / 3 per million tokens

GLM-4.7-Flash

LightFree

Free model, providing inclusive experience, extending the general capabilities of the GLM-4.7 base

Context: 200K Max Output: 128K
Free

GLM-4.6

Context enhanced to 200K, proficient in advanced coding, complex reasoning, and tool invocation

Context: 200K Max Output: 128K
¥1 / 5 per million tokens

GLM-4.5

A balanced performance and cost general model, suitable for dialogue, analysis, and document processing

Context: 128K Max Output: 96K
¥2 / 8 per million tokens

GLM-4.5-Air

Light

High-performance and cost-effective lightweight model, with stable performance in inference, coding, and agent tasks

Context: 128K Max Output: 96K
¥0.8 / 2 per million tokens

GLM-4.5-AirX

Light

High-performance and cost-effective ultra-fast version, suitable for business scenarios with low latency and high responsiveness

Context: 128K Max Output: 96K
¥1 / 4 per million tokens

GLM-4.5-Flash

LightFree

Free model, supports deep thinking mode, supports 128K context processing

Context: 128K Max Output: 96K
Free

GLM-4-Long

Long Ctx

Supports 1M context length, designed for processing long texts and memory-intensive tasks

Context: 1M Max Output: 4K
¥1 per million tokens

GLM-4-Plus

Legacy

Fourth-generation flagship model, strong comprehensive capabilities (old version)

Context: 128K Max Output: 4K
¥5 per million tokens

GLM-4-Air

LightLegacy

Lightweight model, fast speed and low cost (old version)

Context: 128K Max Output: 4K
¥0.5 per million tokens

GLM-4-AirX

LightLegacy

Lightweight Fast Version, Faster Response Speed (Old Version)

Context: 128K Max Output: 4K
¥1 per million tokens

GLM-4-Assistant

AgentLegacy

Agent-Specific Model, Optimized for Agent Scenarios (Old Version)

Context: 128K Max Output: 4K
¥5 per million tokens

GLM-4-FlashX-250414

Light

Flash Enhanced High-Speed Version, Fast Inference Speed, Suitable for High-Concurrency Call Scenarios

Context: 128K Max Output: 16K
¥0.1 per million tokens

GLM-4-Flash-250414

LightFree

Free Model, Supports Long Context Processing, Suitable for Multilingual Understanding and Tool Invocation Scenarios

Context: 128K Max Output: 16K
Free

GLM-4-0520

LegacyDeprecated

Old Version Flagship Text Model, Discontinued on December 30, 2025

Context: 128K Max Output: 4K
¥100 per million tokens

Deep Reasoning

3

GLM-Z1-Air

Deprecated

Deep Inference Model, High-Performance Version, retired on November 15, 2025

Context: 128K Max Output: 16K
¥0.5 per million tokens

GLM-Z1-AirX

Deprecated

Deep Inference Model, Ultra-Fast Version, retired on November 15, 2025

Context: 32K Max Output: 16K
¥5 per million tokens

GLM-Z1-FlashX

Deprecated

Deep Inference Model, Fast and Cost-Effective Version, retired on November 15, 2025

Context: 128K Max Output: 16K
¥0.1 per million tokens

Vision & Multimodal

12

GLM-5V-Turbo

FlagshipNew

First Multimodal Agent Base Model, balancing visual understanding, reasoning, and code generation, supporting multimodal inputs of images, videos, and text

Context: 200K Max Output: 128K
¥5 / 22 per million tokens

GLM-4.6V

Enhanced visual reasoning capability, natively supporting tool calls and long contexts, with more stable frontend code cloning effects

Context: 128K Max Output: 32K
¥1 / 3 per million tokens

GLM-4.6V-FlashX

Light

High-performance visual reasoning model, supporting tool invocation and long context

Context: 128K Max Output: 32K
¥0.15 / 1.5 per million tokens

GLM-4.6V-Flash

LightFree

Free visual reasoning model, supporting tool invocation and long context, with flexible switching of thinking mode

Context: 128K Max Output: 32K
Free

GLM-4.5V

Legacy

Visual understanding model, supporting input of image and video files

Context: 64K Max Output:
¥2 / 6 per million tokens

GLM-4.1V-Thinking-FlashX

Light

Lightweight visual reasoning model, proficient in understanding complex scenes and multi-step analysis

Context: 64K Max Output: 16K
¥2 per million tokens

GLM-4.1V-Thinking-Flash

LightFree

Free visual reasoning model, proficient in understanding complex scenes and multi-step analysis

Context: 64K Max Output: 16K
Free

GLM-4V-Plus-0111

Legacy

Old version of image-text video recognition model (old version)

Context: 8K Max Output:
¥4 per million tokens

GLM-4V-Flash

LightFreeLegacy

Free image understanding model with basic multimodal question-answering capabilities

Context: 16K Max Output: 1K
Free

GLM-4V

Legacy

First-generation multimodal visual model (old version)

Context: 2K Max Output:
¥5 per million tokens

GLM-OCR

OCR

Lightweight image-text parsing model, balancing high precision and efficiency in document understanding, supporting complex layout parsing

Context: Max Output:
¥0.01 每千tokens

AutoGLM-Phone

AgentNew

Mobile intelligent assistant framework, supporting natural language to complete App operation tasks, covering the complete set of mobile operation commands

Context: 20K Max Output: 2K
¥0.04 每次

Image Generation

5

GLM-Image

FlagshipNew

Flagship image generation model, stronger in complex instruction following and knowledge-intensive generation, outstanding in text rendering, especially for Chinese characters

Context: 多分辨率 Max Output:
¥0.2 每次

CogView-4

General image generation model, high-quality generation, rich and diverse style expression, more complete image details

Context: 多分辨率 Max Output:
¥0.06 每次

CogView-3-Plus

Enhanced image generation model, improved image quality, supports more artistic styles

Context: 多分辨率 Max Output:
¥0.1 每次

CogView-3

Legacy

General image generation model, supports multiple resolutions and styles (old version)

Context: 多分辨率 Max Output:
¥0.06 每次

CogView-3-Flash

LightFree

Free image generation model, flexible creative generation, fast generation speed

Context: 多分辨率 Max Output:
Free

Video Generation

5

CogVideoX-3

FlagshipNew

High-intelligence flagship video model, with clear image quality, stronger command compliance and physical simulation, supporting generation of first and last frames

Context: 多分辨率 Max Output:
¥1 每次

Vidu Q1

New

High-quality video generation model, with clear image quality, smooth transitions, richer style expression, and can reduce image distortion

Context: 多分辨率 Max Output:
¥1 每次

Vidu 2

High-speed and low-cost video model, with fast generation speed, natural connection of first and last frames, and stronger consistency with multiple reference images

Context: 多分辨率 Max Output:
¥0.5 每次

CogVideoX-2

Standard video generation model, supporting multi-resolution output

Context: 多分辨率 Max Output:
¥0.5 每次

CogVideoX-Flash

LightFree

Free video generation model, supporting AI sound effects, 4K image quality, and 60fps, with a maximum video length of 10 seconds

Context: 多分辨率 Max Output:
Free

Voice & Audio

5

GLM-TTS

TTS

Voice synthesis model, supporting ultra-human voice generation and emotional expression, providing streaming and non-streaming interfaces

Context: Max Output:
¥0.5 每千字符

GLM-TTS-Clone

TTS

Voice cloning model, capable of quickly generating similar voice in 3 seconds, supporting delicate emotional expression

Context: Max Output:
¥0.2 每秒

GLM-ASR-2512

ASR

High-precision speech recognition model, with low character error rate, supporting custom vocabulary and various dialects

Context: Max Output:
¥0.06 每分钟

GLM-4-Voice

Realtime

Real-time speech dialogue model, supporting Chinese and English speech understanding and generation, adjustable for emotion, tone, and speed

Context: Max Output:
¥80 per million tokens

GLM-Realtime

RealtimeNew

Real-time audio-video model, supporting video calls and long-term dialogue memory, cross-text, audio, and video real-time inference

Context: Max Output:
¥0.04 每分钟

Code

1

CodeGeeX-4

Code completion model, suitable for code auto-completion and development assistance, improving coding efficiency

Context: 128K Max Output: 32K
¥0.1 per million tokens

Embedding

2

Embedding-3

Third-generation vector model, used for semantic retrieval enhancement, clustering, topic modeling, and classification, significantly improving the effect

Context: 8K Max Output:
¥0.5 per million tokens

Embedding-2

Legacy

Second-generation vector model, used for semantic retrieval enhancement, clustering, and classification (old version)

Context: 8K Max Output:
¥0.5 per million tokens

Rerank

1

Rerank

Text reordering model, calculates text relevance score, optimizes recall result sorting and matching effect

Context: 4K Max Output:
¥0.8 per million tokens

Character

2

CharGLM-4

Animate conversation model, suitable for emotional companionship and virtual character interaction, supporting more natural character expression

Context: 8K Max Output: 4K
¥1 per million tokens

Emohaa

Psychological and emotional support model, with professional counseling and emotional guidance capabilities, helping users understand emotions and cope with problems

Context: 8K Max Output: 4K
¥15 per million tokens

Open Source

2

GLM-4-9B

Open Source

Open-source 9B parameter model, supports 128K context, can be locally deployed, MIT license

Context: 128K Max Output:
Free & open source

ChatGLM3-6B

Open SourceLegacy

Third-generation open-source 6B dialogue model, the first generation ChatGLM series, MIT license

Context: 32K Max Output:
Free & open source