

ModelGLM 5.3 FlashX by z-ai
Text+image+video->text
Free credits included. No credit card required.
GLM-5.3-FlashX is the high-speed variant of Z.ai's GLM-5.3-Flash, delivering inference speeds of up to 200 tokens/s. It is the first natively multimodal model in the GLM-5 series, built on a 320B-total/18B-active Mixture-of-Experts architecture with a novel hybrid sparse and linear attention design that reduces attention computation and KV cache costs significantly. The model supports a 1M-token context window with native image and video input, and excels at visual coding, agentic workflows, Office document generation, and long-context tasks. It is released under an MIT license with open weights on Hugging Face.