GLM 5.3 FlashX
GLM-5.3-FlashX is the high-speed variant of Z.ai's GLM-5.3-Flash, a native multimodal model delivering inference speeds of up to 200 tokens/s. Built on the same hybrid sparse and linear attention architecture...
Specifications
Modalities
Supported Parameters
Max Output Tokens
131,072Reasoning Configuration
- Default
- Thinking on (effort: max)
- Selectable effort levels
- maxhighlow
- Turning thinking off
- Not possible (always thinks)
Data policy
- Prompt retention
- Unknown — we could not confirm
- Training
- Not used for training
"Unknown" does not mean "safe". It means we could not confirm it.
Whether a provider trains on prompts is a declared value from our terms with them. Retention is determined from the upstream listing for every host this model can reach. Neither is guessed.
Code Examples
curl https://api.fastmetal.ai/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer YOUR_API_KEY" \
-d '{
"model": "glm-5.3-flashx",
"messages": [{"role": "user", "content": "Hello!"}]
}'How GLM 5.3 FlashX actually answers
Real responses to our standard prompts, recorded on FastMetal.
Count the number of 'r's in 'strawberry'
Count the number of 'r's in 'strawberry'. Explain your reasoning step by step.
Debug This Error
I'm getting the following error in my Node.js application: TypeError: Cannot read properties of undefined (reading 'map') at UserList (/app/components/UserList.js:12:25) at renderWithHooks (/app/node_modules/rea…
Code Review
Please review the following Python function and suggest improvements for readability, performance, and best practices: def get_data(url, retries=3): import requests import time for i in range(retries):…
Frequently asked questions
- How much does the GLM 5.3 FlashX API cost?
- On FastMetal, GLM 5.3 FlashX is billed per token in yen: ¥66.12 per 1M input tokens and ¥223.38 per 1M output tokens, before tax. Usage is drawn from a prepaid balance; there is no subscription or monthly fee.
- Can I call GLM 5.3 FlashX with the OpenAI SDK?
- Yes. Point base_url at https://api.fastmetal.ai/v1 and pass "glm-5.3-flashx" as the model; existing OpenAI-style code works unchanged, including streaming, tool calls and structured output.
- What is the context window of GLM 5.3 FlashX?
- 1,048,576 tokens, shared between the prompt and the response.
- What do I need to try GLM 5.3 FlashX?
- Create an account and add credit; the model is then available both in the browser chat and over the API. There is no contract or minimum spend.
Try GLM 5.3 FlashX right now
GLM 5.3 FlashX is available on FastMetal through one API key. Start in the browser, or call it from the OpenAI SDK.