Should you use zai-glm-4.7?
zai-glm-4.7 is listed for chat workloads and supports a 128K context window.
Check Cerebras's current endpoint status before integrating this listing. Use the alternatives table if the provider listing is unavailable.
GLM-4.7 is Z.AI's (Zhipu AI) latest-generation bilingual model, offered free through Cerebras Cloud on its WSE inference hardware.
https://api.cerebras.ai/v1 zai-glm-4.7 zai-glm-4.7 is listed for chat workloads and supports a 128K context window.
Check Cerebras's current endpoint status before integrating this listing. Use the alternatives table if the provider listing is unavailable.
We only found this glm listing on Cerebras in the current catalog. Use the related models section for nearby alternatives.
| Provider | Model listing | Access | Context | API | Limits |
|---|---|---|---|---|---|
| | zai-glm-4.7 | Check provider | 128K | OpenAI-style | 10 RPM, 100 RPD, 1M TPD |
| Model | Provider | Context | Access |
|---|---|---|---|
| Llama 3.1 70B | Cerebras | 131K | Free tier |
| zai-glm-4.7 (deprecated Aug 2026) | Cerebras | 131K | Check provider |
| llama-3.3-70b | Cerebras | 128K | Check provider |
| qwen-3-235b-a22b-instruct-2507 | Cerebras | 131K | Check provider |
| qwen-3-32b | Cerebras | 131K | Check provider |
zai-glm-4.7 is tagged for chat in this catalog and works with OpenAI-compatible client libraries.
zai-glm-4.7 appears in the free model catalog for Cerebras, but its current endpoint availability should be confirmed with the provider before use.
The model ID shown in this catalog is zai-glm-4.7.
The listed free tier limit is 10 RPM, 100 RPD, 1M TPD. Limits can change per account tier, so confirm against the provider dashboard.
The listed context window is 128K tokens with up to 8K output tokens.
GLM-4.7 is Z.AI's (Zhipu AI) latest-generation bilingual model, offered free through Cerebras Cloud on its WSE inference hardware. With 128K context, OpenAI-compatible API, and no credit card requirement, it is a strong choice for developers who need Chinese-English bilingual performance or want an alternative to the Llama/Qwen families. The free tier is more constrained than some Cerebras endpoints — 100 requests per day at 10 RPM with a 1M token daily cap — so it is best used for evaluation, comparison testing, or low-volume bilingual tasks rather than continuous production traffic.
For API keys, setup steps, and provider-level limits, see the Cerebras provider page.