Early access: features, availability, and pricing may change as Kyara Intelligence evolves.

All models
Below baseline Reasoning

Z.ai: GLM 5.3 Flash

Z.AI's fast multimodal reasoner for responsive, million-context RP.

Model ID z-ai/glm-5.3-flash
Quick start

Drop this model into any OpenAI-compatible client by setting the base URL, your API key, and the model ID below.

curl https://api.kyara-intelligence.com/v1/chat/completions \
  -H "Authorization: Bearer $KYARA_INTELLIGENCE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "z-ai/glm-5.3-flash",
    "messages": [
      { "role": "user", "content": "Hello!" }
    ]
  }'
See the full API reference
Overview

Z.AI's GLM-5.3-Flash is a native multimodal reasoning model built for efficient long-horizon work. Its 320B-parameter MoE activates just 18B parameters per token, pairing low latency with strong coding and agentic performance. A hybrid sparse-and-linear attention architecture keeps its 1M token context window precise and affordable, while image and video understanding support visually grounded scenes. With selectable low, high, or max reasoning effort, it can move naturally between quick dialogue and deeper narrative planning.

Specifications
Context / Memory 🌌 Extreme · 1M tokens
Cost vs baseline Below baseline
Input
0.2x
Output
0.2x
Input
1.4M
cr / M tokens
Output
4.5M
cr / M tokens

Multiples are relative to the catalog median; bars are scaled to the most expensive model.

Model ID
z-ai/glm-5.3-flash
Reasoning
Available
Status
Current
Upstream moderation

All models on Kyara Intelligence are accessed through their original API providers (Mistral, Z.AI/GLM, xAI, DeepSeek, and others). Each provider's content policies and moderation apply to all requests.

Kyara routes API calls without modifying inputs or outputs, and does not log prompt or response content. Users are responsible for complying with the relevant provider's terms of service in addition to Kyara's terms.