Early access: features, availability, and pricing may change as Kyara Intelligence evolves.

All models
Below baseline Reasoning Obsolete

DeepSeek: DeepSeek V4 Flash 0731

A refreshed V4 Flash with sharper reasoning for fast, coherent long-context RP.

This is a legacy model, superseded by newer releases. It remains fully available to use.

Model ID deepseek/deepseek-v4-flash-0731
Quick start

Drop this model into any OpenAI-compatible client by setting the base URL, your API key, and the model ID below.

curl https://api.kyara-intelligence.com/v1/chat/completions \
  -H "Authorization: Bearer $KYARA_INTELLIGENCE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deepseek/deepseek-v4-flash-0731",
    "messages": [
      { "role": "user", "content": "Hello!" }
    ]
  }'
See the full API reference
Overview

A re-post-trained revision of DeepSeek's V4 Flash, built on the same sparse 284B-parameter MoE with 13B active. This refresh sharpens reasoning and instruction-following, keeping character voices consistent and plot logic tight even in demanding, fast-moving scenes. Its massive 1M-token context window holds entire story arcs, lore, and callbacks in view, delivering quick, coherent premium responses at a budget-friendly price.

Specifications
Context / Memory 🌌 Extreme · 1M tokens
Cost vs baseline Below baseline
Input
0.2x
Output
0.1x
Input
1.6M
cr / M tokens
Output
3.2M
cr / M tokens

Multiples are relative to the catalog median; bars are scaled to the most expensive model.

Model ID
deepseek/deepseek-v4-flash-0731
Reasoning
Available
Status
Obsolete
Upstream moderation

All models on Kyara Intelligence are accessed through their original API providers (Mistral, Z.AI/GLM, xAI, DeepSeek, and others). Each provider's content policies and moderation apply to all requests.

Kyara routes API calls without modifying inputs or outputs, and does not log prompt or response content. Users are responsible for complying with the relevant provider's terms of service in addition to Kyara's terms.