Kimi Code membership guide

Kimi Code is a developer-focused benefit within the Kimi membership plan, providing high-performance AI coding capabilities. You can use this benefit through Kimi Code CLI, Claude Code, Roo Code, and other supported tools.

Key advantages

AdvantageDescription
Broad CompatibilityWorks with Kimi Code CLI, Claude Code, Roo Code, and other mainstream coding agents
Standard / HighSpeed TiersThe same model at two speeds — HighSpeed delivers roughly 5–6× the output speed of Standard and switches on demand
Ultra-Fast ResponsesGeneration speeds up to 100 tokens/s, significantly boosting coding efficiency
High-Frequency ConcurrencyApproximately 300–1,200 requests per 5-hour window (depending on your plan), with up to 30 concurrent streams

Quick start

Choose the path that fits your situation:

  • New Users: Go to kimi.com/code, sign in, and subscribe to a Coding Plan.
  • Existing Subscribers: Access the console to manage your API Keys and start using Kimi Code.

Obtaining an API key

  1. Sign in to the Kimi Console.
  2. Navigate to the API Keys page.
  3. Click Create New API Key.
  4. Copy and securely store your API Key (it is only displayed once upon creation).

Do not share your API Key with others or commit it to public code repositories.

One-click login

In Kimi Code CLI, you can use the /login command for quick authorization without manually copying an API Key:

Text

The system will automatically complete device authorization and account binding — the entire process takes just a few seconds.

Device management

  • Each account can be used across multiple devices.
  • Device authorizations that have been inactive for 30 days will automatically expire; you will need to run /login again to re-authorize.
  • You can view and manage authorized devices in the console.

How to switch models

The HighSpeed model is now available. Kimi Code offers two tiers — Standard and HighSpeed — built on the same model with identical coding ability, sharing the same Base URL, API Key, and membership benefits. HighSpeed delivers roughly 5–6× the output speed of Standard, so when you want instant responses and fast iteration, a one-click switch gives you a smoother coding experience. The main differences:

ItemStandardHighSpeed
Model IDkimi-for-codingkimi-for-coding-highspeed
Output speedBaseline~5–6× faster than Standard
Credit usageBaseline~3× that of Standard
Coding abilityFullSame as Standard
Best forEveryday coding tasksInstant responses, fast iteration
MembershipAvailable to all Kimi Code membersRequires an Allegretto plan or above

Ways to switch to the target model:

  • Official Kimi Code CLI: type /model in a session to switch directly between Standard and HighSpeed — no config changes needed.
  • Kimi Code for VS Code: pick the target model from the dropdown menu in the input bar; if HighSpeed isn't listed yet, restart VS Code or reinstall the extension.
  • Third-party tools: set the tool's Model ID to the target model; all other settings stay the same. For where to find it in each tool, see Using in Third-Party Coding Agents.
  • Stable model IDs: Both IDs are stable identifiers; the backend updates the model they map to as models improve, with no client config changes needed.
  • Enter it exactly: The high-speed ID must be kimi-for-coding-highspeed. If mistyped or set to another value, the request silently falls back to the standard kimi-for-coding — no error, but no acceleration either.
  • 401 without access: If your plan doesn't include HighSpeed access, calling it returns a 401; upgrade to Allegretto or above.

Why doesn't the whole task feel 5–6× faster? "5–6×" refers to model output speed (how fast text/code is generated). A coding task's total time is made up of "model output + tool calls (reading/writing files, running commands, web lookups, etc.) + script execution" — how long tool calls and script execution take depends on your project and commands, and HighSpeed doesn't change that part. So if the overall task doesn't feel 5–6× faster, it's usually because tool calls / script execution took up most of that turn, not because model generation slowed down.