VocalFuse is a Fuse Intelligence product.

KIMI CODE

Kimi Code CLI: The MIT Terminal Agent Without a Meter

Kimi Code CLI is Moonshot AI’s terminal coding agent — the successor to kimi-cli, rewritten in TypeScript and shipped as a single MIT-licensed binary that installs with one command and needs no Node.js. It reads and edits code, runs shell commands, fetches web pages, dispatches coder / explore / plan subagents in isolated contexts, speaks the Agent Client Protocol for Zed and JetBrains, and is genuinely model-agnostic: Anthropic-, OpenAI-, and Google-compatible endpoints all work through config. The catch every comparison buries: the MIT tool is free, the tokens are not — memberships run $19–$199/month or the K2.7 Code API bills $0.95 input / $4.00 output per 1M tokens, inside three overlapping quota windows. This page is the corrected record — plus the practical setup: bring-your-own-key, local endpoints, or run Kimi Code alongside Claude Code, Codex, and Cursor as widgets in VibeFuse, the free Windows AI harness — no harness meter, ever.

Why developers search for Kimi Code in 2026

MIT, single binary, model-agnostic

The install is one line — a curl/PowerShell script that ships a compiled binary with a millisecond-start TUI (no Node.js, no PATH gymnastics), or npm install -g @moonshot-ai/kimi-code on Node 24.15+. Version 0.42.0 (September 2026) pulls 33,000+ weekly npm downloads. MCP servers are configured conversationally via /mcp-config, lifecycle hooks gate risky tool calls, and the skills/MCP/data-source marketplace surfaces each install’s trust level up front. Model-agnostic means Anthropic, OpenAI, and Google-compatible endpoints all work through config — Kimi’s models are the default, not the lock-in.

The kimi-cli succession

The original kimi-cli — the Python/Apache-2.0 project that passed 10,000 GitHub stars — froze at v1.49.0, and its own README says Kimi CLI is evolving into Kimi Code CLI, with configuration and session auto-migration on install. Homebrew marked the old formula deprecated with a disable date of January 17, 2027. Plenty of 2026 tutorials still point at the dead repo. If you are evaluating today: install kimi-code, run /login, pick Kimi Code OAuth or a Kimi Platform API key — or wire in a provider you already pay for.

K2.7 Code: the price-per-token play

Moonshot’s pricing documentation (September 2026): kimi-k2.7-code bills $0.95 per 1M input tokens on cache miss ($0.19 on cache hit) and $4.00 per 1M output, with a 262,144-token context. The HighSpeed variant — the same model at roughly 180 tokens/s, up to 260 in short contexts — doubles both rates to $1.90/$8.00. The Batch API runs at 60% of standard ($0.57/$2.40 per 1M), the cheapest published route for bulk work. At $4.00 per 1M output, K2.7 Code runs roughly a fifth of Opus-class output pricing — and Moonshot reports it cuts thinking-token usage ~30% versus K2.6.

The quota reality

Three overlapping controls govern Kimi Code, and none is published as a simple prompt count: a shared monthly membership credit pool, a Kimi Code weekly allowance (refreshed every seven days, no rollover), and a separate usage window the help center describes as 5 hours per week. Community tooling that parses Kimi’s own BillingService endpoint reads the same cap as 200 requests per 5 hours on every tier, with weekly request quotas around 1,024 / 2,048 / 7,168 by plan. Burn any layer and Kimi Code pauses until that layer resets — unused quota never carries over.

What Kimi Code costs now (September 2026)

Path Cost What you get Fine print
The CLI itself $0, MIT Full agent: subagents, ACP, MCP, hooks, skills marketplace, video input Single-binary install script or npm; v0.42.0, several releases a week
Kimi membership $19 / $39 / $99 / $199 per month Weekly-refreshed Kimi Code allowances at 1x / 5x / 15x / 30x, multi-device login, HighSpeed on upper tiers English pricing page tiers; the yuan help-center ladder (¥49/¥99/¥199/¥699) uses different tier names; no rollover anywhere
Kimi Platform API Token-priced $0.95 cache-miss / $0.19 cache-hit input, $4.00 output per 1M (K2.7 Code, 262K context) HighSpeed $1.90/$8.00; Batch API 60% ($0.57/$2.40); gateways track ~$0.68–$0.74/$3.40–$3.50
BYOK / local endpoints Your provider’s rate Point the CLI at Anthropic-, OpenAI-, or Google-compatible servers — or your own machine Model-agnostic config is the design; Moonshot models stay the default, not the lock-in
VibeFuse harness $0, forever Runs Kimi Code next to Claude Code, Codex, Gemini CLI, and Cursor as widgets over your repos Your vendor billing, your model choice per task, no harness meter

Verified against Moonshot’s K2.7 Code pricing documentation, the kimi.com membership surfaces, the MoonshotAI GitHub org, and the npm registry (September 2026). The tier names differ between the English pricing page and the help center — re-check your account’s entitlements before citing any quota number.

Kimi Code vs Claude Code vs the harness shape

What matters Kimi Code CLI Claude Code VibeFuse
License MIT CLI; Kimi model weights closed Proprietary Free download, free license key
Default model Kimi K2.7 Code (always-thinking) Claude Opus/Sonnet only Whatever agent widget you run — mixable per task
Install Single binary, no Node.js required via script npm Windows installer + free license key
Model switching Anthropic/OpenAI/Google-compatible endpoints via config Anthropic endpoint (or the env-var swap to third-party models) One canvas, five official CLIs, each on its own vendor account
Free tier CLI free; tokens need a membership ($19/mo entry) or API key None (Pro $20/mo entry) Harness costs nothing; your vendor plans set the volume
Benchmarks 81.1% MCPMark Verified (vendor-run); no SWE-bench Verified submission as of mid-2026 Frontier reasoning lead on verified suites (Opus-class) N/A — the harness runs the same agents you would benchmark
Rate shape Weekly allowance + 5-hour window + monthly credit pool, none rolling over 5-hour session windows on subscription plans No meter at the harness layer — your vendors set the limits
Where code lives Your machine; Moonshot cloud sees context unless you BYOK or go local Anthropic cloud Your machine — nothing uploads by default
Economy Memberships + pay-as-you-go API; MIT ecosystem Subscription or API Open-source marketplace: sell widgets, skills, styling packs for revenue share

Be honest: where Kimi Code still trails

Credit where due first: Kimi Code is the most complete MIT-licensed terminal agent shipping in 2026 — a single binary with subagents, lifecycle hooks, ACP editor support, conversational MCP config, a trust-labeled install marketplace, and video input, releasing several times a week. The honest gaps: the headline model numbers are vendor-run — K2.7 Code had not been submitted to independent suites like SWE-bench Verified or Terminal-Bench as of mid-2026, and the 81.1% MCPMark figure comes from Moonshot-side runs (third-party cross-harness comparisons carry the usual harness caveats; the base Kimi K2 reads a competitive-but-not-leading ~53.7 on LiveCodeBench v6). The model always thinks — thinking mode cannot be disabled and sampling parameters are fixed (temperature 1.0, top-p 0.95), so every request carries a reasoning budget you pay for. And the quota shape bites: three overlapping windows, none rolling over, so a burned weekly allowance parks the tool for days. If you want the deepest verified-quality ecosystem, Claude Code is still the safer premium default; if you want an open, hackable, cheap-per-token agent you can point anywhere, Kimi Code is the strongest MIT entry in the class — and inside VibeFuse it runs next to every other vendor’s official CLI, so a quota wall on one widget never stops the canvas.

Point it anywhere

The model-agnostic config is the quiet superpower: edit the provider block and Kimi Code speaks to Anthropic-compatible, OpenAI-compatible, or Google-compatible endpoints — including a local server on your own machine. That makes it the cheapest way to keep one terminal-agent muscle memory while rotating the model underneath it: K2.7 Code for bulk loops at $4.00 per 1M output, a frontier model for the hard reviews, a local endpoint for air-gapped work. Your hooks, MCP servers, and skills carry over across all of them.

Manage the weekly wall

The three-control quota shape (weekly allowance, 5-hour window, monthly credit pool) means a hard day of agent work can park Kimi Code until reset. The VibeFuse pattern fixes the workflow, not the quota: run Kimi Code as one widget on the canvas, and when its allowance burns, route the task to Claude Code, Codex, or Cursor in the same session — each billing its own vendor, none waiting on Kimi’s seven-day refresh. The harness never bills a request and your code never uploads by default.

When to pick which

Cost-sensitive, high-volume agent loops: Kimi Code on the Platform API or a membership tier — output tokens run roughly a fifth of Opus-class pricing. Verified frontier quality and the deepest plugin ecosystem: Claude Code. Refusal to be locked to either: both, as widgets in VibeFuse — Kimi for the meter-heavy bulk, Claude for the high-stakes review, on one free canvas. They are different meters for different jobs, not rivals for one.

Also compare the Qwen Code record (the other big 2026 lab CLI), the Gemini CLI shutdown record, the Claude Code tutorial, the Codex CLI vs Claude Code breakdown, the Cursor CLI alternative, the AI agent harness guide, and the best vibe coding tools.

Kimi Code FAQ

What is Kimi Code CLI?

Kimi Code CLI is Moonshot AI's terminal coding agent - the successor to kimi-cli, rewritten in TypeScript and shipped as a single MIT-licensed binary (a one-line install script for macOS/Linux/Windows, or npm install -g @moonshot-ai/kimi-code; the script path needs no Node.js). It reads and edits code, runs shell commands, searches and fetches web pages, and plans multi-step tasks from feedback. Built-in coder, explore, and plan subagents run in isolated context windows, MCP servers are configured conversationally via /mcp-config, lifecycle hooks gate risky tool calls, and kimi acp speaks the Agent Client Protocol so Zed and JetBrains can drive a session. It defaults to Kimi K2.7 Code but is model-agnostic: OpenAI-compatible, Anthropic-compatible, and Google-compatible endpoints all work through config.

Is Kimi Code free?

The CLI is free and open source (MIT license, v0.42.0 as of September 2026 with 33,000+ weekly npm downloads) - but like every 2026 agent CLI, the tokens are not. Three routes: (1) a Kimi membership subscription, from \$19/mo on the English pricing page, which bundles weekly Kimi Code quotas; (2) a Kimi Platform API key billed per token (\$0.95 input cache-miss / \$4.00 output per 1M on K2.7 Code); (3) bring your own provider - the CLI accepts Anthropic-, OpenAI-, and Google-compatible endpoints, so you can point it at any model you already pay for or a local server. Installing Kimi Code also auto-migrates configuration and sessions from the old kimi-cli.

How much does the Kimi K2.7 Code API cost?

Per Moonshot's pricing documentation (September 2026): kimi-k2.7-code bills \$0.95 per 1M input tokens on cache miss (\$0.19 on cache hit) and \$4.00 per 1M output tokens, with a 262,144-token context window. The HighSpeed variant - the same model at roughly 180 tokens/s output, up to 260 in short contexts - doubles both rates: \$1.90 cache-miss input / \$8.00 output. The Batch API runs at 60% of standard pricing (\$0.57 / \$2.40 per 1M), the cheapest published route for bulk work. Third-party gateways track K2.7 Code slightly under list (about \$0.68-\$0.74 input / \$3.40-\$3.50 output). At \$4.00 per 1M output tokens, K2.7 Code runs roughly one-fifth of frontier Opus-class output pricing - the reason long autonomous agent loops are its natural workload.

What do Kimi membership plans include for coding?

The English pricing page (August 2026) lists four tiers at \$19, \$39, \$99, and \$199 per month with relative Kimi Code allowances of 1x, 5x, 15x, and 30x; the help center surfaces the same ladder in yuan (\u00a549 / \u00a599 / \u00a5199 / \u00a5699) under different tier names (Andante/Moderato/Allegretto/Allegro vs Moderato/Allegretto/Allegro/Vivace) - check the entitlements inside your own account rather than assuming a name maps across surfaces. Kimi Code is available to all paid members (model id kimi-for-coding); the HighSpeed tier requires an upper-tier plan. A free Adagio tier exists for general Kimi use, but Kimi Code quota is a paid-membership feature, and third-party reviewers reported membership signups hitting a waitlist in August 2026.

What are the actual Kimi Code rate limits?

Three overlapping controls govern Kimi Code, and none is published as a simple prompt count: (1) a shared monthly membership credit pool that all Kimi features draw from, refreshed monthly with no rollover; (2) a Kimi Code weekly allowance refreshed every seven days from your subscription date, also with no rollover; (3) a separate Kimi Code usage window the help center describes as 5 hours per week and community tooling that parses Kimi's own BillingService endpoint reads as a 200-requests-per-5-hours rolling cap on every tier. The same tooling reports weekly request quotas of about 1,024 (\u00a549 tier), 2,048 (\u00a599), and 7,168 (\u00a5199). Burn any layer and Kimi Code pauses until that layer resets; unused quota never carries over.

What happened to kimi-cli?

Moonshot wound it down in favor of Kimi Code CLI. The legacy Python/Apache-2.0 project (10,000+ GitHub stars) froze at v1.49.0, and its own README states that Kimi CLI is evolving into Kimi Code CLI and that installing the new CLI auto-migrates your configuration and sessions. The successor is TypeScript, MIT-licensed, ships as a single binary with a millisecond-start TUI, and releases several times a week. Homebrew has marked the kimi-cli formula deprecated with a disable date of January 17, 2027. If you are evaluating today, install kimi-code - the old repo remains documented but is not where development happens.

How good is K2.7 Code compared to Claude?

Moonshot reports K2.7 Code cuts thinking-token usage by roughly 30% versus K2.6 while improving agentic performance about 10%, and third-party MCPMark Verified reads put it at 81.1% against a reported 76.4% for Claude Opus 4.8 on the same benchmark. The honest caveat: K2.7 Code had not been submitted to independent suites like SWE-bench Verified or Terminal-Bench as of mid-2026, and most headline numbers are vendor-run or cross-harness comparisons - treat them as directional. Claude keeps the verified-quality and ecosystem edge; Kimi's edge is price per token (\$4.00 per 1M output is a fraction of Opus-class pricing) plus an MIT-licensed, model-agnostic agent wrapped around it. Note K2.7 Code always thinks - thinking mode cannot be disabled and sampling parameters are fixed.

What is VibeFuse's take on Kimi Code?

Kimi Code is a natural fit for the VibeFuse canvas: a free Windows AI harness where every agent runs as a widget over your real repositories. VibeFuse charges \$0 - no harness meter, no weekly allowance, no 5-hour window - your Kimi Platform key (or any other provider's) does the token billing, per widget and per task. That pairing fixes Kimi Code's sharpest limitation: quota burn stays isolated to one widget while your other agents keep working, and when you hit the weekly wall you can route the task to a different vendor's CLI on the same canvas instead of waiting for a reset. VibeFuse also runs an open-source marketplace where creators sell widgets, skills, and styling packs for revenue share - pair it with VocalFuse for local voice dictation while you build.

Run Kimi Code — and everything else — on one free canvas

Download VibeFuse free, register for a free license key, and run Kimi Code alongside Claude Code, Codex CLI, Gemini CLI, and Cursor as widgets over your repositories on Windows — $0 harness, no usage meter, your vendor accounts, code stays local. Pair it with VocalFuse voice if you dictate while you build.

Explore related AI note taking guides