Head-to-head showdown: GLM-5 (Z.ai) ($1.00 in / $3.20 out per 1M) vs Kimi K2.7 Code (OpenRouter) ($0.67 in / $3.40 out per 1M). Kimi K2.7 Code (OpenRouter) is 1.0× cheaper across standard token mixes, with 200K vs 262.1K context windows.
glm-5 · zai
kimi-k2.7-code · moonshot
Benchmark
| Workload Scenario | GLM-5 (Z.ai) | Kimi K2.7 Code (OpenRouter) | Price Delta |
|---|---|---|---|
| 1M input tokens (raw text) | $1.00 | $0.67 | +$0.33 |
| 1M output tokens (generation) | $3.20 | $3.40 | −$0.20 |
| 1M tokens · 70% input / 30% output mix | $1.66 | $1.489 | +$0.171 |
| Standard chat turn (4K in / 800 out, 50% cached) | $0.00496 | $0.00444 | +$0.00052 |
| Monthly scale (10K requests / day) | $1,488.00 | $1,332.00 | +$156.00 |
Negative difference = GLM-5 (Z.ai) is cheaper. Positive = Kimi K2.7 Code (OpenRouter) is cheaper.
Scaling Curve
Total cost of a token volume at 70/30 input/output split (uncached). Log-log scale.
Verdict
Kimi K2.7 Code (OpenRouter) has the cheaper input rate, while GLM-5 (Z.ai) has the cheaper output rate. The crossover happens when output makes up about 62% of your total tokens.
Below that share (retrieval, summarization, extraction — lots of context in, little text out) Kimi K2.7 Code (OpenRouter) is cheaper. Above it (generation, translation, coding — long completions) GLM-5 (Z.ai) wins.
FAQ
On input, Kimi K2.7 Code (OpenRouter) is cheaper ($0.67/M vs $1.00/M). On output, GLM-5 (Z.ai) is cheaper ($3.20/M vs $3.40/M). For workloads where more than 62% of tokens are output, the output-cheaper model wins overall.
A chat-style request (4,000 input + 800 output tokens, 50% cached) costs $0.00496 on GLM-5 (Z.ai) and $0.00444 on Kimi K2.7 Code (OpenRouter) — Kimi K2.7 Code (OpenRouter) is 1.1× more expensive for that workload.
GLM-5 (Z.ai) supports 200,000 tokens (131.1K max output); Kimi K2.7 Code (OpenRouter) supports 262,144 (235.9K max output). Kimi K2.7 Code (OpenRouter) fits 1.3× more context, which matters for long documents and agents.