- List price (verify vendors): DeepSeek V4 Flash $0.14 / $0.28 · Luna $0.20 / $1.20 · DeepSeek V4 Pro $0.435 / $0.87 per 1M
- Flash still wins raw $/token — especially on output (28¢ vs Luna’s $1.20)
- Luna wins product fit for OpenAI-shaped agents: vision, Responses/tool surface, “Luna without Plus,” US-vendor compliance stories
- We can prove the switch: both models are in our hosted pool; Auto Mode default moved from DeepSeek → Luna after the Jul 30 cut
- On CloudAxis credits: DeepSeek ≈ 1.5× · Luna ≈ 2.5× — Luna burns more credits per token, by design
This is the comparison most “cheapest LLM” pages dodge: after OpenAI cut Luna 80%, does it kill DeepSeek? No. It redraws when you pick each. Full Luna card: complete guide.
Price table (API list)
Confirm on OpenAI pricing and DeepSeek’s current API pricing page before you lock a budget. Numbers operators quote as of late July 2026:
- DeepSeek V4 Flash (
deepseek-v4-flash) — $0.14 input / $0.28 output · cached input ~$0.0028 · ~1M context - GPT-5.6 Luna (
gpt-5.6-luna) — $0.20 / $1.20 · cached $0.02 · ~1.05M context (after Jul 30 −80% cut) - DeepSeek V4 Pro (
deepseek-v4-pro) — $0.435 / $0.87 · cached ~$0.003625 · ~1M context
On a balanced 2:1 input:output mix, Flash is still the bargain. Luna’s output rate is the gap that shows up in chatty agent loops. Pro sits between them on input but still undercuts Luna on output. Model this on our cost calculator or the live pricing tracker.
What we actually run (proof)
Unlike aggregator blogs, we host both. DeepSeek V4 Flash/Pro and GPT-5.6 Luna are in the CloudyBot/CloudAxis OpenClaw pool (tier 1 / Free-accessible). Through mid-2026, Auto Mode’s cheap default was DeepSeek-shaped. After Luna’s cut made a current-generation OpenAI model free-tier-viable, Auto Mode standardized on Luna — no picker, every plan including Free. See Auto Mode policy and the Luna landing.
Internal credit multipliers (how we bill AI credits, not vendor invoices): DeepSeek V4 Flash/Pro ≈ 1.5×; Luna ≈ 2.5×. So even on our meters, Luna is the dearer “cheap” lane — we did not switch to save pennies.
Capability tradeoffs that matter for agents
Beyond the price sheet:
- Vision / multimodal: Luna accepts image input on the OpenAI surface; DeepSeek V4 Flash is text-first in most hosted configs — matters for screenshot-heavy browser agents
- Tool APIs: Luna inherits GPT-5.6 Responses tooling (incl. programmatic tool calling). DeepSeek stays OpenAI-compatible chat/completions with its own thinking modes — excellent for volume, different ecosystem
- Vendor / compliance posture: some buyers need a US frontier vendor on the invoice; others optimize purely for $/token and open-weight optionality
- Long context: both advertise ~1M windows; Luna’s usable recall can degrade well before 1M (see the Luna guide caveats). Do not assume either is “perfect recall at 1M”
- Alias gotchas: bare
gpt-5.6is Sol, not Luna — bill-shock post. DeepSeek retireddeepseek-chat/deepseek-reasoneraliases on Jul 24, 2026 — pindeepseek-v4-flash/pro
Worked cost sketch (same token budget)
Illustrative only — your tool loops dominate. Assume 10M input + 5M output tokens/month, no cache hits:
- DeepSeek V4 Flash: 10×$0.14 + 5×$0.28 = $2.80
- GPT-5.6 Luna: 10×$0.20 + 5×$1.20 = $8.00
- DeepSeek V4 Pro: 10×$0.435 + 5×$0.87 = $8.70
Flash wins the spreadsheet. Luna and Pro land in a similar monthly band on this mix — Luna cheaper on input, Pro cheaper on output. Cache-heavy DeepSeek workloads widen Flash’s lead further (cached input ~$0.0028 vs Luna’s $0.02).
Verdict: which to use
Default rules we use internally:
- Pick DeepSeek V4 Flash when raw $/token and cache-friendly volume are the only scoreboard — classification, extraction, high-QPS chat, cost-floor experiments
- Pick DeepSeek V4 Pro when you want DeepSeek’s stronger tier without jumping to Western flagship prices
- Pick GPT-5.6 Luna when you want OpenAI-native agents (vision, Responses tools), ChatGPT-adjacent quality at the cheap GPT-5.6 tier, or Luna without Plus inside a hosted OS
- On CloudAxis: you do not pick — Auto Mode is Luna. That is a product decision (outcome over catalog), not a claim that Luna is the cheapest API row
Related reading
Cluster links:
Frequently asked questions
Is GPT-5.6 Luna cheaper than DeepSeek V4 Flash?
No on typical list rates. Flash is about $0.14/$0.28 per 1M vs Luna’s $0.20/$1.20 after the July 30 cut. Flash especially wins on output tokens.
Why did CloudAxis Auto Mode switch from DeepSeek to Luna?
After Luna’s 80% cut, a current-generation OpenAI model became free-tier-viable for hosted agents. Auto Mode favors outcome and OpenAI-shaped tooling over winning the absolute cheapest API row. Plans still have no model picker.
Should I use DeepSeek V4 Pro instead of Luna?
Often on pure price/quality within the DeepSeek family. Pro is roughly $0.435/$0.87 — competitive with Luna’s monthly band on mixed workloads, with DeepSeek’s own strengths. Choose by vendor needs and evals, not slogans.
Do you still host DeepSeek?
Yes. DeepSeek V4 Flash and Pro remain in the hosted pool (Free-tier accessible). Auto Mode itself is fixed to GPT-5.6 Luna.
Which model ID should API builders pin?
For Luna: gpt-5.6-luna (not gpt-5.6 — that is Sol). For DeepSeek: deepseek-v4-flash or deepseek-v4-pro. Legacy deepseek-chat / deepseek-reasoner aliases were retired July 24, 2026.
GPT-5.6 Luna complete guide · Luna on CloudAxis Free · Sol vs Terra vs Luna · AI cost calculator · What it costs to run AI agents