Hosted Model Pricing for Agents

Compare the frontier models you can assign to agents inside the CloudAxis Cloud OS. Rates, context windows, and notes tuned for long-running agent workloads.

ModelInput / 1MOutput / 1MContextNotes for agents
Claude 3.5 Sonnet$3$15200KExcellent balance for research + synthesis agents
Claude 3 Opus$15$75200KMax quality for final reports or complex reasoning
GPT-4o$2.50$10128KFast and capable for browser and workflow agents
DeepSeek V3 / R1$0.14 / $0.55$0.28 / $2.19128K+Best price/performance for high-volume or monitoring agents
Moonshot$1$4200K–1MGreat for very long file or research contexts

Prices are indicative list rates for hosted access inside CloudAxis. Confirm current rates before high-volume production use. All models are available without you managing API keys.

Choosing models for collaborating agents

In the Cloud OS you don't have to pick one model for everything. Assign a fast, cheap model to high-volume monitoring or data extraction agents, and reserve a higher-quality model for the final synthesis or customer-facing output step. This reference helps you make those trade-offs with real numbers. For routing logic per agent role, read choosing and routing models. When you compare subscription tiers against hosted-model caps, our Lindy Plus Pro and Max plan pricing for 2026 breakdown sits beside CloudAxis Growth math for assistant-platform buyers.

5 use cases

  • Cost-optimize a 24/7 price monitor using DeepSeek, then escalate interesting findings to a Sonnet agent for analysis. For the latest DeepSeek V3 and R1 per-million token rates, see the DeepSeek per-token rates on CloudyBot when you budget high-volume monitoring agents.
  • Use Moonshot for ingesting large research reports or codebases, then hand off to GPT-4o for action planning.
  • Run cheap daily digests on a lightweight model, reserve Opus for the weekly executive summary.
  • Browser-heavy workflows often benefit from GPT-4o's speed on tool use steps.
  • Long-running file processing pipelines can stay economical on DeepSeek while still producing high-quality output when needed.

Frequently asked questions

The table shows the underlying hosted rates. Your effective cost inside a CloudAxis plan depends on the plan tier, included credits, and volume. Use the cost estimator tool to model your actual spend.
Yes — that's a core feature of the OS. Each agent in a workflow or on the home screen can have its own model preference.
Vision and tool-specific pricing are handled inside the platform. This reference focuses on the text model rates most relevant for planning agent text workflows.
It is a maintained reference. For production budgeting, always cross-check the current public pricing pages of the model providers.

For provider-specific rates — including Gemini and Google's hosted API tiers — see the Google model pricing reference on CloudyBot, updated alongside this CloudAxis hosted table.

Teams routing retrieval-heavy agents sometimes budget Cohere Command models separately — compare the Cohere Command API pricing reference on CloudyBot alongside the hosted rates in the table above.

Open-weight Llama routes are another line item when you assign long-context research specialists — see the Meta Llama list pricing reference on CloudyBot for per-million token rates before you scale daily duties.

The estimator helps you model volume. This table helps you choose which model(s) to use for each role in the volume. Use both together.
Built by CloudAxis — free Agent OS tools for operators. All tools · Put work on autopilot →