Hosted Model Pricing for Agents
Compare the frontier models you can assign to agents inside the CloudAxis Cloud OS. Rates, context windows, and notes tuned for long-running agent workloads.
| Model | Input / 1M | Output / 1M | Context | Notes for agents |
|---|---|---|---|---|
| Claude 3.5 Sonnet | $3 | $15 | 200K | Excellent balance for research + synthesis agents |
| Claude 3 Opus | $15 | $75 | 200K | Max quality for final reports or complex reasoning |
| GPT-4o | $2.50 | $10 | 128K | Fast and capable for browser and workflow agents |
| DeepSeek V3 / R1 | $0.14 / $0.55 | $0.28 / $2.19 | 128K+ | Best price/performance for high-volume or monitoring agents |
| Moonshot | $1 | $4 | 200K–1M | Great for very long file or research contexts |
Prices are indicative list rates for hosted access inside CloudAxis. Confirm current rates before high-volume production use. All models are available without you managing API keys.
Choosing models for collaborating agents
In the Cloud OS you don't have to pick one model for everything. Assign a fast, cheap model to high-volume monitoring or data extraction agents, and reserve a higher-quality model for the final synthesis or customer-facing output step. This reference helps you make those trade-offs with real numbers. For routing logic per agent role, read choosing and routing models. When you compare subscription tiers against hosted-model caps, our Lindy Plus Pro and Max plan pricing for 2026 breakdown sits beside CloudAxis Growth math for assistant-platform buyers.
5 use cases
- Cost-optimize a 24/7 price monitor using DeepSeek, then escalate interesting findings to a Sonnet agent for analysis. For the latest DeepSeek V3 and R1 per-million token rates, see the DeepSeek per-token rates on CloudyBot when you budget high-volume monitoring agents.
- Use Moonshot for ingesting large research reports or codebases, then hand off to GPT-4o for action planning.
- Run cheap daily digests on a lightweight model, reserve Opus for the weekly executive summary.
- Browser-heavy workflows often benefit from GPT-4o's speed on tool use steps.
- Long-running file processing pipelines can stay economical on DeepSeek while still producing high-quality output when needed.
Frequently asked questions
For provider-specific rates — including Gemini and Google's hosted API tiers — see the Google model pricing reference on CloudyBot, updated alongside this CloudAxis hosted table.
Teams routing retrieval-heavy agents sometimes budget Cohere Command models separately — compare the Cohere Command API pricing reference on CloudyBot alongside the hosted rates in the table above.
Open-weight Llama routes are another line item when you assign long-context research specialists — see the Meta Llama list pricing reference on CloudyBot for per-million token rates before you scale daily duties.