Parasail
Pricing
We believe that democratizing access to compute is the first step to enabling the next wave of AI breakthroughs for all. That's why we're offering the best prices on the latest hardware and models, all delivered through reliable solutions that fit your needs as you scale.
Our approach is simple — by building relationships with leading cloud providers, we unlock a massive abundance of compute and the best industry prices with a scalable, reliable experience. We embrace open source and its many advantages to distribute the best solutions and innovation to everyone.
All New Accounts: Will be charged every $25 of spend, more information can be found here on billing and the ability to increase the threshold to trigger billing.
Select Your Instance Type
Choose the best infrastructure for your needs—whether you’re running real-time inference, dedicated high-performance GPUs, or large-scale batch jobs.
Serverless
Price-per-token to ensure you only pay for the compute you actually use, offering flexibility and cost-efficiency at scale.
View Serverless Pricing
Dedicated
Hourly pricing for dedicated GPUs gives you predictable costs and full control over high-performance resources tailored to your workload.
View Dedicated Pricing
Batch
Optimize batch processing with our fast and low-cost API using your preferred models and embeddings.
View Batch Pricing
Batch
Our self-service batch processing API offers the best pricing and speeds for your largest jobs. We identify a unique configuration of our fleet, including spot instances, to deliver the best value for your needs. We apply a discount to our serverless pricing for using our batch API as well as an additional discount for cached tokens.
Get started now with Batch
Price per million tokens
| Model Size | Input | Output | Cache Read |
|---|---|---|---|
| 0 - 4.1B params, FP8 | $0.03 / MTok | $0.05 / MTok | $0.01 / MTok |
| 4.1B - 8.1B params, FP8 | $0.04 / MTok | $0.09 / MTok | $0.02 / MTok |
| 8.1B - 16.1B params, FP8 | $0.05 / MTok | $0.14 / MTok | $0.03 / MTok |
| 16.1B - 21.1B params, FP8 | $0.07 / MTok | $0.22 / MTok | $0.03 / MTok |
| 21.1B - 41.1B params, FP8 | $0.10 / MTok | $0.32 / MTok | $0.05 / MTok |
| 41.1B - 80.1B params, FP8 | $0.18 / MTok | $0.50 / MTok | $0.09 / MTok |
| 80.1B - 150.1B params, FP8 | $0.25 / MTok | $0.68 / MTok | $0.12 / MTok |
| 150.1B - 250.1B params, FP8 | $0.32 / MTok | $0.86 / MTok | $0.20 / MTok |
| 250.1B - 500B params, FP8 | $0.57 / MTok | $1.55 / MTok | $0.29 / MTok |
| 500B+ params, FP8 | $0.74 / MTok | $1.93 / MTok | $0.37 / MTok |
Models with special batch pricing
| Model | Input | Output | Cache Read |
|---|---|---|---|
| Qwen/Qwen3-235B-A22B-Instruct-2507-FP8 | $0.40 / MTok | $0.40 / MTok | - |
| Qwen/Qwen3.5-397B-A17B-FP8 | $0.25 / MTok | $1.80 / MTok | - |
| deepseek-ai/DeepSeek-R1-0528 | $2.00 / MTok | $2.00 / MTok | - |
| google/gemma-4-31B-it | $0.07 / MTok | $0.20 / MTok | - |
Need enterprise-grade support? Contact us to learn more about our enterprise offering.