readme.md

For the complete documentation index, see llms.txt. Markdown versions of documentation pages are available by appending .md to page URLs; this page is available as Markdown.

Welcome

Parasail is a cloud GPU platform for AI inference. We aggregate top AI hardware providers to give you scalable, on-demand access to powerful computing resources—without costly hardware investments or lengthy contracts.

We offer three ways to run inference, each designed for a different workload pattern:

Type Description Link
Serverless On-demand, pay-per-token inference for popular open-source models. No setup required—just call the API. Link
Dedicated Instances Private GPU endpoints with full control over model, hardware, and scaling. Ideal for production workloads. Link
Batch Processing Process millions of inferences at 50% off serverless pricing. OpenAI-compatible batch API. Link

Not sure where to start?

Pick a product based on your workload:

If you... Use Why
Want to call a popular open-source model with no setup Serverless Pay-per-token, instant, OpenAI-compatible
Need a specific or private model, predictable latency, or production SLAs Dedicated Instances Private GPU endpoint with full control over model, hardware, and scaling
Process large volumes offline where real-time isn't required Batch 50% off serverless pricing, up to millions of requests

Find your path

Choose the option that best matches what you're trying to do:

Type Description Link
Chat and text generation Build chatbots, assistants, and text generation pipelines Link
RAG and embeddings Build retrieval-augmented generation and vector search systems Link
Batch at scale Process large volumes of prompts, images, or text offline Link
Agents and tool calling Build agentic workflows with function calling and multi-step reasoning Link

Quickstart

Make your first API call in under two minutes:

Type Description Link
Serverless quickstart Call a serverless model with the OpenAI SDK Link
Dedicated quickstart Deploy your own model on a dedicated GPU Link
Batch quickstart Submit a batch job in five lines of Python Link

API reference

Parasail's API is fully OpenAI-compatible. Point your existing OpenAI SDK at our base URL and start making requests.

See the API Reference for full details on authentication, endpoints, and parameters.

For AI coding agents

Building on Parasail on behalf of a user? Read llms.txt for a flat index of every page, then see For AI Coding Agents for per-product setup notes. In short: use the OpenAI SDK against https://api.parasail.io/v1 with the PARASAIL_API_KEY environment variable—every endpoint is OpenAI-compatible.

Explore

Type Description Link
Guides Step-by-step tutorials for common tasks Link
Billing Pricing and billing details Link
Operations Production readiness, retries, limits, and quotas Link
Security Security, privacy, Trust Center, and account controls Link