Models — Parasail

Featured models

reasoningZ.AI | Released: 6-26

GLM-5.2
Z.AI’s next-gen model for agentic engineering with sustained execution on long-horizon tasks.
1M tokens · Text

reasoningZ.AI | Released: 4-26

GLM-5.1
Strong reasoning and coding model with a 1M context window at competitive pricing.
1M tokens · Text

reasoningMiniMax | Released: 5-26

MiniMax M3
Highly efficient MoE model delivering strong reasoning at exceptionally low cost.
1M tokens · Text

reasoningDeepSeek | Released: 5-26

DeepSeek V4 Pro
Open model for reasoning and coding workloads.
128K tokens · Text

codingMoonshot | Released: 5-26

Kimi K2.7 Code
Specialized coding model optimized for software engineering and agentic coding workflows.
256K tokens · Text

visionAlibaba | Released: 5-26

Qwen3-VL 235B-A22B
Large vision-language model with strong document understanding and visual reasoning.
256K tokens · Text + Vision

Recently added

compactAlibaba | Released: 6-26

Qwen3.6 35B-A3B
Efficient MoE model with 3B active parameters — great price-to-performance.
128K tokens · Text

reasoningNVIDIA | Released: 6-26

Nemotron 3 Ultra 550B
NVIDIA’s flagship open model optimized for enterprise reasoning workloads.
128K tokens · Text

reasoningZ.AI | Released: 6-26

GLM-5.2
Z.AI’s next-gen model for agentic engineering with sustained execution on long-horizon tasks.
1M tokens · Text

Explore the library

reasoningZ.AI | Released: 6-26

GLM-5.2
Z.AI’s next-gen model for agentic engineering with sustained execution on long-horizon tasks.
1M tokens · Text

reasoningZ.AI | Released: 4-26

GLM-5.1
Strong reasoning and coding model with a 1M context window at competitive pricing.
1M tokens · Text

reasoningZ.AI | Released: 2-26

GLM-5
Z.AI’s open model for reasoning and agentic workflows.
128K tokens · Text

reasoningMiniMax | Released: 5-26

MiniMax M3
Highly efficient MoE model delivering strong reasoning at exceptionally low cost.
1M tokens · Text

reasoningMiniMax | Released: 3-26

MiniMax M2.5
Cost-effective large model with strong general capabilities and long context.
1M tokens · Text

reasoningDeepSeek | Released: 5-26

DeepSeek V4 Pro
Open model for reasoning and coding workloads.
128K tokens · Text

compactDeepSeek | Released: 5-26

DeepSeek V4 Flash
Ultra-fast, ultra-cheap reasoning model for high-throughput workloads.
128K tokens · Text

reasoningMoonshot | Released: 4-26

Kimi K2.6
Strong reasoning model with excellent long-context understanding.
256K tokens · Text

reasoningAlibaba | Released: 5-26

Qwen3.5 397B-A17B
Large-scale MoE model with frontier reasoning and multilingual capabilities.
256K tokens · Text

reasoningAlibaba | Released: 7-25

Qwen3 235B-A22B
Strong open MoE model with excellent coding and reasoning at low cost.
256K tokens · Text

reasoningTrinity | Released: 4-26

Trinity Large (Thinking)
Reasoning-focused model with extended thinking for complex problem-solving.
128K tokens · Text

reasoningNVIDIA | Released: 6-26

Nemotron 3 Ultra 550B
NVIDIA’s flagship open model optimized for enterprise reasoning workloads.
128K tokens · Text

reasoningMeta | Released: 4-25

Llama 4 Maverick
Meta’s multimodal Llama model with native image understanding.
1M tokens · Text + Vision

generalMeta | Released: 12-24

Llama 3.3 70B
Meta’s reliable workhorse — great quality-to-cost ratio for production.
128K tokens · Text

codingAlibaba | Released: 5-26

Qwen3-Coder-Next
Code-specialized model with strong fill-in-the-middle and agentic coding support.
256K tokens · Text

visionAlibaba | Released: 5-26

Qwen3-VL 8B
Compact vision-language model — fast and affordable for image understanding at scale.
128K tokens · Text + Vision

visionAlibaba | Released: 1-25

Qwen2.5-VL 72B
Proven vision-language model for OCR, document QA, and visual grounding.
128K tokens · Text + Vision

compactAlibaba | Released: 6-26

Qwen3.6 35B-A3B
Efficient MoE model with 3B active parameters — great price-to-performance.
128K tokens · Text

compactAlibaba | Released: 5-26

Qwen3.5 35B-A3B
Compact MoE model with strong reasoning at minimal cost per token.
128K tokens · Text

compactAlibaba | Released: 4-26

Qwen3-Next 80B
Balanced model with strong general capabilities at a competitive price point.
128K tokens · Text

compactMistral | Released: 3-26

Mistral Small 3.2 24B
Mistral’s efficient model — excellent for high-volume, latency-sensitive workloads.
128K tokens · Text

compactGoogle | Released: 3-26

Gemma 4 26B-A4B
Google’s efficient MoE Gemma model with 4B active parameters.
128K tokens · Text

compactGoogle | Released: 3-26

Gemma 4 31B
Google’s latest open model with strong general performance and broad language support.
128K tokens · Text

compactGoogle | Released: 3-25

Gemma 3 27B
Proven compact model with multimodal capabilities and strong multilingual support.
128K tokens · Text

compactOpenAI | Released: 8-25

gpt-oss-120b
OpenAI’s open-weight model with strong reasoning at a low cost.
128K tokens · Text

compactOpenAI | Released: 8-25

gpt-oss-120b (Fast)
Optimized fast tier of gpt-oss-120b with lower latency for real-time workloads.
128K tokens · Text

compactOpenAI | Released: 8-25

gpt-oss-20b
Compact open-weight model from OpenAI — ultra-cheap, fits on a single GPU.
128K tokens · Text

compactXiaomi | Released: 4-26

MiMo v2.5
Xiaomi’s efficient reasoning model with strong coding capabilities.
128K tokens · Text

compactSkyfall | Released: 2-26

Skyfall 31B v4.2
Fine-tuned model optimized for creative writing and conversational tasks.
128K tokens · Text

compactSkyfall | Released: 2-25

Skyfall 36B v2
Fine-tuned model with enhanced instruction-following and creative output.
128K tokens · Text

compactCydonia | Released: 8-25

Cydonia 24B v4.1
Fine-tuned model with enhanced roleplay and character consistency.
32K tokens · Text

embeddingsBAAI | Released: 1-24

BGE-M3
Multilingual, multi-functional embedding model supporting 100+ languages.
8K tokens · Embedding

audioResemble | Released: 4-26

Resemble TTS (English)
High-quality text-to-speech model with natural-sounding English voices.
— · Audio

agentsByteDance | Released: 4-25

UI-TARS 1.5 7B
GUI agent model that can operate computer interfaces autonomously.
128K tokens · Text + Vision

Find the right model for your workload

Start with free credits — no credit card. Call any model in minutes with an OpenAI-compatible API.