Models — Parasail
Featured models
reasoningZ.AI | Released: 6-26
GLM-5.2
Z.AI’s next-gen model for agentic engineering with sustained execution on long-horizon tasks.
1M tokens · Text
reasoningZ.AI | Released: 4-26
GLM-5.1
Strong reasoning and coding model with a 1M context window at competitive pricing.
1M tokens · Text
reasoningMiniMax | Released: 5-26
MiniMax M3
Highly efficient MoE model delivering strong reasoning at exceptionally low cost.
1M tokens · Text
reasoningDeepSeek | Released: 5-26
DeepSeek V4 Pro
Open model for reasoning and coding workloads.
128K tokens · Text
codingMoonshot | Released: 5-26
Kimi K2.7 Code
Specialized coding model optimized for software engineering and agentic coding workflows.
256K tokens · Text
visionAlibaba | Released: 5-26
Qwen3-VL 235B-A22B
Large vision-language model with strong document understanding and visual reasoning.
256K tokens · Text + Vision
Recently added
compactAlibaba | Released: 6-26
Qwen3.6 35B-A3B
Efficient MoE model with 3B active parameters — great price-to-performance.
128K tokens · Text
reasoningNVIDIA | Released: 6-26
Nemotron 3 Ultra 550B
NVIDIA’s flagship open model optimized for enterprise reasoning workloads.
128K tokens · Text
reasoningZ.AI | Released: 6-26
GLM-5.2
Z.AI’s next-gen model for agentic engineering with sustained execution on long-horizon tasks.
1M tokens · Text
Explore the library
reasoningZ.AI | Released: 6-26
GLM-5.2
Z.AI’s next-gen model for agentic engineering with sustained execution on long-horizon tasks.
1M tokens · Text
reasoningZ.AI | Released: 4-26
GLM-5.1
Strong reasoning and coding model with a 1M context window at competitive pricing.
1M tokens · Text
reasoningZ.AI | Released: 2-26
GLM-5
Z.AI’s open model for reasoning and agentic workflows.
128K tokens · Text
reasoningMiniMax | Released: 5-26
MiniMax M3
Highly efficient MoE model delivering strong reasoning at exceptionally low cost.
1M tokens · Text
reasoningMiniMax | Released: 3-26
MiniMax M2.5
Cost-effective large model with strong general capabilities and long context.
1M tokens · Text
reasoningDeepSeek | Released: 5-26
DeepSeek V4 Pro
Open model for reasoning and coding workloads.
128K tokens · Text
compactDeepSeek | Released: 5-26
DeepSeek V4 Flash
Ultra-fast, ultra-cheap reasoning model for high-throughput workloads.
128K tokens · Text
reasoningMoonshot | Released: 4-26
Kimi K2.6
Strong reasoning model with excellent long-context understanding.
256K tokens · Text
reasoningAlibaba | Released: 5-26
Qwen3.5 397B-A17B
Large-scale MoE model with frontier reasoning and multilingual capabilities.
256K tokens · Text
reasoningAlibaba | Released: 7-25
Qwen3 235B-A22B
Strong open MoE model with excellent coding and reasoning at low cost.
256K tokens · Text
reasoningTrinity | Released: 4-26
Trinity Large (Thinking)
Reasoning-focused model with extended thinking for complex problem-solving.
128K tokens · Text
reasoningNVIDIA | Released: 6-26
Nemotron 3 Ultra 550B
NVIDIA’s flagship open model optimized for enterprise reasoning workloads.
128K tokens · Text
reasoningMeta | Released: 4-25
Llama 4 Maverick
Meta’s multimodal Llama model with native image understanding.
1M tokens · Text + Vision
generalMeta | Released: 12-24
Llama 3.3 70B
Meta’s reliable workhorse — great quality-to-cost ratio for production.
128K tokens · Text
codingAlibaba | Released: 5-26
Qwen3-Coder-Next
Code-specialized model with strong fill-in-the-middle and agentic coding support.
256K tokens · Text
visionAlibaba | Released: 5-26
Qwen3-VL 8B
Compact vision-language model — fast and affordable for image understanding at scale.
128K tokens · Text + Vision
visionAlibaba | Released: 1-25
Qwen2.5-VL 72B
Proven vision-language model for OCR, document QA, and visual grounding.
128K tokens · Text + Vision
compactAlibaba | Released: 6-26
Qwen3.6 35B-A3B
Efficient MoE model with 3B active parameters — great price-to-performance.
128K tokens · Text
compactAlibaba | Released: 5-26
Qwen3.5 35B-A3B
Compact MoE model with strong reasoning at minimal cost per token.
128K tokens · Text
compactAlibaba | Released: 4-26
Qwen3-Next 80B
Balanced model with strong general capabilities at a competitive price point.
128K tokens · Text
compactMistral | Released: 3-26
Mistral Small 3.2 24B
Mistral’s efficient model — excellent for high-volume, latency-sensitive workloads.
128K tokens · Text
compactGoogle | Released: 3-26
Gemma 4 26B-A4B
Google’s efficient MoE Gemma model with 4B active parameters.
128K tokens · Text
compactGoogle | Released: 3-26
Gemma 4 31B
Google’s latest open model with strong general performance and broad language support.
128K tokens · Text
compactGoogle | Released: 3-25
Gemma 3 27B
Proven compact model with multimodal capabilities and strong multilingual support.
128K tokens · Text
compactOpenAI | Released: 8-25
gpt-oss-120b
OpenAI’s open-weight model with strong reasoning at a low cost.
128K tokens · Text
compactOpenAI | Released: 8-25
gpt-oss-120b (Fast)
Optimized fast tier of gpt-oss-120b with lower latency for real-time workloads.
128K tokens · Text
compactOpenAI | Released: 8-25
gpt-oss-20b
Compact open-weight model from OpenAI — ultra-cheap, fits on a single GPU.
128K tokens · Text
compactXiaomi | Released: 4-26
MiMo v2.5
Xiaomi’s efficient reasoning model with strong coding capabilities.
128K tokens · Text
compactSkyfall | Released: 2-26
Skyfall 31B v4.2
Fine-tuned model optimized for creative writing and conversational tasks.
128K tokens · Text
compactSkyfall | Released: 2-25
Skyfall 36B v2
Fine-tuned model with enhanced instruction-following and creative output.
128K tokens · Text
compactCydonia | Released: 8-25
Cydonia 24B v4.1
Fine-tuned model with enhanced roleplay and character consistency.
32K tokens · Text
embeddingsBAAI | Released: 1-24
BGE-M3
Multilingual, multi-functional embedding model supporting 100+ languages.
8K tokens · Embedding
audioResemble | Released: 4-26
Resemble TTS (English)
High-quality text-to-speech model with natural-sounding English voices.
— · Audio
agentsByteDance | Released: 4-25
UI-TARS 1.5 7B
GUI agent model that can operate computer interfaces autonomously.
128K tokens · Text + Vision
Find the right model for your workload
Start with free credits — no credit card. Call any model in minutes with an OpenAI-compatible API.