## Featured models

### reasoningZ.AI | Released: 6-26  
**GLM-5.2**  
Z.AI’s next-gen model for agentic engineering with sustained execution on long-horizon tasks.  
1M tokens · Text

### reasoningZ.AI | Released: 4-26  
**GLM-5.1**  
Strong reasoning and coding model with a 1M context window at competitive pricing.  
1M tokens · Text

### reasoningMiniMax | Released: 5-26  
**MiniMax M3**  
Highly efficient MoE model delivering strong reasoning at exceptionally low cost.  
1M tokens · Text

### reasoningDeepSeek | Released: 5-26  
**DeepSeek V4 Pro**  
Open model for reasoning and coding workloads.  
128K tokens · Text

### codingMoonshot | Released: 5-26  
**Kimi K2.7 Code**  
Specialized coding model optimized for software engineering and agentic coding workflows.  
256K tokens · Text

### visionAlibaba | Released: 5-26  
**Qwen3-VL 235B-A22B**  
Large vision-language model with strong document understanding and visual reasoning.  
256K tokens · Text + Vision

## Recently added

### compactAlibaba | Released: 6-26  
**Qwen3.6 35B-A3B**  
Efficient MoE model with 3B active parameters — great price-to-performance.  
128K tokens · Text

### reasoningNVIDIA | Released: 6-26  
**Nemotron 3 Ultra 550B**  
NVIDIA’s flagship open model optimized for enterprise reasoning workloads.  
128K tokens · Text

### reasoningZ.AI | Released: 6-26  
**GLM-5.2**  
Z.AI’s next-gen model for agentic engineering with sustained execution on long-horizon tasks.  
1M tokens · Text

## Explore the library

### reasoningZ.AI | Released: 6-26  
**GLM-5.2**  
Z.AI’s next-gen model for agentic engineering with sustained execution on long-horizon tasks.  
1M tokens · Text

### reasoningZ.AI | Released: 4-26  
**GLM-5.1**  
Strong reasoning and coding model with a 1M context window at competitive pricing.  
1M tokens · Text

### reasoningZ.AI | Released: 2-26  
**GLM-5**  
Z.AI’s open model for reasoning and agentic workflows.  
128K tokens · Text

### reasoningMiniMax | Released: 5-26  
**MiniMax M3**  
Highly efficient MoE model delivering strong reasoning at exceptionally low cost.  
1M tokens · Text

### reasoningMiniMax | Released: 3-26  
**MiniMax M2.5**  
Cost-effective large model with strong general capabilities and long context.  
1M tokens · Text

### reasoningDeepSeek | Released: 5-26  
**DeepSeek V4 Pro**  
Open model for reasoning and coding workloads.  
128K tokens · Text

### compactDeepSeek | Released: 5-26  
**DeepSeek V4 Flash**  
Ultra-fast, ultra-cheap reasoning model for high-throughput workloads.  
128K tokens · Text

### reasoningMoonshot | Released: 4-26  
**Kimi K2.6**  
Strong reasoning model with excellent long-context understanding.  
256K tokens · Text

### reasoningAlibaba | Released: 5-26  
**Qwen3.5 397B-A17B**  
Large-scale MoE model with frontier reasoning and multilingual capabilities.  
256K tokens · Text

### reasoningAlibaba | Released: 7-25  
**Qwen3 235B-A22B**  
Strong open MoE model with excellent coding and reasoning at low cost.  
256K tokens · Text

### reasoningTrinity | Released: 4-26  
**Trinity Large (Thinking)**  
Reasoning-focused model with extended thinking for complex problem-solving.  
128K tokens · Text

### reasoningNVIDIA | Released: 6-26  
**Nemotron 3 Ultra 550B**  
NVIDIA’s flagship open model optimized for enterprise reasoning workloads.  
128K tokens · Text

### reasoningMeta | Released: 4-25  
**Llama 4 Maverick**  
Meta’s multimodal Llama model with native image understanding.  
1M tokens · Text + Vision

### generalMeta | Released: 12-24  
**Llama 3.3 70B**  
Meta’s reliable workhorse — great quality-to-cost ratio for production.  
128K tokens · Text

### codingAlibaba | Released: 5-26  
**Qwen3-Coder-Next**  
Code-specialized model with strong fill-in-the-middle and agentic coding support.  
256K tokens · Text

### visionAlibaba | Released: 5-26  
**Qwen3-VL 8B**  
Compact vision-language model — fast and affordable for image understanding at scale.  
128K tokens · Text + Vision

### visionAlibaba | Released: 1-25  
**Qwen2.5-VL 72B**  
Proven vision-language model for OCR, document QA, and visual grounding.  
128K tokens · Text + Vision

### compactAlibaba | Released: 6-26  
**Qwen3.6 35B-A3B**  
Efficient MoE model with 3B active parameters — great price-to-performance.  
128K tokens · Text

### compactAlibaba | Released: 5-26  
**Qwen3.5 35B-A3B**  
Compact MoE model with strong reasoning at minimal cost per token.  
128K tokens · Text

### compactAlibaba | Released: 4-26  
**Qwen3-Next 80B**  
Balanced model with strong general capabilities at a competitive price point.  
128K tokens · Text

### compactMistral | Released: 3-26  
**Mistral Small 3.2 24B**  
Mistral’s efficient model — excellent for high-volume, latency-sensitive workloads.  
128K tokens · Text

### compactGoogle | Released: 3-26  
**Gemma 4 26B-A4B**  
Google’s efficient MoE Gemma model with 4B active parameters.  
128K tokens · Text

### compactGoogle | Released: 3-26  
**Gemma 4 31B**  
Google’s latest open model with strong general performance and broad language support.  
128K tokens · Text

### compactGoogle | Released: 3-25  
**Gemma 3 27B**  
Proven compact model with multimodal capabilities and strong multilingual support.  
128K tokens · Text

### compactOpenAI | Released: 8-25  
**gpt-oss-120b**  
OpenAI’s open-weight model with strong reasoning at a low cost.  
128K tokens · Text

### compactOpenAI | Released: 8-25  
**gpt-oss-120b (Fast)**  
Optimized fast tier of gpt-oss-120b with lower latency for real-time workloads.  
128K tokens · Text

### compactOpenAI | Released: 8-25  
**gpt-oss-20b**  
Compact open-weight model from OpenAI — ultra-cheap, fits on a single GPU.  
128K tokens · Text

### compactXiaomi | Released: 4-26  
**MiMo v2.5**  
Xiaomi’s efficient reasoning model with strong coding capabilities.  
128K tokens · Text

### compactSkyfall | Released: 2-26  
**Skyfall 31B v4.2**  
Fine-tuned model optimized for creative writing and conversational tasks.  
128K tokens · Text

### compactSkyfall | Released: 2-25  
**Skyfall 36B v2**  
Fine-tuned model with enhanced instruction-following and creative output.  
128K tokens · Text

### compactCydonia | Released: 8-25  
**Cydonia 24B v4.1**  
Fine-tuned model with enhanced roleplay and character consistency.  
32K tokens · Text

### embeddingsBAAI | Released: 1-24  
**BGE-M3**  
Multilingual, multi-functional embedding model supporting 100+ languages.  
8K tokens · Embedding

### audioResemble | Released: 4-26  
**Resemble TTS (English)**  
High-quality text-to-speech model with natural-sounding English voices.  
— · Audio

### agentsByteDance | Released: 4-25  
**UI-TARS 1.5 7B**  
GUI agent model that can operate computer interfaces autonomously.  
128K tokens · Text + Vision

## Find the right model for your workload

Start with free credits — no credit card. Call any model in minutes with an OpenAI-compatible API.
