## Specs & substance

### Spec sheet

- **Developed by:** Google
- **Model family:** Gemma
- **Category:** compact
- **Modality:** Text
- **Context window:** 128K tokens
- **Architecture:** MoE 4B active
- **Version:** 4
- **License:** Gemma License
- **Pricing:** $0.13 in $0.40 out · $0.05 cache
- **Released:** Mar 2026
- **Endpoint:** parasail-gemma-4-26b-a4b-it

Google’s efficient MoE Gemma model with 4B active parameters.

Gemma 4 26B-A4B is part of the Gemma family by Google, categorized as a **compact** model. It supports a 128K tokens context window and is built on a MoE 4B active architecture.

Key strengths: efficient, fast. Parasail serves Gemma 4 26B-A4B on a global fleet of current-gen GPUs behind a single OpenAI-compatible endpoint — with per-token pricing, no minimums, and dedicated capacity options when you need guaranteed throughput.

## Integrate

### Drop-in via the OpenAI SDK

Point any OpenAI-compatible chat client at Parasail and change the model name. That's it.

#### Python

```python
from openai import OpenAI

client = OpenAI(
    api_key="YOUR_API_KEY",
    base_url="https://api.parasail.io/v1"
)

response = client.chat.completions.create(
    model="parasail-gemma-4-26b-a4b-it",
    messages=[
        {"role": "user", "content": "Hello, what can you do?"}
    ],
    stream=True,
    stream_options={"include_usage": True},
    top_p=1,
    max_tokens=1000,
    temperature=1
)

for chunk in response:
    if chunk.choices and chunk.choices[0].delta.content is not None:
        print(chunk.choices[0].delta.content, end="", flush=True)
```

#### Bash

```bash
curl https://api.parasail.io/v1/chat/completions \
  -H "Authorization: Bearer $PARASAIL_API_KEY" \ 
  -H "Content-Type: application/json" \ 
  -d '{
    "model": "parasail-gemma-4-26b-a4b-it",
    "messages": [{"role": "user", "content": "Hello, what can you do?"}],
    "stream": true,
    "max_tokens": 1000
  }'
```

## Explore the library

- **Model DeepSeek V4 Flash**  
  Ultra-fast, ultra-cheap reasoning model for high-throughput workloads.  
  **$0.14/M in**
- **Model Qwen3.6 35B-A3B**  
  Efficient MoE model with 3B active parameters — great price-to-performance.  
  **$0.15/M in**
- **Model Qwen3.5 35B-A3B**  
  Compact MoE model with strong reasoning at minimal cost per token.  
  **$0.15/M in**
