## Specs & substance

### Spec sheet

- **Developed by**: Z.AI  
- **Model family**: GLM  
- **Use case**: Large language  
- **Modality**: LLM  
- **Context window**: 1M tokens  
- **Architecture**: MoE + DSA  
- **Version**: 5.2  
- **License**: MIT  
- **Pricing**: $1.40 in $4.40 out · $0.26 cache  
- **Released**: Jun 2026

GLM-5.2 is Z.AI’s next-generation model for agentic engineering, designed to sustain autonomous execution across hours-long tasks. Its MoE + DSA architecture supports a 1M context window for long-horizon coding workflows that require continuous iteration at production scale.

Parasail serves GLM-5.2 on a global fleet of current-gen GPUs behind a single OpenAI-compatible endpoint — with per-token pricing, no minimums, and dedicated capacity options when you need guaranteed throughput.

## Integrate

### Drop-in via the OpenAI SDK

Point any OpenAI-compatible client at Parasail and change the model name. That's it.

**Python**:  
```python
from openai import OpenAI

client = OpenAI(
    api_key="YOUR_API_KEY",
    base_url="https://api.parasail.io/v1"
)

response = client.chat.completions.create(
    model="zai-org/GLM-5.2",
    messages=[\
        {"role": "user", "content": "Implement Hello World in Python"}\
    ],
    stream=True,
    stream_options={"include_usage": True},
    top_p=1,
    max_tokens=1000,
    temperature=1
)

for chunk in response:
    if chunk.choices and chunk.choices[0].delta.content is not None:
        print(chunk.choices[0].delta.content, end="", flush=True)
```

**bash**:  
```bash
curl https://api.parasail.io/v1/chat/completions \
  -H "Authorization: Bearer $PARASAIL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "zai-org/GLM-5.2",
    "messages": [{"role": "user", "content": "Implement Hello World in Python"}],
    "stream": true,
    "max_tokens": 1000
  }'
```

**json**:  
```json
{
  "id": "chatcmpl-143",
  "object": "chat.completion",
  "created": 1741224586,
  "model": "zai-org/GLM-5.2",
  "choices": [\
    {\
      "index": 0,\
      "finish_reason": "stop",\
      "message": {\
        "role": "assistant",\
        "content": "print(\"Hello, World!\")"\
      }\
    }\
  ],
  "usage": {
    "prompt_tokens": 38,
    "completion_tokens": 12,
    "total_tokens": 50
  }
}
```
