## Specs & substance

### Spec sheet

Developed by NVIDIA  
Model family Nemotron  
Category flagship  
Modality Text  
Context window 128K tokens  
Architecture MoE NVFP4  
Version 3 Ultra  
License NVIDIA Open Model License  
Pricing $0.50 in $2.50 out · $0.10 cache  
Released Jun 2026  
Endpoint parasail-nemotron-3-ultra-550b-nvfp4

NVIDIA’s flagship open model optimized for enterprise reasoning workloads.

Nemotron 3 Ultra 550B is part of the Nemotron family by NVIDIA, categorized as a **flagship** model. It supports a 128K tokens context window and is built on a MoE NVFP4 architecture.

Key strengths: reasoning, enterprise. Parasail serves Nemotron 3 Ultra 550B on a global fleet of current-gen GPUs behind a single OpenAI-compatible endpoint — with per-token pricing, no minimums, and dedicated capacity options when you need guaranteed throughput.

## Drop-in via the OpenAI SDK

Point any OpenAI-compatible chat client at Parasail and change the model name. That's it.

### Python

```python
from openai import OpenAI

client = OpenAI(
    api_key="YOUR_API_KEY",
    base_url="https://api.parasail.io/v1"
)

response = client.chat.completions.create(
    model="parasail-nemotron-3-ultra-550b-nvfp4",
    messages=[
        {"role": "user", "content": "Hello, what can you do?"}
    ],
    stream=True,
    stream_options={"include_usage": True},
    top_p=1,
    max_tokens=1000,
    temperature=1
)

for chunk in response:
    if chunk.choices and chunk.choices[0].delta.content is not None:
        print(chunk.choices[0].delta.content, end="", flush=True)
```

### Bash

```bash
curl https://api.parasail.io/v1/chat/completions \
  -H "Authorization: Bearer $PARASAIL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "parasail-nemotron-3-ultra-550b-nvfp4",
    "messages": [{"role": "user", "content": "Hello, what can you do?"}],
    "stream": true,
    "max_tokens": 1000
  }'
```

## Explore the library

### Model **GLM-5.2** 
Z.AI’s next-gen model for agentic engineering with sustained execution on long-horizon tasks.  
**$1.40/M in**  
### Model **GLM-5.1** 
Strong reasoning and coding model with a 1M context window at competitive pricing.  
**$1.40/M in**  
### Model **GLM-5** 
Z.AI’s open model for reasoning and agentic workflows.  
**$1.00/M in**

## Run Nemotron 3 Ultra 550B in the platform

Start with free credits — no credit card. Call any model in minutes.
