Trinity Large (Thinking) — Trinity | Parasail

Specs & substance

Spec sheet

Reasoning-focused model with extended thinking for complex problem-solving.

Trinity Large (Thinking) is part of the Trinity family by Trinity, categorized as a flagship model. It supports a 128K tokens context window and is built on a Dense architecture.

Key strengths: reasoning. Parasail serves Trinity Large (Thinking) on a global fleet of current-gen GPUs behind a single OpenAI-compatible endpoint — with per-token pricing, no minimums, and dedicated capacity options when you need guaranteed throughput.

Drop-in via the OpenAI SDK

Point any OpenAI-compatible chat client at Parasail and change the model name. That's it.

Python

from openai import OpenAI

client = OpenAI(
    api_key="YOUR_API_KEY",
    base_url="https://api.parasail.io/v1"
)

response = client.chat.completions.create(
    model="parasail-trinity-large-thinking",
    messages=[\
        {"role": "user", "content": "Hello, what can you do?"}\
    ],
    stream=True,
    stream_options={"include_usage": True},
    top_p=1,
    max_tokens=1000,
    temperature=1
)

for chunk in response:
    if chunk.choices and chunk.choices[0].delta.content is not None:
        print(chunk.choices[0].delta.content, end="", flush=True)

Bash

curl https://api.parasail.io/v1/chat/completions \
  -H "Authorization: Bearer $PARASAIL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{\
    "model": "parasail-trinity-large-thinking",\
    "messages": [{"role": "user", "content": "Hello, what can you do?"}],\
    "stream": true,\
    "max_tokens": 1000\
  }'