Parameters | Parasail

For the complete documentation index, see llms.txt. This page is also available as Markdown.

Parasail supports all parameters that vLLM supports. These parameters control the randomness, diversity, and length of model outputs.

Sampling parameters

temperature (default: 1.0)

Controls randomness in token selection.

top_p (Nucleus Sampling) (default: 1.0)

Controls the probability mass of token selection.

top_k (default: -1, disabled)

Limits token selection to the top k most probable tokens.

max_tokens (default: None)

Sets the maximum number of tokens to generate. Helps prevent excessively long responses.

repetition_penalty (default: 1.0)

Penalizes repeated tokens to avoid looping responses. Common values: 1.1 to 1.2.

presence_penalty (default: 0.0)

Increases the likelihood of introducing new tokens. Useful for making outputs more diverse.

frequency_penalty (default: 0.0)

Penalizes tokens that have appeared frequently. Helps prevent excessive repetition of common words.

seed (default: None)

Sets a fixed seed for reproducible results. Useful for debugging or deterministic sampling.

How parameters work together

Example

from openai import OpenAI

client = OpenAI(
    base_url="https://api.parasail.io/v1",
    api_key="<PARASAIL_API_KEY>"
)

response = client.chat.completions.create(
    model="parasail-deepseek-r1",
    messages=[{"role": "user", "content": "Write a creative story about a robot."}],
    max_completion_tokens=1000,
    temperature=0.7,
    top_p=0.1,
    repetition_penalty=1.1,
    extra_body={"top_k": 50}
)

print(response.choices[0].message.content)

Next steps