Qwen3-VL 8B — Alibaba | Parasail
Specs & substance
Spec sheet
- Developed by: Alibaba
- Model family: Qwen
- Category: vision
- Modality: Text + Vision
- Context window: 128K tokens
- Architecture: Dense VL
- Version: VL 8B
- License: Apache 2.0
- Pricing: $0.25 in $0.75 out · $0.12 cache
- Released: May 2026
- Endpoint: parasail-qwen3vl-8b-instruct
Compact vision-language model — fast and affordable for image understanding at scale.
Qwen3-VL 8B is part of the Qwen family by Alibaba, categorized as a vision model. It supports a 128K tokens context window and is built on a Dense VL architecture.
Key strengths: vision, efficient. Parasail serves Qwen3-VL 8B on a global fleet of current-gen GPUs behind a single OpenAI-compatible endpoint — with per-token pricing, no minimums, and dedicated capacity options when you need guaranteed throughput.
Drop-in via the OpenAI SDK
Point any OpenAI-compatible chat client at Parasail and change the model name. That's it.
Python
from openai import OpenAI
client = OpenAI(
api_key="YOUR_API_KEY",
base_url="https://api.parasail.io/v1"
)
response = client.chat.completions.create(
model="parasail-qwen3vl-8b-instruct",
messages=[
{"role": "user", "content": "Hello, what can you do?"}
],
stream=True,
stream_options={"include_usage": True},
top_p=1,
max_tokens=1000,
temperature=1
)
for chunk in response:
if chunk.choices and chunk.choices[0].delta.content is not None:
print(chunk.choices[0].delta.content, end="", flush=True)
Bash
curl https://api.parasail.io/v1/chat/completions \
-H "Authorization: Bearer $PARASAIL_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "parasail-qwen3vl-8b-instruct",
"messages": [{"role": "user", "content": "Hello, what can you do?"}],
"stream": true,
"max_tokens": 1000
}'
Explore the library
Model Qwen3.5 397B-A17B
Large-scale MoE model with frontier reasoning and multilingual capabilities.
$0.50/M inModel Qwen3 235B-A22B
Strong open MoE model with excellent coding and reasoning at low cost.
$0.14/M inModel Qwen3-Coder-Next
Code-specialized model with strong fill-in-the-middle and agentic coding support.
$0.12/M in
Run Qwen3-VL 8B in the platform
Start with free credits — no credit card. Call any model in minutes.