Skip to main content

Qwen (Alibaba Cloud)

Uses SSE streaming with OpenAI-compatible tool-calling format. Supports vision (base64 images) and reasoning content. Qwen offers one of the widest model ranges of any provider — from the ultra-cheap Qwen 3.5 Flash at 0.06/0.06/0.24 per MTok to the frontier Qwen 3.8 Max. All Qwen3+ models support three reasoning modes (None, High, Max) and up to 1M context. The dedicated Qwen3 Coder Plus model is tuned for code generation tasks. Since v1.0.235, connecting Qwen auto-selects Qwen 3.8 Max — the new flagship, and the first Max tier to add vision, at a lower price than the 3.7 Max it replaces. Best for: Cost-efficient agentic workflows, code generation, multilingual tasks, and workloads that benefit from a wide selection of price/performance tiers.
Qwen 3.5 Flash is one of the cheapest reasoning-capable models available — at 0.06/0.06/0.24 per MTok with 1M context, it’s significantly cheaper than DeepSeek V4 Flash while still supporting full reasoning modes. Great for high-volume tasks where cost matters.

Getting an API Key

  1. Go to qwencloud.com
  2. Sign up or log in
  3. Navigate to API Keys and create a new key
  4. Paste it into Wolffish → Settings → Models → Qwen

Models

Reasoning modes

How this model reasons is set in the model card beside the chat input — hover the model switch to preview the card, click to pin it open. The top chip row, labelled Thinking, is the control: pick a chip and it applies from your next message. Two separate ideas combine here:

Thinking — whether the model reasons

  • Off — the model answers immediately. Fastest and cheapest; ideal for simple, direct tasks.
  • On (the chip reads Normal) — the model first works through the problem in a dedicated reasoning pass before replying. Slower and uses more tokens, but markedly more accurate on multi-step, logical, or ambiguous tasks.

Effort — how hard it thinks

Only effort-capable models expose this; it applies once thinking is on.
  • High — standard reasoning depth. The right default for most agentic work.
  • Max — the model reasons longer and deeper for the hardest problems. More tokens and latency in exchange for higher quality on complex work.

The chip row

The active chip is tinted; the rest sit quiet. Each model shows only the chips it genuinely supports — a model that always reasons has no Off, a model with a single mode shows that one chip inert, and a model that can’t reason at all replaces the row with “Reasoning is not supported by this model.” Wolffish remembers your choice per model. (Before v1.0.236 this was a colour-coded brain button on the composer; it moved into the model card, beside the Single/Workflow row, so every model knob sits in one panel.) On Qwen: qwen3.x models support Off / High / Max (effort via a thinking-token budget). qwq and qvq reason always-on (locked on). Legacy qwen-max/plus/turbo/flash don’t reason.

Cost Comparison

Qwen spans a wide price range, competing at every tier:
Start with Qwen 3.5 Flash for high-volume tasks, Qwen 3.7 Plus for general agentic work, or Qwen 3.8 Max when you need frontier reasoning or vision. The dedicated Qwen3 Coder Plus model is a good pick for code-heavy workflows at a budget price.