Skip to main content

Z.ai (Zhipu GLM)

OpenAI-compatible SSE streaming with tool-calling. Reasoning streams as reasoning_content, thinking is controlled with thinking: { type } (plus reasoning_effort on GLM-5), and prompt-cache hits are reported under usage.prompt_tokens_details.cached_tokens — the same wire format as Kimi. GLM reasoning is per-model — GLM-5 models expose real effort tiers (Off / High / Max via reasoning_effort), while GLM-4.x are a simple on/off toggle (any mode other than none enables reasoning, none disables it). GLM-5.2 also offers a 1M-token context window, the largest in Z.ai’s lineup. Best for: Cost-efficient agentic work, long-context workflows (GLM-5.2), and vision-capable tasks. GLM models are strong all-rounders for tool chains and code at budget-tier pricing.

Getting an API Key

  1. Go to z.ai
  2. Sign up or log in
  3. Open API Keys and create a new key
  4. Paste it into Wolffish → Settings → Models → Z.ai

Models

Reasoning modes

How this model reasons is set in the model card beside the chat input — hover the model switch to preview the card, click to pin it open. The top chip row, labelled Thinking, is the control: pick a chip and it applies from your next message. Two separate ideas combine here:

Thinking — whether the model reasons

  • Off — the model answers immediately. Fastest and cheapest; ideal for simple, direct tasks.
  • On (the chip reads Normal) — the model first works through the problem in a dedicated reasoning pass before replying. Slower and uses more tokens, but markedly more accurate on multi-step, logical, or ambiguous tasks.

Effort — how hard it thinks

Only effort-capable models expose this; it applies once thinking is on.
  • High — standard reasoning depth. The right default for most agentic work.
  • Max — the model reasons longer and deeper for the hardest problems. More tokens and latency in exchange for higher quality on complex work.

The chip row

The active chip is tinted; the rest sit quiet. Each model shows only the chips it genuinely supports — a model that always reasons has no Off, a model with a single mode shows that one chip inert, and a model that can’t reason at all replaces the row with “Reasoning is not supported by this model.” Wolffish remembers your choice per model. (Before v1.0.236 this was a colour-coded brain button on the composer; it moved into the model card, beside the Single/Workflow row, so every model knob sits in one panel.) On Z.ai: GLM-5 models support Off / High / Max (genuine effort tiers). GLM-4.x are a simple On / Off toggle with no effort control.