Skip to main content

Stepfun (Always-On Reasoning)

Uses SSE streaming with OpenAI-compatible tool-calling format. Supports vision (base64 images) and reasoning content. Stepfun’s Step-3 models reason on every request — there’s no toggle to disable thinking. Reasoning tokens count toward the completion budget, so the model balances thinking depth against output length automatically. This makes Stepfun a straightforward choice when you always want reasoning without managing mode settings. Best for: Reasoning-heavy tasks, workloads where you always want the model to think, and vision-capable workflows.

Getting an API Key

  1. Go to platform.stepfun.ai
  2. Sign up or log in
  3. Navigate to API Keys and create a new key
  4. Paste it into Wolffish → Settings → Models → Stepfun

Models

Step-3 models always reason — the enable_thinking parameter is accepted but ignored. Reasoning tokens count toward the completion token budget. All models support tool calling with no hard tool-count limit.
Stepfun’s context window is 128K — smaller than the 1M offered by DeepSeek, Qwen, or xAI. If your workflows regularly exceed 128K tokens of context, consider a provider with a larger window.

Reasoning modes

How this model reasons is set in the model card beside the chat input — hover the model switch to preview the card, click to pin it open. The top chip row, labelled Thinking, is the control: pick a chip and it applies from your next message. Two separate ideas combine here:

Thinking — whether the model reasons

  • Off — the model answers immediately. Fastest and cheapest; ideal for simple, direct tasks.
  • On (the chip reads Normal) — the model first works through the problem in a dedicated reasoning pass before replying. Slower and uses more tokens, but markedly more accurate on multi-step, logical, or ambiguous tasks.

Effort — how hard it thinks

Only effort-capable models expose this; it applies once thinking is on.
  • High — standard reasoning depth. The right default for most agentic work.
  • Max — the model reasons longer and deeper for the hardest problems. More tokens and latency in exchange for higher quality on complex work.

The chip row

The active chip is tinted; the rest sit quiet. Each model shows only the chips it genuinely supports — a model that always reasons has no Off, a model with a single mode shows that one chip inert, and a model that can’t reason at all replaces the row with “Reasoning is not supported by this model.” Wolffish remembers your choice per model. (Before v1.0.236 this was a colour-coded brain button on the composer; it moved into the model card, beside the Single/Workflow row, so every model knob sits in one panel.) On Stepfun: Step models always reason — the button stays locked on. There’s no off switch and no effort control.