8.7.11. dasLLAMA-10 — Thinking Models
A thinking model works through the question before answering, and the reply’s
wire format carries that reasoning span: Qwen3 and GLM wrap it in
<think>...</think>, gpt-oss speaks Harmony channels, gemma-4 a thought
channel. This tutorial splits replies into reasoning and answer — buffered and
live-streamed — and then turns thinking off.
Run it with a thinking-family GGUF (Qwen3-0.6B-Q8_0 works):
daslang.exe -jit tutorials/dasLLAMA/10_thinking.das -- path/to/Qwen3-0.6B-Q8_0.gguf
8.7.11.1. The reply carries its reasoning
respond streams the raw pieces — reasoning markers included — while
chat.history stores the reply already reasoning-stripped (tutorial 02).
Thinking burns token budget, so raise max_new past its 256 default:
var chat = create_chat(m, "You are a concise assistant.", 1024l)
add_user(chat, "Which is larger: 17 * 24 or 20 * 20? One sentence.")
let full = respond(m, chat, SamplingParams()) $(_piece) => true // buffered: collect only
// full starts "<think> ..." — the raw wire; history stores only the answer
8.7.11.2. split_reasoning: the buffered split
split_reasoning splits a complete reply at the family’s reasoning
boundary. Both halves come back stripped when a reasoning span was found; a
reply with no reasoning passes through untouched — so it is safe to run on
every reply, thinking model or not.
let sp = split_reasoning(chat, full)
print("reasoning: {sp.reasoning}\n") // the model's work: 17*24 = 408, 20*20 = 400 ...
print("content: {sp.content}\n") // the answer: 17 * 24 is larger than 20 * 20.
8.7.11.3. Streaming the split: think_feed / think_finish
A live UI wants the reasoning marked while it streams, not after. The
ThinkStream splitter turns each streamed piece into (reasoning, content)
deltas, holding partial markers across piece boundaries. Create it for the
turn, feed every piece, and flush at the end — an unclosed reasoning span (the
token budget cut mid-thought) classifies as reasoning:
var chat2 = create_chat(m, "You are a concise assistant.", 1024l)
add_user(chat2, "What is the capital of Australia? One sentence.")
var ts = make_think_stream(chat2)
respond(m, chat2, SamplingParams()) $(piece) {
var r = ""
var c = ""
think_feed(ts, piece, r, c) // overwritten per piece; either may be empty
fprint(fstdout(), c) // show the answer live; dim or hide r as you like
fflush(fstdout())
return true
}
var tail_r = ""
var tail_c = ""
think_finish(ts, tail_r, tail_c)
On a model with no reasoning format the stream is a pass-through: every piece comes back as content. And a small model sometimes ends a later turn inside its think block — the flush then returns the tail as reasoning, and the content half stays empty; the tutorial file shows the pattern that handles both endings.
When the reply is already whole — a captured respond return, a server’s
non-streaming route — think_drain runs the same splitter in one call:
feed, finish, and the strip rule together:
var ts2 = make_think_stream(chat2)
let full = respond(m, chat2, SamplingParams()) $(piece) {
return true // capture only
}
let sp = think_drain(ts2, full) // ThinkSplit: .reasoning / .content
8.7.11.4. Turning thinking off
Hybrid families answer directly when the turn opens with the template’s empty
think block. set_thinking(false) renders exactly that (a no-op for models
with no think specials in the vocabulary); the default is on.
effective_stop_ids is the merged view of every stop in force — the
template’s stops plus the thinking-off extras while thinking is off:
set_thinking(chat, false)
let stops <- effective_stop_ids(chat)
add_user(chat, "And 12 * 12 vs 11 * 13? One sentence.")
// respond(...) now answers directly — no <think> span on the wire
A loop that samples the tokens itself splits that list in two.
turn_stop_ids are the template’s own stops and end the turn outright;
make_nothink_guard carries the thinking-off extras, and
nothink_stop_here decides each token: a channel marker before the reply’s
first content piece is a leading thought the reply matcher splits (gemma-4’s
E-series opens a media turn with one even in instruct mode), a marker after
content ends the turn:
let turn_stops <- turn_stop_ids(chat)
var guard <- make_nothink_guard(chat)
let prompt <- render_turn(m, chat)
generate(m, chat.session, prompt, SamplingParams(), 96l) $(id, piece) {
return false if (find_index(turn_stops, id) >= 0)
return false if (nothink_stop_here(guard, id, piece))
print("{piece}")
return true
}
8.7.11.5. A reasoning budget
A server caps how long a reply may think with the request’s thinking_budget:
once the reply has spent that many tokens inside its reasoning span, the
scheduler writes the family’s own close for it and the answer follows. The
scheduler reads tokens, not text, so the span’s bounds travel as token marks.
think_budget_marks builds them for the family the chat is on: the sequence a
reply writes to open the span (Qwen3’s <think>; gemma-4’s <|channel>
followed by the word thought; gpt-oss’s channel mark followed by
analysis), the one token the model writes to leave it, and the tokens the
budget forces at the cut - a symmetric family’s close special after the recipe’s
“Considering the limited time…” sentence, each followed by the template’s
blank line - and the markers that end the turn should the model re-open its span
after the cut. A family with no reasoning span returns empty marks, and a budget
is ignored on it:
var marks <- think_budget_marks(m, chat)
if (!empty(marks.close)) {
print("open: {marks.open} close token: {marks.close_tok} forced: {decode(m, marks.close)}\n")
}
See also
Full source: tutorials/dasLLAMA/10_thinking.das
Next tutorial: dasLLAMA-11 — Tool Calling
Chat basics and the transcript: dasLLAMA-02 — Chat and Templates