Wolffish Sleeps On It
Memory records what happened. The feedback loop records what worked at the tool level. Reflection closes the last gap: every night, Wolffish reviews its own finished conversations the way a careful colleague reviews their day — what was attempted, how it actually ended, what worked well enough to repeat, what failed and why, and what you showed it about how you like things done. The lessons distill into a playbook that rides into every future conversation. The strongest signal in that review is yours — not a grade, but your own words in the transcript: what you appreciated, what you corrected, what you asked it never to do again.Location
reviewed.json entries are pruned after 14 days.
The Nightly Review
The nightly job looks back over the last 7 days and reviews every conversation — automations and procedure runs included — that has been quiet for at least the configured window (default 12 hours). A conversation that gets continued after its review is simply re-reviewed the next night. Each conversation gets one review call on your configured Brain (reasoning off, billed like the other utility side-calls, never against the conversation’s own meter). The result is an honest block appended to the day file:score: line is the model’s own honest self-score — held to what actually happened rather than effort, with your own words in the transcript as the strongest evidence: a turn that annoyed you, wasted work, or failed silently scores low even if every tool “succeeded”. It stays internal to reflection, never shown in chat.
The Playbook
The same nightly run folds new review blocks intoplaybook.md — a living document rewritten whole, never an append-only log. It has exactly five sections:
user-said (you stated it) or inferred (the review concluded it). The merge rules keep it trustworthy:
- Newest evidence wins a contradiction — a re-learned preference replaces the stale one instead of sitting beside it.
- Inferred entries decay after 30 days unless reinforced; what you said yourself doesn’t fade.
- Hard ceiling of 6,000 characters — the playbook must stay a briefing, not an archive.
- Every rewrite keeps the previous version as
playbook.md.bak, so a bad night is one copy away from undone.
Your Words Are Its Ground Truth
Earlier versions asked you to grade every answer on a 0–10 rating bar — in the app, on the phone, or by texting a bare number to Telegram or WhatsApp. Since v1.0.260 that scoring is retired everywhere: the bar, the channel invitations, the bare-number votes, and the terminal’s/rate are all gone, along with their switches. Grading every answer was homework, and the numbers told the review less than what you already say naturally.
What reflection reads instead is the conversation itself — the strongest evidence there is:
- Appreciation — “perfect, exactly that” marks a win worth repeating, along with what triggered it.
- Corrections — “no, round to one decimal” becomes a durable preference the moment you say it.
- Frustration — a redo, a complaint, an abandoned thread scores a conversation honestly without you filing anything.
The Monthly Deep Reflection
Nightly accumulation has a failure mode: self-taught bad habits. Once a month (the 1st, one hour after the nightly run), a deep reflection takes the opposite stance and attacks what the nights built up — stale rules, sweeping conclusions drawn from one incident, entries the evidence has turned against. It audits the playbook and the five knowledge files, reading up to a month of reflection day files as evidence, and rewrites only what changed — knowledge rewrites keep a.bak beside the original, same as the playbook.
It skips itself until reflection has actually produced something to audit.
Schedule and Settings
Everything lives in Settings → Knowledge → Reflection (the Knowledge page now has two tabs: Compaction and Reflection):- Nightly reflection — the hour it runs (default 3 AM).
- Review after quiet for — how long a conversation must sit untouched before it counts as settled: 1, 2, 3, 6, 12 (default), 24, or 48 hours.
- Run now buttons for both the nightly review and the deep reflection, plus last-run cards showing when each ran, how long it took, the tokens in and out, and what it produced.
config.json:
Reflection is core — there is no off switch, only the hour. Like compaction, it runs through the brainstem’s scheduler: a laptop asleep at 3 AM simply runs its review on the next launch (missed fires are swept 90 seconds after startup and every 30 minutes after that). Both runs also appear on the Automations page as “Nightly reflection” and “Deep reflection”.