SPEC-0016: AI-editorialized journal
- Capability: journal
- Source packages:
internal/journal(journal.go),internal/store(journal.go,journal_calendar.go,schema.gov11),internal/config(config.goDefaultDigestPrompt),internal/cli(journal.go),internal/web(journal.go,templates/journal.html,templates/home.html) - Related ADRs: 📝 ADR-0023, 📝 ADR-0011, 📝 ADR-0010
- Related specs: SPEC-0005 — the sibling extraction feature
Overview
msgbrowse builds a day-by-day journal over the local archive in two layers. The
mechanical layer is a deterministic per-day rollup (counts and top senders)
derived entirely from messages with no LLM and no network egress; it is always
built. The digest layer is one prose summary per day written by the
configured chat model, cached in SQLite and keyed by (day, model, prompt_version); it is the only network egress and is gated on
journal.digest_enabled. Both layers bucket messages by UTC calendar day,
honor journal.exclude_conversations before any content is assembled, and are
re-ingest-safe (day-keyed, no foreign key to messages).
Phase 2 upgraded the digest from a single prose paragraph to a structured
editorial object — the LLM now returns a JSON document (summary, people,
themes, mood, timed highlights, standout media, notable links) parsed tolerantly
and cached alongside a plain-text summary fallback, with mood denormalized into
its own column. The read surface changed to match: the flat, keyset-paginated day
list (REQ-0016-010) is superseded by a mood-tinted month-calendar navigator
with an editorial day card and headline stat tiles (REQ-0016-014). Phase 2
requirements (REQ-0016-011 through REQ-0016-016) add to the Phase 1 requirements
below; none of the two-layer, UTC-bucketing, cache-key, or privacy invariants
(REQ-0016-001 through REQ-0016-009) changed.
Requirements
REQ-0016-001: Mechanical day journal, deterministic and egress-free
The mechanical journal MUST be a per-day rollup — message count, conversation
count, per-source counts, and top senders — derived solely from the local
messages table, with NO LLM call and NO network egress. It MUST be built on
every msgbrowse journal run regardless of journal.digest_enabled. Rebuilding
it MUST be idempotent: a re-run over unchanged messages MUST produce the same
journal_days rows (an upsert keyed on day), never duplicates.
Scenario: Mechanical build never egresses
- Given an imported archive and
journal.digest_enabled = false - When the user runs
msgbrowse journal - Then per-day
journal_daysrows are written frommessagesalone, no call is made tollm.base_url, and re-running produces identical rows.
REQ-0016-002: UTC day bucketing
A day MUST be derived as substr(ts,1,10) (equivalently
date(ts_unix,'unixepoch')). Because messages.ts_unix is the wall-clock string
parsed as UTC, day bucketing MUST NOT apply any timezone conversion; using
'localtime' (which would double-shift the already-UTC value and misfile
messages across midnight) is prohibited. The mechanical rollup and the transcript
assembled for a digest MUST use the same UTC day definition.
Scenario: A late-night message stays on its own UTC day
- Given a message with
ts = '2026-07-01 23:30:00'(ts_unix= that instant as UTC) - When the journal is built
- Then it is counted under day
2026-07-01, and no'localtime'conversion shifts it into an adjacent day.
REQ-0016-003: Opt-in digest as the single egress
The LLM digest MUST be produced only when journal.digest_enabled is true, and
each digest MUST be the only network egress — one llm.Chat call per day to
llm.base_url using llm.chat_model. Import and serve MUST NOT trigger digest
generation. A per-day digest call MUST use a longer per-call context timeout
(~180s) than the facts default, because a full day's transcript can be large.
Scenario: Digest is explicit and gated
- Given
journal.digest_enabled = trueand days with no cached digest - When the user runs
msgbrowse journal - Then each eligible day is summarized via
llm.base_url; runningsignal-importorservealone never calls the LLM, and settingdigest_enabled = falseproduces the mechanical journal with zero LLM calls.
REQ-0016-004: Digest cache keyed by day, model, and prompt version
Each digest MUST be cached in journal_digests (PRIMARY KEY day) alongside the
model and prompt_version that produced it, where prompt_version = sha256hex(lower(trim(effective DigestPrompt))) (the same normalization recipe as
factHash). A day whose cached (model, prompt_version) matches the current
effective values MUST be skipped with no LLM call. Editing journal.digest_prompt
or switching llm.chat_model MUST make previously-digested days eligible again on
the next run. There MUST be no foreign key from journal_digests (or
journal_days) to messages.
Scenario: Prompt change invalidates the cache
- Given a day already digested under prompt version
P1and modelM - When
journal.digest_promptis edited (yielding versionP2) andmsgbrowse journalruns again - Then that day is re-digested and its row is updated to
(M, P2); with the prompt and model unchanged instead, the day is skipped with no LLM call.
REQ-0016-005: One digest per day
The digest MUST be a single cross-thread summary per calendar day (matching
config.DefaultDigestPrompt, "summarizing one day"), keyed on day alone.
Per-conversation digests are OUT OF SCOPE for this capability.
Scenario: Multi-thread day yields one digest
- Given a day with messages across several conversations
- When the digest for that day is generated
- Then exactly one
journal_digestsrow is written for that day, summarizing the day as a whole.
REQ-0016-006: Honor the exclude list before assembly
journal.exclude_conversations MUST be applied during day enumeration and
transcript assembly, before any message content is gathered, so an excluded
conversation's content never reaches the assembled transcript or the LLM. This
MUST NOT affect a day's inclusion when other, non-excluded conversations have
messages that day.
Scenario: Excluded thread is never sent
- Given a conversation named in
journal.exclude_conversationswith messages on a given day - When that day's digest is generated
- Then the excluded conversation's content is absent from the transcript sent to
llm.base_url, while non-excluded threads from the same day are summarized normally.
REQ-0016-007: CLI flags and run semantics
msgbrowse journal MUST expose --since YYYY-MM-DD, --backfill,
--regenerate, and --dry-run. A default run MUST be incremental: it
digests every day whose cached digest is absent or stale by (model, prompt_version). --backfill MUST apply the same eligibility across all history
with no day cap. --regenerate MUST wipe all cached digests and rebuild them.
--since MUST set a day floor (process only days on or after the given date).
Scenario: Incremental default skips current days
- Given some days already digested under the current model and prompt and some not
- When
msgbrowse journalruns with no flags - Then only days lacking a current digest are sent to the LLM; already-current days are skipped.
Scenario: Regenerate rebuilds everything
- Given a fully-digested history
- When
msgbrowse journal --regenerateruns - Then all cached digests are wiped and re-generated for every eligible day.
REQ-0016-008: Per-run day cap with remaining count
An incremental run MUST cap the number of days it digests at
journal.max_days_per_run (0 = unbounded) so a scheduled run has a bounded cost,
and MUST report how many eligible days remain after the cap so the operator can
schedule follow-up runs. --backfill MUST ignore the cap.
Scenario: Cap bounds a cron run and reports the remainder
- Given 50 days eligible for a digest and
journal.max_days_per_run = 20 - When
msgbrowse journalruns - Then 20 days are digested and the run reports 30 days remaining; a subsequent run digests the next batch.
REQ-0016-009: Dry-run makes zero LLM calls
--dry-run MUST make NO llm.Chat calls. It MUST enumerate the days that lack a
current digest and estimate input tokens with a local len(runes)/4 heuristic.
It MUST NOT print a dollar/cost figure — the codebase has no tokenizer or price
table and does not decode the provider's usage object, so a real cost estimate
is explicitly future work.
Scenario: Dry-run reports days and a token estimate only
- Given several days lacking a current digest
- When
msgbrowse journal --dry-runruns - Then it prints the eligible-day count and an approximate input-token total (
runes/4), makes no network call, and prints no dollar amount.
REQ-0016-010: Journal web page and home quick-link
Phase 2 (superseded presentation): the flat, keyset-paginated day list this requirement describes has been replaced by the mood-tinted month-calendar navigator specified in REQ-0016-014. The load-bearing invariants below carry forward unchanged — a day renders its cached digest when present and falls back to the mechanical rollup otherwise, an empty archive renders an empty state, the home page offers a fourth quick-link tile to
/journal, and rendering the page issues no LLM call. Only the day-navigation surface changed (keyset list → calendar grid + editorial day card); the scenarios below record the Phase 1 behavior.
GET /journal MUST list days newest-first, showing each day's digest when one is
cached and falling back to the mechanical rollup when it is not, with keyset
pagination over days and an empty state when no days exist. The home page MUST
offer a fourth quick-link icon tile linking to /journal. Rendering the page
MUST NOT trigger any LLM call.
Scenario: Days render newest-first with fallback
- Given a day with a cached digest and an earlier day with only a mechanical rollup
- When the user opens
/journal - Then the digested day appears first showing its prose, the earlier day shows its counts and top senders, older days page in via keyset pagination, and no LLM call is made.
Scenario: Empty state
- Given an archive with no imported messages
- When the user opens
/journal - Then the page renders an empty state rather than an error or a blank list.
REQ-0016-011: Structured editorial digest with tolerant parsing
The digest MUST be produced as a single JSON object with the keys summary
(string), people (string array), themes (string array), mood (enum string),
highlights (array of {text, time} objects), standout_media (string array),
and notable_links (string array), matching config.DefaultDigestPrompt. The
response MUST be parsed tolerantly: the object is extracted from the first
{ to the last } (tolerating markdown fences or surrounding prose), and every
field is coerced rather than rejected — each list item is trimmed and empty items
are dropped, an unknown mood is coerced to "neutral", and a malformed HH:MM
highlight time is blanked ("") while its text is kept. A response that
carries no JSON object or whose summary is empty MUST be treated exactly
like an empty body: the day is skipped-and-logged (counted as skipped, no row
written) so a re-run retries it, and it MUST NOT wedge the resumable run. The
validated digest MUST be re-canonicalized to JSON before it is stored.
Scenario: Fenced, partly-malformed response is coerced, not rejected
- Given an LLM response wrapping the JSON object in markdown fences with a blank
peopleentry, an unknownmood, and one highlight whosetimeis notHH:MM - When the digest is parsed
- Then the object is extracted from the fences, the empty
peopleitem is dropped,moodbecomes"neutral", the bad highlight time is blanked while its text is retained, and the canonicalized digest is stored.
Scenario: No-JSON or empty-summary response is skipped, not persisted
- Given an LLM response containing no JSON object (or a JSON object with an empty
summary) - When that day is digested
- Then no
journal_digestsrow is written, the day is logged and counted as skipped, and a subsequent run treats the day as still eligible.
REQ-0016-012: Mood is a fixed enum, coerced and denormalized
mood MUST be one of a fixed allowlist — journal.Moods = upbeat, neutral,
quiet, tense — kept in sync with config.DefaultDigestPrompt. Any value
outside the allowlist (including an empty or unrecognized string) MUST be coerced
to "neutral" at parse time. The coerced mood MUST be denormalized into the
journal_digests.mood column so the calendar and stat reads can tint day cells
without unmarshaling the structured blob per day.
Scenario: Unknown mood degrades to neutral
- Given a model that returns
"mood": "chaotic" - When the digest is parsed and stored
- Then
journal_digests.moodis written as"neutral", and the calendar reads the column directly without parsingstructured.
REQ-0016-013: Structured columns via in-place schema amendment and prompt-version re-derivation
journal_digests MUST carry a structured column (the canonical JSON of the
validated digest) and a mood column, both defaulting to '' so a prose-only or
legacy row reads cleanly; body MUST remain the plain-text summary used as the
fallback and empty-response guard. Because these columns were added by amending
schemaV11 in place (safe only while v11 is unmerged/unshipped), and because
journal_digests is a derived cache with no source-of-truth data, no schema
version bump and no --regenerate are required to populate them: editing
config.DefaultDigestPrompt changes prompt_version, which makes every existing
prose-era digest eligible again and re-derives it in structured form on the next
run — a free data migration. A database already at user_version = 11 from before
the amendment MUST be understood NOT to retroactively gain the new columns (the
CREATE TABLE IF NOT EXISTS is a no-op on an existing table); the documented
remedy is to recreate the derived-cache database (it is rebuildable from
messages), not to hand-alter the table.
Scenario: Prompt-version bump re-derives prose rows as structured
- Given days digested before Phase 2 (prose
body, emptystructured/mood) and aDefaultDigestPromptedited to demand the JSON object - When
msgbrowse journalruns against a database that has the amendedjournal_digestsshape - Then each affected day is re-digested under the new
prompt_version, and its row is updated in place with canonicalstructuredJSON and a denormalizedmood, with no schema bump and no--regenerate.
REQ-0016-014: Mood-tinted month-calendar navigator with editorial day card and stat tiles
GET /journal MUST render a mood-tinted month calendar navigator: year tabs
(newest first), a fixed month grid whose present days are tinted by that day's
mood and annotated with a message-count subscript, previous/next-month controls,
and a mood legend. It MUST render an editorial day card for the selected day
(summary, timed highlights, people, themes, standout media, notable links) and a
row of stat tiles (days-with-entries, longest streak, most-active weekday,
peak hour). Navigation MUST be by boosted query parameters (?year, ?month,
?day) with no client-side state. The bare /journal MUST open on the newest
day's card; a year tab (?year with no month) MUST open on that year's latest
active month, never an empty January. When the selected day has no structured
digest, the card MUST fall back to the prose body, then to the mechanical
top-senders. Rendering MUST issue no LLM call, and an archive with no journal MUST
render the empty state.
Scenario: Bare landing opens the newest day
- Given a built journal whose newest day is
2026-07-11 - When the user opens
/journalwith no query parameters - Then the calendar shows July 2026,
2026-07-11is selected and its editorial card is rendered, and no LLM call is made.
Scenario: Year tab opens the year's latest active month
- Given a year whose activity begins in March
- When the user clicks that
?yeartab (no month specified) - Then the calendar opens on that year's most recent active month, not January, and no empty grid is shown.
Scenario: Day cell tinted by mood with a count subscript
- Given a day with a cached digest whose mood is
tenseand 42 messages - When the month grid renders
- Then that day's cell carries the
tensemood tint and a42count subscript and links to?day=for that date; days without content are inert blanks.
REQ-0016-015: Calendar read surface, UTC-bucketed and exclude-honoring
The calendar/stat reads MUST be served by dedicated store methods:
JournalMonth (≤31 rows off journal_days joined to journal_digests.mood),
JournalStats, GetJournalDay (rollup joined with its digest, structured and
mood included), LatestJournalDay and LatestJournalDayInYear, and
JournalYears. JournalStats MUST derive days-with-entries and longest streak
from the journal_days key set in Go (longestStreak, adjacency by date
arithmetic so month/year rollovers count), and most-active weekday and peak hour
as argmax GROUP BY reads over messages (journalArgmax), a year bounded by a
sargable ts_unix range (not a date(ts_unix) wrap). All day bucketing MUST
be UTC. Year 0 means all-time (a full scan), but the web layer MUST always pass
a concrete year. JournalStats MUST honor journal.exclude_conversations — the
same denylist journal_days was built with — so excluded threads never inflate
the stat tiles.
Scenario: Streak counts a month rollover
- Given journal days
2026-01-30,2026-01-31,2026-02-01 - When
JournalStatscomputes the longest streak - Then the streak is 3, because adjacency is date arithmetic, not string succession.
Scenario: Stats honor the exclude denylist
- Given
journal.exclude_conversationsnaming a high-volume thread - When the stat tiles are computed for a year
- Then the excluded thread's messages are absent from the weekday/peak-hour argmax and from the day/streak counts, matching the journal the day rollups were built with.
REQ-0016-016: Untrusted structured fields are escaped; links resolve against archive facts, never against model strings
Every structured-digest field is model-derived and MUST be treated as untrusted
output. All rendered fields (summary, highlights, people, themes, standout media,
notable links) MUST be emitted through html/template auto-escaping. Mood tints
MUST be applied as CSS classes keyed by the fixed mood enum (cal-day--<mood>),
never an attacker-supplied class or inline style. The page MUST remain compatible
with the site CSP (style-src 'self', no 'unsafe-inline'): no style=
attribute may carry model-derived values.
A notable link MAY render as an anchor only when its normalized URL matches a
row in links for that day; on a match the anchor's href MUST be the
stored URL, never the model's string, re-checked server-side against an
http/https allowlist. A javascript:, data:, or otherwise non-http(s)
value is prohibited as an href without exception — matching cannot rescue it,
because it will never match anyway. On any miss the link renders as plain text.
A person chip MAY link to /contact/{id} only when the name resolves to exactly
one contact who participated in that day's conversations; absent or ambiguous
names render as plain chips. Resolution happens in the handler, not the template:
the template only branches on whether a resolved destination exists.
Why matching-not-filtering is the rule. The original requirement was "text only, never a raw
href" - a blanket refusal that also blocked real affordances (issue #371). Filtering hostile values instead would rebuild the web's URL-parsing minefield one CVE at a time. Matching inverts the burden: the archive's own record of what it observed is the allowlist, so a hallucinated or injected entry has nothing to resolve to and is inert by construction rather than by inspection. The display text always stays the model's escaped string - a matched link must not be able to relabel itself as something else - while the href comes from stored fact.
Scenario: A malicious digest field cannot inject markup or a live link
- Given a digest whose
summarycontains<script>and whosenotable_linkscontains ajavascript:URL - When the day card renders
- Then the
<script>is HTML-escaped as inert text, the link is shown as escaped text with no clickablehref, no inlinestyle=is emitted, and the mood tint is a fixedcal-day--<mood>class.
Scenario: An unobserved but plausible link stays inert
- Given a digest naming
https://plausible.example/solar-deal, a well-formed URL never seen in the archive - When the day card renders
- Then the entry renders as escaped text with no anchor at all - no hover affordance implies clickability.
Scenario: An observed link opens externally from the desktop shell
- Given a digest link whose normalized URL matches a stored
linksrow for that day - When the reader clicks it
- Then the anchor points at the stored URL and routes through the desktop external-open path (
desktop.jsto/desktop/open-url) withrel="noopener noreferrer"; it never navigates the shell webview itself.
Scenario: People chips resolve to participants only
- Given a digest person name equal to exactly one contact who messaged that day, a second name shared by two contacts that day, and a third name present in no conversation
- When the day card renders
- Then the first chip links to
/contact/{id}, and the ambiguous and absent names render as plain chips.
REQ-0016-017: The Journal page is a reading surface — no pipeline status, ever
GET /journal is a reading surface. It MUST NOT render any journal-pipeline
machinery: digest coverage counts or percentages, a build progress bar, the
configured chat model, the built-through date, run history, run error strings, or
Build / Rebuild controls. Those belong on the Settings surface that owns the
journal pipeline (REQ-0004-010). The same prohibition applies to the semantic
index — no index card of any kind renders on /journal.
The single permitted exception is a one-line, non-interactive progress note shown only while a build is actively running, linking to the Settings tab. It MUST NOT carry a coverage percentage, a model name, a history table, error text, or any control.
Why this is a prohibition and not a preference. This requirement exists because the build card was repeatedly re-added to
/journal— reasonably, since nothing forbade it and issue #274 positively required the Journal page to "keep its Build / Rebuild controls". Rendering an eight-row table of failed LLM runs on top of the archive's most narrative surface is the failure mode. Anyone tempted to place pipeline status here again should change this requirement first, in a PR that argues the case — not add the markup and leave the requirement contradicted.
Scenario: A partially built journal renders no diagnostics
- Given a journal with 3,591 of 3,920 days digested and several failed runs recorded
- When the user opens
/journal - Then the page renders the calendar and the selected day's editorial card only, with no coverage figure, no progress bar, no run history, no error strings, and no Build or Rebuild control.
Scenario: An active build shows at most a one-line note
- Given a journal build in progress
- When the user opens
/journal - Then at most a single-line progress note linking to the Settings journal tab is shown, carrying no percentage, model name, history, or control.
Related Artifacts
Direct relationships declared in YAML frontmatter (per the SDD plugin's ADR-0023 / SPEC-0018 frontmatter-graph conventions). Run /sdd:graph chain SPEC-0016 for the transitive view.