Specification
A knowledge base is a directory. Everything below describes what is in it.
strauss-kb schema emits JSON Schema generated from the Zod schemas that
enforce it, so it cannot drift from what a write will accept.
Bundle layout
<kb>/
<type>.<slug>.md records
INDEX.md index derived, store-owned
log.jsonl history primary, append-only
.gitattributes merge store-owned, written on first write
.index.sqlite search derived, gitignored
The default base is .strauss/kb, relative to the working directory;
--bundle PATH addresses any other.
INDEX.md and log.jsonl are both store-owned, and differ in kind:
INDEX.md | log.jsonl | |
|---|---|---|
| Nature | derived, recomputable | primary — events nothing else holds |
| Write | full regenerate | append |
| Repair | rebuilt when it disagrees | malformed lines reported |
| If lost | reconstructed free | gone |
The store excludes both from record listings and repairs the index on read, which holds only because it is the sole accessor.
INDEX.md
A projection, one line per record, sorted by concept id:
- [Title](fact.cache-key.md) — fact · accepted · tags: cache · A region-less key serves the wrong data.
It carries description, not just the title, so a reader can decide what is
worth opening. No lock is needed: two writers compute the same function.
log.jsonl
One JSON object per line, appended with O_APPEND:
{
"at": "2026-08-16T09:14:00Z",
"by": "agent",
"operation": "write",
"conceptId": "decision.cursor-v2"
}
| Field | Required | Meaning |
|---|---|---|
at | yes | ISO timestamp |
by | yes | the actor, from STRAUSS_KB_ACTOR |
operation | yes | e.g. write, verify:refused |
conceptId | yes | the record acted on |
target | no | the operation's other end: a second id for supersession, the other base's absolute path for promote-in/promote-out |
reason | no | why, where the operation demands one — anchor-set does |
anchors | no | what anchor-set changed: { op, from?, to? } per pointer, op one of move, add, drop |
Unknown keys are kept, not rejected: one base is read by every version that
touches the repository, so the log has to read forward — under a strict schema
the first version to add a field turned its own entries into malformed for
every older reader. A line missing a required field, or carrying an at that is
not an ISO-8601 UTC datetime, is still malformed. Writes go through a strict
schema, since this package controls what it appends.
reason and anchors arrived in 0.1.22, when the read schema was still strict.
A CLI or MCP server built before it reports every anchor-set line as
malformed and drops it from entries. Malformed lines are reported, never
repaired, so nothing is lost — but rebuild or update before reading a base for
its audit trail.
::: Malformed lines are reported with their 1-based
position and never rewritten. Reads are sorted by at and deduplicated on
exact equality over the whole parsed entry.
.gitattributes and cross-worktree writes
Git's line-level merge is wrong for a file both sides only append to, so the first call that appends a log line writes a merge driver, and marks every store-owned file generated so a review of a committed base is a review of its records:
log.jsonl text eol=lf merge=union linguist-generated=true
INDEX.md linguist-generated=true
.index.sqlite linguist-generated=true
union is built in, so the attribute alone is enough, and eol=lf pins line
endings regardless of core.autocrlf. Each attribute is checked separately, so
a base written before this block grew gains only the lines it lacks and a value
already set — merge=ours, -linguist-generated — is left alone. The step is
best-effort and never fails the mutation that triggered it.
git merge, not to GitHubGitHub computes pull request merges through its own service, which does not read
.gitattributes merge-driver declarations, so its merge button can leave
conflict markers in log.jsonl. Reads skip every marker line — <<<<<<<,
=======, >>>>>>>, and diff3's ||||||| base section — keep both sides'
entries, and warn once for the file.
Records
Identity
The filename is the identity: fact.auth-retries.md has concept id
fact.auth-retries, the bundle-relative path with .md removed.
slug ^[a-z0-9]+(?:-[a-z0-9]+)*$
concept id ^[a-z0-9]+(?:-[a-z0-9]+)*\.[a-z0-9]+(?:-[a-z0-9]+)*$
Both halves are kebab-case, enforced at the entry point because concept ids are interpolated into markdown links unescaped.
Frontmatter
Records are OKF v0.2 concepts: type is the only required key, everything else
is optional, and unknown keys are kept rather than stripped.
OKF keys
| Key | Type | Meaning |
|---|---|---|
type | string | the only always-required key |
title | string | one line, in the reader's terms |
description | string | what breaks if this is wrong |
resource | string | a path this concept names |
tags | string[] | free-text labels, filterable — see below |
sources | Source[] | material the record draws on |
generated | { by, at } | who wrote it, and when |
verified | { by, at }[] | the append-only trail of checks |
stale_after | string | the date this record stops being trusted |
A source is { id, resource, title?, author?, last_modified? }, id and
resource required; footnotes key to id.
tags are selectable: kb_list, kb_query and kb_catalog take a tags
array (CLI --tag, repeatable) and return the records carrying every tag in it,
and a kb_context profile takes excludeTags to keep tagged records out of the
injected block without unpinning the base. Selection runs after adjudication, so
a tag never changes a record's standing or the replacement it names. Matching is
exact and the vocabulary is not enforced, so an unknown tag returns nothing
(Architecture).
strauss extensions, namespaced so a later OKF version cannot collide:
| Key | Type | Meaning |
|---|---|---|
strauss_status | enum | see Standing; defaults to draft |
strauss_supersedes | string[] | ids this record replaces |
strauss_superseded_by | string | the id that replaced this one |
strauss_anchors | Anchor[] | where it attaches in code — see Anchors |
strauss_links | Link[] | typed causal edges — see below |
strauss_verify | string[] | checks that would confirm the record still holds |
strauss_answered | { by, at } | who resolved an open question, and when |
strauss_assumption | boolean | the claim has no source |
strauss_materiality | enum | blocking, important, non-blocking |
strauss_confidence | enum | low, medium, high |
strauss_owner | string | a name |
Anchors
strauss_anchors:
- file: src/kb-store.ts
symbol: KbStore.setStatus
hash: sha256:9f2c…
lines: 24
resolved_at: 2026-08-16T09:14:00Z
resolver: tree-sitter
strict() — file is required and nothing outside this table is accepted:
| Field | Required | Meaning |
|---|---|---|
file | yes | the repo-relative path the concept names |
symbol | no | a symbol within it; absent means the file |
span | no | { start, end }, 1-based and inclusive |
side | no | old or new; absent reads as new |
hash | no | sha256:<64 hex> over the anchored text |
hash_kind | no | raw or ast; absent reads as raw |
lines | no | the line count that hash was taken over |
resolved_at | no | ISO timestamp of the last resolution |
resolver | no | tree-sitter, regex or span; absent reads as regex |
repo | no | which repository; absent means the base's own |
ref | no | git rev the evidence was taken at |
Anchors stay symbolic because they are written while the code is still moving;
once it settles, a resolution pass stamps hash, hash_kind, lines,
resolved_at, and resolver. Those five are measured; repo, ref, span
and side are author-owned and never stamped. A tree-sitter stamp hashes the
span's normalised token stream — comments dropped, whitespace collapsed — and
records hash_kind: "ast", so reformatting the anchored code is not drift. A
raw hash keeps comparing raw text; the two kinds are never compared to each
other. CRLF is normalized to LF before hashing, and lines is what lets a drift
report say how much changed.
Prefer a symbol: it survives reformatting and moves. Two claims have no symbol to hang on, and each gets its own address.
- A
spannames lines in a file no resolver can name a symbol in — YAML, SQL, Markdown, JSON. It is hashed exactly as written, alwaysraw: a slice is not a syntactic unit. An anchor names a symbol or a span, never both;kb_validatereports two addresses, a backwards range, or anasthash over a span. side: "old"with arefnames code as it was at that rev — what a refactor removed — read withgit cat-file blob <ref>:<file>, never from the working tree;kb_validatereports a missingref. Committed bytes cannot drift: the anchor reportsmatchuntil the rev changes and is never searched for moves. A rev this clone lacks isref-unavailable— unchecked, notgone— so a shallow checkout says nothing about the old side.kb_doctorcounts old-side anchors on a line of their own.
Drift
An anchor carrying a hash can be re-resolved and compared. Four states:
| State | Meaning |
|---|---|
stamped | had no hash; one was just written |
match | the code still hashes to what was recorded |
drifted | it resolves, and hashes to something else |
unresolved | it no longer resolves at all |
unresolved carries a reason: file-missing, symbol-not-found,
symbol-ambiguous, span-out-of-range, ref-unreadable, ref-unavailable,
resolver-unavailable, outside-repo, file-too-large, file-unreadable,
remote-unreachable, ref-not-found, repo-unauthorized,
default-branch-unknown, ref-invalid, or repo-invalid. drifted carries
resolver-changed when the resolver changed and the code did not.
Drift classes
Drift is also classified, so a reader only sees what a machine cannot settle:
| Class | Meaning |
|---|---|
moved | the stored hash resolves at another file or symbol — same code |
cosmetic | old and new spans are one token stream; only formatting changed |
gone | the file, symbol, span or ref is no longer there to read |
changed | everything else |
gone and changed are settled by the hash comparison itself and appear on
every read path. moved needs a repository-wide search and cosmetic needs the
committed text, so both are computed on demand by
reassess and
doctor --drifted. cosmetic needs a grammar, so
the regex resolver never reports it — there the class is changed. A span is
searched for moved by sliding its recorded line count over its own file, since
it names no definition to look for elsewhere. An old-side anchor is neither
searched nor diffed: committed bytes cannot move or be reformatted.
Moving a pointer
file, symbol, span, side, repo and ref are the anchor's address;
the five measured fields are its baseline. A refactor that renames the anchored
symbol, or extracts part of it, leaves a record pointing at code that is no
longer there — and nothing mechanical can say which new symbol replaces it.
anchor-set is the reader's answer: the new
set of anchors, and a required reason.
A record's first write and anchor-set share one check: no two anchors at one
address. Beyond that the set is taken as given. An anchor keeps its baseline by
carrying its hash and the rest of its stamp forward; carried under a new
file or symbol, the drift the rename hid is reported the moment the pointer
resolves, and clearing it is still
anchor-resolve --rebaseline. Whether the
code behind a pointer was read is the caller's claim, and the anchor-set log
entry records who made it and why — it is never a verified[] event.
What drift does and does not see
A hash over one span answers one question, which bounds the whole mechanism:
- A body that changes under the same name drifts. That is the case the hash exists for.
- An unchanged caller does not inherit an unanchored helper's change: the caller's own span still hashes the same, so the helper needs its own anchor to be watched.
- A runtime change — an environment variable, a feature flag, a config value read at startup — changes no bytes and so drifts nothing. A risk about one stays open after the code drift clears; only a reading closes it.
The reassessment packet
Per record that still needs a reading: its title, why, and the claim section
of its type; per anchor the class and, with --with-diff, a unified diff of the
old span against the new; and the record's
impact set, because a fact that stopped holding
invalidates what depends on it. The old span is recovered with git show <ref>:<file>, or from the last commit touching that path before resolved_at;
neither resolving makes the diff unrecoverable. Diffs are capped per anchor
and per packet.
Each packet names a type-based default — fact, constraint and contract
anchored to changed or gone code read as presumed invalidated; decision
and risk as rationale may survive, check. It is a starting point, not a
verdict: no drift path writes verified[] or changes standing.
Symbol resolution
A symbol resolves through a chain, and tree-sitter answers when its tags query defines the symbol:
| Resolver | Covers |
|---|---|
tree-sitter | the 20 languages with both a grammar and a definitions query |
regex | every other extension, and any symbol the tags query does not define |
| whole-file | an anchor with no symbol |
Upstream tags queries define functions, classes, methods, interfaces, traits,
structs and modules — but not constants, type aliases, TypeScript enums or
class fields. A symbol the query does not define falls through to regex, and
the anchor records resolver: regex. An ambiguous AST match
(symbol-ambiguous) and a grammar that will not load (resolver-unavailable)
never fall through: one would be settled by guessing, the other would trade a
precise span for a guessed one.
Which languages those are is data, not code. A language pack is a WASM grammar
and its definitions query pinned together at one package@version;
grammars/packs.json lists the 30 packs and where each part comes from, and is
the only file a human edits. pnpm grammars pin resolves both parts, proves
them — every pack's WASM must load under the installed web-tree-sitter with
every other pack resident, and each query must compile against its own — and
writes grammars/manifest.json, which carries the URL,
hash and file extensions the runtime reads. A part that is missing or will not load fails the pin
rather than shipping a language that would report itself unavailable, and
pnpm grammars check re-proves every pack weekly against the real CDN. An
extension whose pack has no tags query stays with the regex heuristic, as
before the resolver existed.
Definitions are whatever upstream's tags query captures as @definition.*,
named by its @name. A dotted symbol (KbStore.setStatus) resolves to the
definition whose enclosing chain matches — a Go method through its receiver, a
Rust function through its impl block. A bare symbol must match exactly one
definition; two make it symbol-ambiguous, except that a signature loses to
the implementation of the same symbol. Only declarations count: a symbol that
appears only in a call is no tree-sitter match, and goes down the chain.
Neither half of a pack is shipped with the package: the grammar and the tags
query both download from jsDelivr on first use, are verified against the sha256
pinned in grammars/manifest.json, and are cached under
~/.strauss/grammars/<language>/ — 49 MB of WASM in every install, for a
feature most installs never reach, is a bad trade. A grammar
that cannot be obtained or verified is resolver-unavailable, never a throw
and never a silent fall back to regex, whose span for the same symbol is a
different hash. --offline and STRAUSS_KB_GRAMMARS=off use the cache without
fetching; the report names the grammar and the repair.
The regex resolver ranks candidates by shape, scopes a dotted symbol to the nearest parent above it, and captures by brace depth or Python indentation. It runs where no grammar applies, and for symbols no grammar defines.
Because the two resolvers span code differently, an anchor stamped by one and
re-resolved by the other can hash differently over unchanged code. That is
reported as drifted with reason resolver-changed, and
anchor-resolve --rebaseline accepts it;
nothing is ever restamped silently.
Drift is computed on read, never stored. load
and query re-resolve hash-carrying anchors against
the working tree and attach a warning:
{
"kind": "drifted",
"anchors": [
{ "file": "src/kb-store.ts", "symbol": "KbStore.setStatus", "diffSize": 6 }
]
}
diffSize is null when the anchor recorded no line count. Anchors with no
hash are never read, and any failure degrades to no drift information rather
than failing the read. When --repo-root is omitted and every checked anchor
comes back missing, the finding is discarded.
Anchors in another repository
An anchor whose repo is not this root's origin is read from that
repository's remote, never from a local checkout — a checkout is one person's
possibly stale view of it. Resolution fetches into a bare cache at
~/.strauss/repo-cache/<host>/<org>/<name>.git (STRAUSS_KB_REPO_CACHE
overrides), one git fetch --depth 1 per (repo, rev) per run, and reads the
blob with git cat-file. Authentication is git's own. A fetch that hangs is
cut off after 30s (STRAUSS_KB_FETCH_TIMEOUT_MS).
With a ref, the evidence is checked at that commit and "current" is the
remote's default branch:
| State | Meaning |
|---|---|
matches-ref | the hash holds at the ref, and on the default branch |
drifted-from-ref | the hash does not hold at the commit the record names |
drifted-on-default | it holds at the ref, and the default branch moved past |
Without a ref there is one state, against the default branch. Only a full URL
can be fetched from, so validate warns on a short repo; records are never
rewritten. load and query never fetch — they read the cache, and report
anything they could not check as unchecked.
repo and ref are record data, and both reach git argv, where an argv array
stops the shell but not git's own option parsing. So both are checked before git
sees them: a ref must be an option-free, range-free name git's own
check-ref-format accepts (ref-invalid otherwise), and a repo must be an
https, ssh, or git remote carrying no password (repo-invalid otherwise —
ext::, file://, and plaintext http:// are all refused).
STRAUSS_KB_REPO_PROTOCOLS widens that list and exists for the test suite, not
for production.
Typed causal links
strauss_links carries directed, typed edges. Every edge reads source →
target and lives on the source's frontmatter, so
{ target: fact.b, rel: depends_on } on record A says A needs B.
strauss_links:
- { target: fact.index-on-created-at, rel: depends_on }
- { target: requirement.stable-ordering, rel: satisfies }
The vocabulary is closed — eight rels, and nothing else may be written:
rel | Meaning | Dependant |
|---|---|---|
depends_on | the source needs the target to hold | source |
constrains | the source bounds what the target may do | target |
informs | the source shaped the target without binding it | target |
blocks | the target cannot proceed until the source is settled | target |
invalidates | the source makes the target no longer hold | target |
verified_by | the target is the check that confirms the source | source |
satisfies | the source discharges the target's requirement | source |
related_to | a pointer worth following, with no claim of dependence | — |
The dependant column is load-bearing: dependence does not follow the
direction of the edge, so impact follows each rel
in whichever direction its dependence runs, and nothing propagates along
related_to. Supersession is a lifecycle rather than a rel.
Tolerant read, strict write. The schema keeps rel a plain string so an
unknown rel stays readable; composeRecord refuses to write one, and
validate turns a stored one into an error. target need
not resolve. The write path caps links at 64 and refuses a self-link. Each
link also renders into the body as one sentence from a fixed per-rel template —
Depends on [fact.b](fact.b.md).
Body
Section headings come from the record's type (see Record types) and are ordered; one the type does not define is rejected, and one left empty is omitted rather than stubbed. A markdown link in the body renders an edge rather than being one — see Edges.
---
type: decision
title: Compare-and-swap rather than a lock
description: A stale lock hold blocks every later writer.
generated: { by: agent, at: 2026-08-16T09:14:00Z }
verified: []
strauss_status: accepted
strauss_anchors:
- { file: src/kb-store.ts, symbol: KbStore.setStatus }
---
## Decision
Read-modify-write checks a content digest immediately before publishing.
## Rejected
A lock file, which adds a stale-hold failure mode.
Record types
The types differ only in what their body answers and where they start in the lifecycle.
| Type | Purpose | Sections | Initial status |
|---|---|---|---|
fact | Observed or sourced fact | Claim · Evidence · Implication | accepted |
requirement | Required behavior or outcome | Claim · Evidence · Implication | proposed |
constraint | Limitation, boundary, policy, or restriction | Claim · Evidence · Implication | accepted |
decision | Chosen or proposed direction | Decision · Rationale · Rejected · Impact | accepted |
assumption | Unsourced working assumption | Claim · Why we think so · What would falsify it | draft |
open-question | Question needing resolution | Question · Why it matters · Default assumption | open |
risk | Something that can go wrong | Risk · Why it matters · Mitigation · Verification | open |
contract | API, data, event, schema, or permission contract | Contract · Producer · Consumer · Compatibility | proposed |
flow | Sequence, lifecycle, or state behavior | Flow · Trigger · Steps · Failure modes | accepted |
affected-system | Component, service, package, or external system | System · How it is affected · Blast radius | accepted |
source-note | Extracted note from source material | Note · Where it came from | accepted |
strauss-kb types prints this table from the code. An unrecognised OKF type
is a note, not a failure.
Standing and adjudication
OKF answers "is this still true?"; standing answers "is this settled, and does it still apply?". Seven statuses map onto five standings:
strauss_status | Standing | Warning attached |
|---|---|---|
accepted | current | — |
resolved | current | — |
draft | unsettled | unsettled |
proposed | unsettled | unsettled |
open | open | unresolved-question |
rejected | rejected | rejected |
superseded | superseded | superseded, with the resolved heads |
Adjudication attaches standing to every hit and never drops one: a filtered
result set is invisible. query drops a superseded record only when its
replacement is also in the results.
Warnings
| Warning | Meaning |
|---|---|
rejected | explicitly not adopted — a well-formed assertion of what someone decided not to do |
superseded | replaced; carries by, the surviving heads |
unsettled | draft or proposed |
unresolved-question | says a matter is unresolved — valuable as a result, never as an answer |
broken-chain | strauss_superseded_by names a record not in the bundle |
chain-cycle | the supersession walk revisited a record |
forked-chain | two records claim to replace this one, so every head is reported |
stale | stale_after is in the past |
unverified | verified[] is empty |
drifted | anchored code moved; carries anchors, each with diffSize and any reason |
unchecked | a foreign anchor nothing could reach — neither drift nor a clean match |
Supersession
A record whose meaning changed is superseded, never edited or deleted: one that quietly becomes something else invalidates every reference to it. Both directions are written together:
old.strauss_superseded_by = new
new.strauss_supersedes = [old, …]
Two paths write that pair: supersede <old> <new>, and write /
write-decision carrying supersedes, which publishes the new record first and
then marks each prior record it names.
An id naming a record that does not exist yet is legal, one naming the record's
own id is a no-op, duplicates mark once, and the array is capped at 32. The
write returns { conceptId, action, supersededIds }, where supersededIds
holds only the ids actually marked, so a crash mid-way is reported by
validate rather than silent.
The one deletion
sweep is the single exception: review-tagged
records in a terminal status are never traced, so git history is their archive.
It deletes only records carrying the tag it was given and sitting in
resolved, rejected or superseded, refuses without a tag, keeps any record
a surviving record still points at — by typed link or by supersession — and logs
each deletion as sweep.
Chain resolution happens on read
Walking to the head is done at read time, following both pointers: a stored
head would have to be rewritten on every ancestor whenever a chain grows. A
cycle terminates with chain-cycle, a fork reports every head, and a missing
replacement is broken-chain.
Verification
verified[] is a record's append-only trail of checks: OKF's { by, at } actor
stamp, plus a required note on entries verify writes. Prior entries are
spread forward untouched, and the schema still reads the noteless OKF shape.
It holds judgments only. anchor-resolve never writes it: an anchor's hash
and resolved_at are the mechanical evidence. verify refuses the actor
unknown.
A verifier whose actor equals the record's generated.by, compared
case-insensitively, is refused unless the actor is human:-prefixed: a
generator re-reading its own output is not an independent check. The refusal is
logged as verify:refused, and human: is an honor-system label. Adjudication
today reports only the unverified warning.
Writes
Records are staged to a sibling file and published atomically with link, which
fails when the name is taken — a 409 carrying action: "refused" in its
details. rename is used only when the caller passes overwrite.
Read-modify-write (status, answer) checks a content digest immediately
before publishing; see
Architecture.
Write input
write takes a type and this object; write-decision takes the same minus
sections, plus alternative and impact. The schema is .strict(), so
unknown keys are rejected.
| Field | Required | Meaning |
|---|---|---|
slug | yes | the second half of the concept id |
title | yes | one line → OKF title |
why | yes | what breaks if this is wrong → OKF description |
sections | no | keyed by the type's section headings; unknown ones rejected |
anchors | no | { file, symbol? }[] |
sources | no | { id, resource, title?, author?, last_modified? }[] |
assumption | no | true when no source exists |
stale_after | no | YYYY-MM-DD, and a real date |
verify | no | checks that would confirm this still holds |
tags | no | free-text labels |
relatedConceptIds | no | stored as related_to links, and rendered as prose |
links | no | typed edges { target, rel }, max 64; a self-link is refused |
supersedes | no | ids this record replaces, max 32 |
materiality | no | blocking | important | non-blocking |
confidence | no | low | medium | high |
owner | no | a name |
Every written record gets generated: { by, at } and verified: [].
no-decision
decision.none is the explicit claim that a piece of work had nothing to
decide, so a workflow gate can ask "did you answer?" rather than "did you write
a decision?". It is idempotent, and "what was decided" never returns it.
Validation rules
Per-record shape is enforced on every read, so validate covers only what a
single record cannot see; a problem it reports means someone hand-edited a file.
| Check | Severity | Reported when |
|---|---|---|
type | error | the type is not a known type — a note, since OKF permits any |
superseded_by | error | a superseded record names no replacement, or a missing one |
backlink | error | the replacement does not list this record in strauss_supersedes |
supersedes | error | a named target is missing, or is not marked superseded |
link_rel | error | a rel outside the closed vocabulary |
link_target | error | a target that is not a well-formed concept id |
link_target | warning | a well-formed target not in the bundle, or a record linking to itself |
assumption | error | strauss_assumption set and sources non-empty |
Each problem is { check, conceptId, note, severity }, and only errors fail
the check — the CLI exits 1 on at least one. The line between the two is
whether time can fix it: a link to a record that does not exist yet is the
ordinary state of a base being written. These checks live here rather than in
the schema because a file that fails to parse is skipped by listings.
Retrieval
Three axes decide whether a record answers a question, and only one is a search problem:
| Axis | Source | Question |
|---|---|---|
| Relevance | BM25 where an index exists, else substring | does this match? |
| Standing | strauss_status, the supersession chain | is this still what we hold? |
| Freshness | stale_after, verified[] | has anyone confirmed it lately? |
@tobilu/qmd is an optional peer dependency providing BM25 over a
.index.sqlite per base, rebuilt when a record is newer than the index. Absent
— the default — query falls back to a substring scan over concept ids, titles,
descriptions, and bodies, and only recall degrades: on a twenty-record base,
eight of nine probe queries returned what substring returned.
Budgets and refusals
load refuses rather than truncating past a ceiling, because a truncated
base is indistinguishable from a complete one; pack and context do the same
at their own budgets.
| Ceiling | Default | Held against |
|---|---|---|
--budget / budgetTokens | 25,000 | the estimated size of what is handed back |
The comparison is strictly greater, so a base at the budget loads.
{
"loaded": false,
"recordCount": 62,
"approxTokens": 31000,
"budgetTokens": 25000,
"message": "Refusing to load this base whole: 31000 tokens is past the 25,000-token budget. …"
}
message names the budget value, the next calls, and both escape hatches. A
successful load reports recordCount, budgetTokens, and tokensLoaded;
nulls mark that all was used.
The load digest
Every load result, refused or not, carries a digest: one SHA-256 over the
records it would hand back. Each current record contributes
<conceptId>:current:<hash of its canonical recomposed markdown> and each
superseded stub <conceptId>:superseded:<hash of the stub>, sorted, joined,
hashed again — any change flips it. A hook or kb_stamp (SAA-719) detects
change without loading. It also drives
cache-stable placement.
The digest hashes each record's canonical recomposed form, not its on-disk bytes, and the parser does not normalize the body's line endings — so a record authored with CRLF digests differently from the same record authored with LF.
Superseded records come back as stubs — { conceptId, title, supersededBy, at } — not bodies, because over a long session a body outlives its qualifier.
trace still reaches the content by id.
Edges
Four edge kinds connect records in one bundle:
| Kind | Two records are neighbours when | Directed |
|---|---|---|
typed-link | one declares a strauss_links entry naming the other | yes |
supersession | either direction of a supersession pair | no |
anchor | they share a code anchor | no |
source | they share a source | no |
typed-link is the edge a record itself makes; the other three are
symmetric. pack walks all four with the whole rel vocabulary including
related_to. trace walks the same four narrowed to the causal rels:
related_to would flood a timeline.
Prose is not walked. A markdown link in a body is the rendering of an
edge, never the edge. validate is the one
reader of a body.
A pair connected by two rels comes back with both in via, and an unknown rel
is never traversed. The inbound half of a typed edge is answered
by backlinks and
impact.