The Layered Memory Architecture

Status legend: ✅ shipped · 🔭 designed · 🧪 spike (being de-risked). The engram capture-and-store surface described in Memory & Recall is shipped today. The layered architecture on this page is the design Pinard is building toward; treat 🔭 items as roadmap, not current behavior.

The memory layer is organized as a stack. Each layer has one job, and the layers above depend only on the interface below them — so any single layer can be swapped without breaking agents.

A cellar cutaway with one rooftop scope landscape and exactly six interior floors: curation, ontology, knowledge graph, recall, one store-of-record floor, and capture. L6 · Scope & promotion L5 · Curation & wiki L4 · Ontology L3 · Knowledge graph L2 · Recall L1 · Store of record L0 · Capture Portability & embeddings
The memory cellar, bottom to top. Each level has one responsibility and exposes an interface to the level above. The side mechanisms represent portability and embeddings across the stack.
  1. L0Capture ✅ — curated lessons, teaching episodes, and passive observations enter through Engram.
  2. L1Store of record 🔭 — documents, graph edges, and vectors share one SurrealDB store.
  3. L2Recall ✅ — typed query/fetch intent and compact boot injection.
  4. L3Knowledge graph 🧪 — temporal graph value is evaluated before commitment.
  5. L4Ontology 🔭 — stable pinard-core plus specialized repository domains.
  6. L5Curation & wiki ✅ — a git-tracked, human-readable OKF artifact synchronized with the store.
  7. L6Scope & promotion ✅ — curated knowledge rises from vigne to vignoble to global.

The design through-line is signal over noise: capture curated knowledge instead of dumping transcripts, keep it in one multi-model store instead of a zoo of databases, and treat memory as a versioned, portable artifact you can ship with an agent.

Layer 0 · Capture — Engram (✅ shipped)

Knowledge enters through Engram, not raw transcripts. Three inputs:

  • /lesson <text> ✅ — a one-shot pinned fact or rule, high-confidence and eligible for promotion (e.g. “always use a fix:/feat: commit prefix”). Available in conductor and vendangeur.
  • /teaching ✅ — a mode: while active (with a visible status-line indicator, auto-off at session end) the session is captured as a teaching-episode. Supports retroactive capture via --all or --from <duration>. Conductor only.
  • Passive capture ✅ — session summaries and curated mem_save observations, so lessons are never purely manual.

Using Engram as the write path removes most extraction noise: it already produces typed, scoped, curated observations.

Layer 1 · Store of record — SurrealDB 🔭

A single multi-model store replaces the original zoo (FalkorDB + Qdrant + git-YAML). SurrealDB holds documents, graph edges (RELATE), and vectors (HNSW) in one engine, and runs either as a central server or as an embedded, file-based database — the duality that makes portable subsets possible (see below). It is a single static binary, which fits the HPC “no Docker” constraint. Agents never touch it directly; all access is through Layer 2.

Layer 2 · Recall — a typed intent API 🔭

The memory service exposes a small, stable API over NATS request-reply — not raw queries — so the store stays swappable:

  • recall — semantic / vector neighbors,
  • lookup — lexical / full-text,
  • trace — knowledge-graph traversal.

It is fail-open with a ~3s timeout: if memory is slow or down, the agent proceeds without it rather than blocking. Humans get a separate direct read-only path for exploration.

Layer 3 · Knowledge graph — Graphiti, time-boxed 🧪

A temporal knowledge graph gives two hard things cheaply: bi-temporal validity (facts expire) and LLM entity/edge extraction. Rather than commit to it permanently, Pinard runs Graphiti (on FalkorDB) as an evaluation sidecar, fed one-way from the store of record (SurrealDB → jsonl → Graphiti). A Phase-2 decision then picks one of: keep Graphiti, build native temporal-KG in SurrealDB, or adopt Spectron (an upcoming Graphiti-like temporal KG on SurrealDB) — the preferred long-term slot because it needs no store migration.

Layer 4 · Ontology — layered 🔭

Knowledge is typed by a two-layer ontology: a small, stable pinard-core (agent-operational concepts that mirror babysitter primitives) plus a per-repo domain layer that subclasses it. See The Ontology.

Layer 5 · Curation & wiki ✅

A self-evolving git-tracked OKF wiki ⇄ SurrealDB turns the typed memory graph into curated, human-readable pages, with an ontology gardener that proposes structural extensions via human-reviewed MRs. See Teaching & Curation.

Layer 6 · Scope & promotion ✅

Knowledge is stored at the finest grain (a vigne) and rolled up — vigne → vignoble → global — with curate-on-promote ensuring only synthesized wiki_doc entries rise (never raw entities). Vignoble-shared results sync to wiki/_shared/ for human visibility. See Scope & Promotion.

Cross-cutting

  • Embeddings — Rosetta 🔭: an OpenAI-compatible endpoint (qwen3-emb-0.6b, 1024-dim). The same endpoint embeds both writes and queries, guaranteeing vector comparability, with no token/login dance.
  • Portability 🔭: the central store can be subset by scope into an embedded SurrealDB file — so a pinard agent = harness + babysitter process + memory, all three versioned together. See The Ontology.

Status & de-risking

Shipped (Layers 0, 5, 6): the WikiCurator (outbound, per-vigne namespacing, LLM-synthesized summaries), WikiSyncer (inbound), OntologyGardener, the vignoble OKF bundle scaffold, /lesson, /teaching (with retroactive modes), curate-on-promote (only wiki rises, vignoble-shared sync-out), and boot injection v2 (compact manifest with drill-down via recall(fetch=<ref>)) are all live.

The remaining in-flight work is two 🧪 spikes: (1) Engram → SurrealDB curated ingestion and recall quality on 1024-d vectors (Layer 1 as the full store of record), and (2) SurrealDB → jsonl → Graphiti, to judge the temporal-KG value before committing (Layer 3). Until those land, the shipped engram path (Memory & Recall) remains the primary agent-facing write path.