The Ontology & Portable Memory

Status: โœ… shipped (core ontology + domain extension + CLI). The Go ontology registry replaced the Python pinard-core package. The broader multi-model store design is ๐Ÿ”ญ designed โ€” see The Layered Memory Architecture.

For a fleet’s knowledge to be queryable and portable, it has to be typed. Pinard types memory with a layered ontology and ships it as versioned, portable subsets.

A sketched grapevine with a stable shared trunk, distinct domain branches, young learned shoots, one human-reviewed graft toward the core, and a three-compartment portable case holding runtime, process, and memory. L1 ยท pinard-core L2 ยท repository ontology Learned types Human-reviewed promotion Portable agent ยท one versioned unit
Stable core, specialized branches, deliberate promotion. Repository ontologies extend the shared operational vocabulary; learned types remain local unless recurrence and human review justify promotion.
  • L1pinard-core โ€” the small, stable operational trunk shared by every agent.
  • L2Repository ontology โ€” domain concepts branch from and subclass the core.
  • โ†‘Learned โ†’ promoted โ€” only a reviewed, proven domain type is grafted toward core.
  • โ–ฃPortable agent โ€” runtime, process.js, and scoped memory are versioned together.

Why layered

A single flat ontology forces a false choice: either it is generic and useless, or it bakes one domain’s specifics (say, GWAS/HPC terms) into what should be shared by everyone. Pinard splits it into two layers instead.

Layer 1 โ€” core ontology

Repo-agnostic, agent-operational concepts that mirror the primitives of a semi-deterministic loop. Encoded as declarative YAML embedded in the Pinard binary (internal/ontology/core.yaml) โ€” no Python or external runtime required. Small and stable, versioned centrally in Pinard:

  • Entities: task / step, verdict / decision, gate (a breakpoint), action, diagnosis, log_pattern, environment_condition, artifact.
  • Edges: DependsOn, Produces, Consumes, IndicatesProblem, ResolvedBy, RequiresCondition, TriggersDecision.

Layer 2 โ€” per-repo domain

Each repository defines its own ontology that subclasses core and lives next to its process.js. For example, a GWAS pipeline repo might define:

  • SlurmJob is-a Task/execution
  • ShardThresholdDecision is-a Decision
  • GWASStudy is-a Artifact
  • ProvenanceRecord

Granularity is per-repo by default (a per-process override only where a process genuinely diverges; per-agent is too granular). A data-pipeline mid-layer between core and domain is deferred but intended.

Writing a domain ontology file โœ…

A domain ontology is a YAML file conforming to the Pinard meta-schema. Drop it in your vigne’s pinard/ontology/ directory (see Domain extension loading below) and the ingester picks it up automatically โ€” no Pinard rebuild required.

domain: my-pipeline
version: "1.0.0"
group_ids:
  - my-pipeline-build

entities:
  pipeline_job:
    description: "A batch job submitted by the pipeline."
    is_a: task           # inherits task properties from core
    properties:
      job_id:
        type: integer
        description: "Job ID assigned at submission."
      partition:
        type: string
        description: "Compute partition the job ran on."

edges:
  SubmittedTo:
    pairs:
      - [pipeline_job, environment_condition]

File envelope fields:

FieldRequiredNotes
domainโœ“Short identifier for your domain
versionโœ“Semver X.Y.Z
entitiesโœ“Map of role-name โ†’ entity definition
edgesโœ“Map of edge-name โ†’ {pairs: [[src, tgt], โ€ฆ]}
group_idsโ€”Which group IDs this domain applies to (empty = applies to all)
suppressedโ€”List of core entity/edge names to remove from this composition

Each entity entry:

  • description โ€” required
  • is_a โ€” core role to inherit properties from
  • properties โ€” map of property-name โ†’ JSON Schema fragment (type, description, format, minimum, maximum, enum, items)
  • required โ€” list of required property names

Domain extension loading

The ingester discovers domain files at startup from:

  1. PINARD_ONTOLOGY_DIRS โ€” colon-separated list of directories scanned for *.yaml / *.yml / *.json files (for k8s / CI use).
  2. <vignoble>/pinard/ontology/*.{yaml,yml,json} โ€” auto-discovered when VIGNOBLE_DIR is set (the default in a local vignoble).

Missing directories are non-fatal (logged warning). Invalid files are skipped with a warning; valid siblings still load.

Composition semantics

Compose(group_id) = core + domain(group_id) โˆ’ suppressed:

  • is_a inheritance โ€” a domain entity with is_a: task merges task’s core properties under its own (domain properties win on key collision).
  • Edge extension โ€” domain edges sharing a name with a core edge append their pairs rather than replacing the core pairs.
  • Suppressed โ€” role or edge names in suppressed are excluded from the result.

CLI tools โœ…

Two new aoc subcommands help domain authors validate and inspect their ontology:

# Validate a domain file โ€” exit 0 on success, non-zero on error
aoc ontology validate path/to/my-pipeline.yaml
# โœ“ my-pipeline.yaml is valid

# Inspect the composed result for a group_id
aoc ontology inspect --group-id my-pipeline-build
# Composed ontology for group_id="my-pipeline-build" (core 1.0.0 + domain my-pipeline@1.0.0):
# Entity roles:  task  step  verdict  decision  gate  action โ€ฆ  pipeline_job  โ€ฆ
# Edge types:    DependsOn (3 pairs)  โ€ฆ  SubmittedTo (1 pairs)

Use aoc ontology validate as a CI gate in domain repos to catch schema errors early.

Lifecycle: prescribed โ†’ learned โ†’ promoted

Types are not frozen. They move through a lifecycle:

  • prescribed โ€” declared up front (core, or a repo’s domain);
  • learned โ€” new types emerge during /teaching sessions;
  • promoted โ€” a proven domain type is elevated toward core.

Promotion domain โ†’ core is human-gated (a git PR), and a suppressed_types list retires types that stop earning their place. This is the same recurrence-plus-review pattern used for rule and scope promotion.

Portable memory subsets

This is the payoff of a single, embeddable store of record. The central memory (SurrealDB server) can be subset โ€” scoped to a repo or pipeline โ€” into a specialized embedded SurrealDB file that an agent loads locally.

That makes a pinard agent a self-contained, reproducible artifact:

 pinard agent  =  harness  +  babysitter process  +  memory subset
                  (the runtime) (the loop, process.js) (embedded, scoped store)

All three are versioned together (e.g. in ExoHub), so an agent-centric data pipeline is reproducible and auditable: you can rebuild exactly the agent, its loop, and the knowledge it had at a given version.

Version-stamping

Every portable subset version-stamps both the pinard-core ontology version and the domain ontology version it was extracted under. A migration policy governs what happens when core or domain versions change under an already-shipped subset โ€” new, load-bearing scope that portability introduces.

Next