03 — Guardian stack · Deploying

Deploying a guardian, from an empty host to a running stack

A guardian is one stack of services behind one port, set up by a single script from two inputs: the realm credential VeilNet issues, and a settings file. This page is what the host needs, what setup does, the sign-in and compliance choices, installing with no network, the FIPS position, backup and upgrades.

On this page

Requirements

The guardian ships a stack, not a host. The host's hardening, its clock, its encrypted storage and its log collection stay the deployer's, and the stack cannot check most of them, so a stack that starts cleanly has proved none of them.

NeedWhat it takes
A container hostDocker with its Compose plugin, and openssl and python3, which setup uses.
The anchor binariesSupplied beside the stack rather than built by it, and baked into its API's image. The control tool must be the locked-down build: one able to mint a realm root is refused when the stack is built.
MemoryAbout 2.7 GB reserved across the services, with a 10 GB ceiling, plus the host's own overhead. The limits are there to contain a fault in one service, not to ration normal load.
DiskRoom for ninety days of metrics, an amount that grows with the fleet, on the same encrypted storage as everything else.
Encrypted volumesThree volumes hold data that must be encrypted at rest: the database, the audit records, and the identity service's own data. The host's volume encryption is what meets that, and the stack cannot see whether it is there.
One port443/tcp, the front door, on an address the host's firewall decides. Nothing else is published. The address need not be public, or even in DNS.
A host name, used exactlyOperators' browsers and every node must reach the guardian by exactly the name it was set up with, because sign-in and every manifest are built from it. Change it later and the machines already commissioned go on renewing against the old name until each is given a fresh manifest.
Trust in its certificateBy default the guardian's certificate comes from its own internal authority, which must be trusted on every operator's machine and every node, or browsers warn and nodes refuse to renew. A certificate of your own, or one from a public ACME authority, are the alternatives.
Accurate timeSign-in tokens and certificates both depend on it. Drift shows up as intermittent sign-in failures that are hard to trace.

Inputs and setup

Setup takes two inputs, both the deployer's, and asks no questions, so the same steps are an unattended install.

  1. 01
    Put the realm credential in place

    The one file that is yours to provide, copied into the deployment's secrets directory. Setup refuses to run without it, and checks it before changing anything: its format, its kind and every field it must carry. It warns when the credential allows no sub-realms, so that is learned before a tree is designed rather than after.

  2. 02
    Fill in the settings

    The example settings file, filled in. The host name has no default and is the one setting that must be given. The compliance profile, the sign-in mode, where the certificate comes from and whether machines report telemetry all have defaults, each explained where it is set.

  3. 03
    Run setup
    $ ./setup.sh

    It checks every setting and the credential, and names every problem at once before it changes anything. Then it writes the deployment's secrets, its internal certificate authority and its settings into the secrets directory, builds the stack, starts it, and waits for every service to be healthy.

  • The secrets directory is the deployment. Everything the containers take from the host is written there, the settings included, so a backup of it is a backup of the configuration.
  • Re-running setup is always safe. It fills in what is missing, re-derives what follows from the settings, such as a certificate that no longer names the host, and leaves everything else byte for byte as it was. To change a setting, edit the copy in the secrets directory and run setup again: the example file is read only on the first run.
  • --no-start fills the secrets directory and stops before building or starting anything.

After the first run

Sign in to the identity service's administration as the bootstrap administrator, whose password setup writes to the secrets directory and never prints. Create the real break-glass administrator, check that it works, and disable the bootstrap account: leaving it enabled is the most likely finding in a first assessment. Then ship the audit records to your SIEM, an integration of your own, exercise the break-glass path with your identity provider unreachable, and commission the genesis tier.

Sign-in and compliance profiles

Who may operate the guardian is decided by its identity service, in one of two modes chosen in the settings.

ModeWho holds the accounts
Standalone, the defaultThe identity service itself, with a one-time code required on every account. For sites with no identity provider of their own, and for air-gapped ones.
BrokeredYour organisation's identity provider, which the identity service federates to. The guardian never holds a password, and a change upstream takes effect at the next sign-in. Your groups map to the guardian's roles, and what each role may do stays the guardian's decision.

The four roles, admin, manager, auditor and user, are a ladder: each contains the one below. So grant each person one role, or in brokered mode one group. What each role can do.

The compliance profile sets the values that differ between jurisdictions: us-dod, au-ism or uk-mod. The architecture, the images and the controls are the same in all three. It is one system built to the FIPS bar, configured three ways.

SettingUS DoDAustralian ISMUK MODControl
Session idle timeout15 minutes30 minutes30 minutesAC-11
Session maximum11 hours12 hours12 hoursAC-12
Password minimum length151414IA-5
Password history588IA-5
Password maximum age60 days365 daysNoneIA-5
Lockout threshold3 attempts5 attempts5 attemptsAC-7
Audit retention, hot90 days90 days90 daysAU-11
  • The UK profile sets no maximum password age on purpose: NCSC guidance is against forced periodic expiry, and leans on multi-factor sign-in and monitoring instead.
  • The Australian profile favours length over churn, with a long maximum age and a deeper history.
  • Audit records kept beyond the profile's retention are your SIEM's to keep.

Air-gapped install

The stack runs with no outbound network, but building it needs one: the build pulls base images and the sources it compiles. So it is built on a connected host, and carried across.

  1. 01
    Build on a connected host, and save the images

    Every image the stack runs, saved to one archive. Include the FIPS base only where the deployment opts into it.

  2. 02
    Generate an SBOM for each image

    At build time, to travel with the images. An air-gapped site cannot regenerate one later.

  3. 03
    Carry the images across, and load them
  4. 04
    Run setup without building
    $ ./setup.sh --skip-build

    It builds nothing and pulls nothing, starts the stack from the images already loaded, and refuses, naming each one, if any is missing.

  • The anchor binaries travel inside the API's image, built in on the connected host. Nothing is fetched on the target, which is why the release digests the realm credential carries are there for an operator to read rather than acted on.
  • The realm credential travels separately and by hand, with the secrets rather than with the images: it carries the realm's key. For an offline deployment it is issued without a renewal address, so the guardian never calls out, and a fresh one is handed over before it lapses. Online, offline or expired.
  • Choose a certificate that needs no network, the guardian's own authority or your own, and a sign-in that stays inside: standalone, or an identity provider on the same network.
  • For smartcard sign-in on an isolated network, stage a revocation list locally.

FIPS

How the guardian's application layer uses FIPS 140-3 validated modules, and what that asks of the host. The host kernel's FIPS posture is the deployer's, and is not claimed.

  • The identity service always runs in strict FIPS mode, with a validated provider of its own. Password hashing and token signing happen there, and the hashing is PBKDF2, the one approved option, at 210,000 iterations.
  • The rest of the application layer runs inside a validated module on the FIPS base image, built once as a release job and named in the settings. Setup checks the image is there.
  • The validated provider is pinned at FIPS 140-3, certificate #4985, valid to March 2030.

Test on a FIPS-mode host before calling the deployment ready.

Backup and restore

Two things are backed up, and kept apart. The database, as a full dump: it holds both the identity service's store and the guardian's own record of realms, machines and renewals. And the secrets directory, separately and under different custody, because it is the whole of the deployment's configuration: the settings, the realm credential, every generated secret and the internal certificate authority. A backup holding both the database and the keys that protect it is a single point of compromise.

Restoring

  1. 01
    Check out the stack on the new host, with the anchor binaries beside it
  2. 02
    Copy the secrets directory back
  3. 03
    Run setup

    It finds the settings and every secret in place, overwrites none of them, and brings the stack up. The example settings file is not read.

  4. 04
    Restore the database dump

Upgrades

The guardian's own records migrate in place. Each upgrade's migrations run once, in order, when the API starts, and keep every row already there: every realm, every machine and every sealed value survives, so nothing in the fleet is re-commissioned.

The checklist

  1. 01
    Read the release notes

    The identity service's especially: its own migrations are one-way.

  2. 02
    Back up the database volume and the secrets directory
  3. 03
    Update the pinned digest of each image that changed

    The digest, not just the tag.

  4. 04
    Refresh the anchor binaries if they changed

    The stack will not build without them, and still refuses a control tool able to mint a realm root.

  5. 05
    Rebuild, and bring the services up one at a time

    Check each is healthy before the next. The identity service migrates its own store on its first start after an upgrade, which can take a while on a large realm.

  6. 06
    Generate the SBOMs again