2026-09-21
The Stateful Node: Architectural Vision for Sovereign Edge Infrastructure
A reference architecture for running stateful workloads on sovereign edge infrastructure without sacrificing cattle-not-pets principles. Design patterns for the four-tier storage hierarchy, workload-aware caching, and automated recovery.
Platform & Infrastructure · architecture vision — implementation in progress · ~9 min read · companion: The Second Life of Computers · cited by the Sovereign Fleet series
Where the Sovereign Edge Stops
The sovereign edge architecture handles stateless workloads well. Workers, KV, D1, R2 — those primitives can build a surprising amount from scratch. What they cannot run is an existing application: the database that needs POSIX filesystem semantics, the search index that needs low-latency block I/O, the dependency graph that nobody is going to rewrite as a serverless function.
Those workloads have to live somewhere. The question this post answers is: where can they run that is sovereign, reliable and cost-effective — without rebuilding the vendor lock-in the rest of the architecture exists to escape?
The Principle: A Recoverable Pet, Not True Cattle
The stateful node is the compatibility layer that completes the sovereign edge. It rests on three principles:
- Location-agnostic. Whether the machine sits in a garage, a data centre or a cloud region is irrelevant. The architecture cares only about reliability — can it survive failure and recover automatically?
- Sovereign by design. No vendor lock-in. Provider-agnostic abstractions, so moving from one provider to another — or on-premises — does not change the architecture.
- Cattle, with one honest exception. Machine death should be a five-minute inconvenience, not a disaster. For stateful data the honest description is narrower: this is a recoverable pet — 5–30 minutes of recovery, not true cattle. Saying so is the point; the alternative is claiming a property the architecture does not have.
The Architecture: A Four-Tier Storage Hierarchy
Each tier solves a different problem, and the boundary between them is the whole design.
| Tier | What it is | What it is for |
|---|---|---|
| 1 · Edge | Globally distributed, HTTP-accessible, no filesystem semantics | The stateless application |
| 2 · Network-attached block | Detachable block storage (a cloud provider's volumes) | The data. Survives machine death; POSIX semantics; ~0.5–1 ms |
| 3 · Local SSD cache | bcache or lvmcache on local NVMe, transparent to the application | Performance. Local SSD is roughly an order of magnitude faster than the network volume |
| 4 · Object backup | Encrypted, deduplicated backups to object storage | Recovery from volume loss; point-in-time restore |
The critical design decision is the order. Cache sits between the application and the network volume — application → cache → (miss) → volume — never application → volume → cache. Get that backwards and the cache stops being a cache.
The Five Load-Bearing Decisions
Each is a trade-off made deliberately, and each has a failure mode attached.
1 · Durability: detachable volumes, disposable system disk
Problem: local disk dies with the machine, and then the data is a pet. Decision: network-attached volumes hold the data; the system disk holds only the OS and is reproducible from a script.
- Machine death (≈5 minutes): provision a new machine, detach the volume from the dead one, attach it to the new one, start containers.
- Volume loss (≈30 minutes): provision a new machine, create a fresh volume, restore from the object backup, start containers.
The system disk is disposable; the data volume is precious. Quarterly recovery drills are what keep that claim true.
2 · Performance: workload-aware caching, with the safe mode as the default
Local SSD cache in front of the network volume, configured per workload. The applications and kernel both cache, and they are not interchangeable:
| Layer | Tool | Purpose |
|---|---|---|
| Application | Database buffer pool | Pages the database manages itself |
| Application | Redis | Shared in-memory cache |
| Kernel | bcache | Block-level cache in front of the network volume |
bcache's mode is a data-loss decision, not a tuning knob:
| Mode | Behaviour | Risk |
|---|---|---|
| writethrough | Writes go to cache and backing volume together | Default. No data loss if the cache device dies; slower writes |
| writeback | Writes land in cache first, flushed later | Dangerous. Cache failure means data loss — only with a UPS, monitoring, and an explicit decision to accept it |
| writearound | Writes bypass the cache entirely | Safe; cache serves reads only. Good for write-heavy workloads |
3 · The agent: an orchestrator, not an inventor
A small agent on each machine configures workloads and reports health. Deliberately, it is deterministic code using proven tools — not an AI runtime and not a homegrown algorithm engine, because an operator has to be able to explain its behaviour at 3 a.m.
- Always on: health reporting, heartbeat to the control plane, service discovery, workload-profile parsing.
- Toggleable: workload-specific cache configuration, backup orchestration, restarting failed services, network monitoring.
- Explicitly not in scope: custom caching algorithms, predictive prefetching, cache federation, model-based tuning, network-level optimisation (that belongs to the kernel and the tunnel, not the agent).
4 · Sovereignty: one interface over every provider
Providers have different APIs. Hardcode one and the architecture is locked to it. All machine and volume operations go through a small provider-abstraction layer, so no provisioning script contains provider-specific code:
# The shape of the abstraction (illustrative)
create_volume() {
case "$CLOUD_PROVIDER" in
digitalocean) doctl compute volume create "$1" --size "$2" --region "$3" ;;
aws) aws ec2 create-volume --size "$2" --region "$3" --volume-type gp3 ;;
gcp) gcloud compute disks create "$1" --size="$2GB" --zone="$3" ;;
esac
}
The rule is simple to state and easy to break: no provider-specific call outside this layer. Sovereignty here is not a policy statement; it is the property that you can leave.
5 · The HTTP-only boundary: the edge never sees storage
The edge layer never knows block storage exists. Edge and node communicate over HTTP through a secure tunnel:
Edge worker Stateful node
│ HTTP: POST /api/db/query │
│ ─────────────────────────────────►│ internally: local SSD cache,
│ │ network volume, containers
│ HTTP: 200 OK { results: […] } │
│ ◄─────────────────────────────────│
│ (the worker knows nothing │
│ about storage) │
Separation of concerns, in one boundary: the edge is stateless and global, the node is stateful and regional, and neither knows the other's internals. Cross that boundary with a storage detail and you have started rebuilding the cloud you left.
Where AI Fits — and Where It Must Not
This architecture is deliberately unexciting about AI, and that is the position worth stating in a collection about AI-assisted work.
AI accelerates the building, not the running. The agent's runtime is deterministic on purpose; generated code is held to the same review and gate as anything else. The failure mode to avoid is an agent whose behaviour at 3 a.m. no operator can explain — an "inventor" where an orchestrator was needed.
The collection's thesis is load-bearing here. AI amplifies whatever structure it is given: inside clear boundaries it produces more of that clarity, and inside a tangle it produces more tangle. A sovereign fleet is what the fundamentals look like when the stakes are physical — a machine that fails, a volume that detaches, a restore that either works or does not. And the honest limits below are the part that makes delegating any of it defensible.
What This Is, and What It Is Not
This is: a reference architecture for stateful workloads on sovereign infrastructure; single-node per workload (not HA, not distributed); a recoverable pet (5–30 minute recovery); regional, because volumes are zone-bound; appropriate where 99.9% availability is enough.
This is not: a replacement for a managed database when you need a 99.99% SLA; multi-node high availability (no clustering, no leader election); zero-RPO (recovery point is the last backup, not the last transaction); serverless — you are managing the machine.
Use it when you need POSIX semantics, you want sovereignty from cloud vendors, you can tolerate a 5–30 minute recovery window, and you have the operational bandwidth to run infrastructure. Do not use it when any of those is false — the architecture is honest about that trade, and so should you be before adopting it.
The Cost Question
What a node costs is a fair question, and it deserves a precise answer rather than a confident one.
| Component | Planning figure |
|---|---|
| Droplet, 1 vCPU / 1 GB (Sydney) | $6 / month |
| Block volume, 100 GB | $10 / month |
| Object backup storage (~20 GB compressed) | $0.30 / month |
| Object backup egress | $0 (zero-egress design) |
| Marginal cost per machine | ~$16.30 / month |
These are the author's own working figures — a planning model, not an audited invoice. They are marginal cost per node, and the total cost of ownership is larger: the control plane must run somewhere, the node's egress to the edge is not included, operational time is the real cost, and system-disk and configuration backups are separate. What is not measured yet: p95 latency under production-shaped load, per-node support labour, egress and transit, and the hardware replacement rate. Those are the lines that would turn the model into evidence, and none of them is going to be invented here.
The Open Problems
Two parts of this architecture are design, not implementation, and it would be dishonest to present them as finished.
Database backup consistency. A file-level snapshot of a live database can capture an inconsistent state. The options under evaluation are the database's own continuous archiving plus a snapshot filesystem — the shape a PostgreSQL workload would use: base backup plus write-ahead-log archiving for point-in-time recovery, a filesystem snapshot for generic block volumes. Target RPO is one hour with continuous archiving (24 with nightly snapshots); target RTO is five minutes for a volume reattach, up to thirty for a full restore.
The security model. The node may run untrusted customer workloads, so compromise must be contained to that node — no lateral movement to the control plane, no access to other nodes' volumes. The layers are defined and the controls are not yet chosen: encryption at rest with its key management, how the node receives database credentials, how the agent authenticates to the control plane, container isolation level, and least-privilege provider tokens. A security review is a precondition of production use, not a follow-up.
The Orchestrator's Takeaway
Run stateful workloads where you control the exit — data on a detachable volume, the system disk disposable, the cache in front of the volume, and the edge kept ignorant of storage. Then say what it really is: a recoverable pet with a 5–30 minute recovery window, not cattle. The honesty is what makes it safe to rely on.
Where to Read Next
- The Second Life of Computers — the case for the fleet, and the pattern this node completes.
- The Edge-Native Stack — what the edge layer relocates back into the application.
- The Sovereign Fleet series (in preparation) documents the build layer by layer, and cites this architecture rather than repeating it.
For the full method, see The DevOps Engineer's Guide to Effective AI Usage.