Platform & Infrastructure

2026-09-26

The Edge-Native Stack: What Moves Back into the Application

Edge-native platforms do not delete infrastructure — they relocate it. What you now own, what the platform still handles, and what an edge database actually promises.


Platform & Infrastructure · manifesto · ~14 min read · companion: AI Parallel Development and the Verification Bottleneck

Every few years a platform arrives with the same promise: the infrastructure problem is over. Deploy a function to a global network, bind it to a database, an object store, a queue, and a model, and the operational surface that once justified a platform team, an SRE rotation, and a network team simply disappears.

That promise is half true, and the false half is the one that decides whether you sleep. The infrastructure does not disappear. It moves. Work that a load balancer, an operations rotation, and a network appliance used to do on your behalf becomes code you write, state you own, and failures you are on call for.

The edge-native stack is a relocation of responsibility, not a deletion of it. This post is a map of where each piece goes, and an honest account of the pieces that land on your desk.

Two terms carry the argument, so define them before they are used.

A binding is a declaration in configuration that becomes a direct, in-runtime call to a managed service. You write the call; the platform supplies the address, the credentials, the connection, and the placement. A binding is not a network client. There is no endpoint to configure, no connection pool to size, no service discovery to run, and no access policy to hand-write between the two.

An edge function is a unit of compute that the platform deploys to every node in its network, so the same code answers a request in Paris and in São Paulo. Anycast addressing — one IP address announced from many locations — lets the internet's routing protocol deliver each user to the nearest node without a global load balancer in the path.

The gain from those two primitives is real and large. Everything after this paragraph is the bill.

The Infrastructure Problem Does Not Disappear. It Moves.

For two decades the central challenge of building software was infrastructure. You provisioned servers, configured networks, installed databases, tuned connection pools, orchestrated containers, and scaled clusters. That difficulty is why large engineering organisations exist: the platform team, the SRE team, the security team, and the network team were all answers to a genuinely hard operational problem.

Provisioning really is ending. Nobody starting today should stand up a subnet to ship a product.

But the difficulty was never only in the provisioning. It was in the responsibilities that provisioning implied — keeping a backend healthy, failing over when it is not, filtering hostile traffic, holding state somewhere durable, and knowing whether any of it is true. A managed edge platform absorbs some of those responsibilities, and hands the rest back to you in a different form.

ResponsibilityTraditional homeEdge-native home
Global traffic distributionLoad balancer applianceThe platform's network
Origin health checkingLoad balancer health checksCode in your edge function
Failover decisionTarget groups and policiesCode in your edge function
WAF rulesWAF configurationRequest inspection in your function
Session affinityLoad balancer stickinessState you place in a store
Backups and restoreManaged database consoleWhatever you build
Schema migrationsManaged database toolingWhatever you build
ObservabilityAn agent on the hostStructured output from your code

Read the right-hand column as a job description. Some of it is easier than the left. None of it is gone.

That table is the thesis of this post. The sections that follow take the rows apart one at a time, then turn to the two questions a migration plan usually skips: what consistency an edge database actually promises, and what any of this costs.

What the Edge Stack Actually Removes

The edge-native architecture does not require every service to be replaced. It requires every service to be reachable from the edge, ideally through a binding that respects the runtime's constraints. Three tiers are worth separating.

Edge-native services run in the same runtime and the same network as the function that calls them. A relational store, an object store, a key-value store, a queue, and a model endpoint are reached through bindings rather than over the public internet. They are the default for new work, because the operational surface is a declaration rather than a deployment.

Managed services with edge access live in a regional data centre and are reached from the edge through a pooling or caching layer, or through an HTTP-based driver. They offer capabilities that edge-native services do not yet match: full PostgreSQL feature sets, mature tooling, extensions, and larger datasets.

Specialised services — GPU clusters, data warehouses, compliance-bound workloads — stay where they are and are reached by API call. They are still part of the architecture; they are simply not edge-native.

TierWhere it runsHow it is reachedUse it when
Edge-nativeSame network as the functionBindingThe need is ordinary and latency-sensitive
Managed with edge accessRegional data centrePooling layer or HTTP driverYou need features the edge does not have
SpecialisedIts own infrastructureAPI callThe workload cannot move

The design principle is narrow and worth stating plainly: use edge-native services where they suffice, and connect to managed services where they provide something the edge does not. The same command vocabulary and the same data contract can sit over either, which is what keeps the choice reversible.

One consequence is easy to miss. When a function reads from a bound store, writes to bound object storage, and publishes to a bound queue, it is not making three network requests to three vendors. Those hops — the ones that made microservice architectures expensive to operate — are gone. That is the real removal, and it is a removal of hops, not of responsibilities.

The Edge Router: The Network Becomes Programmable Logic

The traditional stack needed a dedicated network layer: a load balancer to distribute traffic, an ingress controller to route requests to services, a WAF to filter hostile traffic, and a firewall to control access. Each was a device or a managed service with its own configuration, its own failure modes, and its own review process.

The edge-native replacement is one idea: the edge function as a programmable request router. Instead of configuring a load balancer, you write a function that inspects the request and decides what to do with it. The platform distributes that function to every node.

A minimal router looks like this in outline:

export default {
  async fetch(request, env) {
    const url = new URL(request.url);

    // Security: reject unauthenticated admin traffic
    if (url.pathname.startsWith('/admin') && !isAuthorised(request)) {
      return new Response('Forbidden', { status: 403 });
    }

    // Routing: API and static assets go to different backends
    if (url.pathname.startsWith('/api/')) {
      return fetch(new Request(API_ORIGIN + url.pathname, request));
    }
    if (url.pathname.startsWith('/static/')) {
      return env.ASSETS.fetch(request);
    }

    const target = await pickHealthyOrigin();
    return fetch(new Request(target + url.pathname, request));
  }
};

That one function replaces the load balancer's routing, the ingress controller's path mapping, and the WAF's request filtering. The difference is not that it is smaller. It is that it is code: reviewable, testable, versioned, and diffable in the same repository as everything else.

Traditional componentWhat it didWhat replaces it
Load balancerDistributed traffic across originsTarget selection inside the router
Ingress controllerMapped paths to servicesURL inspection and forwarding
WAFFiltered hostile HTTP trafficRequest inspection in the function
FirewallRestricted network accessPlatform sandbox isolation
VPC and subnetsSegmented the networkNot needed in the same form
TLS terminationDecrypted HTTPS at the edgeThe platform, automatically

What moves back into the application is the logic: which paths are public, which headers are trusted, what a canary looks like, when a request is suspicious. Configuration becomes code, and code becomes something a reviewer can refuse.

High Availability: The Platform Takes Global Failover, You Take Origin Failover

This is the question a sceptical reader should ask first. If the load balancer is gone, what happens when something fails?

The honest answer is that a traditional load balancer did two different jobs, and they now live in two different places.

Global failover moves to the platform. Because the function is deployed to every node, there is no single appliance to fail. Anycast routing sends each user to the nearest healthy node, and if a node or region is unavailable, the network routes around it. DDoS absorption, TLS certificates, and automatic scaling come with the same model. You do not configure them, and in the ordinary case you do not see them.

Origin failover does not move. The platform has no idea whether the backend your function forwards to is healthy. That responsibility moves into your router, as code:

const ORIGINS = [PRIMARY_ORIGIN, SECONDARY_ORIGIN, TERTIARY_ORIGIN];

async function healthyOrigins() {
  const checks = await Promise.all(ORIGINS.map(async (origin) => {
    try {
      const res = await fetch(origin + '/health', { signal: AbortSignal.timeout(2000) });
      return res.ok;
    } catch {
      return false;
    }
  }));
  return ORIGINS.filter((_, i) => checks[i]);
}

A traditional load balancer's health check does the same thing. The difference is where it lives and who owns it. Yours can query a database, check a dependency, apply a weighted or regional policy, serve a degraded read-only response, and ramp traffic back gradually after recovery — logic that a fixed target group cannot express.

CapabilityLoad balancer applianceRouter function
Health checkFixed HTTP or TCP probeAny logic you can write
Failover decisionPreconfigured target groupWeighted, canary, or regional
Degraded modeStatic error pageCached or read-only response
RecoveryRe-add when the probe passesGradual ramp with validation
Multi-provider failoverManual or DNS-basedPer-request, in code

The same trade recurs across the whole stack: more control, more responsibility, and fewer appliances. The engineer loses the comfort of a device that "just works". The engineer gains failover logic that can be tested, reviewed, and changed on a Tuesday afternoon.

"Edge SQL" Is Not One Category

The phrase edge SQL covers at least two architectures with different promises, and the difference decides whether your application is correct.

The first is a per-region, embedded-style database — a SQLite-compatible engine placed close to the user, sometimes as a single node in a region, sometimes with read replicas. Writes usually go to one primary. Read-after-write holds on the primary but not necessarily on a replica, and not necessarily for the next request if that lands in another region. Conflict handling is either absent or explicitly the application's concern.

The second is a globally replicated database that synchronises writes across regions. Here the interesting properties are not latency but the promises: the read-after-write window, the conflict window and its resolution rule, and the failover semantics when the primary region is unreachable.

Those are not implementation details. They are the contract your application is written against, and they differ from provider to provider.

QuestionPer-region or embeddedGlobally replicated
Read-after-writeOn the primary; replicas lagA stated window; verify it
Concurrent writesSingle writer, or app-managedA conflict rule, or app-managed
FailoverA new primary; what is lost?Stated promotion semantics
Latency storyLocal reads, possibly remote writesLocal reads and writes
Best fitPer-user or per-tenant dataShared global datasets

The question to put to any provider is not "is it fast?" It is: what does a write promise, to whom, for how long, and what happens to a write in flight when a region fails? Ask for the failover semantics in writing. If the answer is a diagram rather than a sentence, treat the sentence as unverified.

Consistency also has a second-order consequence. A per-region database with one location per user is straightforward to restore, because each region's data is small and self-contained. A globally replicated database must be restored consistently across regions, which is a harder problem and a drill you should run before you need it.

The practical split is common in production systems: an edge-native store for session state and per-tenant data, a managed relational database for the shared core, both called from the same function. That is not a compromise. It is choosing the promise each dataset needs.

The Relocation Ledger: What You Now Own

Here is the part a migration plan tends to leave out. Moving compute to the edge removes a set of appliances and adds a set of responsibilities. The responsibilities are not harder than what they replaced, but they are yours, and they are invisible until they fail.

Backups. A managed database console used to give you a restore button. With an edge-native store you must decide what is backed up, where the copy lives, how long it is retained, and how a restore is verified. A backup nobody has restored is a hope, not a recovery plan.

Schema migrations. The migration tooling that came with a managed relational database may not exist for an edge-native store. Migrations become a script you own, a version you track, and a check that proves the schema on every region matches what the code expects.

WAF logic. Rules that arrived as a curated ruleset become request-inspection code. This is a genuine improvement in precision and a genuine increase in surface area: you now own the false positives as well as the attacks.

Health checks and failover. As above, these are code. Code has tests; code has bugs. A failover path that has never been exercised under load is a hypothesis.

Session state. A load balancer could pin a user to a server. An edge function cannot, because the next request may land anywhere. State moves to a binding or to a signed client token, and the engineer must think about state explicitly rather than relying on stickiness.

Observability. There is no host on which to install an agent. Structured output from your code, correlated across requests and regions, is the only signal you will have. If it is not emitted deliberately, it does not exist.

None of this is an argument against the edge-native model. The hops really are gone, the provisioning really is gone, and the operational surface really is smaller in the ordinary case. It is an argument against the belief that the surface reaches zero. The infrastructure problem did not disappear; it took up residence in your application, where your tests and your on-call rotation are what keep it true.

How that relocated ownership interacts with AI-generated code — where the verification burden lands, and why it becomes the binding constraint — is the subject of the companion post, AI Parallel Development and the Verification Bottleneck.

The Cost Model We Have Not Built Yet

The last row of the ledger is the one no architecture diagram shows: what it costs.

The sibling post in this shelf, The Second Life of Computers, takes the same position: it lays out what a fleet costs to run and what has to be measured — a working model, not an audited invoice. This post goes less far than that. We have not produced a cost model for the edge-native control plane described here, and we are not going to invent one. What follows is the measurement plan, not the measurement.

A useful cost model needs five things, each of which is a line an operator can produce from a bill and a benchmark:

What to measureWhy it mattersWhere it comes from
Monthly control-plane line itemsThe fixed cost of the platform, before trafficProvider invoice, itemised
Egress and transitEdge architectures move data between regions and providers; this is where surprise bills liveProvider invoice, per region
Request and operation countsThe unit that scales with successInvoice plus a workload profile
p95 and p99 latency, per regionThe number a user experiences; averages hide the tailA benchmark on production-shaped traffic
Restore and failover drill timeThe real cost of the responsibilities aboveA rehearsal, timed

Two habits keep the model honest. First, measure p95 rather than mean latency, because the tail is what users feel and it is what the failover thresholds inside your router should be based on. Second, price the drill: the time to restore a database and fail over an origin is a real operational cost, and it appears on no invoice.

Until those numbers exist, treat any statement about edge economics — including ours — as a hypothesis. The architecture can be argued on its own terms. Its price cannot be argued into existence; it has to be measured.

The Orchestrator's Takeaway

Adopt the edge-native stack for what it genuinely removes — provisioning, network hops, appliances — and staff for what it moves into your code: health checks, failover, WAF logic, state, migrations, backups, and observability. Budget the relocation, not just the migration. The teams that struggle are not the ones that moved compute to the edge; they are the ones that did not notice the responsibilities arriving with it.

Appendix: Glossary

TermMeaning
BindingA declaration in configuration that becomes a direct, in-runtime call to a managed service
Edge functionA unit of compute deployed to every node of a provider's network
Edge-native serviceA managed service built for the edge runtime and reached by binding
Pooling layerA managed proxy that caches connections and results between edge functions and regional databases
AnycastAnnouncing one IP address from many locations so routing sends each user to the nearest node
Edge routerAn edge function that performs routing, security, and failover decisions
Read-after-writeThe guarantee that a read following a write observes that write
Failover semanticsWhat a system promises about data and traffic when a primary becomes unavailable

Platform Engineering
Software Architecture
Infrastructure