Skip to content

Systems architect: designing for the second year

A systems architect decides where the boundaries go, and writes down what each boundary costs. These are the principles I apply, the trade-offs I accept, and the checks I run before a design is allowed to become code.

What I mean by systems architect

Architecture is the set of decisions that are expensive to change. My job is to find those decisions early, make them deliberately, and keep everything else cheap to change. In practice that means owning the whole path of a request: browser or device, API, domain rules, database, background workers, network, deployment and recovery.

Tip: before choosing a technology, list the five things that would be most painful to change in two years (tenancy model, money type, identifier scheme, sync model, deployment target). Design those first; pick frameworks last.

Start from constraints and quality attributes

I begin with measurable constraints rather than diagrams: how many concurrent users, what a minute of downtime costs, what must work offline, what an auditor will ask for, who operates the system. Each becomes a quality attribute with a number or a testable statement.

  • Availability: a restaurant counter must keep billing when the internet drops. That single sentence produced the offline-first design in OrderRestro and KinetiRx.
  • Consistency: an invoice must change stock, ledger, tax and numbering atomically, so those live in one transaction boundary in Rechvix.
  • Auditability: every state change must be explainable later, which dictates append-only logs and audit rows written inside the same transaction.
  • Operability: if one person must run it, it needs one deployable unit, one place to read logs and a restore that has been tested.
  • Changeability: modules with narrow interfaces so a new domain can be added without touching the others.

Monolith, modular monolith or services

I default to a modular monolith: one deployable process, strict internal module boundaries, and interfaces between modules that could become network calls later. It keeps transactions simple and operations cheap, and it keeps the option of extraction open. I split a service out only when it has a different scaling profile, a different failure domain, a different release cadence or a different owner.

Tip: enforce module boundaries with tooling, not goodwill. In Go, keep domain packages free of HTTP and database imports and check it in CI; in TypeScript, use lint rules for import boundaries. A boundary that is not enforced is a suggestion.

Microservices solve organisational scaling problems. For a small team they mostly add distributed transactions, version skew and an on-call burden. I treat each extra network hop as a cost that needs a named benefit.

Tenancy and isolation

Multi-tenancy is decided in the first migration. Every tenant-owned table carries an organisation identifier, enforced by PostgreSQL row-level security and again in the repository layer, so one forgotten WHERE clause cannot leak data. Branch and warehouse scoping sit on top of that as permissions. Isolation levels, from lowest to highest cost:

  • Shared schema with row-level security: cheapest, best for many small tenants, needs discipline in every query.
  • Schema per tenant: stronger separation and per-tenant migrations, more operational overhead.
  • Database or deployment per tenant: strongest isolation and simplest compliance story, highest cost; right for large or regulated customers.

The reasoning, and why row-level security alone is not enough, is in RLS is not your only tenant boundary.

Data architecture and consistency

  • Make the database the system of record and put invariants in it: foreign keys, check constraints, unique indexes, balanced journals.
  • Store money as a decimal or integer minor units, never a float. Keep a currency with every amount.
  • Model stock and balances as an append-only movement ledger and derive the balance, so history cannot be silently overwritten.
  • Use idempotency keys on every write that a client might retry, so a flaky network produces one invoice, not two.
  • Use the transactional outbox pattern to publish events: write the event in the same transaction as the change, then deliver it from a worker.
  • Prefer UUIDv7 or other time-ordered identifiers for primary keys: they are index-friendly and can be generated offline.

Tip: for each table, write the invariant it protects in one sentence. If you cannot, the table is probably two concepts. If the invariant is only enforced in application code, ask what happens when a second service or a manual script writes to it.

I use event sourcing selectively, for domains where the history is the product (stock, ledgers), and plain state tables elsewhere. See what is event sourcing.

Offline-first and synchronisation

When the network is unreliable, the local device is the primary writer and the server reconciles. That forces explicit decisions that online-only systems can postpone: client-generated identifiers, a durable local queue of operations, conflict rules per entity, and server-authoritative handling of money and stock. I choose the simplest sync that is correct: one-way server-sent events when updates only flow one way, WebSockets when the kitchen display needs instant push, and an operation log with replay when devices write while disconnected.

Designing for failure

  • Timeouts on every outbound call, with retries that use exponential backoff and jitter, and a circuit breaker around anything that can stay down.
  • Put slow or external work (e-Invoice APIs, notifications, report generation) behind a queue and a worker, never inside a request.
  • Make every consumer idempotent; assume at-least-once delivery.
  • Degrade by feature, not by system: if tax-portal submission is down, invoicing continues and submission is retried.
  • Write down the failure mode of each dependency: what the user sees, what is retried, what pages someone.

Tip: run a pre-mortem. Assume the system failed in production six months from now and list why. The top three reasons become design requirements or monitoring alerts.

Security by structure

Security is a property of the architecture, not a layer added at the end. I design in depth: authentication at the edge, authorisation as granular permissions checked in the service layer, tenant isolation in the database, secrets outside the code, least-privilege database roles (the application role cannot bypass row-level security), and audit logs the application cannot edit. See Cybersecurity.

Observability and operations

  • Structured logs with a request identifier that follows a request across API, worker and database.
  • Metrics for the four golden signals: latency, traffic, errors, saturation. Alert on symptoms users feel, not on every cause.
  • Health and readiness endpoints that check real dependencies.
  • Backups with a documented, rehearsed restore, and stated recovery point and recovery time objectives.
  • Migrations that are backward compatible for one release (expand, migrate, contract), so deploys and rollbacks never need downtime.

The operational side is covered in DevOps and Infrastructure.

Performance and capacity

I measure before optimising. The usual order is: fix the query and its index, remove the N+1, add caching only with a clear invalidation rule, then scale out. Most business systems are limited by database access patterns, not by CPU, so I read query plans with EXPLAIN ANALYZE, keep transactions short, and use connection pooling deliberately. A load test against realistic data volumes belongs before launch, not after the first complaint.

Recording decisions

Every significant choice gets a short architecture decision record: context, decision, alternatives considered and consequences. The consequences section matters most, because it lists what was given up and what contains the damage. Months later it lets someone else change the decision safely.

# ADR-007: Stock as an append-only movement ledger

Status: accepted
Context: balances drifted when documents were edited after posting.
Decision: store movements only; balances are a rebuildable projection.
Alternatives: mutable balance column (rejected: no history).
Consequences: +audit trail, +rebuild on demand; -extra read cost,
  mitigated by a projection table refreshed in the same transaction.

A checklist I run before build

  • Are the five hardest-to-change decisions made and written down?
  • Does every invariant live in the database, not only in code?
  • Is tenant isolation enforced at two independent layers?
  • Is every external call time-limited, retried safely and idempotent?
  • What happens when the network, the database, or a third party is down?
  • Can one person deploy, observe, back up and restore it?
  • Is the restore tested, and are the changelog and known gaps honest?

Have something in mind?

Let’s build something useful.

Tell me about the idea, product, or workflow you’re working through.

Tap to say hello