Skip to content
Working reference · English & Spanish

Software architecture atlas

The map of decisions that get made before the first line of code: which patterns exist, where each one lives, what it costs and when not to use it. It is not a list of things to add — it is the vocabulary for arguing about a design before implementing it.

129

Patterns and principles

7

Families

6

Stations on the path

10

Questions before you code

What it is
A catalogue of the patterns that exist, what each one solves, what it costs and when not to use it.
What it is for
So that the design decision is explicit before the code, instead of being made by omission.
How to read it
Find the station where it hurts, go into the family, read the card. Then answer the ten questions.
What it is not
A list of things to add. Half the atlas exists to justify not using the other half.
The map

Where each pattern lives

Almost everything that enters a system travels the same six stations. Every pattern belongs to one of them: if you know which station hurts, you know which family to look in.

The path of an external eventfrom a third party’s webhook to the data already indexedpushesacceptshands offwritescalls out01Originwebhook (push)polling / long pollingCDC from the databasebatch load02EdgeAPI gatewayrate limitingauthentication / signatureschema validationdedupe by event idload shedding03Bufferdurable queuetransactional outboxdead letter queuebackpressureclaim check04Workworker poolretry + jittersaga / compensationdurable executionfan-out / fan-inreconciliation05Statedatabase per serviceCQRS / projectioncache-aside06Egresstimeoutcircuit breakerbulkheadoutbound rate limitcut across all six: observability (trace, metric, log)feature flags · idempotence · versioned contracts
The stations are not optional: an event crosses all of them, even the empty ones. An empty station is not a station that does not exist — it is a decision nobody made, and it only shows up under load.
The ladder

Coupling: every rung buys durability and charges operations

The “should we make it async?” argument is almost never binary. There are six rungs, and the design work is picking which one — not climbing to the top because it sounds modern.

↑ what survives a crashwhat it costs to operate →THE DURABILITY LINEto the left, the work lives in memory and dies with the process01Directsynchronous call02+ timeout andretry with jitter03+ circuit breaker04Durable work queue05Publishedevent (pub / sub)06Durable workflow
The dotted line is the only boundary that really matters: to its left the work lives in one process’s memory and disappears with a restart, an OOM or a redeploy. Climbing one rung too far costs real operations; staying below when the business cannot tolerate losing the event costs an incident.
RungWhat survivesWhat you payWhen it is the right answer
01Direct synchronous callNothing. If the other side goes down, you go down with it.Nothing. It is the simplest thing there is.A cheap read, the caller needs the answer now, and losing the operation does not hurt.
02+ timeout and retry with jitterA hiccup of a few seconds on the other side.p95 latency and extra load. Without jitter, you create a stampede.A normally healthy dependency with transient failures. It demands idempotence on the other side.
03+ circuit breakerA long outage on the other side without it dragging you down.One more state to tune, to observe and to explain.You call a third party that can degrade: a search, payments or shipping provider.
04Durable work queueA process restart and the traffic peak.A broker to operate, a DLQ to watch, and ordering stops being guaranteed.The work cannot be lost and the caller does not need the result. The typical case of a webhook intake.
05Published event (pub / sub)On top of that: the emitter stops knowing who listens, so adding consumers never touches it.Eventual consistency and versioned contracts, taken seriously.Several parties interested in the same fact, and none of them should hold up the producer.
06Durable workflowThe whole process with its state, for hours or days, across restarts.A cluster, dedicated workers, workflow versioning, a learning curve.Long processes, with compensable steps and per-step retries. Where agent orchestration lives.
Family 01

Shape of the system — where do I cut?

Decides what ships together and what ships apart. It is the most expensive decision in the atlas to reverse: cutting it wrong costs years, not sprints.

Monolith

Everything in one deployment. One transaction, one log, one deploy: the right answer until the domain or the team stops fitting in a single head.

It gets abandoned too early, and almost always out of fashion.

Modular monolith

One deployment, but with hard internal boundaries between modules: nobody imports the neighbour’s insides.

If the boundaries are not enforced by tooling, they dissolve within six months.

Microservices

One service per business capability, with its own database and its own deployment cycle.

It pays for the network, for distributed observability and for eventual consistency. Without teams to operate them, it is cost with no benefit.

Bounded context (DDD)

The edge of the model is the edge of the service: an “order” in sales and an “order” in logistics are two different things, and that is fine.

Ports and adapters

The domain defines the interfaces; infrastructure implements them. Swapping Postgres for something else never touches the logic.

On a small CRUD it adds three layers for nothing.

Clean / Onion

Dependencies always point inward. The core does not know a framework exists.

Layers (n-tier)

Presentation, application, data. Simple, and everybody already knows it.

A single business change touches all three layers: cohesion ends up spread thin.

Vertical slice

Organise by whole use case instead of by technical layer. A feature lives in one folder.

Microkernel / plugin

A minimal core plus loadable extensions. It earns its keep when the variability is per client or per channel.

Pipes and filters

Chained stages that transform a stream, each one ignorant of the rest. A transcoding or indexing pipeline is exactly this.

Serverless / FaaS

The unit of deployment is the function, and it scales to zero.

Cold starts, time limits and a bill that surprises you under sustained traffic.

Cell-based architecture

Complete copies of the stack per group of customers. The blast radius of a failure is the cell, not the whole system.

The most expensive item on the list: every cell is its own infrastructure, deployment and monitoring.

Micro-frontends

The front end split by domain, each part with its own team and its own deploy.

It duplicates the bundle and breaks visual consistency unless there is a design system.

SOA with a bus

Large services talking through a central bus that also transforms and routes.

The bus becomes the new monolith, and that is where all the logic nobody wanted to place ends up living.

Space-based

A replicated shared-memory grid, with no database on the critical path. For extreme, predictable peaks.

Family 02

How they talk to each other — who waits for whom?

Once the system is cut, every conversation needs a shape. Almost every distributed incident is born here: somebody was waiting for an answer that never came.

Request / response

REST or gRPC: the caller blocks until the answer arrives. Easy to reason about and easy to debug.

Chaining six synchronous hops multiplies the odds of failure — it does not average them.

Work queue

Point to point: a producer drops the task, one consumer out of a pool picks it up. It absorbs peaks by design.

Publish / subscribe

The emitter announces a fact and does not know who is listening. Adding consumers never touches the producer.

Nobody knows who depends on what until you break the contract.

Event log

The event persists and can be re-read: new consumers can start from the beginning, and a slow one does not block the rest.

Orchestration

One component dictates the order of the steps. Easy to see, easy to audit, easy to fix.

The orchestrator accumulates everybody else’s business logic.

Choreography

Every service reacts to events, with no central brain. Maximum decoupling.

Nobody can draw the whole flow, and debugging turns into archaeology.

Saga

A long business transaction split into local steps, each one with its own compensation to undo it.

Compensating is not rolling back: the world already saw the intermediate state.

API gateway

A single front door holding auth, rate limiting, routing and edge observability.

Backend for frontend

One backend per client type, assembling exactly the response that screen needs.

Sidecar

A companion process that supplies cross-cutting capabilities (TLS, metrics, proxying) without touching the service’s code.

Ambassador

The sidecar that speaks outward on your behalf: it concentrates the retries, timeouts and breakers of outbound calls.

Service mesh

Sidecars plus a control plane: mTLS, traffic policy and retries declared outside the code.

A whole piece of infrastructure. You do not adopt it for three services.

Anti-corruption layer

A translator at the edge that converts the foreign model into yours, so the vendor’s or the ERP’s vocabulary never leaks into the domain.

Strangler fig

The new thing eats routes off the old one behind the same edge, until the old one has no traffic left and gets switched off.

Webhooks

The third party pushes the fact to you the moment it happens. Cheap and real time.

You do not control the rate. With no limit and no dedupe, the third party decides your capacity.

Change data capture

Changes from the database log get published as events, without touching the code that writes.

It couples consumers to the physical schema of the table.

Claim check

A reference travels in the message, not the heavy payload: the content stays in a store and the consumer fetches it if it needs it.

Contract testing

The contract between two services is verified in CI instead of being trusted. It breaks the PR, not production.

Gateway aggregation

The edge joins several internal calls into one response, so the client does not make six round trips.

Polling and long polling

When there is no webhook: you set the pace, and that alone is free flow control.

Latency against cost, and you have to keep a cursor that never gets lost.

Family 03

Data — who owns it, and what may be lost?

This is where you decide which truth is single and which truth is allowed to lag. Most of the strange production bugs are an implicit answer to this question.

Database per service

A single service writes a table; everyone else asks over an API. The schema stops being a public contract.

Shared database

Several services reading and writing the same schema. Blazing fast today.

Any migration becomes a coordinated deployment of everybody. It is the distributed monolith through the back door.

CQRS

The model you write with and the model you read with are different, and they evolve separately.

Event sourcing

State is not stored: it is derived from the sequence of facts. The complete, auditable history comes free.

Versioning old events is forever. You do not adopt it “just in case”.

Materialised view

A precomputed, denormalised read that answers the expensive query cheaply.

Every view needs a written answer to “who rebuilds it, and when?”.

Cache-aside

I read the cache; if it is not there I go to the source and fill it. The default caching pattern.

Invalidation is the problem, not the filling.

Write-through / write-behind

The write goes through the cache: synchronous (consistent and slow) or deferred (fast and at risk of loss).

Sharding

Split by key — customer, region — so that no single instance carries everything.

Choosing the key badly is paid for with a data migration, not with a refactor.

Read replicas

Scale reads by separating them from writes.

Replication lag means a user does not see what they have just saved.

Polyglot persistence

Each kind of data in the engine that suits it: relational for transactions, document for flexible, index for search.

Every extra engine is one more backup, one more monitor and one more expertise.

Transactional outbox

The event is written in the same transaction as the state, and a relay publishes it afterwards. It solves “I saved it but never told anyone”.

Idempotent consumer

The receiver keeps a record of the ids it already processed, so receiving the same message twice does nothing the second time.

Idempotency key

The caller names the operation, so retrying it is safe. Without this, rung 02 of the ladder duplicates charges.

Optimistic locking (etag / version)

I write saying which version I read; if it changed, I am rejected. Two concurrent edits stop overwriting each other in silence.

Distributed lock

One process at a time over a resource.

It needs a TTL and an owner: a lock with no expiry is an outage waiting for its turn.

Snapshot

A periodic cut of the state, so you do not re-read ten years of events on every start.

Two-phase commit

An atomic transaction across two systems.

Almost always the wrong answer between services: it blocks, it does not scale and it fails ugly. The alternative is a saga.

Eventual consistency

The data converges; it does not match instantly. It is the right answer every time the business tolerates it — and it tolerates it more often than we think.

Append-only and audit

Nothing gets overwritten: things get appended. Investigating an incident stops depending on somebody having logged the right thing.

Separate analytical plane

Reports are not computed over the operational database. A heavy query stops being able to take the product down.

Family 04

Resilience and flow control — what happens when 100× shows up?

All of these patterns answer the same question from different angles: what do I do with the work I cannot take right now? There are four honest answers — queue it, reject it, degrade it, or fall over. Not deciding is choosing the fourth.

Timeout

Nothing waits forever. It is the cheapest pattern in the atlas and the one most often missing.

The client’s timeout has to be shorter than its caller’s, or the whole chain hangs anyway.

Retry with backoff and jitter

Wait longer each time, with random noise, so that nobody comes back all at once.

Without jitter a thousand clients retry in the same millisecond and you rebuild the outage yourself.

Retry budget

A global ceiling: if more than N % of the traffic is retries, they get cut off. It stops the resilience layer from being the attack.

Circuit breaker

After N failures it opens the circuit and fails fast without touching the network; later it lets one probe through to see whether the other side is back.

Bulkhead

Watertight compartments: pools, queues or pods separated by kind of work, so a noisy neighbour cannot sink the ship.

Rate limiting

How much I let in per unit of time. Token bucket tolerates bursts, leaky bucket smooths, sliding window is the fairest and the most expensive.

Per-tenant quota

The limit applies per customer, not globally. A single one stops being able to eat everybody’s capacity.

A global limit protects the service and does not protect the customers from each other.

Load shedding

When I cannot keep up, I throw the least valuable work overboard on purpose and protect the critical path. Rejecting fast is better service than accepting and dying.

Backpressure

Telling whoever is pushing to ease off, instead of accepting and piling up in memory until the OOM.

Queue-based load levelling

The queue absorbs the peak and the consumer works at a steady pace. It turns a capacity problem into a latency one.

Dead letter queue

Whatever failed N times leaves the line: it is neither lost nor blocking the rest, and it stays visible enough to fix.

Deduplication

The same fact twice does not count twice. Without this, any retry from the provider multiplies your work.

Request coalescing

A thousand identical requests arriving together are resolved with a single call to the origin. It kills the cold-cache stampede.

Graceful degradation

Answering something useful when the ideal is not available: the stale cached value, the unsorted list, the generic response.

Fail closed / fail open

When the check cannot run: do I let it through or block it? Both are valid; what is unacceptable is for it to be an accident.

A credential block that fails open is not blocking anything at all.

Health checks

Liveness says whether to restart me; readiness whether I can take traffic. Confusing the two causes cascading restarts.

Graceful shutdown

Stop accepting, finish what is in flight, close. Without this, every deploy loses the work in progress.

Failover

A replica ready to take over. It is worth exactly as much as the last drill that was actually run.

Chaos engineering

Breaking it on purpose during office hours, so as not to discover it on a Sunday.

Family 05

Work over time — who does it, and what if it dies halfway?

Everything that is not resolved inside the request lives here. One question orders the whole family: if the process restarts right now, does anybody find out, and does anybody finish it?

Consumer pool

Several workers compete for the same queue; scaling is adding workers. The workhorse of asynchronous processing.

Durable execution

The workflow’s state is persisted step by step: it survives a worker restart and carries on where it was, with per-activity retries. Temporal is the well-known implementation.

In-memory background work

Fire the task inside the same web process and answer 200. It looks like a queue and it is not.

It dies with the pod, it has no retry and no visibility. It is the origin of most silent losses of work.

Scheduler / agent / supervisor

Somebody watches whatever was left half done and requeues or compensates it. The safety net under everything else.

Fan-out / fan-in

Split one job into N parallel ones and wait for them all to come back. Careful with multiplying the load on whoever is downstream.

Reconciliation loop

Periodically compare the desired state against the real one and correct the difference. It is the only thing that saves you when every event was lost.

Per-entity debounce

Twenty changes to the same product in a minute collapse into a single reindex.

Measure it before dismissing it: it is common for most reindexes to change nothing, and that is pure work to throw away.

Batch and micro-batch

Group N operations into one call. It improves throughput enormously and makes individual latency worse.

Priority queue

The urgent thing does not wait behind last night’s bulk load.

Without a service floor, low priority never gets processed at all.

Leader election

A single owner among several replicas for the tasks that cannot be duplicated.

Idempotent cron

A scheduled job that can run twice without harm. Because sooner or later it will run twice.

Cursor with checkpoint

Save how far I got, so I can resume without re-reading everything or skipping anything.

Time windows

Aggregate events by fixed or sliding window. The basis of any metric over a continuous stream.

Compensating sweep

A job that looks for whatever was left inconsistent and fixes it. Less elegant than a perfect design and far cheaper.

Family 06

Delivery and operations — how do I turn it on, turn it off and find out?

A pattern you cannot switch off without deploying, or observe without SSH-ing in, is not finished. This family is the non-negotiable half of any design.

Deploy ≠ release

The code arrives switched off and gets turned on separately. It separates technical risk from product risk.

Feature flags

Typed, declared flags: release and experiment with an expiry date, killswitch and operational permanent.

Killswitch

Turn the feature off in seconds without deploying. The first thing anyone looks for in an incident and the last thing anyone implements.

Canary

Send 1 % of the traffic to the new version and watch the metrics before going any further.

Blue-green

Two complete environments and one traffic switch. Going back is instant.

Rolling update

Replace pods a few at a time. It demands that two versions coexist speaking the same contract.

Shadow traffic

Copy real traffic to the new version without using its answer. You test under real load without putting anyone at risk.

Expand / contract

A schema migration in three steps: add the new, write to both, retire the old. Never a destructive change in a single deploy.

Dual write and backfill

Write to both sides while an idempotent, resumable job fills in the history.

Contract versioning

Add optional fields, never change the meaning of an existing one. Old consumers stay alive.

A literal in a payload may be acting as protocol for somebody else.

Observability: all three

Traces (where), metrics (how much), logs (what it said). Two of the three is not enough.

A silent rejection shows up on no dashboard.

SLO and error budget

An agreed number for how much is allowed to fail. It turns “it feels slow” into a decision with a threshold.

Symptom-based alerting

Alert on what the user suffers, not on CPU. A new metric with no alert is a dashboard nobody looks at.

Infrastructure as code

The environment is rebuilt from the repo. If an infrastructure change is not in a diff, it does not exist.

Runbook

What to look at and what to do, written before the incident. During the incident nobody designs anything.

Family 07

Principles — what applies even when you choose no pattern at all

Patterns get chosen; principles get respected. These are the ones that come up again and again in code review.

SRP

One class, one reason to change. If two business areas touch it, it is two classes.

OCP

Open to extension, closed to modification. Adding a case should not require editing everybody’s match statement.

LSP

An implementation must be able to replace its interface with no surprises for whoever uses it.

ISP

Several small interfaces beat one fat one that forces you to implement what you do not use.

DIP

Depend on abstractions, not on the concrete implementation. It is the engine behind ports and adapters.

Cohesion and coupling

What changes together lives together; what does not, gets separated. It summarises half the list above.

Separation of concerns

Business, transport and persistence do not share a function.

KISS

The simplest solution that solves today’s real problem.

YAGNI

Do not build it until it is needed. Almost all the flexibility anyone anticipates never gets used.

DRY, carefully

Do not repeat knowledge. Two similar pieces of code that change for different reasons are not duplication.

Composition over inheritance

Assembling behaviour out of pieces ages better than inheriting it.

Dependency injection

Receive what you need instead of constructing it. It is what makes the domain testable.

Fail fast

Validate at the edge and blow up early, with a message that says what was missing.

Least surprise

It should behave the way the name promises. True for functions and for endpoints alike.

Immutability

What does not change cannot produce a race condition.

Idempotence

Doing it twice gives the same result as doing it once. It is the prerequisite of almost everything distributed.

ACID

Atomic, consistent, isolated, durable. What a transaction gives you inside one database.

BASE

Basically available, soft state, eventually consistent. What you are left with once you cross services.

CAP

With the network partitioned you pick consistency or availability. Choosing is not optional.

PACELC

And with no partition you are still choosing: latency or consistency. The honest version of CAP.

12-factor

Config per environment, stateless processes, logs to stdout, dev/prod parity.

Conway’s law

The system ends up shaped like the org chart. If you do not like the cut, look at the teams.

The 8 fallacies

The network is not reliable, nor zero-latency, nor infinite in bandwidth, nor secure, and the topology is not fixed.

Law of Demeter

Talk to your neighbours, not to your neighbours’ friends.

The limit principle

Every shared resource — memory, connections, a queue — needs an explicit ceiling. Whatever is unbounded runs out.

Least operational astonishment

If understanding what happened requires reading the code, a log or a metric is missing.

Antipatterns

Traps: patterns you choose without noticing

An antipattern is almost never chosen on purpose. It shows up by omission, when nobody asked the question. These are the ones that cost the most.

Distributed monolith

Separate services that still have to be deployed together, because they share a schema or a contract. You pay the price of microservices with none of the benefits.

Integration through the database

One service reads another one’s table “because it is faster”. It turns every migration into a negotiation between teams.

Background work without durability

Answer 200 and process in memory. One redeploy, one OOM or one restart and the work evaporated without anyone finding out.

Long synchronous chain

Six hops where each one waits for the next. Availabilities multiply: six services at 99.9 % give you 99.4 %.

Retry with no ceiling and no jitter

The stampede you build yourself. A service that was recovering falls over again with the first wave of retries.

Queue with no dead letter queue

A poison message retries forever and blocks everything behind it. A healthy queue dies because of a single one.

Webhook without deduplication

Every provider resends when in doubt. Without dedupe, their retry policy becomes your compute bill.

Limit with no dashboard and no alert

A rate limit that rejects in silence protects the service and hides from the team that somebody was left outside. The invisible 429 is a failure that does not exist until a customer calls.

Resource with no ceiling

An in-memory queue, a connection pool or a list that grows without limit. It always ends up in the same place.

Cache with no invalidation plan

It gets added to fix latency and turns into a parallel source of truth that nobody knows when to expire.

God service

The one that has to be touched on every feature. It is a monolith with network latency.

Pattern chosen for the résumé

Kafka because it is Kafka, microservices because they are microservices. The question that undoes it is always the same: which measured problem does it solve?

The checklist

The ten questions, before you open the editor

Ten written answers, one or two lines each. If an answer is “I don’t know”, measuring it is the first thing to do.

  1. 01How much comes in, at the peak?Events per second in the worst minute, not the monthly average. Without measured volume, everything else is opinion.
  2. 02What happens if the one next door does not answer?Timeout, retry with jitter, breaker or degradation. Picking one is mandatory; “nothing happens” is not an answer.
  3. 03What happens if the same message arrives twice?If the answer is not “nothing”, idempotence is missing. And it will arrive twice.
  4. 04Where does the work live if the process dies halfway?If it lives in RAM, it is lost. Look at the ladder and decide which rung you are standing on.
  5. 05Does this have to be synchronous?Who is waiting for the answer, and what do they do with it? If nobody looks at it, the caller should not be waiting.
  6. 06Who owns this piece of data?A single service writes; the rest ask. If two write, there is a conflict waiting for its turn.
  7. 07How do I find out that it broke?Trace, metric, alert — and a dashboard where what gets rejected is visible, not only what gets processed.
  8. 08How do I switch it off without deploying?A declared feature flag or killswitch. A user-visible change with no flag does not ship.
  9. 09How do I roll it back?An expand-contract migration, a backward-compatible contract, a previous version that still understands the new data.
  10. 10What is the simplest thing that works?And if the honest answer is the monolith with a queue, then it is the monolith with a queue.

The goal is not ceremony. It is that the architecture conversation happens before the code and stays written down, so that six months from now someone can read why the system is the way it is.

None of these ten questions is expensive to answer. All of them are ruinously expensive to answer late.

In practice

From atlas to practice

At YaVendio I built ya-architecture-design on top of this atlas: a Claude Code skill in the company’s internal engineering harness. It declares a change’s architecture level before any code exists, then runs only the analysis that level demands.

The four levels

N0
One repo, one process, no contract or network change, undone with a revert.
Nothing runs. The level and its reason, in one line.
N1
Touches a contract, a schema or an existing call.
Four questions in the ticket: who reads it today, is it backward compatible, how do I roll it back, how do I find out it broke.
N2
Adds or changes a call between processes, takes traffic nobody controls, changes data another service reads, changes throughput in either direction, adds an external dependency, touches money or conversation state on the write path, or was born from an incident.
Measure, pick a rung on the ladder, answer the ten questions, write the ADR before the code, and update the C4 model if the shape of the system changed.
N?
Not enough written down to decide.
Stop and ask. A ticket that says nothing has not said that the change is small.

What it takes from this atlas

The default is inverted: the skill declares the level, and an engineer who disagrees overrules it with one line that stays in the ticket. A deterministic pre-filter can raise a level, never lower it.