Early in my career I’d hear a term in a design review, nod, and Google it under the table. Plenty of us do this. The vocabulary sounds scarier than the ideas behind it.

So here’s my plain-English version of 49 terms. I’ve grouped them by the kind of problem each one solves, not alphabetically, because I find that’s easier to remember.

Part 1: Handling Growth and Traffic

1. Scalability. Can your system survive getting popular? That’s all it means. The catch is that it’s only as good as its weakest piece. Beef up the web servers and the database quietly becomes the new bottleneck.

2. Vertical scaling. Buy a bigger machine. It’s the easy fix because nothing about your code changes, but you eventually hit the biggest machine money can buy, and it’s still a single point of failure.

3. Horizontal scaling. Buy more machines instead of a bigger one. This is how the big companies do it. Losing one box is a minor dip, not a disaster.

4. Stateless servers. The rule that makes #3 work. If a server remembers things about you between requests, you’re stuck going back to that exact server. Keep memory out of the servers and any of them can handle any request.

5. Load balancer. The traffic cop in front of your servers. It spreads requests around and skips any server that stops responding. Users see one website, even though a dozen machines are behind it.

6. Reverse proxy. A middleman that takes requests on behalf of your servers. It can cache, compress, and handle HTTPS. A load balancer is really just a reverse proxy with a narrower job.

7. API gateway. One front door for all your backend services. Login checks, rate limits, and routing happen there once, instead of every service reinventing them. Just make sure the front door doesn’t become the thing that falls over.

8. DNS. The phone book of the internet, turning yoursite.com into an IP address. It can also steer users to different servers by location. The downside is that it’s slow to update, since everyone caches the answers.

9. CDN. Copies of your images, videos, and scripts stored in cities around the world, so a user in Tokyo doesn’t fetch a file from Virginia. The risk is that if the original changes, old copies can linger.

10. Latency. How long you wait for one response. It’s the “feels snappy or feels slow” number, and it’s usually a sum of network time, server time, and database time. Find the biggest chunk first.

11. Throughput. How much total work gets done per second. This is different from latency. A system can answer one request lightning-fast yet only handle one at a time, which means great latency and terrible throughput.

Part 2: Caching

12. Cache. A small, fast layer holding things you’ve needed recently, so you don’t redo slow work. The price is that it can serve outdated data.

13. Cache-aside. The everyday default. Check the cache first. On a miss, ask the database, then save the answer in the cache for next time. If the cache dies, the app still works, just slower.

14. Write-through. Every write updates the cache and the database together. The data always matches, but each write takes a bit longer.

15. Write-behind. Write to the cache and let the database catch up later. It’s fast, but if the cache crashes before it syncs, you lose data. Fine for analytics, risky for money.

16. Consistent hashing. A clever way to decide which server holds which data, so that adding or removing a server only moves a small slice of keys instead of reshuffling nearly everything.

17. Hash ring and virtual nodes. The usual way to picture #16: servers sit on a circle, and each key goes to the next server clockwise. Giving each server several spots on the circle (virtual nodes) spreads the load more evenly.

Part 3: Storing Data

18. Database. Where data lives after the request is over. It’s also the choice you’ll regret most if you get it wrong, because everything else gets built on top of it.

19. SQL (relational) database. Tidy tables, with rows linked together by relationships. It’s great when your data is structured and you need reliable transactions. Scaling writes across many machines is where it gets painful.

20. NoSQL database. An umbrella term for key-value, document, graph, and other non-table databases. Each gives up some SQL features for scale or flexibility. Pick based on how you’ll read the data, not on what’s trendy.

21. ACID. The four promises of a safe transaction. It’s all-or-nothing, it leaves the data valid, concurrent transactions don’t trip over each other, and it survives a crash once saved. Think bank transfers: you never want money to leave one account and not arrive in the other.

22. Index. Like the index at the back of a textbook. Lookups jump straight to the answer instead of scanning every row. But every write now has extra bookkeeping, so only index what you actually search by.

23. Replication. Keeping copies of your data on several machines. Usually one machine takes writes and the others serve reads. Remember that copies can lag behind for a moment.

24. Sharding. Splitting one huge dataset across many databases. Picking the shard key is the whole game: choose badly and one shard does all the work while the rest sit idle.

25. Data partitioning. The bigger umbrella over sharding. You could also split by date, with one table per month, so old data is easy to archive and queries scan less.

26. Object storage. A giant, cheap, durable bucket for files like photos, videos, and backups. S3 is the famous one. Don’t treat it like a database. It’s built for whole files, not tiny quick lookups.

27. Event sourcing. Instead of saving “balance: $100,” you save every deposit and withdrawal and work out the balance from the history. You get a full audit trail and the ability to rewind. Reading the current state means replaying events, so people use snapshots to speed it up.

Part 4: Talking Between Services

28. Message queue. A waiting line for tasks. The sender drops off a message and moves on, and the receiver handles it when it can. It smooths out spikes, but you don’t get the result immediately.

29. Pub/sub. One announcement, many listeners. When “order placed” fires, inventory, emails, and analytics can each react independently, and the order service never needs to know who’s listening.

30. Dead-letter queue. A parking lot for messages that keep failing. Instead of one broken message jamming the line forever, it gets set aside so a human can look at it later.

31. Backpressure. The receiver telling the sender “slow down, I’m drowning.” Without it, queues grow endlessly until something runs out of memory and crashes.

32. WebSocket. A connection that stays open so both sides can talk any time. It’s great for chat, multiplayer games, and live collaboration, though keeping thousands of long-lived connections organized takes extra infrastructure.

33. Server-Sent Events (SSE). A simpler cousin where only the server talks, and the client just listens. It’s perfect for notifications, live scores, and streaming AI responses, with automatic reconnection built in.

Part 5: Life Across Many Machines

34. Distributed system. Many computers cooperating over a network. You gain power, but you also inherit unreliable networks, partial failures, and the headache of keeping copies in agreement.

35. CAP theorem. When the network splits, you can’t have everything. Because partitions happen in real life, you’re really choosing between consistency (be correct, even if you refuse some requests) and availability (keep answering, even if the answer might be slightly stale).

36. Strong consistency. Every read sees the latest write, no matter which machine answers. It costs speed, and it’s worth it for things like bank balances or the last concert ticket.

37. Eventual consistency. Copies may disagree briefly, then catch up. It’s faster and more available, and it’s fine for like counts or follower numbers, where nobody is harmed by being a few seconds behind.

38. Consensus. Getting a group of machines to agree on one answer even when some fail. Raft and Paxos are the famous methods. They require a majority, which stops two machines from both thinking they’re in charge.

39. Leader election. How a group picks one machine to coordinate things, and how it picks a replacement when that one dies. There’s a short awkward gap during the switch where leader-only work pauses.

40. Idempotency. An action you can repeat safely with the same outcome. Reading a page twice changes nothing. Submitting an order twice might create two orders. It matters because networks lose messages, so retries are unavoidable.

41. Idempotency key. A unique ID the client attaches to a request. If the server sees the same ID again, it returns the earlier result instead of doing the work twice. It’s how payment APIs avoid double charges.

42. Two-phase commit. A “can everyone do this?” vote, followed by “okay, go” or “abort.” It makes a transaction span multiple databases, but if the coordinator crashes mid-way, everyone is stuck holding locks.

43. Saga pattern. The more popular alternative to #42. Do the steps one at a time, and if a later step fails, run “undo” actions for the earlier ones, such as refunding a payment or releasing reserved stock.

44. Clock skew. Different computers disagree about what time it is, even if only by milliseconds. That’s enough to scramble the order of events if you trust timestamps.

45. Vector clock. A workaround for #44. Instead of real time, each machine keeps counters that track what it has seen, which lets the system tell whether one event could have caused another, or whether the two happened independently.

Part 6: Staying Alive Under Pressure

46. Circuit breaker. Like the one in your house. If a dependency keeps failing, you stop calling it for a while, fail fast, and test it again later. That keeps one sick service from dragging down everything that depends on it.

47. Rate limiting. A cap on how many requests one client gets per time window. The common approach is a token bucket: tokens refill steadily, each request spends one, and short bursts are allowed while the long-run average is controlled.

48. Load shedding. When you’re overwhelmed, deliberately turn some requests away so the rest get served properly. Better to serve 80% well than 100% terribly, and drop the least important work first.

49. Bloom filter. A tiny structure that answers “have I seen this?” It’s never wrong when it says “no,” but sometimes it says “maybe” when the real answer is no. That’s a great trade when you just want to skip pointless disk lookups.

How I Keep Them Straight

I stopped trying to memorize definitions. When a term comes up, I ask one thing: what would go wrong without this? No load balancer, and one server melts. No idempotency, and customers get charged twice. No backpressure, and memory runs out. Once you can name the pain, the term sticks.

Which of these tripped you up the most? Tell me in the comments and I’ll write a follow-up on i