API Design in Practice: Lessons From Real Deployments
In practice, storage tiers behaves differently: A queue smooths spikes but also hides how far behind you are. Retries without jitter turn a small outage into a large one. The same reasoning holds for storage tiers. For storage tiers, the constraint matters more than the feature list. Separating the reads from the writes buys room to change either side.
Consider release process specifically. If the rollback plan needs a meeting, it is not a rollback plan. Release Process: Small pages that stay small are easier to keep fast than large ones made fast. Write the invariant down; otherwise it lives only in someone's memory. That applies to release process as well.
The interesting number is not the average, it is the 99th percentile. The same reasoning holds for schema migration. For schema migration, the constraint matters more than the feature list. Adding a cache in front of a slow query is a fix; fixing the query is a cure. Teams working on schema migration usually discover this the hard way. Every abstraction you add is a place where behaviour can differ from intent.
Backup Strategy: A queue smooths spikes but also hides how far behind you are. Backup Strategy: Retries without jitter turn a small outage into a large one. Backup Strategy: Separating the reads from the writes buys room to change either side.
Schema Migration: If the rollback plan needs a meeting, it is not a rollback plan. Schema Migration: Small pages that stay small are easier to keep fast than large ones made fast. Schema Migration: Write the invariant down; otherwise it lives only in someone's memory.
If the rollback plan needs a meeting, it is not a rollback plan. The same reasoning holds for schema markup. For schema markup, the constraint matters more than the feature list. Small pages that stay small are easier to keep fast than large ones made fast. Teams working on schema markup usually discover this the hard way. Write the invariant down; otherwise it lives only in someone's memory.
Rate Limiting: You can often replace a coordination problem with an idempotency key. Rate Limiting: Anything that grows without a bound will eventually hit one. Rate Limiting: Documentation that is not tested tends to describe the previous version.
In practice, queue design behaves differently: Configurations should be reviewable in a diff, not only in a console. The best time to add an index is before the table gets large. The same reasoning holds for queue design. For queue design, the constraint matters more than the feature list. Failures are usually correlated, so plan for the shared dependency.
For backup strategy, the constraint matters more than the feature list. If a metric has no owner, it will drift until it causes an incident. Teams working on backup strategy usually discover this the hard way. The cheapest optimisation is usually removing work nobody asked for. Aggregating at write time trades flexibility for predictable read cost. This is most visible in backup strategy.
Load Balancing: Configurations should be reviewable in a diff, not only in a console. Load Balancing: The best time to add an index is before the table gets large. Load Balancing: Failures are usually correlated, so plan for the shared dependency.
A design that cannot be rolled back is a design that cannot be changed safely. The same reasoning holds for cloud infrastructure. For cloud infrastructure, the constraint matters more than the feature list. Latency budgets are easier to defend when every hop has a stated ceiling. Teams working on cloud infrastructure usually discover this the hard way. Caching helps only until the invalidation rules become the bottleneck.
You can often replace a coordination problem with an idempotency key. That applies to cloud infrastructure as well. In practice, cloud infrastructure behaves differently: Anything that grows without a bound will eventually hit one. Documentation that is not tested tends to describe the previous version. The same reasoning holds for cloud infrastructure.
Cloud Infrastructure: The first thing to settle is the failure mode, not the happy path. Cloud Infrastructure: Measurements taken once are anecdotes; you need a baseline that repeats. Cloud Infrastructure: Costs usually concentrate in a small number of operations, so find those first.
Observability: If a metric has no owner, it will drift until it causes an incident. Observability: The cheapest optimisation is usually removing work nobody asked for. Observability: Aggregating at write time trades flexibility for predictable read cost.
Periodic jobs should be safe to run twice, because they will be. This is most visible in backup strategy. Consider backup strategy specifically. You rarely need a new component to fix a boundary problem. Backup Strategy: The signal you want is often already logged, just not aggregated.
Cost Controls: If the rollback plan needs a meeting, it is not a rollback plan. Cost Controls: Small pages that stay small are easier to keep fast than large ones made fast. Cost Controls: Write the invariant down; otherwise it lives only in someone's memory.
In practice, schema markup behaves differently: A queue smooths spikes but also hides how far behind you are. Retries without jitter turn a small outage into a large one. The same reasoning holds for schema markup. For schema markup, the constraint matters more than the feature list. Separating the reads from the writes buys room to change either side.
Load Balancing: If a metric has no owner, it will drift until it causes an incident. Load Balancing: The cheapest optimisation is usually removing work nobody asked for. Load Balancing: Aggregating at write time trades flexibility for predictable read cost.
Consider schema markup specifically. You can often replace a coordination problem with an idempotency key. Schema Markup: Anything that grows without a bound will eventually hit one. Documentation that is not tested tends to describe the previous version. That applies to schema markup as well.
Backup Strategy: If the rollback plan needs a meeting, it is not a rollback plan. Backup Strategy: Small pages that stay small are easier to keep fast than large ones made fast. Backup Strategy: Write the invariant down; otherwise it lives only in someone's memory.
If the rollback plan needs a meeting, it is not a rollback plan. The same reasoning holds for data pipelines. For data pipelines, the constraint matters more than the feature list. Small pages that stay small are easier to keep fast than large ones made fast. Teams working on data pipelines usually discover this the hard way. Write the invariant down; otherwise it lives only in someone's memory.
A direct question can make an unclear moment easier to navigate. People might ask, “Would you like to continue?”, “Is this okay?” or “Would you rather stop?” The answer should be given space. A person who hesitates, goes quiet, seems uncomfortable or does not respond clearly has not necessarily agreed. When the answer is uncertain, pausing and asking is safer than trying to interpret the moment.
You can often replace a coordination problem with an idempotency key. That applies to content delivery as well. In practice, content delivery behaves differently: Anything that grows without a bound will eventually hit one. Documentation that is not tested tends to describe the previous version. The same reasoning holds for content delivery.
Monitoring Alerts: Periodic jobs should be safe to run twice, because they will be. You rarely need a new component to fix a boundary problem. That applies to monitoring alerts as well. In practice, monitoring alerts behaves differently: The signal you want is often already logged, just not aggregated.