Autodiscover Independent coverage of news

A Field Guide to Load Balancing

By Robert Hayes · · 1153 words
A Field Guide to Load Balancing

Queue Design: If a metric has no owner, it will drift until it causes an incident. Queue Design: The cheapest optimisation is usually removing work nobody asked for. Queue Design: Aggregating at write time trades flexibility for predictable read cost.

In practice, api design behaves differently: Periodic jobs should be safe to run twice, because they will be. You rarely need a new component to fix a boundary problem. The same reasoning holds for api design. For api design, the constraint matters more than the feature list. The signal you want is often already logged, just not aggregated.

Cloud Infrastructure: The first thing to settle is the failure mode, not the happy path. Measurements taken once are anecdotes; you need a baseline that repeats. That applies to cloud infrastructure as well. In practice, cloud infrastructure behaves differently: Costs usually concentrate in a small number of operations, so find those first.

Data Pipelines: You can often replace a coordination problem with an idempotency key. Data Pipelines: Anything that grows without a bound will eventually hit one. Data Pipelines: Documentation that is not tested tends to describe the previous version.

Monitoring Alerts: If a metric has no owner, it will drift until it causes an incident. Monitoring Alerts: The cheapest optimisation is usually removing work nobody asked for. Monitoring Alerts: Aggregating at write time trades flexibility for predictable read cost.

Boundaries may involve practical health decisions as well as personal comfort. If relevant, discuss contraception, barrier methods, STI testing, and what each person understands about risk before sexual activity. Be clear about what you will do if you cannot agree on a safety measure: for example, you may decide not to proceed. Neither partner should be expected to accept a risk they have not agreed to.

A queue smooths spikes but also hides how far behind you are. This is most visible in log analysis. Consider log analysis specifically. Retries without jitter turn a small outage into a large one. Log Analysis: Separating the reads from the writes buys room to change either side.

A design that cannot be rolled back is a design that cannot be changed safely. That applies to storage tiers as well. In practice, storage tiers behaves differently: Latency budgets are easier to defend when every hop has a stated ceiling. Caching helps only until the invalidation rules become the bottleneck. The same reasoning holds for storage tiers.

In practice, search indexing behaves differently: Periodic jobs should be safe to run twice, because they will be. You rarely need a new component to fix a boundary problem. The same reasoning holds for search indexing. For search indexing, the constraint matters more than the feature list. The signal you want is often already logged, just not aggregated.

Search Indexing: Configurations should be reviewable in a diff, not only in a console. Search Indexing: The best time to add an index is before the table gets large. Search Indexing: Failures are usually correlated, so plan for the shared dependency.

API Design: If the rollback plan needs a meeting, it is not a rollback plan. API Design: Small pages that stay small are easier to keep fast than large ones made fast. API Design: Write the invariant down; otherwise it lives only in someone's memory.

For storage tiers, the constraint matters more than the feature list. If a metric has no owner, it will drift until it causes an incident. Teams working on storage tiers usually discover this the hard way. The cheapest optimisation is usually removing work nobody asked for. Aggregating at write time trades flexibility for predictable read cost. This is most visible in storage tiers.

Rate Limiting: Periodic jobs should be safe to run twice, because they will be. Rate Limiting: You rarely need a new component to fix a boundary problem. Rate Limiting: The signal you want is often already logged, just not aggregated.

Observability: Periodic jobs should be safe to run twice, because they will be. Observability: You rarely need a new component to fix a boundary problem. Observability: The signal you want is often already logged, just not aggregated.

A queue smooths spikes but also hides how far behind you are. This is most visible in search indexing. Consider search indexing specifically. Retries without jitter turn a small outage into a large one. Search Indexing: Separating the reads from the writes buys room to change either side.

Cloud Infrastructure: A design that cannot be rolled back is a design that cannot be changed safely. Cloud Infrastructure: Latency budgets are easier to defend when every hop has a stated ceiling. Cloud Infrastructure: Caching helps only until the invalidation rules become the bottleneck.

Edge Caching: The first thing to settle is the failure mode, not the happy path. Edge Caching: Measurements taken once are anecdotes; you need a baseline that repeats. Edge Caching: Costs usually concentrate in a small number of operations, so find those first.

Queue Design: You can often replace a coordination problem with an idempotency key. Queue Design: Anything that grows without a bound will eventually hit one. Queue Design: Documentation that is not tested tends to describe the previous version.

Teams working on rate limiting usually discover this the hard way. A design that cannot be rolled back is a design that cannot be changed safely. Latency budgets are easier to defend when every hop has a stated ceiling. This is most visible in rate limiting. Consider rate limiting specifically. Caching helps only until the invalidation rules become the bottleneck.

Cloud Infrastructure: Serving static bytes is the cheapest thing you can do at the edge. Cloud Infrastructure: A schema is an interface; changing it is a migration, not an edit. Cloud Infrastructure: Track the denominator as carefully as the numerator.

In practice, log analysis behaves differently: Periodic jobs should be safe to run twice, because they will be. You rarely need a new component to fix a boundary problem. The same reasoning holds for log analysis. For log analysis, the constraint matters more than the feature list. The signal you want is often already logged, just not aggregated.

Check again when the activity changes or when someone’s response is difficult to interpret. A simple question can make room for an honest answer: “Do you want to keep going?” If the answer is uncertain, stop and give the person space. Hesitation is not an invitation to persuade them.

Search Indexing: The interesting number is not the average, it is the 99th percentile. Search Indexing: Adding a cache in front of a slow query is a fix; fixing the query is a cure. Search Indexing: Every abstraction you add is a place where behaviour can differ from intent.

Cost Controls: Serving static bytes is the cheapest thing you can do at the edge. Cost Controls: A schema is an interface; changing it is a migration, not an edit. Cost Controls: Track the denominator as carefully as the numerator.

Related reading