Autodiscover Independent coverage of news

Cloud Infrastructure Compared: What Actually Matters

By Robert Hayes · · 1221 words
Cloud Infrastructure Compared: What Actually Matters

Backup Strategy: Periodic jobs should be safe to run twice, because they will be. Backup Strategy: You rarely need a new component to fix a boundary problem. Backup Strategy: The signal you want is often already logged, just not aggregated.

Schema Migration: The first thing to settle is the failure mode, not the happy path. Schema Migration: Measurements taken once are anecdotes; you need a baseline that repeats. Schema Migration: Costs usually concentrate in a small number of operations, so find those first.

Teams working on observability usually discover this the hard way. If the rollback plan needs a meeting, it is not a rollback plan. Small pages that stay small are easier to keep fast than large ones made fast. This is most visible in observability. Consider observability specifically. Write the invariant down; otherwise it lives only in someone's memory.

In practice, cost controls behaves differently: A queue smooths spikes but also hides how far behind you are. Retries without jitter turn a small outage into a large one. The same reasoning holds for cost controls. For cost controls, the constraint matters more than the feature list. Separating the reads from the writes buys room to change either side.

If the rollback plan needs a meeting, it is not a rollback plan. The same reasoning holds for cost controls. For cost controls, the constraint matters more than the feature list. Small pages that stay small are easier to keep fast than large ones made fast. Teams working on cost controls usually discover this the hard way. Write the invariant down; otherwise it lives only in someone's memory.

Teams working on monitoring alerts usually discover this the hard way. If the rollback plan needs a meeting, it is not a rollback plan. Small pages that stay small are easier to keep fast than large ones made fast. This is most visible in monitoring alerts. Consider monitoring alerts specifically. Write the invariant down; otherwise it lives only in someone's memory.

Load Balancing: The interesting number is not the average, it is the 99th percentile. Load Balancing: Adding a cache in front of a slow query is a fix; fixing the query is a cure. Load Balancing: Every abstraction you add is a place where behaviour can differ from intent.

Storage Tiers: A design that cannot be rolled back is a design that cannot be changed safely. Storage Tiers: Latency budgets are easier to defend when every hop has a stated ceiling. Storage Tiers: Caching helps only until the invalidation rules become the bottleneck.

Configurations should be reviewable in a diff, not only in a console. This is most visible in access control. Consider access control specifically. The best time to add an index is before the table gets large. Access Control: Failures are usually correlated, so plan for the shared dependency.

Edge Caching: Periodic jobs should be safe to run twice, because they will be. Edge Caching: You rarely need a new component to fix a boundary problem. Edge Caching: The signal you want is often already logged, just not aggregated.

A design that cannot be rolled back is a design that cannot be changed safely. That applies to backup strategy as well. In practice, backup strategy behaves differently: Latency budgets are easier to defend when every hop has a stated ceiling. Caching helps only until the invalidation rules become the bottleneck. The same reasoning holds for backup strategy.

Release Process: Serving static bytes is the cheapest thing you can do at the edge. Release Process: A schema is an interface; changing it is a migration, not an edit. Release Process: Track the denominator as carefully as the numerator.

Pay attention to the conditions around the conversation. A substantial power difference, financial dependence or fear of someone’s reaction can make it harder to speak openly. These circumstances do not automatically determine a legal outcome, but they are reasons to take extra care and avoid pressuring the other person. Give them time and a genuine opportunity to say no.

If a metric has no owner, it will drift until it causes an incident. This is most visible in observability. Consider observability specifically. The cheapest optimisation is usually removing work nobody asked for. Observability: Aggregating at write time trades flexibility for predictable read cost.

API Design: The first thing to settle is the failure mode, not the happy path. API Design: Measurements taken once are anecdotes; you need a baseline that repeats. API Design: Costs usually concentrate in a small number of operations, so find those first.

API Design: A queue smooths spikes but also hides how far behind you are. API Design: Retries without jitter turn a small outage into a large one. API Design: Separating the reads from the writes buys room to change either side.

Boundaries may involve practical health decisions as well as personal comfort. If relevant, discuss contraception, barrier methods, STI testing, and what each person understands about risk before sexual activity. Be clear about what you will do if you cannot agree on a safety measure: for example, you may decide not to proceed. Neither partner should be expected to accept a risk they have not agreed to.

In practice, log analysis behaves differently: Periodic jobs should be safe to run twice, because they will be. You rarely need a new component to fix a boundary problem. The same reasoning holds for log analysis. For log analysis, the constraint matters more than the feature list. The signal you want is often already logged, just not aggregated.

For load balancing, the constraint matters more than the feature list. The first thing to settle is the failure mode, not the happy path. Teams working on load balancing usually discover this the hard way. Measurements taken once are anecdotes; you need a baseline that repeats. Costs usually concentrate in a small number of operations, so find those first. This is most visible in load balancing.

Observability: A design that cannot be rolled back is a design that cannot be changed safely. Observability: Latency budgets are easier to defend when every hop has a stated ceiling. Observability: Caching helps only until the invalidation rules become the bottleneck.

If a metric has no owner, it will drift until it causes an incident. This is most visible in queue design. Consider queue design specifically. The cheapest optimisation is usually removing work nobody asked for. Queue Design: Aggregating at write time trades flexibility for predictable read cost.

Monitoring Alerts: You can often replace a coordination problem with an idempotency key. Monitoring Alerts: Anything that grows without a bound will eventually hit one. Monitoring Alerts: Documentation that is not tested tends to describe the previous version.

Observability: Periodic jobs should be safe to run twice, because they will be. You rarely need a new component to fix a boundary problem. That applies to observability as well. In practice, observability behaves differently: The signal you want is often already logged, just not aggregated.

For rate limiting, the constraint matters more than the feature list. Configurations should be reviewable in a diff, not only in a console. Teams working on rate limiting usually discover this the hard way. The best time to add an index is before the table gets large. Failures are usually correlated, so plan for the shared dependency. This is most visible in rate limiting.

Related reading