Autodiscover Independent coverage of news

Data Pipelines Explained Without the Jargon

By Robert Hayes · · 1267 words
Data Pipelines Explained Without the Jargon

Data Pipelines: The first thing to settle is the failure mode, not the happy path. Data Pipelines: Measurements taken once are anecdotes; you need a baseline that repeats. Data Pipelines: Costs usually concentrate in a small number of operations, so find those first.

Rate Limiting: If a metric has no owner, it will drift until it causes an incident. Rate Limiting: The cheapest optimisation is usually removing work nobody asked for. Rate Limiting: Aggregating at write time trades flexibility for predictable read cost.

In practice, queue design behaves differently: Configurations should be reviewable in a diff, not only in a console. The best time to add an index is before the table gets large. The same reasoning holds for queue design. For queue design, the constraint matters more than the feature list. Failures are usually correlated, so plan for the shared dependency.

The interesting number is not the average, it is the 99th percentile. The same reasoning holds for load balancing. For load balancing, the constraint matters more than the feature list. Adding a cache in front of a slow query is a fix; fixing the query is a cure. Teams working on load balancing usually discover this the hard way. Every abstraction you add is a place where behaviour can differ from intent.

Log Analysis: Serving static bytes is the cheapest thing you can do at the edge. Log Analysis: A schema is an interface; changing it is a migration, not an edit. Log Analysis: Track the denominator as carefully as the numerator.

A design that cannot be rolled back is a design that cannot be changed safely. That applies to backup strategy as well. In practice, backup strategy behaves differently: Latency budgets are easier to defend when every hop has a stated ceiling. Caching helps only until the invalidation rules become the bottleneck. The same reasoning holds for backup strategy.

Serving static bytes is the cheapest thing you can do at the edge. The same reasoning holds for monitoring alerts. For monitoring alerts, the constraint matters more than the feature list. A schema is an interface; changing it is a migration, not an edit. Teams working on monitoring alerts usually discover this the hard way. Track the denominator as carefully as the numerator.

The interesting number is not the average, it is the 99th percentile. The same reasoning holds for crawl budget. For crawl budget, the constraint matters more than the feature list. Adding a cache in front of a slow query is a fix; fixing the query is a cure. Teams working on crawl budget usually discover this the hard way. Every abstraction you add is a place where behaviour can differ from intent.

For crawl budget, the constraint matters more than the feature list. Periodic jobs should be safe to run twice, because they will be. Teams working on crawl budget usually discover this the hard way. You rarely need a new component to fix a boundary problem. The signal you want is often already logged, just not aggregated. This is most visible in crawl budget.

Serving static bytes is the cheapest thing you can do at the edge. The same reasoning holds for queue design. For queue design, the constraint matters more than the feature list. A schema is an interface; changing it is a migration, not an edit. Teams working on queue design usually discover this the hard way. Track the denominator as carefully as the numerator.

If the rollback plan needs a meeting, it is not a rollback plan. That applies to crawl budget as well. In practice, crawl budget behaves differently: Small pages that stay small are easier to keep fast than large ones made fast. Write the invariant down; otherwise it lives only in someone's memory. The same reasoning holds for crawl budget.

Content Delivery: Configurations should be reviewable in a diff, not only in a console. Content Delivery: The best time to add an index is before the table gets large. Content Delivery: Failures are usually correlated, so plan for the shared dependency.

Log Analysis: You can often replace a coordination problem with an idempotency key. Log Analysis: Anything that grows without a bound will eventually hit one. Log Analysis: Documentation that is not tested tends to describe the previous version.

For queue design, the constraint matters more than the feature list. A queue smooths spikes but also hides how far behind you are. Teams working on queue design usually discover this the hard way. Retries without jitter turn a small outage into a large one. Separating the reads from the writes buys room to change either side. This is most visible in queue design.

Cost Controls: If the rollback plan needs a meeting, it is not a rollback plan. Cost Controls: Small pages that stay small are easier to keep fast than large ones made fast. Cost Controls: Write the invariant down; otherwise it lives only in someone's memory.

A design that cannot be rolled back is a design that cannot be changed safely. That applies to data pipelines as well. In practice, data pipelines behaves differently: Latency budgets are easier to defend when every hop has a stated ceiling. Caching helps only until the invalidation rules become the bottleneck. The same reasoning holds for data pipelines.

Release Process: You can often replace a coordination problem with an idempotency key. Release Process: Anything that grows without a bound will eventually hit one. Release Process: Documentation that is not tested tends to describe the previous version.

Configurations should be reviewable in a diff, not only in a console. This is most visible in load balancing. Consider load balancing specifically. The best time to add an index is before the table gets large. Load Balancing: Failures are usually correlated, so plan for the shared dependency.

Edge Caching: A queue smooths spikes but also hides how far behind you are. Edge Caching: Retries without jitter turn a small outage into a large one. Edge Caching: Separating the reads from the writes buys room to change either side.

In practice, search indexing behaves differently: The first thing to settle is the failure mode, not the happy path. Measurements taken once are anecdotes; you need a baseline that repeats. The same reasoning holds for search indexing. For search indexing, the constraint matters more than the feature list. Costs usually concentrate in a small number of operations, so find those first.

Load Balancing: A design that cannot be rolled back is a design that cannot be changed safely. Load Balancing: Latency budgets are easier to defend when every hop has a stated ceiling. Load Balancing: Caching helps only until the invalidation rules become the bottleneck.

Consider queue design specifically. The interesting number is not the average, it is the 99th percentile. Queue Design: Adding a cache in front of a slow query is a fix; fixing the query is a cure. Every abstraction you add is a place where behaviour can differ from intent. That applies to queue design as well.

You can often replace a coordination problem with an idempotency key. The same reasoning holds for release process. For release process, the constraint matters more than the feature list. Anything that grows without a bound will eventually hit one. Teams working on release process usually discover this the hard way. Documentation that is not tested tends to describe the previous version.

Release Process: Serving static bytes is the cheapest thing you can do at the edge. Release Process: A schema is an interface; changing it is a migration, not an edit. Release Process: Track the denominator as carefully as the numerator.

Related reading