In the previous post, we explored the foundational trade-offs in data systems architecture. While Chapter 1 set the stage, Chapter 2 of Designing Data-Intensive Applications (Second Edition), "Defining Nonfunctional Requirements," gets into the "how" of building systems that aren't just functional, but actually work in the messy reality of the real world.
Most developers are driven by functional requirements, the buttons, screens, and specific operations the software must perform. However, the authors argue that nonfunctional requirements, performance, reliability, scalability, and maintainability, are just as vital. An app that has all the right features but is unbearably slow or frequently crashes might as well not exist.
A Case Study in Scale: The Social Network Timeline
To ground these abstract concepts, the chapter uses a case study of a social network (style of X/Twitter). Imagine a service where users make 5,800 posts per second on average, spiking to 150,000.
The challenge isn't just storing the posts, but delivering home timelines, a feed of recent posts from everyone a user follows. A simple SQL join to fetch this data for millions of active users is prohibitively expensive. The solution often involves materialization: precomputing each user’s timeline into a cache so it’s ready the moment they log in. This introduces a "fan-out" problem where a single post by a celebrity must be "delivered" to millions of follower mailboxes, showcasing the intense write-load trade-offs required for fast reads.
1. Reliability: Surviving the Inevitable
Reliability means a system continues to work correctly even when things go wrong. The book distinguishes between faults (one component deviating from its spec) and failures (the whole system stopping).
- Hardware Faults: In large-scale systems, disks failing or power going out are not accidents; they are part of normal operations. Redundancy is the first line of defense.
- Software Faults: These are often systematic and harder to deal with than hardware faults because they are correlated across nodes (e.g., a "retry storm" that overwhelms a recovering system).
- Human Error: Humans are creative but unpredictable. The authors advocate for blameless postmortems, where the focus is on fixing the system's incentives and priorities rather than blaming individuals.
2. Scalability: It’s Not a One-Dimensional Label
The authors warn against saying "X scales" or "Y doesn't scale". Scalability is about having options to cope with increased load.
- Measuring Load: You must first define your load using metrics like throughput (requests per second) or response time.
- Vertical vs. Horizontal: You can scale up (more powerful machines) or out (more smaller machines). The shared-nothing architecture (horizontal scaling) has gained popularity for its potential to scale linearly, though it introduces significant distributed systems complexity.
3. Performance: The Power of Percentiles
When describing performance, looking at the "average" response time is often misleading because it hides the experience of users in the "tail". The authors emphasize percentiles (like p95 or p99). High percentiles are critical because tail latency amplification can occur: if a single end-user request requires multiple backend calls, just one slow backend call can slow down the entire user experience.
4. Maintainability: The Long Game
- Software often outlives its original creators. To minimize the pain of legacy systems, we must design for:
- Operability: Making it easy for teams to keep the system running smoothly.
- Simplicity: Managing complexity by using the right abstractions.
- Evolvability: Ensuring the system can adapt to changing requirements or new legal contexts like GDPR.
Conclusion
Chapter 2 serves as a reminder that the "invisible" parts of an application, the parts the user doesn't see until they break, are what define its success. By using well-understood building blocks and focusing on these core nonfunctional pillars, we can build systems that don't just work today, but continue to work as they grow and change.
Reference
Designing Data-Intensive Applications: The Big Ideas Behind Reliable, Scalable, and Maintainable Systems
Book by Martin Kleppmann and Chris Riccomini
This blog post is a summary of my personal notes and understanding from reading "Designing Data-Intensive Applications" by Martin Kleppmann. All credit for the original ideas belongs to the author.
Comments
Post a Comment