Books Deep-Read

Great CS Books, Distilled — BigCat's Bookshelf

> one chapter, one diagram · read this page ≈ understand the chapter
DDIA — Designing Data-Intensive Applications— Martin Kleppmann · 2017
Part I · Foundations of Data Systems
Ch 1Reliable, Scalable, Maintainable — set three yardsticks and make "is this system good?" measurable with load parameters and response-time percentilesKleppmann · 2017 Ch 2Data Models & Query Languages — what relational, document and graph are each good at, and why declarative beats imperativeKleppmann · 2017 Ch 3Storage & Retrieval — how databases store and find data: LSM-trees vs B-trees, row-oriented OLTP vs columnar analyticsKleppmann · 2017 Ch 4Encoding & Evolution — how data is serialized and how schemas evolve with forward/backward compatibility, no downtimeKleppmann · 2017
Part II · Distributed Data
Ch 5Replication — one dataset on many machines: single-leader, multi-leader, leaderless, and the pitfalls of replication lagKleppmann · 2017 Ch 6Partitioning — splitting a dataset too big for one node: by key range or by hash, hotspots and rebalancingKleppmann · 2017 Ch 7Transactions — what ACID actually guarantees, which anomalies weak isolation lets through, and why write skew is the trap to rememberKleppmann · 2017 Ch 8The Trouble with Distributed Systems — partial failures, unreliable clocks, network and process pausesKleppmann · 2017 Ch 9Consistency & Consensus — linearizability, total-order broadcast and consensus, what to really know after CAPKleppmann · 2017
Part III · Derived Data
Ch 10Batch Processing — MapReduce's tally-then-total, one distributed sort turning huge immutable input into derived results, recompute on failureKleppmann · 2017 Ch 11Stream Processing — process a never-ending event stream one at a time on a replayable log (Kafka); stream-table duality turns the database inside out, with time and watermarks as the hard partKleppmann · 2017 Ch 12The Future of Data Systems — unbundle the database and reassemble it around one ordered event log; correctness comes end-to-end, not from the layers belowKleppmann · 2017
Continuous Delivery— Humble & Farley · 2010
Part I · Foundations
Ch 1The Problem of Delivering Software — manual deploys and release-day hell; turn releasing into a repeatable pipelineHumble & Farley · 2010 Ch 2Configuration Management — everything under version control, one artifact for every environment, no manual changes to machinesHumble & Farley · 2010 Ch 3Continuous Integration — merge to trunk every day and stop the line when it goes red; CI is a practice, not a toolHumble & Farley · 2010 Ch 4Implementing a Testing Strategy — four quadrants: automate what repeats, leave curiosity to peopleHumble & Farley · 2010
Part II · The Deployment Pipeline
Ch 5Anatomy of the Deployment Pipeline — one automated road from commit to production: build once, stop the line when redHumble & Farley · 2010 Ch 7The Commit Stage — a trustworthy red or green within ten minutes; that budget governs people, not machinesHumble & Farley · 2010 Ch 8Automated Acceptance Testing — separate what to verify from how to click it, or the suite becomes unaffordableHumble & Farley · 2010 Ch 10Deploying and Releasing Applications — blue-green and canary make going live revocable; the wall is never the code, it's the dataHumble & Farley · 2010
Part III · The Delivery Ecosystem
Ch 11Managing Infrastructure and Environments — only automation builds or changes a machine, from definitions in version control; edit one by hand and it is unreproducibleHumble & Farley · 2010 Ch 12Managing Data — code swaps back to the previous version, data does not; non-destructive migrations and a compatibility window make releases revocable againHumble & Farley · 2010 Ch 13Components and Dependencies — you split to get the build back under ten minutes; the price is losing one-commit-one-verification, bought back with pinned versions and an integration pipelineHumble & Farley · 2010
SRE — Site Reliability Engineering— Google · 2016
Ch 1Introduction — ask a software engineer to design an ops team, then pin it down with a 50% engineering floor and an error budgetGoogle · 2016 Ch 3Embracing Risk — 100% is the wrong target; treat 1 minus the SLO as a budget you spend, ship while it lasts, freeze when it's goneGoogle · 2016 Ch 4Service Level Objectives — pick the right thing to measure before arguing about nines; a good SLO names its window and measurement pointGoogle · 2016 Ch 5Eliminating Toil — the test isn't whether you hate the work but whether it grows linearly; 50% is a ceiling, not a targetGoogle · 2016 Ch 6Monitoring Distributed Systems — the four golden signals as a minimum sufficient set; never page for work a script could doGoogle · 2016 Ch 8Release Engineering — shipping fast is the result; reproducibility is the foundation, and config needs its own version tooGoogle · 2016 Ch 9Simplicity — reliability is capped by what you can still understand; code is a liability, and the best change is a negative oneGoogle · 2016 Ch 12Effective Troubleshooting — a craft, not a gift: stop the bleeding first, cut from the middle, then ask where the effort goesGoogle · 2016 Ch 15Postmortem Culture — chase the conditions, not the person; the output isn't a document, it's action items tracked to closedGoogle · 2016 Ch 21Handling Overload — stop measuring capacity in QPS; take on less on purpose and output holds at the ceilingGoogle · 2016 Ch 22Addressing Cascading Failures — failure breeds; removing the trigger won't bring the system back, you have to cut the loopGoogle · 2016 Ch 23Managing Critical State — heartbeat-and-timeout doesn't avoid consensus, it implements it wrong; take this one off the shelfGoogle · 2016 Ch 25Data Processing Pipelines — the enemy isn't data volume, it's the periodic schedule itself: idle troughs, herd peaks, and a middle you can't monitorGoogle · 2016 Ch 26Data Integrity — replicas don't protect your data, they just sync the bad delete faster; three lines of defense doGoogle · 2016
Accelerate— Forsgren, Humble, Kim · 2018
Ch 1Accelerate — speed and stability aren't a trade-off; improve by finding your weakest capability, not by climbing maturity levelsForsgren et al. · 2018 Ch 2Measuring Performance — stop counting effort; count four outcomes that constrain each other so no side can be won by wrecking the otherForsgren et al. · 2018 Ch 3Measuring and Changing Culture — stop asking what the culture is like; watch whether bad news reaches the person who needs itForsgren et al. · 2018 Ch 4Technical Practices — continuous delivery isn't a pipeline you buy; it's keeping trunk shippable, and who writes the tests separates teams more than whether tests existForsgren et al. · 2018 Ch 5Architecture — stop arguing monolith vs microservices; ask two questions: can you test alone, can you ship aloneForsgren et al. · 2018 Ch 6Integrating Infosec — security isn't a door before release; it's paving the road so the right thing is easier than going aroundForsgren et al. · 2018 Ch 9Making Work Sustainable — don't ask if the team is tired; ask what day and hour they last deployed. Burnout is an output of the environmentForsgren et al. · 2018
Ch 11Leaders and Managers — how transformational leadership amplifies technical practiceForsgren et al. · 2018
Database Internals— Alex Petrov · 2019
Part I · Storage Engines
Ch 1Introduction and Overview — DBMS architecture, row vs column stores, memory vs diskPetrov · 2019
Ch 2B-Tree Basics — why disk-based databases keep choosing B-TreesPetrov · 2019
Ch 3File Formats — pages, cells and slotted pages — how bytes are laid out on diskPetrov · 2019
Ch 4Implementing B-Trees — splits and merges, concurrency and page management in practicePetrov · 2019
Ch 5Transaction Processing and Recovery — write-ahead logging, the page cache and ARIES-style recoveryPetrov · 2019
Ch 6B-Tree Variants — copy-on-write, lazy B-Trees, FD-Trees and Bw-TreesPetrov · 2019
Ch 7Log-Structured Storage — LSM-Trees: how write-optimized storage actually worksPetrov · 2019
Part II · Distributed Systems
Ch 8Introduction and Overview — the core problems and abstractions of distributed systemsPetrov · 2019
Ch 9Failure Detection — deciding a node is dead: heartbeats and φ-accrual detectorsPetrov · 2019
Ch 10Leader Election — election algorithms and guarding against split brainPetrov · 2019
Ch 11Replication and Consistency — the spectrum of consistency models, quorums and CRDTsPetrov · 2019
Ch 12Anti-Entropy and Dissemination — gossip protocols, Merkle trees and read repairPetrov · 2019
Ch 13Distributed Transactions — 2PC/3PC, Percolator and CalvinPetrov · 2019
Ch 14Consensus — Paxos, Raft, Zab and Byzantine consensusPetrov · 2019
Fundamentals of Data Engineering— Reis & Housley · 2022
Part I · Foundation and Building Blocks
Ch 1Data Engineering Described — what a data engineer actually does, and where data science beginsReis & Housley · 2022
Ch 2The Data Engineering Lifecycle — ingestion, storage, transformation, serving — plus the undercurrentsReis & Housley · 2022
Ch 3Designing Good Data Architecture — architecture principles and trade-offs, lakehouse vs warehouseReis & Housley · 2022
Ch 4Choosing Technologies Across the Lifecycle — build vs buy, open source vs cloud, and what it costsReis & Housley · 2022
Part II · The Lifecycle in Depth
Ch 5Data Generation in Source Systems — where data comes from: databases, APIs and change data captureReis & Housley · 2022
Ch 6Storage — object storage, columnar formats and the warehouse/lakehouse storage layerReis & Housley · 2022
Ch 7Ingestion — batch vs streaming ingestion, ETL vs ELTReis & Housley · 2022
Ch 8Queries, Modeling, and Transformation — SQL and query engines, data modeling, dbt-style transformationReis & Housley · 2022
Ch 9Serving Data for Analytics and ML — BI, analytics and feeding features to modelsReis & Housley · 2022
Part III · Security and the Future
Ch 10Security and Privacy — the security responsibilities a data engineer actually ownsReis & Housley · 2022
A Philosophy of Software Design— John Ousterhout · 2021
Ch 1It's All About Complexity — the whole goal of software design is holding complexity downOusterhout · 2021
Ch 2The Nature of Complexity — three symptoms: change amplification, cognitive load, unknown unknownsOusterhout · 2021
Ch 3Working Code Isn't Enough — tactical vs strategic programming — shipping fast compounds the debtOusterhout · 2021
Ch 4Modules Should Be Deep — small interface over thick implementation; shallow modules breed complexityOusterhout · 2021
Ch 5Information Hiding and Leakage — what to hide, what counts as leakage, and why temporal decomposition traps youOusterhout · 2021
Ch 6General-Purpose Modules are Deeper — a somewhat general interface ends up simpler and easier to useOusterhout · 2021
Ch 7Different Layer, Different Abstraction — pass-through methods and decorator sprawl signal broken layeringOusterhout · 2021
Ch 8Pull Complexity Downward — better hard inside the module than hard for every callerOusterhout · 2021
Ch 9Better Together or Better Apart? — when to join and when to split — the test is not line countOusterhout · 2021
Ch 10Define Errors Out of Existence — design the exception away instead of catching it everywhereOusterhout · 2021
Ch 11Design it Twice — your first design is rarely the best — force out a second oneOusterhout · 2021
Ch 13Comments Describe What Isn't Obvious — good comments add what code cannot say, not restate itOusterhout · 2021
Ch 14Choosing Names — a name is the smallest abstraction; a vague one signals a design problemOusterhout · 2021
Ch 18Code Should be Obvious — obviousness is judged by the reader, never by the authorOusterhout · 2021
Ch 20Designing for Performance — clean and fast usually align — and when to trade clarity for speedOusterhout · 2021
Software Engineering at Google— Winters, Manshreck, Wright · 2020
Part I · Thesis
Ch 1What Is Software Engineering? — programming is code; engineering is code that survives time and scale — plus Hyrum's LawGoogle · 2020
Part III · Processes
Ch 8Style Guides and Rules — why rules exist, how to set them, and how to enforce them automaticallyGoogle · 2020
Ch 9Code Review — how Google reviews, and what it really buys beyond catching bugsGoogle · 2020
Ch 10Documentation — treat documentation as code and maintain it the same wayGoogle · 2020
Ch 11Testing Overview — why test at all, how tests are sized, and the payoff modelGoogle · 2020
Ch 12Unit Testing — maintainable unit tests: assert behavior, not implementationGoogle · 2020
Ch 13Test Doubles — mocks, stubs and fakes — the trade-offs and the cost of overusing themGoogle · 2020
Ch 14Larger Testing — integration and end-to-end tests that don't turn brittleGoogle · 2020
Ch 15Deprecation — how to retire a widely depended-on system without breaking everyoneGoogle · 2020
Part IV · Tools
Ch 16Version Control and Branch Management — why Google chose a monorepo and trunk-based developmentGoogle · 2020
Ch 18Build Systems and Build Philosophy — artifact-based builds, reproducibility and remote caching — the Bazel wayGoogle · 2020
Ch 20Static Analysis — what makes static analysis actually get adopted by developersGoogle · 2020
Ch 21Dependency Management — diamond dependencies, the limits of semver, and living at HEADGoogle · 2020
Ch 22Large-Scale Changes — safely making one sweeping change across millions of filesGoogle · 2020
Ch 23Continuous Integration — what CI looks like, and costs, at very large scaleGoogle · 2020
Ch 24Continuous Delivery — small, frequent, reversible — turning release into a non-eventGoogle · 2020
Ch 25Compute as a Service — from Borg to managed compute: treating machines as an abstract resourceGoogle · 2020
Fundamentals of Software Architecture— Richards & Ford · 2020
Ch 1Introduction — what architecture is, what architects do, and why there's no standard definitionRichards & Ford · 2020
Part I · Foundations
Ch 2Architectural Thinking — architecture vs design, breadth over depth, and thinking in trade-offsRichards & Ford · 2020
Ch 3Modularity — measuring cohesion and coupling: connascence, abstractness, distance from the main sequenceRichards & Ford · 2020
Ch 4Architecture Characteristics Defined — what the -ilities actually mean, implicit vs explicit characteristicsRichards & Ford · 2020
Ch 5Identifying Architectural Characteristics — extracting the few that matter from domain concerns and requirementsRichards & Ford · 2020
Ch 6Measuring and Governing — making characteristics measurable and guarding them with fitness functionsRichards & Ford · 2020
Ch 7Scope of Architecture Characteristics — the architecture quantum — scope is not the whole systemRichards & Ford · 2020
Ch 8Component-Based Thinking — how to partition components, at what granularity, aligned to the domainRichards & Ford · 2020
Part II · Architecture Styles
Ch 10Layered Architecture Style — the default monolith: benefits, layers of isolation, and the sinkhole anti-patternRichards & Ford · 2020
Ch 11Pipeline Architecture Style — pipes and filters — the classic skeleton for ETL and data processingRichards & Ford · 2020
Ch 12Microkernel Architecture Style — core system plus plug-ins — the usual shape for product softwareRichards & Ford · 2020
Ch 13Service-Based Architecture Style — coarse-grained services with a shared database — the pragmatic middle groundRichards & Ford · 2020
Ch 14Event-Driven Architecture Style — mediator vs broker topologies, and what asynchrony buys and costsRichards & Ford · 2020
Ch 15Space-Based Architecture Style — removing the database bottleneck with in-memory grids for extreme concurrencyRichards & Ford · 2020
Ch 17Microservices Architecture — bounded contexts, the granularity trap, data isolation and communicationRichards & Ford · 2020
Ch 18Choosing the Appropriate Style — deriving the style backwards from the architecture characteristicsRichards & Ford · 2020
Part III · Techniques
Ch 19Architecture Decisions — ADRs: recording the decision and the reasoning behind itRichards & Ford · 2020
Ch 20Analyzing Architecture Risk — risk matrices, risk storming and continuous assessmentRichards & Ford · 2020