RSS

What build systems taught us about Terraform automation

Infrastructure Engineering Terraform OCaml

The problem

Terrateam automates Terraform operations in pull requests. You open a PR that touches infrastructure, we run a plan, you review it, you merge, we apply. Simple enough in concept. For a single Terraform root module, it's totally manageable.

But nobody has a single root module. Not in production. Not at scale. In the real world, you have hundreds of root modules, thousands of workspaces. You have dependencies between them. Your application can't deploy until your database exists, your database can't deploy until your VPC exists. And you need to handle concurrent modifications safely.

We built a system to handle this called Abb_flow. It worked, but adding new functionality was difficult. The architecture made changes harder than they should have been.

We also had a configuration problem that will be familiar to anyone who's operated infrastructure at scale. Our workflow definitions were doing double duty, specifying both what to do (init, plan, apply) and where to do it (dev, staging, prod, us-east-1, eu-west-2). This meant workflow definitions were proliferating. The same twelve lines of YAML, copy-pasted with slight variations. Something had to give.

The paper that changed our perspective

There's a common attitude in our industry toward academic computer science: "That's interesting, but I have a ticket due Thursday." There's some truth to that tension. But we've learned that the academics are often right, and ignoring their work means rediscovering things the hard way.

In 2018, Andrey Mokhov, Neil Mitchell, and Simon Peyton Jones published a paper called "Build Systems à la Carte." It examines build systems like Make, Shake, Bazel, Buck, and Excel (yes, Excel is a build system) and asks: what are the fundamental dimensions along which these systems vary?

They found two:

  1. How does the system determine what needs to be rebuilt? They call this the "rebuilder," covering concepts like minimality and early cutoff
  2. How does the system schedule tasks? Topological order? Restarting? Suspending?

Two dimensions. And suddenly Make, Excel, and Bazel aren't three completely different things. They're points in a design space. You can reason about them, see the tradeoffs, and make informed decisions about where your system should live.

This is what good academic work does. It gives you a vocabulary and framework for thinking about problems. It's not about implementing their code. It's about applying their ideas.

The realization

Reading this paper, we realized that Terraform automation is a build system.

In retrospect it seems obvious, but we hadn't been thinking about it that way. We'd been thinking about it as a webhook handler that runs shell commands. That's technically accurate, but it misses the underlying structure.

Once you see Terraform automation as a build system, everything snaps into focus:

This is what build systems do. This is what Make has been doing since 1976. Forty-seven years of accumulated wisdom about dependency resolution, incremental computation, and parallel execution, and we'd been ignoring all of it because we thought we were solving a different problem.

We were not solving a different problem.

The implementation

Armed with this mental model, we rebuilt the execution engine. Here are the key concepts:

Work manifests: artifacts, not actions

In our old system, we thought about actions like "run a plan" and "do an apply." In the new system, we think about work manifests, which are immutable descriptions of work to be performed.

This distinction matters. An action is ephemeral. It happens and then it's gone. An artifact is tangible. You can inspect it. You can replay it. You can ask "what was the state of the world when this decision was made?" This is the same insight that made content-addressable storage valuable. When you reify your operations into data, everything gets easier.

Layered execution

The execution model is layered:

start → plan → apply → [next layer] → plan → apply → ... → complete

Within a layer, operations run in parallel (they have no dependencies on each other). Between layers, we enforce ordering. The layers themselves are computed via topological sort of the dependency graph.

This isn't novel. This is how Make works. But it took reading an academic paper about build systems for us to realize we should be doing what Make does.

Stacks: parameterization without proliferation

We introduced a concept called "stacks" to solve the configuration explosion problem. A stack is a named group of workspaces with associated configuration:

stacks:
  names:
    prod:
      tag_query: production
      variables:
        environment: prod
      rules:
        apply_after: [dev]
    dev:
      tag_query: development
      variables:
        environment: dev

Now your workflow definitions can be generic. The stack provides the parameterization. This is separation of concerns, the kind of thing we learned in our first software engineering course, but it's easy to forget when you're deep in YAML configuration.

Dependency rules

Stacks can declare dependencies on other stacks:

That last one is subtle and powerful. If your app stack is modified_by: [database], then any change to database also triggers app. And if you specify modified_by but not plan_after, we infer plan_after from modified_by. Because if you're modified by something, you probably can't plan until it's applied. This is the system being helpful instead of pedantic.

Learning from being wrong

When we originally designed stacks, we had to decide whether stacks should be nestable. Can you have a "prod" stack that contains "prod-database" and "prod-compute" as sub-stacks?

We said no. In our design document (RFD 653), we wrote:

Nested stacks were rejected because their use cases were unclear.

Shortly after shipping stacks to production, we started hearing from users: "We need to group these stacks together. We need hierarchy. We need to say 'all of prod depends on all of dev.'"

So we wrote RFD 725, which began:

When stacks were introduced in RFD 653, nesting and layering stacks were explicitly rejected as an option. However, in real-world usage of stacks, nested stacks are both desirable and necessary for useful workflows.

We went from "use cases unclear" to "desirable and necessary" in a few months. That's how engineering works. You make decisions with incomplete information, you ship, you learn, you adapt. The mistake isn't being wrong. The mistake is not writing down why you made the decision in the first place.

Because we had RFD 653, we could go back and see exactly what we were thinking. We could examine our assumptions and see where they broke down. We couldn't have done that if the decision had been made in a Slack thread that scrolled away.

Write things down. Especially the things you decide not to do.

Takeaways

Read academic papers

Not all of them, but when you're stuck on a hard problem, there's a decent chance someone thought about it years ago. The "Build Systems à la Carte" paper wasn't about Terraform or infrastructure, but its abstractions applied directly to our problem.

Recognize when you're solving a solved problem

We were building a build system and didn't know it. This is why breadth of knowledge matters. This is why you should read about systems outside your immediate domain.

The right abstraction is a force multiplier

Once we saw our problem through the lens of build systems, features fell out naturally. Parallel execution is just topological layers. Cycle detection is validating the dependency graph. Incremental rebuilds are the whole point of a build system. The abstraction didn't just clarify our thinking, it generated solutions.

Humility is useful

We were wrong about nested stacks. We'll be wrong about other things. The goal isn't to be right all the time. The goal is to create systems and processes that let you recover from being wrong, such as design documents, version control, incremental rollouts, and feature flags.

The industry forgets too much

Make is from 1976. The ideas in it are almost fifty years old. And yet we keep building systems that ignore those ideas, and then we rediscover them. We can do better.


Stategraph is open source under MPL-2.0. The code is at github.com/stategraph/stategraph.

The "Build Systems à la Carte" paper is here. It's short, well-written, and might save you from reinventing Make.