Skip to content Skip to footer

Taming Coding Agents Using Chemical Engineering Principles

Kicking And Screaming Into The World Of Engineering

TL;DR

Shapez 2 is a good illustration of how to make software engineering a proper engineering discipline by applying lessons from chemical engineering.

Software Engineering
Software Architecture

Overhead view of a Shapez 2 factory with rows of processing machines connected by conveyor belts.
Shapez 2 Factory

Introduction

I’ve recently gotten addicted to a game called Shapez 2. It’s a game where you build a factory that manufactures shapes. Yes, it is that simple of a game. I find Shapez 2 similar to what I have already discussed about the game Factorio, which is to say that it feels a lot like programming. Shapez 2 takes different shapes and allows you to manipulate those shapes using a series of pieces of manufacturing equipment. In software, we manufacture data instead of shapes, and we use algorithms instead of specific pieces of manufacturing equipment. Also, in my Factorio blog post, I pointed out that chemical engineers are trained to do manufacturing at scale. So, instead of data, chemical engineers manufacture chemicals, and instead of algorithms, they have unit operations.

In this post, I discussed how chemical engineers further built on the concept of unit operations to create fundamental mathematical equations to help them design chemical plants that scale well. These equations are part of a concept chemical engineers call transport phenomena. I also talked about how we, too, can create our own math equations for software engineering. What is important to know for the purposes of this blog post is that both unit operations and transport phenomena made chemical engineering a proper engineering discipline. Unit operations helped to enforce discipline in how to organize manufacturing equipment, and transport phenomena helped to enforce discipline in ensuring that optimal performance was achieved for each piece of equipment.

What does this have to do with AI? I believe that our industry is trying to use outdated tactics with AI and is getting subpar results. We have spent years trying to minimize the lines of code written by downloading library after library or through the heavy use of inheritance or reflection, as is the case with Java Spring and ASP.NET. Like it or not, AI can write code faster than any human. The annoying part is that you’re never quite sure what code it is going to write. As I played Shapez 2, I started to realize that I was essentially building my factory in a way that reflected how I would like to write software in the future using agents. To be fair to Factorio, these concepts also work there, but I am terrible at Factorio, and it is a much harder game to explain in a simple blog post.

The Basics

In Shapez 2, you have pieces of equipment which make specific changes to your shapes. In the example below, you can see how this cutter simply destroys half of this shape.

As you can see, half of the shape gets destroyed, but upstream, the shapes are queuing up behind the cutter. This means that the flow of output shapes does not match the flow of input shapes. To fix this, we just need to know how many shapes a specific machine can handle. As it turns out, you need three cutters to handle one full input belt, as you can see in the example below.

If we were using chemical engineering terms, then we could say that different kinds of equipment are unit operations and the math required to understand how many pieces of equipment are needed to handle a full belt is a part of transport phenomena. If we instead used computer science terms, then we could say that different kinds of equipment are algorithms and the math required to understand how many pieces of equipment are needed to handle a full belt is a part of what I am calling data transport phenomena.

So what’s the implication of all of this? The implication is that what ultimately matters is that we have a system that runs at optimal performance and produces the correct output. In other words, I could have laid out my cutters like this, but the performance and the output are the same.

Worrying about the exact lines of code written by an AI agent is like worrying about how exactly the equipment is laid out in Shapez 2. It does not matter as long as you can fix and extend the software. Personally, when I read the code an AI agent writes, I am verifying that it is, in fact, getting the performance I want and the data I expect. I am not as worried about the exact layout of the code, as that is more subjective. To be more precise, I want the agent to be guaranteed to follow my data flow specification exactly as I want it, but I don’t care what specific lines were written to achieve this. If I know my specifications are followed, then there is no reason to review the code. While this might sound like waterfall, I have never found myself drastically changing the shape of the data or how the data flows in a system that is well-designed. In well-designed systems, I find that you usually add a new entry point to the data flow, like having something triggered by a cron job instead of a button click, or add new kinds of data. The only really turbulent part of the software is how the user interacts with your data. In poorly designed systems, I find there is a rush to ship software as quickly as possible, which means the data is not correctly defined, and the data flows are difficult to follow, which ultimately leads to a rewrite.

Blueprints at Scale

As it turns out, in Shapez 2, you have platforms where you put your equipment and conveyor belts as well as space belts that feed into those platforms. For the purposes of this post, the important thing to know is that a space belt can feed 12 full conveyor belts into a platform. This means we can increase our throughput by 12x, as in the example below.

In Shapez 2, you can make blueprints, so you can save a combination of space belt and platform configurations to essentially copy and paste the combination again and again.

Furthermore, a full space belt can supply 4 platforms, so we can increase our throughput by another 4x by using our platform as a blueprint like so:

This will work well enough, but Shapez 2 would be a very boring game if you only had to cut shapes. What if we have to cut and then do a rotation like this?

This approach works, but now it will be much more difficult to scale my factory if we want to add even more combinations of operations. What could we do instead? Put the rotation equipment on its own platform and make it a blueprint!

Since I designed my individual platforms to run at peak performance, I know that when I chain them together, everything will, in turn, run at peak performance. Sure, I’m going to take up more space, but I am reducing the number of decisions I need to make in order to get my optimal performance.

Blueprint of Blueprints

Chaining platforms works well, but sometimes you have the same kinds of patterns come up again and again to the point where you might make a blueprint of blueprints, as is the case for taking a full shape and splitting it into four pieces.

Again, this further reduces the mental overhead of understanding the system because you are thinking at an even higher level, but still have guarantees of data quality and performance.

Shapes at Scale

Having a strong foundation to build on through the use of optimally designed blueprints means you can design complex shapes at scale like this:

Because each blueprint is well-designed, it is almost trivial to make a complex shape since it simply involves gluing together the right platforms. What if we could do the same for software? What if we had the ability to glue together algorithms, all of which are well-designed in terms of performance? Could we have high-quality software that is written much faster?

Agentic Coding at Scale

To bring it back to software, imagine we have code like this:

type Person = {
  firstName: string;
  lastName: string;
};

function getFirstNames(people: Person[]): string[] {
  const firstNames: string[] = [];

  for (const person of people) {
    firstNames.push(person.firstName);
  }

  return firstNames;
}

const people: Person[] = [
  { firstName: "Ada", lastName: "Lovelace" },
  { firstName: "Alan", lastName: "Turing" },
];

const firstNames = getFirstNames(people);
console.log(firstNames); // ["Ada", "Alan"]

Being a seasoned TypeScript developer, you might decide to simplify the code for readability.

const people = [
  { firstName: "Ada", lastName: "Lovelace" },
  { firstName: "Alan", lastName: "Turing" },
];

const firstNames = people.map((person) => person.firstName);
console.log(firstNames); // ["Ada", "Alan"]

Now, you stand up in front of everyone and say, “Wow! This code is readable!” I agree that this is true, but you still have issues. You now have a function pointer to call. I’m sure the V8 engine probably optimizes the code in fun and interesting ways, but it can only do so much. Really, what the programmer is trying to say is, “I simply want to convert from one kind of data to another and do it in the most optimal way.” Before LLMs, I think the map function approach was the better approach even with a performance hit, but now with LLMs writing our code, I think we have an opportunity to rethink this approach. What if instead we simply had specialized models that are designed to do specific algorithms in the most optimal way? To be honest, you might not even want an LLM that does this part. Really, what you want is a model whose implementation you don’t have to double-check, just as you don’t have to double-check standard library functions. This theoretical model would essentially be like using a blueprint in Shapez 2. You could even chain multiple algorithms together just like you do blueprints in Shapez 2.

const people = [
  { firstName: "Ada", lastName: "Lovelace" },
  { firstName: "Alan", lastName: "Turing" },
  { firstName: "Grace", lastName: "Hopper" },
];

const firstNames = people
  .filter((person) => person.firstName.startsWith("A"))
  .map((person) => person.firstName);

console.log(firstNames); // ["Ada", "Alan"]

In the example above, we now have a filter added to our chain of algorithms. If we again were able to have another model to give the most optimal performance for a filter operation and had a guarantee that it would always successfully filter, then we’d have very efficient code that is doing what we want. Not only is the software doing what we want, but we have a much better chance of scaling for larger systems. This would force the software engineers to adopt standardized algorithms, as is already done with sorting, but for all algorithms. Not standardizing would mean that deterministic models would be nearly impossible to create, which is essentially the mess we’re in now. Since we would need to standardize our algorithms for every kind of operation we might perform on data, there would be an opportunity to start to leverage equations in the form of data transport phenomena in order to ensure you pick the correct series of algorithms to achieve the best performance for your specific situation.

Just like in Shapez 2, there would be nothing stopping us from then having algorithms of algorithms. Potential examples would include an implementation of the HTTP protocol, OAuth 2.0, or serializing JSON. Really, the standardization of components is the way engineers are able to optimize and build without being an expert in every aspect of a system.

The Paradigm Shift

If we start moving towards a paradigm more focused on data flow in software development, then I think we will start to see a clearer distinction between software engineers and computer scientists. I think computer scientists will do research and development of novel algorithms. Computer scientists will research the strengths and weaknesses of a given algorithm as well as determine its governing equations. Software engineers will be taught how to use the equations that the computer scientists discover for specific algorithms and how to glue algorithms together to make complex data pipelines that scale well. I think software engineers will have to make what are effectively construction plans for data flow and will only look at code to verify it is within the desired specification.

If I turn out to be correct, one major implication I see is that we will no longer need to use libraries for our software. The libraries would be replaced by the use of a series of glued-together standardized algorithms that are implemented by some kind of model. The only exception I could see would be operating system calls, as I don’t think regular users will be willing to give up an operating system in the near future.

In my mind, I don’t see why a standardization of algorithms for all aspects of data processing is not possible. With the nonstop clickbait articles about how coding is dead, I think it’s clear software engineers are worried about what kinds of skills will be valued. I think that means we are in the middle of a paradigm shift, and to me, that means we need to focus on the flow of the data and not the lines of code. Coding might soon die, but software engineering is just being born.