Do not turn AI into software

How to build robust AI solutions instead

Uwe Friedrichsen

17 minute read

Rough wooded landscape (seen in Bosnia and Herzegovina)

Do not turn AI into software

I recently did a presentation at Compass AI & Tech Summit in Budapest. The title of my presentation was “The 7 pitfalls of AI”. The presentation was about some still unresolved, yet often neglected issues regarding AI. We still need to find good answers for these issues if we want to arrive where we want to be at the end of our AI transition (if there is an end at all). If we fail to find good answers, we risk ending up in a place where we do not want to be, not being able to leverage the potential of this powerful and fascinating technology.

Confusing AI with humans or software

Two of the pitfalls I presented were misconceptions I often see. Both misconceptions influence the decision-making processes in more or less subtle ways and point in opposite directions. However, I often see people falling for both of them – surprisingly often at once. The one misconception is confusing AI, i.e., agentic AI based on LLMs, with humans. The other misconception is confusing AI with software.

We see both of them all the time in many variants. The most widespread one is the conflation of both ideas, i.e., that AI would be the “best of humans and software united” (which it definitely is not). But we also see many other variants of both misconceptions in the wild.

However, AI is neither human nor software. It is something different. Confusing it with one or the other leads to flawed decisions. This ranges from the idea that AI could replace humans to the expectation that AI could be made as reliable as software. Both are wrong.

AI is not software

There would be a lot to say about anthropomorphizing AI, i.e., attributing it human-like properties. Clarke’s Third Law, “Any sufficiently advanced technology is indistinguishable from magic,” does its part to fuel this misconception. However, as this is not what I want to discuss in this post, I will not dive deeper into this misconception here.

In this post, we will focus on the other misconception: confusing AI with software and its detrimental consequences. Let us start with revisiting the different properties of software and AI.

For many decades, all we had on a computer were software programs. Software has some very specific properties. I discussed many of them in my blog series “Software - It’s not what you think it is”. When it comes to AI, some of them stand out as they massively differentiate themselves from AI. I already discussed them in my blog post “It has never been about code”. Here is the short version:

  • Software runs deterministically. We can reliably repeat a software-based program as often as desired. It always works the same way. We can even prove the correctness of software (although we rarely do).
  • On the other hand, software provides a terribly cumbersome interface. Writing code using a 3rd Generation Language (3GL) can be painful because software is unbelievably stupid. If you do not phrase things exactly in the way the respective programming language expects it, the software created does the wrong thing very reliably. Every developer who got lost in an excruciating debug session knows what I am talking about.
  • Additionally, the user interfaces lean heavily toward the capabilities of the machines. We need to do the things the machine expects exactly in the right order, or the machine will not do what we want it to do. And let me not get started about ambiguity. Software has no idea how to deal with ambiguity. You need to be very precise and totally non-ambiguous, or the machine will not be able to do what you want it to do.

AI, i.e., an AI agent powered by an LLM, has almost the opposite properties:

  • AI provides an interface that leans much more towards the way humans interact with each other. It is very easy to communicate with an AI solution. You do not need to execute a link, button, and menu sequence in the exact right order to get the desired result. You simply tell it in your words what you want, and the AI solution figures out how to do it. It can even deal surprisingly well with ambiguity, which is an inherent part of human language.
  • Developing AI solutions is also much more straightforward than developing software solutions. You basically “program” it in natural language. Again, the AI tries to make sense of it, and surprisingly often it is able to turn a mediocre request into a desired result.
  • The price to pay for this increased convenience is non-deterministic behavior. Due to their functioning principle, LLMs are never fully deterministic, reliable, or secure. Their superpowers regarding their abilities to process natural language are also their kryptonite. They cannot just deal with ambiguity. They are ambiguous (in a certain way).

The new kid in computer town

If we compare these properties, we realize that AI is not software. It is something different. While it lacks the reliability and determinism of software, it is much more approachable from a developer’s and user’s perspective. It makes it a lot easier to develop a solution that runs on a computer. It makes it a lot easier to use a solution that runs on a computer. AI is the new kid in (computer) town, a new tool available on a computer. 1

For decades we only had software as the only tool to turn ideas into something that could be executed by a computer. We are so used to this fact and thus are so focused on software, its design, its creation, its deployment, its operation, that we completely missed that we recently got a new tool: AI, based on AI agents and LLMs.

Trying to turn AI into “software+”

This is probably one of the main reasons that people still conflate AI and software. The new kid in town is still too new. This is probably also the reason why so many people try to make AI solutions as reliable as software solutions. They want the best of both worlds: the reliability of software and the approachability of AI. They want “software+”.

As a consequence, people are always trying more and more elaborate and complicated ways to keep AI from failing, i.e., from answering and acting incorrectly. They want the same reliability from AI as from software:

  • A given request shall always lead to the same result.
  • The result needs to be correct all the time, of course.
  • Even if the request is phrased ambiguously, the AI shall still return the correct result.
  • AI must be able to respond correctly in a reliable way, no matter what the input is.

In short: Most people want what a former colleague of mine called a “DWIM machine”: Do what I mean.

Hmmm …

Exorcising the superpower

The problem with this approach is that if we were able to remove all erroneous behavior, most likely we would also remove the superpower of AI. The superpower is why we want to use AI in the first place: its ability to respond to natural language surprisingly well even if the request is phrased in a vague or ambiguous way. If we would eliminate the non-deterministic behavior of AI, most likely we would also eliminate its superpower. If LLMs were deterministic, they would be a lot weaker at handling human language. We have seen a lot of deterministic approaches to human language in the past. At best, they were “meh”.

Therefore, my conclusion is that if we want perfect predictability and reliability, we need to use software 2. AI solutions do not give us this level of predictability and reliability – at least not without depriving them of their desired properties. Maybe future AI built on a different technological basis will. But the AI of today does not. Therefore, if we want to use AI, we need to accept that failure is inevitable.

Pondering resilience

But we still want to use AI whenever it makes sense. Not all problems require a perfectly reliable solution. However, in many cases we still cannot accept persistent failures. As I pondered this issue while preparing the presentation mentioned above, it reminded me of some of the core ideas of resilience. Those ideas led me to the conclusion that trying to avoid failures completely, i.e., trying to turn AI into a “software+”, is the wrong approach.

In my blog series “The long and winding road towards resilience” I laid out a prototypical journey towards resilience for an IT organization. The journey starts in an environment that does not ponder resilience at all and ends with a fully resilient organization. The journey itself is split into four steps, I call “plateaus” because I use a mountain-climbing metaphor. From these four plateaus, the first two are of particular interest in this context:

  • The first plateau (the “plateau of stability”) describes an IT organization that tries to eliminate errors in its software completely. The motto is Avoid errors at all costs.
  • The second plateau (the “plateau of robustness”) describes an IT organization that accepts that failure is inevitable and thus embraces it. The motto changes to Maximize availability (and embrace failures). 3

Moving from the first to the second plateau starts with the realization that failures are inevitable. In the case of software, the reasons are different from in the case of AI. However, the pattern is the same: Failures happen, no matter how hard we try to avoid them. The realization means that simply trying to avoid errors is not enough. We can reduce the number of errors that will happen, but we cannot completely eliminate them. More importantly, some of them will become failures, i.e., become observable from the outside of the respective system.

To build reliable systems from unreliable components, i.e., components that may fail, we need to change our approach. This change basically consists of four measures:

  • We still try to reduce the number of errors, but we accept that it will not be able to avoid them completely. As a consequence, we do not rely on avoiding errors as the only measure.
  • Additionally, we try to identify errors and failures as soon as possible and recover from them or at least mitigate them. The goal is to maximize availability, which is defined by the relation between MTTF (Mean Time To Failure) and MTTR (Mean Time To Recovery) 4. The bigger MTTF in relation to MTTR, the higher the availability. Accepting that we cannot increase MTTF indefinitely, we additionally focus on reducing MTTR to increase availability.
  • We do not attempt to handle errors and failures at a technical level alone. Sometimes, it is not possible to resolve an error quickly, e.g., if a required service is not available. In such situations, the problem needs to be resolved at a business level by switching to a degraded service level to keep availability high. E.g., if a price cannot be fetched from the pricing engine, the default price from the product catalog is used. However, this is a business-level decision. The business experts need to decide how to respond to an error based on the criticality of the encompassing use case. Therefore, IT does not try to decide on error and failure handling strategies alone anymore but includes the business experts.
  • We actively contain the blast radius of errors to minimize their impact as much as possible. As the system should not fail as a whole, we split it into failure units that are isolated against cascading errors via bulkheads. Bulkheads can be all sorts of measures ranging from complete parameter and return value checks over fallback actions to dedicated thread pools for different connections.

These measures allow us to implement a robust system landscape, i.e., systems that not only try to avoid errors but are also able to recover quickly from them. Side note: This does not yet make us resilient, but robustness is sufficient for this discussion. 5

A different approach to AI

If we ponder these ideas for a moment, we quickly realize that we can (and should) apply them also to AI solutions because we face a similar problem: Failures are inevitable. Therefore, it is not sufficient to work towards a stable AI solution, i.e., one that does not make any errors. Instead, we need to work towards a robust AI solution, i.e., one that can quickly recover from errors:

  • We should try to reduce the number of errors an AI solution makes. However, we should accept that it will never reduce the error rate to zero.
  • We should add error detection and handling capabilities to our AI solutions. If an error occurs, we should be able either to recover from it or at least to mitigate it. Error detection is often located outside the solution itself. This means we need some kind of AI monitoring that allows us to detect errors on a technical as well as a business level. Additionally, we need to design recovery and mitigation flows that we can trigger as needed.
  • As with software, we need to include our business experts to understand how critical errors are from a business point of view, and which mitigation strategies or service level reductions are possible. This is not about improving a harness. This is about the solution itself.
  • We also need to control the blast radius of errors. This means that we need to split up our AI solutions into parts and isolate them from each other. At the moment, monolithic thinking still prevails when it comes to AI solutions. Also, subagents typically implement composition patterns and not delegation patterns, which are needed for working isolation.

This leads to a quite different approach regarding the usage of AI. Instead of trying to implement the perfect harness (which does not exist), we try to create a good harness that reduces the probability of errors. But we are not done there. We also need to look at the solution itself. We need to add external monitoring that detects if something has gone wrong. We implement remediation strategies that are based on the business-level criticality of the error. And we split up the solution to avoid cascading error effects.

Of course, all this is not trivial. E.g., building monitoring and bulkheads for AI solutions are both still very young domains. We still need to learn how to do it. However, if we accept that it is not possible to build error-free AI solutions, I think this is the route we need to take.

Memories, guesses, and apologies

But what about errors that turn into failures and leave the area of our influence? How shall we fix them?

This is a sensible question: How can we handle a failure that already triggered actions which lie outside our reach? We can do recovery or mitigation work inside the boundaries of our influence. But we cannot directly do anything about the effects the failure triggered outside those boundaries.

Pat Helland and Dave Campbell provide the answer in their seminal paper “Building on Quicksand”. Again, the context is different. The paper is about the effects of going from synchronous to asynchronous communication in a distributed system and how to build robust solutions on top of it.

In this context, Helland and Campbell also discussed the problem of what to do if a wrong answer leaves a boundary of influence and potentially triggers wrong actions in downstream systems. In their context, the trigger for the wrong answer was incomplete knowledge due to missing updates – a general problem in distributed systems that send their updates asynchronously. But we face the same problem if a wrong answer is created by an AI agent, and we recognize it only after it has left the boundaries of our influence.

In such a situation, Helland and Campbell suggested sending an “apology”. An apology attempts to trigger a compensating action. It “apologizes” for its failure by sending an additional answer that reverses the effects of a wrong answer. E.g., if we erroneously withdraw money from an account and realize it later, we apologize by depositing the money back to that account (and withdrawing it from the account it was erroneously sent to).

This approach is as simple as powerful. We accept that we cannot prevent failures from happening and thus potentially trigger effects outside our area of influence. If it happens, we do not attempt to “undo” the failure. Instead, we send something that undoes its effects. The failure remains visible. The apology also remains visible, and the effects of both of them cancel each other out. This is the only sensible approach because we do not have the means to erase a failure silently anymore once it has left the boundaries of our control.

Again, it means more effort than simply “upgrading” a harness. But then again, the goal is achieving robustness and not sticking to an illusion, i.e., that we can avoid all errors from happening.

If apologies are not an option

But what if an apology is not an option? E.g., there is a meme where a person shows its AI a picture of a fly agaric and asks it if that mushroom is safe to eat. The AI answers confidently: “Yes, it is safe to eat.” In the next picture, the person is dead, and the AI apologizes that it was wrong: “Sorry, you are completely right. This mushroom is not safe to eat.”

While funny as a meme, it illustrates quite drastically that there are situations where a failure leaving the boundaries of our influence has drastic consequences and cannot be reversed later. Does that mean that apologies do not make any sense?

No, apologies make a lot of sense. We simply cannot use them blindly everywhere. There are situations where we simply cannot afford failures because of the risk – because of its safety impact, its economic impact, its psychological impact, or the like.

This limits what we can solve using AI. In such high-risk settings, AI is simply not an option because we cannot guarantee the absence of failures 6 – at least with the type of AI we currently deal with.

If failures are not an option, we need to use software, not AI. Software is deterministic, and we can prove its correctness if needed.

This is determined at the business level. Additionally, it depends on the respective use case and context. This cannot be decided at a general level, i.e., we cannot decide in general if apologies are an option or not. This must be decided on a case-by-case basis.

Summing up

We discussed that people still often confuse AI and software even if they have very different, partially opposite properties. One effect of this confusion is that people attempt to turn AI into “software+”, i.e., an AI solution that acts completely deterministically and does not make any errors. The problem with this approach is that it is not possible. And even if it were, most likely it would deprive AI of the very properties we use it for in the first place.

The question how we still can build reliable systems from unreliable components leads us to the domain of resilience engineering. Part of resilience engineering is to build robust systems. When building robust systems, we accept that we cannot completely avoid errors and failures. Thus, we build the systems in a way that they can quickly detect and handle errors if they happen, instead of attempting in vain to avoid them completely.

We then added the idea of apologies to deal with failures that we only detect after they have left our area of influence and may have triggered undesired effects downstream. We also discussed that situations exist where such failures leaving our system context are not an option. In such situations, AI solutions as we know them at the moment are simply not an option.

As with software, building robust systems is more effort than simply pretending we could completely avoid errors. However, as in resilience engineering and distributed systems, completely avoiding errors is an illusion with AI systems. Therefore, we better start connecting the dots and apply the ideas we know from the other domains to AI engineering.

I am curious where it will lead us …


  1. A smartphone is also a computer – a computer that is also capable of making a phone call. A smart TV is also a computer, and so on. A computer is defined by its capability of running software – and AI agents today. ↩︎

  2. Let us ignore the fact that perfect reliability is not achievable. But we can get quite close with software, and if something goes wrong then, it usually is not because of the software. ↩︎

  3. If it confuses you that I use both the words “error” and “failure”: Put simply, an error is something that goes wrong in a system but is not (yet) observable from the outside. A failure is an error that has become observable from the outside. Many fault-tolerance techniques try to identify and contain errors before they become failures. ↩︎

  4. MTTF (Mean Time To Failure) describes how long it takes on average until an error occurs. MTTR (Mean Time To Recovery) describes how long it takes on average until an error is resolved. ↩︎

  5. If you are curious what is missing to become resilient, please read the referenced blog series “The long and winding road towards resilience”. ↩︎

  6. It is impossible to guarantee the complete absence of failures. Still, it is possible to guarantee them with software under defined conditions (which can be quite wide). We cannot do the same with AI because of its non-deterministic behavior. ↩︎