Back to Blog

The One Workflow Nobody Modeled

August 30, 2026
WorkflowsOrchestrationDecision-MakingEngineering CultureTooling

Picking a workflow engine turned into a six-week debate with five options and no clear owner — the one workflow nobody thought to model was the decision itself.

The One Workflow Nobody Modeled

The workflow migration I mentioned last time is still going. I've spent the last few weeks doing the unglamorous part — reading wikis, tracing the current implementation, working out what the system actually does versus what the diagrams claim it does. That part I don't mind. It's the part after it that's been interesting, for reasons I didn't expect.

First, a reality check I had to give myself. I used to run a team responsible for workflows ten times more complex than this one, peaking at something like 600,000 runs in a single night. I'm fairly sure that number is higher now. The thing I'm migrating runs about a hundred times a year. That's not a typo — a hundred, annually. And a fair chunk of those hundred are apparently attempts that customers claim don't even work, which is possibly why the number isn't higher. So the throughput conversation, the one I'd normally have first, barely applies here. Good. One less thing to argue about.

The Two Options That Were Supposed to Be the Whole Menu

Going in, there were two paths, and I built out a proposal for one of them: something that follows the company's golden path, cloud native, built and owned by us. I'll say up front I might be biased — it's my proposal — but on paper it holds up. The tradeoff is honest and not small: more effort to build, more to test, more to deploy, and then the part everyone forgets to price in, more to maintain indefinitely. In exchange, we own the thing our customers actually touch, instead of handing that experience to a platform another team maintains on a roadmap we don't influence.

The other path was the default one — the migration target the team that owns the system currently being deprecated was steering everyone toward, which is worth sitting with for a second. It has real integration coverage already built, which makes it the path of least resistance on paper. It also has a reputation, and not from people with an axe to grind — from the people who actually use it, who've built workflows in it themselves. Onboarding and making changes reportedly takes days, sometimes weeks, waiting on approval. And the design is apparently brittle enough that a small change in one place can break things for everyone else on the platform, simultaneously, which is a fairly aggressive way for a shared system to behave.

Two options. Both real, both defensible, both with a downside you could write on a whiteboard in one line. That part I was fine with. That part didn't last.

Then Everyone Had an Idea

The moment this went to internal design discussion, it stopped being two options.

The first new proposal was pure oversimplification: skip a dedicated workflow engine entirely and wire the UI straight into an existing background worker. It sounded lightweight on paper, but only because it ignored the actual requirements—forcing interactive, multi-step API coordination through a batch script that was never meant to handle it. Convenience answered a design question that needed an actual architectural answer.

The second proposal was the inevitable modern classic: slapping AI capabilities onto a problem that didn't need them. The idea was to put an AI agent in front of a chat interface to capture user requests. But because the underlying process had to be strictly deterministic, the proposed solution was to hardcode all the execution logic in standard code underneath the agent. That solved the determinism problem by eliminating the entire reason for using AI in the first place—leaving an agent burning tokens to do what a simple web form does for free, with an audit trail that's twice as hard to debug. It wasn't an orchestration layer; it was a chat window with better manners.

Four options now, up from two. I pushed for us to actually converge — pick two, maybe three, and spare everyone a slide that reads like a lunch menu. What we did instead was pick one, label it "Recommended," and leave the other three up there anyway, because apparently it felt wrong not to show people the whole menu.

The Meeting About the Meeting

We took it to the wider org, expecting a conversation about whether to migrate off a system that's already scheduled to be deprecated. Instead, most of the room wanted to talk about whether the new design would also fix a list of unrelated problems — a list that, to be fair, we'd included in the doc ourselves, because it seemed generous to mention. In hindsight I'd have left it out. Deprecation is a sufficient reason on its own; the new thing doing the same job, or a bit more, was never supposed to be the bar. But once it's written down, it's fair game, and the conversation drifts to it every time.

Weeks of this. Hours of talks. No decision. I kept turning the same question over afterward: was it that we didn't have enough data to choose, or that we had plenty of it and just buried it under too many options for a decision that only needed two? I'm still not sure which one it was.

An Outside Data Point

Around the same time, I sat in on an interview for someone from a different org entirely — different team, different history, no shared context with any of this. They'd worked with the exact same tooling we'd been debating, as a user of it at that other org, and had gone as far as advising their own team there against adopting it. Their reasoning lined up with mine almost exactly, arrived at independently, by someone with zero incentive to agree with me. That's about as close to a second opinion as you get in this line of work, and it landed right where my read already was.

Then the colleague sitting next to me — the same one behind the short-circuit idea — asked the obvious follow-up: so what do you use instead? The candidate named another system entirely, one their team happened to like. That was the whole exchange. No context on their scale, their constraints, or whether any of it maps onto ours — just a name, dropped in passing, in response to a question that wasn't even about us. And that's how we ended up with a fifth option. Not from a proposal, not from a design discussion, not from anyone opening the thing up and reading how it works. From one sentence in an interview for a role that has nothing to do with this migration, because it happened to fit somebody else's use case. If I'm being honest about the standard I'm holding everyone else to, this one hasn't cleared it either — it's just sitting on the list now, which at this point might be the closest thing we have to a group policy.

The One Workflow Nobody Modeled

I spent a whole post arguing that everything is a workflow — that human steps belong in the same graph as automated ones, with inputs, an owner, and a timeout, instead of living invisibly in someone's head. I still believe that. What I didn't notice until now is that the decision-making process I just went through never got that treatment. No one owned it. Past a certain point it stopped being about data at all — someone just needed to make the call, and nobody was willing to be that someone. No timeout, so it just kept running past its own usefulness, one more session at a time.

Everything is a workflow, sure. But a workflow that loops forever isn't a workflow — it's a stall with better branding. It needs an exit condition, same as any other node in the graph: do we move forward with a pick, do we scrap the whole thing, does it all quietly go back to being manual because that was somehow less painful than this. I genuinely don't know which of those it's going to be. That's the part that actually bothers me.

Somewhere between the five options, the scope creep, and an outside opinion that somehow grew the list instead of shrinking it, I think the real answer is closer to "people didn't care enough to converge" than "we lacked data." But I'm honestly not certain, and that uncertainty is the tell. If the process had been modeled the way I keep insisting everything else should be, I'd know. Instead I've got a recommendation, a shelf of alternatives nobody ruled out, and a very good case study for the next post about writing things down.