Back to Blog

Congrats, You Can Be an AI Expert Now Too

August 2, 2026
AIEngineeringSoftware DesignCareer

On the sudden abundance of AI experts and expert coders, and why vibe-coding without the underlying concepts is a debt you pay back with interest.

Congrats, You Can Be an AI Expert Now Too

Somewhere in the last couple of years, two new job titles appeared that nobody applied for: AI expert and expert coder. You didn't get promoted into either. You just woke up one day, opened a chat window, and apparently qualified.

I don't say that to be dismissive. The tools genuinely are that good now. Someone with zero background can describe an app in plain English and watch something real show up on screen twenty minutes later. That's not nothing — it's actually one of the more remarkable things to happen to software in a long time. But there's a gap opening up between being able to produce code and understanding what the code is doing, and it's a gap a lot of people are stepping over without noticing it's there.

Two Titles Nobody Applied For

There are two kinds of person who've actually earned the title "AI expert." One understands how the models actually work: it's statistics dressed up as deliberation, a system generating one token at a time that genuinely doesn't know what it's going to say next until it arrives at that token and the math resolves — there's no draft sitting upstream waiting to be read out loud. The other understands how to work the models — prompting well, managing context, getting a genuinely unreliable tool to behave reliably in production. Both are real, hard-won skills. Neither is what's walking around with the title right now. A few prompts in, a few blog posts read, and suddenly there are strong opinions on model architecture and what will and won't be "solved" in six months — the confidence of expertise, arriving well before the expertise does.

The expert coder version is the one I actually worry about, because it has a blast radius. Vibe-coding — describing what you want and letting the model write it — is a legitimately useful way to prototype. I use it. Most engineers I know use it. The problem isn't the tool. The problem is when someone ships that output believing they've written software, when what they've actually done is accept software. Writing it means you can trace how a request moves through the pieces, what a real user does to it that you didn't expect, what breaks first under load and why. Accepting it means it ran once, for you, on the input you happened to try.

Those are not the same skill.

Understanding What It's Doing

I'm not anti-AI-in-coding. Quite the opposite — I think refusing to use it at this point is its own kind of stubbornness. But there's a non-negotiable line for me: you need to understand what the thing is doing. Not necessarily every line, but the shape of it — what it's touching, why this approach and not another, what happens when the input isn't the happy path.

Anyone can prototype now. Anyone can vibe-code a working demo in an afternoon. That part's democratized, and good riddance to the gatekeeping that used to surround it. But a demo isn't a system, and the gap between them is exactly the stuff that doesn't show up in a quick prompt-and-generate loop: what good design looks like versus bad design, and — the one people skip past fastest — the tradeoffs. Every real design decision is a tradeoff. Consistency versus availability, simplicity versus flexibility, speed of delivery versus speed of change later. AI can generate code that runs. It cannot make the tradeoff for you, because the tradeoff isn't a coding problem, it's a judgment problem, and judgment is the part you can't skip by generating harder.

Skip that step enough times and you get a codebase that works right up until it doesn't, and nobody in the room can explain why, because nobody in the room actually decided anything. Compiling was never the hard part, and it's even less of one now — any model clears that bar without trying. The bar was always sitting underneath it: does this do the correct thing, for reasons someone in the room can name, when the input stops cooperating?

Which brings back the mechanism, because the same misreading shows up here in a costlier form. A reasoning model says "let me think step by step," produces something shaped exactly like a chain of logic, and it takes real effort not to conclude there's a mind in there weighing options. What's actually happening is more tokens. The model generates intermediate tokens that become additional context for the final answer to condition on, and that extra context genuinely does improve results on hard problems — it's a good trick, and it earns its keep. Like most good tricks, it plays better if you don't peek at the mechanism mid-performance.

It is not, however, deliberation, and researchers have been busy finding the seams. Work on trimming "overthinking" has found that cutting a model off halfway through its reasoning trace still lands the correct answer most of the time it would have anyway — an odd property for an explanation to have, if the explanation were the thing doing the deciding. A related paper, titled without much ceremony, "Large Language Models Decide Early and Explain Later", comes at it from the other side: force a model to commit partway through and the answer is frequently already locked in, with a few hundred tokens of reasoning still to go. The reveal happens in the mechanism, not on stage. So when you hand the model a tradeoff, you're not consulting a junior colleague. You're sampling a probability distribution and reading the sample as though it deliberated — which is fine, genuinely, as long as you know that's the transaction.

The AI-Stamping Trap

The same gap shows up somewhere less technical and more expensive: reaching for AI for the sake of being seen reaching for it. Not because it solves anything in particular, but because it needs to be in there — on the slide, in the changelog, in the pitch. AI-powered this, AI-enhanced that, a chatbot bolted onto a feature that was doing fine and had never once been asked a question. Good tooling decisions start with a problem and work forward to the tool that fits it. This one runs the play backwards: start with the tool everyone's excited about, then go find a problem worth mentioning it for.

You can usually spot it with one question: what does this solve that the earlier version didn't? A clear answer means someone did the work. A vague one means the slide came first.

And to be fair to the people asking for it, the enthusiasm is not irrational — it's just often aimed a little wide. An MIT study on enterprise AI adoption found that 95% of generative AI pilots in 2025 hadn't yet delivered measurable P&L impact, despite tens of billions in enterprise investment — not because the models are bad, but because a lot of that spend chased the label rather than a specific, well-scoped problem the label happened to solve.

The Knowledge Base That Demoed Well

This is the version I've watched play out more than once, and it's instructive precisely because the technology is fine — it's the planning around it that goes missing. Someone stands up an AI knowledge base, it demos beautifully on day one against five hand-picked documents, everyone signs off happy. The person building it, in the case I'm thinking of, hadn't yet run into the term RAG — perfectly understandable, since it's recent jargon for anyone outside the field, except it happens to be the name of the exact mechanism doing the retrieving, and skipping past it means skipping past a real design decision. Scale that knowledge base to ten thousand documents, some of which quietly contradict each other, and you run into a well-documented behavior researchers call "lost in the middle": a Stanford-led study found retrieval accuracy is highest when a fact sits at the very start or end of the context and drops noticeably for anything buried in between, a pattern that showed up across several model families, including ones built specifically for long context. Newer models are getting better at this — genuinely, that's good news — but "the next model release will handle it" is a bet, not a plan. Context is also a line item: it behaves like a budget with diminishing returns, and every token that flows through it is a token somebody's paying for. Get the planning right and it's a great tool. Skip it and you've bought something slower and pricier than what it replaced, with wrong answers that now arrive with a confident citation attached.

Often there was a simpler, sturdier option sitting right there the whole time — a lookup table, a proper search index, a rules engine that gives the same answer for the same input every time, for a fraction of the cost and none of the drift. That option rarely loses on merit. It loses because "we built a lookup table" doesn't fit on a slide the way "we're in the AI space" does.

The Determinism Detour

My favorite side quest, and it grows straight out of that same misreading of the mechanism, is the crusade to make AI outputs deterministic. Pin the seed. Clamp the temperature to zero. Wrap the whole thing in retries and validators until it says the same thing twice in a row, at which point everyone celebrates.

I get the impulse — predictability is comforting, and comforting is underrated. But these models are probabilistic by design, and that's not a flaw hiding in there waiting to be patched out; it's the exact quality that lets a model handle a prompt nobody wrote a rule for. Chase determinism hard enough and you end up spending more effort fighting the tool's nature than you would have spent just designing around it — validating outputs, adding guardrails, building a system that expects to be sometimes wrong, rather than one pretending it never is. Often the better move is simpler anyway: if a task truly needs the same output every time, that's usually a sign the task wants a deterministic system, not a probabilistic one wearing a costume.

Best story I've got on this: someone once told me, genuinely pleased with themselves, that they'd cracked it — their AI feature was now fully deterministic, same input, same output, every time. I asked to see the code. Turned out the "AI" step was quietly calling a plain deterministic function under the hood; the model had been retired from the part of the pipeline that actually mattered. Which is a fine engineering choice, to be clear — but it does raise the delightfully awkward question of why the AI was invited to that meeting in the first place. If the fix for "the model isn't reliable enough" is "stop letting the model decide," you haven't tamed anything. You've swapped in a calculator and kept the AI branding on the box.

Where That Leaves Us

None of this is an argument against using AI to code, or against having opinions on a fast-moving field before you've spent a decade in it. It's an argument against skipping the part where you actually understand what you're shipping. The tools got faster. The judgment they require didn't get any cheaper.