BA AI Engineer & Full-Stack TypeScript Developer

About

Twelve years of things that had to stay up

I spent over a decade on production systems for a company that operated entirely online — where an outage isn't a ticket, it's a customer who can't travel. That environment teaches one thing faster than anything else: software only creates value when it fits the business reality behind it.

The shift

From features, to workflows, to systems

Nobody plans this move. It happens because you keep noticing that the feature was never the constraint.

Then

Build the thing on the ticket. Ship it. Next ticket. The work is measured in output, and output is easy to produce.

Now

Understand how the work actually happens, find the point where automation creates leverage, and design around the constraints that are real rather than the ones that are convenient.

I built a lot of well-made features nobody could name a number for. Applied AI is where I went next — it moves the constraint rather than the backlog, and it changes what a small team can do without hiring for it.

Approach

AI is not the product. It's an amplifier inside a system.

Which sounds like a slogan until you watch what happens when someone skips the system part.

What goes wrong

  • An agent wired directly to a database with no boundary, so a prompt change becomes a data-access change.
  • Retrieval with no citations, so nobody can tell a hallucination from a bad chunk.
  • A pipeline with no idempotency, so the retry that was meant to save you sends the email twice.
  • No event contract, so six months later nobody can prove the feature did anything.

What I do instead

  • Typed tools with hard-coded audience scoping — the model cannot ask for data it shouldn't see.
  • Structured output validated server-side; a failed guard throws and retries rather than passing something plausible downstream.
  • Exactly-once side effects, written pending-first so a crash mid-way is recoverable rather than ambiguous.
  • An event taxonomy defined before launch, with bot filtering, so the numbers survive being looked at.

FAQ

Things worth asking before you trust me with this

The ones that matter are about judgement rather than tooling, so those are the ones answered here — including where I'm weakest.

You're an engineer, not an ML researcher. Why does that make you the right person for AI work?

Because almost nothing that goes wrong with a production LLM feature is a model problem. It is an agent that could reach data it should never have seen, a retry that sent the same email twice, a retrieval step nobody can debug because the answer arrived without its sources, or six months of usage that proves nothing because no one agreed what to count.

Those are engineering failures wearing an AI costume. I do applied LLM work — agents, tools, retrieval, structured output and the systems that keep them upright. Not model training or fine-tuning. If you need someone to train a model, that is a different person and I will say so.

What does “AI is an amplifier inside a system” actually mean in practice?

It means the model is never the thing holding the guarantee. Anything with consequences — writing to a CRM, sending an email, quoting a price — is done by ordinary code with explicit rules, or by a person. The model handles the conversation and decides when to call a tool; it does not get to be the authority.

The practical test is what happens when the model misbehaves. If a bad output can only produce a bad sentence, the system is built right. If it can produce a bad row in your database, it isn't, and no amount of prompt engineering fixes that.

How do you decide what not to build?

By writing down what a thing is supposed to move before building it, and being willing to conclude that nothing here will move it. The list of things I turned down is shorter than the list of things I talked a client out of.

The other half is writing reversals down as reversals. When a decision I argued for turns out to be wrong, the record says so, with the date. A document that only contains the good calls is marketing.

Twelve years at one company — doesn't that narrow you?

It narrows the logos and it widens almost everything else. The same site over twelve years meant designing it, rebuilding it twice, measuring it, arguing about it, and eventually being the person who had to live with decisions made six years earlier.

That is the part that does not transfer from short engagements: you rarely get to find out which of your choices aged badly. I also ran a separate client concurrently for a year, so working across two contexts is not theoretical either.

What's the weakest part of your AI skill set right now?

Evaluation. I have guards, typed tools and server-side validation of structured output, so a malformed response cannot pass through as if it were fine. What I do not yet have is a systematic eval harness — a regression suite that tells me whether a change to an agent's instructions made it measurably better or worse.

Cost attribution per conversation and latency instrumentation sit in the same gap. I would rather you hear that from me in the first conversation than discover it in month three.

In public

I write down what I figure out

The part I still like most is finding out how something actually behaves — which usually means building it rather than reading about it. I write those up for the next person who goes looking. The first few are drafted and none are public yet.

Start a conversation