06 · Training as a service

When prompting stops working, teams need a safe path to train their own model.

Once prompting hit its limit, the question became how to make training our own model a safe, normal thing to do instead of a one-off exception. I wrote the rules for it and designed routing that saves the expensive LLM, and people's time, for the rows that need them.

70–85%fewer LLM calls on classification, by design

My roleProduct lead; wrote the rules for training a model
When2026
Domainclassification
StatusIn place, with a written risk sign-off

Prompting has a ceiling. On the classification work it showed up at level two of the taxonomy: the frontier model stopped improving no matter how the prompt was written. The usual answer, keep prompting, burns tokens and still can't give a user a score they can audit. The other answer, training your own model, had no safe path at the firm: no approved place to train one, no agreed risk sign-off, and no rule for whose data could train what.

Tokens add up, too. Every LLM call costs money, and as AI use spreads across the firm, so does the bill. This work is part of a bigger firm push to spend tokens only where they actually help.

I was the product lead. I wrote the rules for training a model here: which teams can do it, with whose data, on which approved infrastructure, and who signs off on the risk before anything runs. I designed the routing: our trained model takes the rows it's sure about, the LLM gets the ones it isn't, and a person sees anything still below the bar. I worked through the data-rights questions with legal and kept each client's data separate. The engineers built and trained the models.

Make the safe way the easy way. Teams kept asking for one-off exceptions to train a model. Instead, I set up one path anyone could follow: approved infrastructure, a risk sign-off written once, each client's data kept separate, and an honest "not sure" instead of a made-up answer when the model isn't confident.

It's a cost decision too. Every row our own model handles is a row we don't pay an LLM for.

  1. rows to classifyclient data, isolated per client
  2. our trained modelhandles most rows on its own
  3. confidence checka threshold per workflow
  4. LLMcalled only when confidence is low
  5. below the threshold"no signal" beats a guess

Our model goes first; the expensive, less predictable LLM runs only where it's needed.

An illustrative chart of how sure our trained model is about each row, from less sure to more sure. Rows it's very sure about, about three quarters, stay with our model. Rows in the middle go to the LLM. The least certain rows go to a person. a person decidesthe LLM takes a lookour model handles it less suremore surehow sure our trained model is about each row illustrative: by design, about three quarters of rows never need the LLM
Every row gets a confidence score. The sure ones stay with our own model, the middle goes to the LLM, and the least certain go to a person.
Built withCustom-trained modelConfidence gatingOur model first, LLM secondApproved infrastructureDocumented risk sign-offPer-client isolation
70–85%
fewer LLM calls on classification, because our model goes firsta design estimate from the share of rows our model can handle; the dollar saving is small at today's volumes, but predictable costs matter
1
safe, approved way to train a model, where there had been none

It got signed off. The path has a written risk sign-off and keeps each client's data separate.

Honestly, the cost savings were a bonus. What mattered was getting the same answer every time, with a score anyone can check.How I describe it

The firm has a path where there was none. Teams that hit the prompting ceiling can train their own model without turning it into a procurement project, and two teams already use the same model.

  • Count tokens before someone makes you. Deciding where each row goes is about trust first, and it saves money too.
  • Reuse is the proof. A model a second team picks up is worth more than a benchmark.