Prompting has a ceiling. On the classification work it showed up at level two of the taxonomy: the frontier model stopped improving no matter how the prompt was written. The usual answer, keep prompting, burns tokens and still can't give a user a score they can audit. The other answer, training your own model, had no safe path at the firm: no approved place to train one, no agreed risk sign-off, and no rule for whose data could train what.
Tokens add up, too. Every LLM call costs money, and as AI use spreads across the firm, so does the bill. This work is part of a bigger firm push to spend tokens only where they actually help.
I was the product lead. I wrote the rules for training a model here: which teams can do it, with whose data, on which approved infrastructure, and who signs off on the risk before anything runs. I designed the routing: our trained model takes the rows it's sure about, the LLM gets the ones it isn't, and a person sees anything still below the bar. I worked through the data-rights questions with legal and kept each client's data separate. The engineers built and trained the models.
Make the safe way the easy way. Teams kept asking for one-off exceptions to train a model. Instead, I set up one path anyone could follow: approved infrastructure, a risk sign-off written once, each client's data kept separate, and an honest "not sure" instead of a made-up answer when the model isn't confident.
It's a cost decision too. Every row our own model handles is a row we don't pay an LLM for.
- rows to classifyclient data, isolated per client
- our trained modelhandles most rows on its own
- confidence checka threshold per workflow
- LLMcalled only when confidence is low
- below the threshold"no signal" beats a guess
Our model goes first; the expensive, less predictable LLM runs only where it's needed.
It got signed off. The path has a written risk sign-off and keeps each client's data separate.
Honestly, the cost savings were a bonus. What mattered was getting the same answer every time, with a score anyone can check.How I describe it
The firm has a path where there was none. Teams that hit the prompting ceiling can train their own model without turning it into a procurement project, and two teams already use the same model.
- Count tokens before someone makes you. Deciding where each row goes is about trust first, and it saves money too.
- Reuse is the proof. A model a second team picks up is worth more than a benchmark.