Benchmarking is the practice's bread and butter, and every study starts with a client mapping its cost data onto our taxonomy. 86% of surveyed clients named it their biggest pain. Analysts copy-pasted between taxonomies for days. Clients skipped the survey rather than do the mapping, which hurt the firm's ability to benchmark. Then four teams asked for help, in four vocabularies.
I was the group product leader from discovery to handoff. I conducted the user research and the client interviews. I turned four requests into one capability, wrote the definition of good with the analysts who'd use it most, and handed ownership to the practice team. The engineers built and trained the model.
When better prompts stopped helping, change the approach. The big general-purpose model did fine at the top level of the taxonomy and got unreliable one level down, no matter how we wrote the prompt. So I made the case for training a smaller model of our own on mappings analysts had already done. It gives the same answer for the same row, with a confidence score at every level. The general model steps in only when ours isn't sure, and an analyst checks anything below the bar.
When trust is the problem, predictable beats clever. Our own model gives a score you can check at every level. A big model grading its own confidence doesn't.What I told the engineers
- client rowscost data in the client's own taxonomy
- our trained modela score at every level of the taxonomy
- LLM fallbackonly for the rows our model isn't sure about
- below the thresholdan analyst decides, with a view of why
- mappedonto our taxonomy, ready for the study
Above the threshold a row maps on its own. The explanation view shows analysts why two similar rows went different ways.
The bar was written before the first demo, with the analysts who'd use it most. About 70% right at level one is useful: the analyst checks the rest and saves half the time. 80–85% is better: the client verifies, and manual mapping is only about 90% right anyway. The argument between an 85% and a 95% bar became a configurable threshold instead of a fight. One live pilot client said the score at every level is what earned their trust.
The practice does. I handed it to the practice team as citizen developers, and they reported the change to their own leadership. Teams beyond the practice can use it in proposals. The path to train our own model became the reusable capability described in the governed model-training case.
- Trust is the product. The time savings got attention. The score at every level is what got adoption.
- A better tool is not adoption. Handing ownership to the practice was the step that made it stick.