03 · Applied machine learning

Taxonomy mapping took clients two and a half weeks. Now it takes two days.

Clients wanted less copying without losing the ability to check the answer. I helped turn four separate tool requests into one capability: a model trained on our own past mappings for the core task, visible confidence at every level, and a person for uncertain rows.

2.5 wk → 2 daysclient effort per taxonomy mapping

My roleGroup product leader, discovery to handoff
WhenAug 2025 – 2026
Usersanalysts and clients of a benchmarking practice
StatusOwned by the practice team; teams beyond the practice can use it in proposals

Benchmarking is the practice's bread and butter, and every study starts with a client mapping its cost data onto our taxonomy. 86% of surveyed clients named it their biggest pain. Analysts copy-pasted between taxonomies for days. Clients skipped the survey rather than do the mapping, which hurt the firm's ability to benchmark. Then four teams asked for help, in four vocabularies.

I was the group product leader from discovery to handoff. I conducted the user research and the client interviews. I turned four requests into one capability, wrote the definition of good with the analysts who'd use it most, and handed ownership to the practice team. The engineers built and trained the model.

When better prompts stopped helping, change the approach. The big general-purpose model did fine at the top level of the taxonomy and got unreliable one level down, no matter how we wrote the prompt. So I made the case for training a smaller model of our own on mappings analysts had already done. It gives the same answer for the same row, with a confidence score at every level. The general model steps in only when ours isn't sure, and an analyst checks anything below the bar.

When trust is the problem, predictable beats clever. Our own model gives a score you can check at every level. A big model grading its own confidence doesn't.What I told the engineers
  1. client rowscost data in the client's own taxonomy
  2. our trained modela score at every level of the taxonomy
  3. LLM fallbackonly for the rows our model isn't sure about
  4. below the thresholdan analyst decides, with a view of why
  5. mappedonto our taxonomy, ready for the study

Above the threshold a row maps on its own. The explanation view shows analysts why two similar rows went different ways.

A made-up row from the explanation view. The client's line "Third-party IT support" gets a suggested mapping with a confidence score at each level of the taxonomy: 0.97 at level one, 0.62 at level two, 0.55 at level three. The bar is 0.80, so the row goes to an analyst, with the reason shown. a client's row "Third-party IT support" level 1 · Technologylevel 2 · Infrastructurelevel 3 · Managed services 0.970.620.55 the bar: 0.80 below the bar at level 2, so an analyst decides why: rows like this split between Infrastructure and Applications
A made-up row: the model is sure at the top level and unsure one level down, so a person makes the call, and can see why.
Built withCustom-trained modelLLM fallbackPer-level confidence thresholdsExplainability viewFeedback backlogHierarchical taxonomy
2.5 wk → 2 days
client effort per mappingmy year-end review; the practice team reported the change to its own leadership
5 → 2 days
analyst effort per mapping
90 min → 12
to process 1,000 entries; process mapping per study 2.5 weeks → 4 daysthe team's year-end summary and a dated win post

The bar was written before the first demo, with the analysts who'd use it most. About 70% right at level one is useful: the analyst checks the rest and saves half the time. 80–85% is better: the client verifies, and manual mapping is only about 90% right anyway. The argument between an 85% and a 95% bar became a configurable threshold instead of a fight. One live pilot client said the score at every level is what earned their trust.

The practice does. I handed it to the practice team as citizen developers, and they reported the change to their own leadership. Teams beyond the practice can use it in proposals. The path to train our own model became the reusable capability described in the governed model-training case.

  • Trust is the product. The time savings got attention. The score at every level is what got adoption.
  • A better tool is not adoption. Handing ownership to the practice was the step that made it stick.