Publish confidence thresholds, not accuracy headlines
The argument: "It's 95% accurate" is unfalsifiable in the way that matters. Publish the curve and a threshold per workflow, and say what happens to the other five percent.
The most dangerous sentence in enterprise AI is "it's 95% accurate." Not because it is false. Because it is unfalsifiable in the way that matters. Accurate on which data? At which level of the taxonomy? Compared to what human baseline? And most of all: what happens to the other five percent?Opening lines
Comes from the taxonomy-classification case and the evaluation table in the open-source intake pipeline. Papers: unified routing and cascading, margin-based cascades, selective classification, judge validation. Publishing first
Three gates: human approval belongs inside the workflow, not in the policy deck
The argument: the gate took an afternoon to build. Deciding what to route through it took months, and I got it wrong first. Oversight has a capacity; spend it where it changes the outcome.
Every responsible-AI policy I have read says the same thing: keep a human in the loop. Every agent I have watched fail in production had one. The policy was satisfied. The human was not in the loop; the human was on a slide.Opening lines
Comes from the AI-intake case and the editorial-workflow case. Papers: the four oversight papers in section B. Publishing second
How to read your own AI portfolio, and find the category your taxonomy will miss
The argument: sort requests by what the model does to information, not by what the requester called it. Count. Then write down what doesn't fit instead of forcing it in; that's where the money is.
I sat down with every AI request my team had received over two years, a few hundred of them, logged by many hands, inconsistently, and tried to answer a question leadership kept asking: what should we build or scale next? The finding is no longer the interesting part. The method is, because it is the part you can run on your own board this week.From the opening
Comes from the AI-portfolio strategy case. Papers: TnT-LLM, the Anthropic Economic Index. Publishing third
Intake as a system: running AI demand like a portfolio
The argument: "How do we prioritize?" is the wrong first question. The first question is how do we see. You can't rank demand you haven't captured in a form that can be compared.
Nobody fills in the form. Or they fill it in badly, and the "problem statement" field says "we want to use AI for our workflow." The demand you can compare is the demand you extracted from a conversation with the person who has the pain.From section 1, "Capture is a product, not a form"
Comes from the AI-intake case and the AI-portfolio strategy case. Papers: RECOVER, LENS, TnT-LLM. Publishing fourth
Let's get loopy: the three loops that replaced my assembly line
The argument: working software is now how a team communicates, decides, and validates. The line became three loops — a human in the loop, a feedback loop, and a continuous improvement loop — and a PM's job moved from deciding what to build to deciding what deserves to ship.
For most of my career, product work ran like an assembly line. Research, then a spec, then design, then engineering, then a launch, and every handoff was a place to wait. That made sense when building was expensive. It isn't anymore.Opening lines, draft
Comes from the Approach page. Papers: section G, the vibe-coding studies and the prompt-to-design trial. The assembly-line and jazz-band framing is credited to Ravi Mehta; the loops and the evidence are mine. Draft · publishing fifth
Oversight has a capacity: what 250 AI requests taught me about putting humans in the loop for real
A field report on making "human in the loop" real: three gates placed at kickoff and handoffs, confidence thresholds published as curves instead of accuracy headlines, and a decision log that governs agents across workflows. From 250-plus enterprise AI requests, with an open-source tool and the 2026 research behind it.
- Where human approval gates belong in an agentic workflow, and why "review everything" reduces safety.
- How to set and publish confidence thresholds per workflow using a risk–coverage curve.
- How a decision log turns per-prompt governance into a reusable component.
Forty-five minutes, or twenty with the demo. Program committees: the CFP-ready abstract, bio, and takeaways are one email away. Available