For a year I'd argued that the way to fix requirements is to encode the discovery method itself, keep the human decision inside the workflow, and write down what good looks like before anyone argues about a demo. The only proof I had was work nobody outside my company could open. Confidential systems make a weak argument.
I designed and wrote all of it, alone and from scratch: no proprietary code, prompts, data, or names. Three fictional discovery calls in three recording formats are the test set.
Write the eval before tuning anything, and publish what it says. I would have spent the whole time improving extraction. The eval said extraction was fine and calibration was broken: my confidence scores weren't calibrated at the top of the range. So the README shows the weak spots instead of an accuracy headline.
- recorded callthree formats, all synthetic
- extract22 fields, each with a quote and a confidence score
- scoreDFV and RICE, plus the requirement nobody said out loud
- three approvalsspeakers, summary, stories
- stories outa structured record and a Jira-ready export
An MCP server exposes read-only tools to an agent client, with allow-listed paths and no reachable credentials.
The eval harness reports field accuracy, faithfulness, coverage, DFV label match, judge agreement with κ, and a risk–coverage curve. It publishes the weak spots: a non-monotone curve and a judge that agreed with the gold labels no better than chance on one metric. You can see where it breaks before you trust it.
Anyone can. The method is inspectable from the outside, so you can run it instead of taking my word for it.
Repository: github.com/christine-q-nguyen/transcript-to-requirements. Its sister project: How I built this portfolio, open source.
- Traceability is the reliable part. Every field quotes its source. Completeness still needs the person, which is why the three approvals exist.
- Publish the failure modes. A README that shows κ = 0 on one judge is worth more than one that claims 95% accuracy.