08 · Open source

The intake pattern, rebuilt from scratch so you can inspect it.

Confidential work is hard to evaluate from the outside, so I rebuilt the intake method in public. You can inspect every field, source quote, confidence score, approval, test, and weak spot instead of taking my word for it.

22 fieldseach with a quote and a confidence score; evals published as-is

My roleDesigned and wrote it, alone
WhenSep 2026
DataSynthetic only; MIT licensed
StatusWorking; evals published

For a year I'd argued that the way to fix requirements is to encode the discovery method itself, keep the human decision inside the workflow, and write down what good looks like before anyone argues about a demo. The only proof I had was work nobody outside my company could open. Confidential systems make a weak argument.

I designed and wrote all of it, alone and from scratch: no proprietary code, prompts, data, or names. Three fictional discovery calls in three recording formats are the test set.

Write the eval before tuning anything, and publish what it says. I would have spent the whole time improving extraction. The eval said extraction was fine and calibration was broken: my confidence scores weren't calibrated at the top of the range. So the README shows the weak spots instead of an accuracy headline.

  1. recorded callthree formats, all synthetic
  2. extract22 fields, each with a quote and a confidence score
  3. scoreDFV and RICE, plus the requirement nobody said out loud
  4. three approvalsspeakers, summary, stories
  5. stories outa structured record and a Jira-ready export

An MCP server exposes read-only tools to an agent client, with allow-listed paths and no reachable credentials.

Built withPython (stdlib core)CLIMCP serverEval harnessRisk–coverage curveJira exportpytest
22
fields, each with the quote it came from and a confidence score
3
points where a person says yes before anything is written
13
tests, and CI re-runs the published eval on every push

The eval harness reports field accuracy, faithfulness, coverage, DFV label match, judge agreement with κ, and a risk–coverage curve. It publishes the weak spots: a non-monotone curve and a judge that agreed with the gold labels no better than chance on one metric. You can see where it breaks before you trust it.

Risk among auto-accepted fields at each confidence threshold: 21% at 0.5, 19% at 0.6, 15% at 0.7, 20% at 0.8, and 30% at 0.9. Those thresholds auto-accept 100%, 71%, 41%, 30%, and 15% of fields. Risk is lowest at 0.7 and climbs again above it. risk among the fields it auto-accepts, by confidence threshold 10%20%30% 21%19%15%20% 30%: the curve climbs again auto-accept from 0.7; everything below goes to a person 0.50.60.70.80.9 takes 100%71%41%30%15% of fields auto-accepted at each threshold · 3 synthetic calls, heuristic baseline
The eval as the README publishes it. Risk bottoms out at 0.7 and climbs again above it, because the baseline over-trusts any sentence with a number in it. So 0.7 is the auto-accept line, and everything below it goes to a person. The same run found a rubric judge that agreed 96% of the time with κ = 0: it never said no, so it can't be trusted yet.

Anyone can. The method is inspectable from the outside, so you can run it instead of taking my word for it.

Repository: github.com/christine-q-nguyen/transcript-to-requirements. Its sister project: How I built this portfolio, open source.

  • Traceability is the reliable part. Every field quotes its source. Completeness still needs the person, which is why the three approvals exist.
  • Publish the failure modes. A README that shows κ = 0 on one judge is worth more than one that claims 95% accuracy.