I read so I can name what I do, and know where it breaks.
Seven habits from the work, each tied to the papers that back it up and the papers that warn about it. Thirty-three papers, none of them decorative: every one changed a threshold, a gate, or a sentence in a PRD.
The habit: classify the turns of a recorded conversation, attach evidence to every claim, then generate structured requirements, and work out the need nobody said out loud. Pieces 01 and 04.
What the papers say: the turn-classify-then-generate architecture has been validated with practicing requirements engineers (RECOVER). Explicit requirements come out at F1 around 84%, and three quarters of the latent ones an LLM infers are judged useful (LENS). Traceability to the transcript is the reliable part, near 99%; completeness against a human is not, around 60% (goal extraction with GPT-4o). Which is exactly why the gates exist.
Two metrics I took from it: requirements faithfulness (are the stories supported by the transcript?) and interview coverage (is the transcript covered by the stories?), the pair Inter2US formalized. The open-source artifact reports both.
- RECOVER — Requirements Elicitation from Conversations
- LENS — LLM-Enabled Needs Discovery from Stakeholder Interviews
- Goal extraction from interview transcripts with GPT-4o
- Inter2US — alignment between elicitation interviews and requirements
The habit: put human approval inside the workflow, then treat reviewer attention as a budget. Route what changes the outcome, not everything. Pieces 01, 03, 04, and the bonus case.
What the papers say: reviewers only moderately agree on what's risky (κ ≈ 0.52), realized safety is an inverted U in the escalation rate, and flooding a reviewer produces rubber-stamping (Oversight Has a Capacity). Current agent designs push humans out of the loop unless approval is designed: which actions, when, at what granularity, with alternatives and batch review, and with counter-measures for skill atrophy (Agents Push Humans Out). Oversight works better as an independent component with intervention conditions, role resolution, and approve / reject / defer semantics than as a line in each prompt (Decoupled HITL). And an execution-boundary control layer took unsafe actions from 88% to zero in adversarial tests (Organizational Control Layer), which is why the agent gateway is governance, not plumbing.
What it changed for me: confidence thresholds became a load-aware escalation policy, not a quality knob. "Flag, never silently correct" became the first ruling in the decision log, because editors keep the judgment muscle that way.
- Oversight Has a Capacity
- AI Agents Push Humans Out of the Loop
- A Decoupled Human-in-the-Loop System for Controlled Autonomy
- Organizational Control Layer
The habit: a small deterministic model first, the LLM only on low confidence, thresholds calibrated on data and published per workflow. Pieces 02 and 04.
What the papers say: cascades with a scoring function and per-stage thresholds cut cost at equal quality (FrugalGPT); the small model's own uncertainty is a good criterion, and thresholds should be calibrated on early data rather than fixed in a meeting (margin-based cascades); whether any routing or cascading scheme works depends mostly on the quality of the estimator (unified routing and cascading); and underneath all of it is the risk–coverage trade-off, abstain below a confidence threshold to guarantee a target error rate (selective classification).
What it changed for me: accuracy varied by dataset, so the estimator had to be validated per dataset, so the thresholds were published per workflow instead of one accuracy headline. And a deterministic encoder's score was preferred over LLM self-report because LLM confidence is calibrated in some settings and not in others.
- FrugalGPT · RouteLLM
- Margin-sampling cascades with dynamic thresholds
- A Unified Approach to Routing and Cascading for LLMs
- Selective Classification for Deep Neural Networks · Language Models (Mostly) Know What They Know
The habit: let the model propose and label a taxonomy, then ship a small classifier; classify demand by what the model does to information, and by whether it augments or automates. Pieces 02 and 06.
What the papers say: an LLM can generate and refine a label taxonomy, then pseudo-label data to train lightweight classifiers that run at scale (TnT-LLM), which is the published version of both the classification platform and the 275-request taxonomy. Usage of assistants concentrates in a few task families and splits roughly 57/43 between augmentation and automation (Anthropic Economic Index), a second axis for prioritizing demand. And a field experiment found the biggest productivity gains go to the least experienced workers (Generative AI at Work), which is what one of the newer editors told us on a live report.
- TnT-LLM — Text Mining at Scale with LLMs
- Which Economic Tasks are Performed with AI? (Anthropic Economic Index)
- Generative AI at Work
The habit: write "good" as binary criteria before the demo; validate the judge before trusting the rubric; promote on thresholds across many runs. Pieces 03 and 04.
What the papers say: tens of thousands of expert-written, binary, weighted criteria graded by a model is the reference design for evaluating open-ended output (HealthBench), and it's the template for the editorial rubrics. Criteria-driven form-filling evaluation correlates better with humans than reference-based metrics (G-Eval). Strong judges reach human-level agreement, but carry position, verbosity, and self-enhancement biases (Judging LLM-as-a-Judge), so you validate the judge against human rulings first. Faithfulness, answer relevance, and context relevance cover retrieval systems (RAGAS).
Where it bit me: the artifact's judge agreed with the human 96% of the time and had κ = 0. It passed everything and missed the one human failure. High agreement with zero kappa is the exact trap the literature warns about, and it's in the README table, not a footnote.
- HealthBench · G-Eval
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena · A Survey on LLM-as-a-Judge
- RAGAS
The habit: least privilege before more agents. Treat tool access, credentials, and injection as the first design constraints, not the last. Piece 05.
What the papers say: LLMs with tool protocols are open to malicious code execution, remote access, and credential theft through tool poisoning and injection (MCP Safety Audit), and the mitigations are the boring ones: authentication, scoped permissions, logging, isolation (Enterprise-Grade Security for MCP). That's why one shared credential across every agent was wave zero, and why the intake skill removing an internal reference before wider sharing mattered.
- MCP Safety Audit
- Enterprise-Grade Security for the Model Context Protocol
The habit: build the first version yourself, in a week; prescribe a straw-man to understand the problem; show a POC instead of holding ninety meetings; keep the friction you can afford. The Approach page, and pieces 01 and 04.
What the papers say: the first empirical study of vibe coding finds expertise is redistributed toward context management, rapid evaluation, and deciding when to take the wheel, not removed (Sarkar & Drosos). A review of 47 sources describes an iterative generate-evaluate-revise loop whose outcomes depend on how the output is evaluated and governed; the evidence is strongest for prototypes and weakest for production (multivocal review), which is the line I draw for myself. A survey of 162 vibe coders finds making software got democratized while verifying it did not, a perception–action gap (From Prompting to Verification); enablement is the second half. A randomized trial of prompt-to-design tools found about 20% shorter task times, with the larger gain for product managers (Does AI Save Time on Product Design?). Prototypes work as the elicitation instrument itself, because people correct a prototype faster than they describe a problem (SERGUI). A Microsoft study of 885 PMs lands on a line I already use: accountability must not be delegated to non-human actors. And the ironies-of-automation literature applied to AI-assisted design names the thing I refuse to give up: de-skilling comes from removing the friction (De-skilling, Cognitive Offloading). Framing, Judging, Steering gives the split a vocabulary I can assess against.
- Vibe coding: programming through conversation with AI
- Vibe Coding in Software Development: A Multivocal Literature Review
- From Prompting to Verification: How Experience Shapes Vibe Coding Practices
- Does AI Save Time on Product Design? A Randomized Controlled Experiment
- Self-Elicitation of Requirements with Automated GUI Prototyping
- Prototyping with Prompts (collaborative software teams)
- Product Manager Practices for Delegating Work to Generative AI
- De-skilling, Cognitive Offloading, and Misplaced Responsibilities · Framing, Judging, Steering
I keep this list because it's more useful than the one above. Calibrating and publishing risk–coverage curves with the data behind them. Oversight-capacity design as explicit parameters: escalation rate, batch size, reviewer budget. Latent-requirement inference as a named step in the output, not a lucky finding. An augmentation-versus-automation field on every demand. Fluency in one eval tool well enough to run and read it, not build it. Each has a proof artifact in the plan, and each is honest about where the evidence stops.
One ID is cited from memory and should be verified before print: the Anthropic Economic Index (2503.04761).
Where
Christine Nguyen · Austin, Texas.
Open to Head of AI Enablement, Director of AI Transformation, and AI adoption leadership. Don't let the titles lock you in: if the work matches my skills, I'm open to it.
If you've read this far, we should probably talk.
Bring me the AI problem your team is still trying to explain.
christineqnguyen@gmail.comClick to copy.