
Field guide
Three readouts from the work I do most. They cover the parts that rarely make the headlines, which happen to be the parts that decide whether AI pays off. Take what helps.
Readout 1
Before you bet on the score
How to tell whether an AI score is ready to drive a business decision.
An AI score is only as useful as the decision it supports. Most teams test the model and skip the harder question, which is whether the score agrees with the people who know the work best. This is the sequence I use, from the business question down to the checks.
Start with the decision
Name the decision each score will drive, such as who gets coaching, which cases get a second look, or where the budget goes. A score with no decision attached is a report, and nobody notices when it drifts.
Build a golden set
Pick a small set of real examples, often fifty to a few hundred, and have your most trusted experts label them. Write down why each label is right. This set becomes the reference everything else is measured against, so a model never writes it.
Calibrate the people first
Before you score a model, score your raters. Have several people label the same examples, compare their answers, and talk through every disagreement. If your experts only agree most of the time, no model will do better than that. Disagreement isn't noise. It shows you where the definition is unclear.
Break the judgment into parts
One overall quality score is hard to trust and impossible to fix. Split it into specific measures a person could check, each tied to evidence in the source. When the model misses, you can see which measure missed and why.
Compare the model with the experts
Measure agreement for each measure on its own, not just overall. Check it by segment, because a model can do well on average and poorly on the cases that matter most. Keep checking after launch, because drift is quiet.
Decide who owns the bar
The business, not the data team, should set how much agreement is enough, based on what a wrong answer costs. Write down the threshold, who reviews it, and how often. That is the heart of AI governance, and it keeps the score honest after the project team moves on.
Ten questions to ask before you trust an AI score
Readout 2
Clean data starts with people
Getting a business ready for AI is mostly a people problem.
Every transformation I've worked on hit the same wall. The data was messy because the business wasn't aligned. The same field meant three things to three teams, and nobody had written down which one was right. Tools don't fix that. People do, with a little structure.
Find the conflicts first
Pick the handful of measures leadership argues about most, and ask each team how they define them. The disagreements are your real starting point.
Name an owner for every critical field
Each field that feeds a decision needs one person who can say what it means. Ownership turns an endless debate into a decision.
Write the dictionary with human context
A useful data dictionary says more than a type and a definition. It says who uses the field, what decision it feeds, what good and bad values look like, and the quirks only the people closest to it know. That context is what makes data usable, by people and by AI.
Settle disagreements in the room
When definitions conflict, put the people who live with the data together and decide. Record the decision and the reason, so the next team doesn't reopen it.
Start where the value is
Don't try to clean the whole warehouse. Start with the data behind the decisions that matter most this year, show the payoff, then widen.
Measure the progress
Track how many critical fields have an owner and a reviewed definition, and how long it takes to answer common questions. Those numbers move before the big wins show up, and they keep sponsors patient.
Aligned data shortens every project that follows. It is also the foundation an AI system needs, because a model can only learn the context your people have actually written down.
Readout 3
Keep the human signal
Why leaning on AI for feedback and definitions can quietly erase your edge.
AI now writes a lot of the notes, feedback and documentation inside companies. That saves time, and it carries a risk most teams don't see yet. When the model writes your record, your record starts to sound like the model.
The drift
A model drafts a definition, someone accepts it, and it becomes the reference. The next draft builds on that one. Over time the organization's own knowledge gets replaced with general knowledge, and the specifics that made you different fade.
Why it matters to strategy
Your advantage lives in what only your people know, the exceptions, the customer patterns, the reasons behind the rules. Generic content can't hold that, and decisions built on it start to look like everyone else's.
The loop to watch
If people use AI to write the labels and feedback that are later used to check the AI, the check becomes circular. Agreement goes up and means less. A suspiciously high match between model and human answers deserves a closer look.
Guardrails that work
None of these slow a team down much, and together they keep the record yours.
- Keep a human-written golden set that no model drafts.
- Mark AI-drafted content until a person has reviewed it.
- Have an owner approve definitions and labels before they become reference.
- Audit a sample of AI-assisted work every cycle.
- Reward specifics. A note that names the customer's actual problem beats a polished summary.
Where AI helps
Use it to draft, summarize, find inconsistencies and flag gaps. Let people decide what's true. The goal is faster people, not absent ones.
AI gets more valuable as your human signal gets clearer. Protect it, and every model you add gets better. Let it erode, and every model you add gets more confident and less right.
Talk it through