Keelson · Blog
Mid-Market AI Workflow Audit Guide
01 · Discovery
What the operator already reads on day one.
The 30-day audit begins with a one-hour read of an operator dashboard the team has already been staring at for a year. Nothing new is installed in week one. The named operator sits down with the metric the CFO already reads on a Monday morning and walks through four artifacts: the one workflow, the one metric, the exceptions queue, and the override log. The entire discovery phase is the act of reading things the customer has already written down. The audit does not begin with a strategy deck.
The named operator opens discovery with a single sentence: claims triage, contract drafting, settlement matching, exceptions cleared per shift. If the conversation ends with “the operations function” or “the AI stack,” the shape is wrong and the audit cannot land in 30 days. A mid-market scope is always one workflow, owned by one team, measureable on one dashboard. The audit before the audit is naming what the workflow actually is.
Claims cycle time, contracts drafted per week, hours reclaimed per close, exceptions cleared per shift, defects per thousand shipped units. The metric decides the quote, the metric decides the handoff, the metric decides whether the engagement is repeatable. Discovery ends when the operator can name the number without opening a ticket, and the CFO can verify it without opening the build. Anything less precise is a wishlist, not a metric, and the audit will not run on it.
Discovery reads the exceptions queue side by side with the override log for the last four weeks. If the queue carries the same shape two weeks running — same field, same upstream feed dropping the same record, same model confidence dipping on the same label — the system is missing a label and the work is a build, not ops. If the queue is genuinely random week to week, the work is ops wear and tear and the right answer is a retainer, not an audit.
The override log is the artifact most operator teams under-use. It is the truth serum of the exception queue: what the human did when the machine’s suggestion disagreed with the operator’s read. Discovery ends with two pages of overrides, sorted by frequency, that already describe the labels the next model needs to ship. The audit reads the words the team has already written and turns them into the metric lock the engagement will be priced against.
02 · Process mapping
The audit document has the same shape, regardless of the workflow.
Weeks two through three turn the discovery artifacts into a single audit document. The document is short, it is the same document every time, and it is what the customer signs before the build begins. Current-state notes, the exception taxonomy, the drift map, the locked metric, and the named escalation contact. Five sections. Every audit lands these five sections or the engagement was not an audit.
Current-state notes are written by the named operator alongside the customer team. They are not a swim lane, they are not a SIPOC, and they are not a vendor process diagram. They are the words the team uses to describe the work as they do it on a Monday morning — in the language the team would use on a Tuesday morning if a new hire was sitting next to them for the day. The point of the notes is to capture what the system already does before anyone proposes what it should do.
The exception taxonomy is the page that turns the exceptions queue into a training set. Every recurring exception is grouped under a label the team already uses in the override log, and every label carries an example trace from the customer’s own environment. The taxonomy is exhaustive enough that the next retraining cycle reads the same document the audit was signed against, which is what makes the AI Ops retainer cheap to renew.
The drift map names the upstream feeds the workflow depends on, the cadence at which each feed changed shape in the last six months, and the downstream metrics that drifted when the upstream feed moved. The map is short and ugly: a list of breakfast-table observations, not a chart. It is what lets the build harden the right places in week seven instead of ten.
The locked metric is the one number on the operator dashboard the customer, the vendor, and the CFO all agree on before the build begins. The named escalation contact is the human who picks up the phone when the metric drifts past the threshold the team set on day fourteen. Together they are the only two lines on the audit document that survive handoff on day 31; the rest is execution detail.
03 · Tool-fit scoring
Reading the vendor pattern when the work hits the build.
By week three, the audit has named the workflow, the metric, and the exception taxonomy. The next conversation is about who actually ships the build. The vendor pattern read takes three looks at the staffing decision and a scorecard the customer can fill out without us in the room. The pattern read is also where the cost of the wrong call becomes visible — on the build method page we document the four signals an operator reads to pick the engagement shape.
A Big Four engagement will staff the audit with the senior partner and the build with the bench. The handoff loses six weeks of context before the first exception hits. The scorecard reads zero on “depth of integration”, “named escalation”, and “metric lock.” The price is the lowest on the table and the cost is the highest across three years.
A freelancer ships well for as long as the freelancer is on the account. When the freelancer’s calendar moves, the runbook walks out the door with them. The scorecard reads high on depth of integration and named escalation and zero on “metric lock” and “post-handoff coverage.” The price is low, the artifact survives only as long as the engagement does.
A single named operator writes the audit, ships the build, writes the handoff document, and is the person on the escalation email for the first month of AI Ops. The scorecard reads high on every row: depth of integration, named escalation, metric lock, post-handoff coverage. The price is fixed against the metric; the runbook survives the operator’s departure because the operator wrote it for the next person, not for themselves.
Scorecard rows: depth of integration into the operator’s own systems, named escalation contact who picks up the phone, lock on a single metric the CFO can verify, post-handoff coverage written into the runbook before the build sits down with the method. The vendor who fails the scorecard on any one row will fail the build on the same row; the scorecard is short on purpose.
04 · 90-day roadmap
The week-by-week milestones that close the audit.
The audit document closes the first two weeks. The next ten weeks are the build, broken into four phases a mid-market operator can read on one page. The product at the end of week twelve is the same product the audit was signed against: a shipped metric, on the operator data, in the operator environment, with the failure cases baked into a runbook the team owns on day 31.
Discovery, process mapping, tool-fit scoring, and the locked metric. The audit document is signed on day fourteen with the metric, the exception taxonomy, the named escalation, and the quote — all on one page.
The named operator ships the system against the locked metric on the customer’s data. Every Friday carries a working drop; every Friday the audit document is the source of truth for what shipped and what slipped.
The drift map earns its keep this month. Every upstream feed the audit named is hardened against the failure modes the package promised: feed goes dark, model confidence crosses the threshold the customer set, the audit log disagrees with the operator’s read of the workbook. Each failure mode picks up the phone.
The handoff document is written for the Monday-morning operator, not the executive. Day 31 the customer owns the runbook, the metric, and the system. Day 31 the operator owns nothing and the metric keeps moving. The artifact is a shipped metric, on the operator data, in the operator environment we document on the method.
05 · What good output looks like
The artifact at the end of week two.
The audit lands one document, signed by the operator, the CFO, and the named vendor operator. It is short enough to read on a Monday morning and structured enough that the build two months later still references it as the source of truth. Five sections, one page, signed before any code is written. The output of a 30-day audit is an artifact you can hand to the next vendor even if you never end up working with the first one.
The lock doc names the workflow, the metric, and the price. The exception taxonomy pages every recurring exception with a label and an example trace. The named escalation contact sits at the top of the runbook with a phone number and a threshold that triggers the call. The drift map lists the upstream feeds that broke in the last six months and the metric that moved when each one shifted. The CFO closes the document by reading the metric, calling it the number the invoice will be measured against, and signing on the cover page. That is the entire output. The build comes after.
Continue reading
Related posts
More working notes for operators building AI systems that hold up after handoff.
AI Ops / Strategy
How mid-market teams scale AI workflows after the first build: strengthen governance, change operator behavior, assign durable ownership, write runbooks and escalation paths, and expand a portfolio without losing metric discipline.
How mid-market operators choose one workflow, establish a baseline they can defend, measure implementation ROI, set pilot exit gates, and hand a production system to the person who will run it on Monday morning.
Governance / AI Ops
A practical operating checklist for mid-market teams: inventory every AI workflow, assign owners, set risk tiers and approval gates, control data and vendors, define human review and exception escalation, log what matters, and review the system monthly and quarterly before drift becomes an incident.
The practical AI ops digest
Working notes on outcome-locked pricing, the 90-day build, and what AI Ops actually does after handoff. No marketing — just the operator notes we send to the people who already read this blog.
Bring us the workflow
Run the 30-day audit on your own workflow.
The self-serve audit walks eight questions an operator can answer in twenty minutes. The output is the same one-page document the named operator produces on the paid engagement — the workflow, the metric, the exception taxonomy, the next step. Or talk to a build operator directly if you already know which metric needs to move.
Replies within one business day.