Choose the job that runs most often, where a mistake is cheapest to catch and the information it needs already exists. That is usually a dull job. Dull is the point.
First AI projects tend to fail in one of two ways. Some pick an impressive workflow that runs twice a quarter, so nobody learns anything for months. Others pick something where a single error reaches a client or an invoice, and the first mistake ends the experiment. Three questions avoid both.
Three questions, scored 1 to 3
How often does it run? Weekly or more scores 3; monthly scores 2; a few times a year scores 1. Frequency is how you get enough repetitions to see whether the workflow is actually better, and it is where the time goes.
How cheap is a mistake to catch? If a person reviews every output before it matters and an error costs a few minutes, score 3. If errors cost money or an awkward conversation, score 2. If an error reaches a client, a regulator or a bank account before anyone sees it, score 1.
Does the context exist, and is it clean? If everything the job needs is written down in one place and current, score 3. If it exists but is scattered or out of date, score 2. If it lives in someone’s head, score 1, and fix that before automating anything.
Multiply the three. A 27 is an obvious first workflow. Anything under 8 should wait, however exciting it sounds.
Match the checking to the stakes
The cost-of-error question is a rough version of the trust tiers described in Trust as Architecture, chapter 4 of Specs In, Software Out by Jhonn Moreno. Tier 1 failures are an annoyance you retry. Tier 2 failures waste time or money. Tier 3 failures cause financial or reputational damage or legal exposure. Tier 4 failures are irreversible or unsafe. The tier sets how much a person has to check. A first workflow should sit in Tier 1 or 2, with a person approving output before it goes anywhere that matters.
The one-page decision template
Fill one of these for each candidate, then compare them side by side.
- Workflow
- One sentence, starting with a verb: “turn client call transcripts into draft tasks”.
- Runs per month
- Counted from your task tool or calendar, not estimated.
- Worst realistic mistake
- What happens if the output is wrong and nobody notices. Then the trust tier, 1 to 4.
- Where the context lives
- The folders, files or records the job needs, and whether each is current.
- Scores
- Frequency, cost of error and context, each 1 to 3. Multiply them.
- Approver
- The named person who checks output before it leaves, and what they check.
- Baseline
- The number today, with its denominator: “3 of 13 actions reached the board”.
- Acceptance test
- Five to ten cases with known answers, run before launch and after every change.
- Review date
- Thirty days out, with the decision it will make: continue, adjust or stop.
Three worked examples
An eight-person marketing agency
Candidates: drafting tasks from client meeting notes, writing monthly reports, and automating client invoicing. Meeting notes run several times a week (3); a person approves each draft task before it reaches the task board (3); the transcripts already exist (3). Score 27. Invoicing runs monthly (2), an error reaches the client’s accounts team (1) and the rates live in a partner’s head (1). Score 2. This is close to what Suncoast Interactive did: its baseline was 3 of 13 agreed actions reaching the task board after one client review, and the first structured run captured 8 of 8, from a different meeting with a different denominator.
A twelve-person management consultancy
Candidates: a weekly client status note from project files, a first draft of proposal pricing, and a research summary before each discovery call. Status notes run weekly per client (3), a partner reads each before it is sent (3), but project files are spread across personal folders (2). Score 18. Pricing drafts run a few times a month (2), a wrong number in a proposal is expensive (1) and past pricing sits in old spreadsheets (2). Score 4. Start with status notes, and the first week of work is moving project files into one place, not writing prompts.
An operations team in a 150-person distributor
Candidates: sorting inbound customer emails into categories with a suggested reply, matching supplier invoices to purchase orders, and approving refunds. Email sorting runs hundreds of times a week (3); an agent suggests and a person sends (3); the categories and reply templates exist but are out of date (2). Score 18. Invoice matching also runs often (3), but a wrong match pays the wrong amount (1) and purchase-order data is clean (3): score 9, a sound second workflow once review is designed. Automatic refunds score 3 × 1 × 2 and wait, because the money leaves before a person looks.
After you choose
Write the baseline before anything is built, run the acceptance test on known answers, and put the review date on the calendar. If the readiness picture is unclear, the AI readiness checklist helps you find the weak area first.
When you do not need this
If no workflow in your business scores above 8, you do not need an AI workflow yet, and buying one will not change the arithmetic. The work is upstream: write down where the information lives, make it current, and decide who checks what. Come back to this page when a frequent job has clean context behind it.