We just finished a five part video series on putting AI to work in a small business. This is the practical companion to it: what we have actually seen survive contact with a real operation, and what has not. The pattern is less about technology than we expected.
Still running
Enquiry triage
Reading inbound forms, pulling out the structured details, scoring and routing. This is the most reliable thing we have deployed. It is frequent, the correct answer is not a matter of opinion, and a mistake is obvious within seconds to the person who receives it.
One client went from a person reading every submission to that person reading a queue that is already sorted. They stopped opening raw forms in week two and have not gone back.
Notes into records
Job notes or call notes turned into a consistent summary on the account. It survives because the person who wrote the notes reads the summary immediately and would notice if it were wrong. The verification is free, because it is a by product of work already happening.
Reconciliation flagging
Comparing two systems that should agree and surfacing what does not. Note that it flags rather than fixes. Everything we have deployed that flags is still running. Some of what we deployed that fixed is not.
Quietly died
The monthly report generator
Beautiful, genuinely useful, ran once a month, and nobody could remember whether last month's had been checked. Once you are not sure, you check it fully, and once you check it fully it has saved you nothing.
Low frequency was the killer. There was never enough repetition to build trust or a habit.
The unsupervised customer email
We advised against it and built a drafting version instead. The client later switched it to send automatically. It worked for a while. Then it sent something confidently wrong to a good account, and the whole category was off the table for a year.
The 70% does not go away because things have gone well for six weeks.
Anything built on unreconciled data
Two of these. Both produced plausible output that quietly disagreed with the finance numbers. Neither failed loudly. They just accumulated doubt until nobody cited them, which is how most of these actually end.
What separated them
Not the model. Not the platform. Three things, consistently.
- Frequency. Daily survived. Monthly did not. Repetition is what builds the trust and the habit, and without both the thing gets skipped and then forgotten.
- Free verification. Everything still running has a human who was going to look at the output anyway. Where checking is a separate chore, checking stops, and shortly after that so does the trust.
- A named owner. Every survivor has a specific person who would complain if it broke. Every casualty was owned by a department.
The workflows that lasted were not the clever ones. They were the ones where being wrong was cheap and someone was already looking.
The thing we changed about how we work
We now ask for the report back mechanism before we quote. Not who is paying, not what the integration looks like. Who will tell us in six weeks whether this is working, and when.
If there is no answer to that, we have learned that the build will technically succeed and practically fail. It is a slightly awkward question to open with. It has been the most reliable predictor we have found.
If you are somewhere in this and want a second opinion on which of your workflows is the right first one, get in touch. That conversation is free and usually short.