Evaluation 8 min read

What AI can see in your customer data (and what it cannot)

What AI can genuinely find in the customer records an SME already holds, what it cannot, and when a spreadsheet honestly beats a model.

Most small businesses hold more customer data than they think. Invoices, quote history, job notes, booking records, email threads, reviews. The analytics pitch says AI will find gold in it: hidden segments, customers about to leave, exactly who to call next. Some of that is real. Some of it only works at a scale most SMEs will never reach, and saying so out loud is rarer than it should be.

This post sets out what the published evidence says AI can actually find in the customer and sales data you already hold: who buys, who leaves, who is worth reaching, and what to say. It is equally specific about what it cannot find, and about the cases where a sorted spreadsheet honestly beats a model.

How much of their customer data do small businesses actually use?

Very little, and far less than large firms. OECD policy research from 2019 reports that across the EU, big data analysis was used by 33% of large firms, 19% of medium firms and 10% of small firms (OECD, Data Analytics in SMEs). A 2020 case study of 53 UK SMEs receiving university analytics support found roughly 5% using predictive analytics at all; the authors observed that most businesses collect and store data but lack the skills to analyse it (Mohamed and Weber, 2020). Small sample, one country, but it matches what I see in Perth engagements: the constraint is almost never access to data. It is that nobody reads it.

That reframes what “AI for customer insight” should mean for a small business. The first win is not a predictive model. It is getting the reading done at all.

Who buys from you: what does segmentation actually need?

For most SMEs, three columns you already have: how recently each customer bought, how often, and how much. That is RFM analysis (recency, frequency, monetary value), and it has been a standard segmentation method for decades. The peer-reviewed literature still treats it as the common baseline while documenting its blind spots: the recency score is noisy (a brand-new customer can score the same as a long-term loyal one), and the model sees only your side of the relationship, not the customer’s experience of it (Scientific Reports, 2024).

Here is the honest part: with a customer base in the hundreds, RFM is a pivot table. You do not need machine learning to learn that 15 customers produce half your revenue, or that a segment of once-regular buyers has gone quiet. Where AI adds something a spreadsheet cannot is the unstructured material sitting around those numbers: what customers actually asked for in their enquiry emails, the themes in your reviews, what the lost quotes had in common. That is reading work, and language models are genuinely good at it.

Who is about to leave: can AI predict churn for a small business?

At research scale, churn prediction works. At typical SME scale, the honest version is churn noticing, not churn prediction. A peer-reviewed 2024 study in Scientific Reports reached 89.6% accuracy predicting telecom customer churn with gradient-boosted models (Scientific Reports, 2024). Results like that are real, and they are also built on conditions most small businesses do not have: a subscription business where “churned” is a clean labelled outcome, and a public dataset with thousands of customer records to learn from.

The sample-size problem is quantified in the methods literature. A peer-reviewed simulation study found that modern techniques such as random forests and neural networks may need more than ten times as many events per variable as plain logistic regression to produce stable results, and were still unstable at 200 events per variable (van der Ploeg, Austin and Steyerberg, 2014). In churn terms, an “event” is a customer who actually left. Use ten customer attributes and that arithmetic wants a couple of thousand recorded departures before the fancier models settle down. A business with 400 customers and 30 known losses is not close, and no tool fixes that.

What is achievable at that scale is simpler and still valuable. Every repeat customer has a rhythm. A customer who ordered every six weeks for two years and has now been silent for five months is a signal you can compute with subtraction. Flagging everyone whose gap since last purchase has stretched well past their own usual gap is arithmetic, not machine learning, and it catches most of what a small business needs churn prediction for. AI earns its place afterwards: reading the last email thread with each flagged customer and summarising what happened before the silence.

Who is worth reaching: is it the customer the model flags?

Not necessarily, and this is where the research is most useful to a small operator. Eva Ascarza’s field-experiment study in the Journal of Marketing Research (peer-reviewed, and winner of the AMA’s 2018 Paul E. Green Award) found that the customers at highest risk of leaving are often not the best targets for retention campaigns; targeting the customers most likely to respond to the intervention reduced churn more than targeting by risk (Ascarza, Retention Futility). Some of your most at-risk customers have already decided, and money spent on them is spent on a goodbye.

The general lesson is about correlation and causation. A model trained on your sales history tells you who resembles the people who bought. It does not tell you what caused them to buy, and it cannot see anyone who never contacted you, because your data only contains the people who showed up. Treat “looks like a past buyer” as a shortlist worth testing, not a verdict.

What to say: where does AI genuinely beat the spreadsheet?

Text. Your won and lost quotes, enquiry emails, review wording, complaint threads and job notes are customer data too, and they are the part a spreadsheet cannot touch. A language model can read every lost quote from the past two years and tell you what the objections had in common, or read your five-star reviews and tell you which phrases your happiest customers use unprompted. Those phrases are usually better marketing copy than anything written from scratch, because they are evidence of what the market already values in you.

This is our own experience as well as the literature’s. DreamCopy, our property marketing tool, assembles campaign copy packs from the listing material an agent already holds, and the transferable lesson from building it was exactly this: the value came from reading and reusing what was already written down, not from predicting anything.

When does a spreadsheet beat a model?

Whenever the data is small and structured and the question is descriptive. Who are my top 20 customers, which segment is going quiet, what is my average gap between first enquiry and first invoice: sort, filter, pivot. A model earns its keep in two situations: when the volume is genuinely large, or when the input is text a human would need days to read. Between those poles, be suspicious of anyone selling prediction to a 300-customer business, and remember that no analysis recovers data you never collected. We have written before about hitting exactly this kind of wall in our own builds, in the invoice verifier post; the pattern of stating where the wall is applies just as much to customer analytics.

What this looks like in practice

A useful reading of an SME’s customer data usually produces four things: a ranked picture of who your revenue actually depends on, a short list of customers whose rhythm has broken, a tested shortlist of who resembles your best buyers, and the actual words your market uses when it is happy. None of that requires new software on day one. It requires someone to sit down with the data you already have and read it properly, with AI doing the reading a human would not get to. That is the substance of our customer targeting and customer insights work, and the wider evidence base behind these judgements is maintained in our State of AI review.

If you want to know what is actually visible in the data your business already holds, start with a conversation: get in touch.


Frequently asked questions

Can AI predict which customers will leave a small business?

Only with more data than most small businesses have. Peer-reviewed churn studies reporting high accuracy are trained on thousands of labelled customer records, usually in subscription businesses. Methods research shows machine learning models need many recorded departures before predictions stabilise. At small-business scale, the reliable version is churn noticing: flagging customers whose gap since their last purchase is well beyond their own usual rhythm, then using AI to read the history behind each flag.

How much data do you need before AI customer analysis is worth it?

For prediction, more than most SMEs hold: simulation research suggests modern models can need over ten times the data of simple statistics, measured in outcomes per attribute, to give stable answers. For reading, far less. If you have a few hundred enquiry emails, quotes or reviews, AI analysis of that text is already worthwhile, because the alternative is a human reading for days or nobody reading at all.

What customer data does a small business already hold that AI can use?

More than most owners expect: invoices and sales records, quote history including the quotes you lost, enquiry emails, booking and job records, reviews and complaint threads. The structured part answers who buys, how often and for how much, often with nothing fancier than a pivot table. The unstructured text is where AI adds real capability, surfacing objections, themes and the exact wording your happiest customers use.

When is a spreadsheet better than an AI model for customer analysis?

When the data is small and structured and the question is descriptive: top customers, quiet segments, average order gaps. Sorting and pivot tables answer those directly, cheaply and transparently. A model earns its place when the data volume is genuinely large or the input is free text. If a vendor proposes predictive modelling on a customer base of a few hundred, ask how many recorded outcomes the model would learn from.

Published 29 August 2026

Perth AI Consulting delivers AI opportunity analysis for small and medium businesses. Start with a conversation.

Prepared by Claude, directed and approved by PAC.

More from Thinking

Evaluation 11 min read

AI in property valuation: the evidence, the design rules, and what it could become

The best Australian evidence on vision AI in valuation measures a different task than the one vendors demo. The findings, and the design rules that follow.

Evaluation 7 min read

Eleven cells moved. Here is what they mean for your business.

Reading the September 2026 State of AI verdict table: what improved, what declined, and what to do differently this quarter.

Evaluation 7 min read

Competitor intelligence for small business: what AI can and cannot see

What AI-assisted competitor intelligence really is for a small business: the public sources worth watching, what they cannot tell you, and the legal line.

Evaluation 10 min read

AI in regulated professional work, Mid-2026

One structure links family law, valuation, and building inspections: a signed document others rely on. How each field's regulator answered the AI question.

Technical 9 min read

The business knowledge base: evidence, risks, and how to build one

What a business knowledge base actually is, what the evidence says it delivers, the security and privacy realities, and how we build one that holds up.

Building 7 min read

What an AI quoting engine actually does

What an AI quoting engine takes in, what it drafts, what the evidence says about accuracy and speed, and why the final price stays with a human.

Adoption 6 min read

Australia's AI adoption gap is bigger than the 12% headline suggests

ABS says 12% of Australian businesses use AI. The real story is 35% of large businesses against 11% of small ones, and the barrier isn't the technology.

Building 7 min read

Why we let AI run the interviews (and why we never let it pretend to be human)

AI-conducted interviews compress weeks of stakeholder discovery into days, standardise what gets asked, and lower the guard that distorts honest answers.

Adoption 14 min read

How AI capability actually moves through a business

The decisive variable in SME AI adoption is the human absorption sequence, not the tooling. A working framework from observation across WA businesses.

Evaluation 7 min read

AHPRA advertising rules for psychologist websites

Recovery stories, 'specialist', 'clinical psychologist', and endorsement titles are where psychology sites breach the National Law. A practical read-through.

Adoption 4 min read

Customer service AI has finally grown up

Chatbots and AI receptionists earned their bad reputation. What changed, and how the mature version answers every call without replacing anyone.

Evaluation 6 min read

Who can use the titles 'Dr', 'Specialist', and 'Surgeon'?

AHPRA restricts 'specialist' and 'surgeon' to specific registrations, and 'Dr' has its own rule. What health practice websites can and cannot claim.

Adoption 5 min read

Your best people hate writing reports

The operators you promote are brilliant at the work and allergic to reporting. A scheduled AI call interviews them, drafts the briefing, they approve it.

Building 6 min read

Your website isn't just for humans anymore

How to build a chatbot that keeps itself up to date, can't leak client information, and won't answer beyond what you've published.

Evaluation 7 min read

Can you show Google reviews on your health practice website?

AHPRA bans clinical testimonials, even true ones, but service reviews are fine. What that means for the Google reviews widget on your practice site.

Evaluation 7 min read

What AHPRA's advertising rules mean for your website

Your practice website is advertising under the National Law. What AHPRA's rules prohibit, who is responsible, and how to check your own site.

Evaluation 8 min read

Is it safe to paste client data into ChatGPT?

Short answer: it depends on one setting, and most people have it wrong. What ChatGPT, Claude and Copilot do with your data, and what the Privacy Act expects.

Evaluation 4 min read

What a good AI audit actually delivers

The audit report named one recommendation specific enough to check, and what the Build that followed looked like: one real engagement, generalised.

Evaluation 7 min read

AI and video, Mid-2026: the models can watch now, not just listen

AI could always transcribe video. It can now read the frames as well, and every hour of footage a business owns becomes something it can question.

Building 7 min read

Case study: a 119-page AML/CTF program in three days

How we built a seven-document AML/CTF compliance pack for a small accounting practice in three days, working from 31 confirmed assumptions.

Building 11 min read

From evidence base to delivery: a production AI methodology

How we delivered 34 evidence-anchored AI briefings to a WA peer-advisory chapter: fact-checked literature review, multi-agent verification, one method.

Technical 9 min read

The six functions of a working AI system

A working AI system is six functions doing six jobs. When all six connect, hallucinations get caught, outputs hold steady, and models become swappable.

Technical 7 min read

Supervised autonomy: the middle path for AI architecture

Between drafts you approve and agents you hope about sits the middle path: an envelope of authorised routine work, supervised, audited, and yours to widen.

Evaluation 5 min read

The state of applied AI in Mid-2026

Our literature review of applied AI in mid-2026: ten capability categories, three fact-check passes, written for operational leaders.

Technical 9 min read

How to design a PHI redaction system for clinical AI

PHI redaction is part of a clinical AI tool's architecture, not a feature you add. What the literature says it should look like, and how we built it.

Building 9 min read

How we built on-device de-identification so AI never sees real names

Most AI privacy is a policy. Ours is architecture: an NER model runs in the browser and strips names before anything leaves the device.

Technical 7 min read

Your agency's clients are about to ask why this costs so much

A solo consultant built in three weeks what your agency quoted twelve for. The client doesn't know why yet. The agencies that survive change what they sell.

Adoption 6 min read

What do you love doing? What do you hate doing?

Ask people what they love doing and what they hate doing, then show them AI is coming for the second list. Why the reframe works, and how it fails.

Technical 7 min read

Why I don't use n8n (and what I do instead)

n8n demos well. But a compelling demo and a reliable production system are different things, and the distance between them is where businesses get hurt.

Technical 10 min read

Your codebase was not built for AI. That's the actual problem.

Amazon's mandatory meeting about AI breaking production is an architecture story: codebases built for human maintainers only, now maintained by AI.

Adoption 4 min read

Your team has AI licences. You don't have an AI system.

Fifteen people, fifteen separate AI accounts, no shared context. The problem isn't the tool; it's the architecture around it. Here's the fix.

Building 7 min read

Your $2,000 day starts the night before: our system keeps you on the tools, not on the phone

Optimised routes overnight, automatic customer notifications, and promises the system keeps or corrects. A scheduling system that protects your daily rate.

Evaluation 4 min read

The fastest way for an executive to get across AI

AI moves faster than any executive can track. One focused conversation, one written report, and a decision you can act on: your time stays on the business.

Building 6 min read

Your IT department will take 18 months. You need this working by next quarter.

Senior leaders know what they need built; the gap is time. A prototype gets the tool working now and hands IT a validated blueprint for later.

Building 8 min read

We built an AI invoice verifier. Here's where it hits a wall.

We built an AI invoice verifier and watched a fake beat a real invoice. Why document analysis alone cannot stop fraud, and the five layers that can.

Building 5 min read

How to build an AI chatbot that doesn't lie to your customers

Woolworths scripted its AI to talk about its mother. The business fix is honesty; the technical fix is architecture that prevents fabrication by design.

Technical 9 min read

Why AI safety features are load-bearing architecture, not political decoration

The 'woke AI' label came from real failures, but they were engineering failures, not safety failures. The difference matters wherever errors have consequences.

Adoption 3 min read

Woolworths' AI told a customer it had a mother. That's a problem.

Woolworths' AI assistant Olive was scripted to talk about its mother and uncle. When callers realised, trust broke instantly. The fix is honesty.

Evaluation 5 min read

Google is no longer the only way your customers find you

Customers now find businesses through ChatGPT, Perplexity, and Gemini. The sites AI cites are structured differently to the sites Google ranks.

Evaluation 6 min read

The personal workflow analysis: what watching a real workday reveals about automation

People describe the work they value, not the work that eats their time. Recording a real workday reveals the automation opportunities interviews miss.

Evaluation 11 min read

An AI audit that starts with your business

How an operations-first AI audit works: what it looks for, how the evidence is collected, what the report contains, and what it tells you to skip.

Building 6 min read

What production AI teaches you that demos never will

The gap between a demo and a working system is where the useful lessons live. Architecture, framing, privacy, adoption: the patterns repeat every time.

Adoption 6 min read

The psychology of why your team won't use AI

You buy the tool, run the demo, and three months later nobody is using it. Five predictable psychological barriers, each with a strategy that works.

Technical 4 min read

Stop telling AI what NOT to do (and what to say instead)

Instructions built on prohibitions make AI cautious and generic. Describing what you want instead transforms the output, and the reason comes from psychology.

Building 5 min read

How we turned generic AI into a specialist: and what that means for your business

Mediocre AI output is rarely the model's fault. Three structural changes that turn the same model from generic to specialist-grade.

Evaluation 6 min read

Your business has 9 customer touchpoints. AI can fix the 6 you're dropping.

You pay to get customers to your door, then lose them to missed follow-up. AI can handle the six touchpoints most businesses drop.

Technical 6 min read

What happens to your data when you press 'Send' on an AI tool

Businesses send customer data to AI tools without knowing what happens during processing. The spectrum of AI privacy is wider than you think.