- Start by telling the AI which decision it is helping you make
- Worked example: the campaign with more leads is not the obvious winner
- Find out what the table does not tell you before moving the budget
- Separate facts, calculations, assumptions and forecasts
- Worked example: does the proposed automation save enough time?
- Test the assumption that could overturn the recommendation
- Worked example: five customer complaints are not a market survey
- What to ask when AI keeps agreeing with you
- Turn a checked recommendation into a job someone can carry out
- Common questions about checking AI business advice
- How much checking is enough?
In this blog post, I'll explain how to check AI business advice before you spend money, change a process or act on a recommendation. It's for small business owners who use AI to think through decisions and want more than a persuasive answer. We'll work through campaign spending, an inbox automation proposal and a customer research problem, so you can check the figures, expose the missing information and decide what deserves a test.
I use AI for a large amount of my work. I also want to know what it opened before it told me something was true. Those two positions fit together perfectly well, although the second does make the first rather less relaxing. My experiment asking one AI to audit another is the longer account of why I now insist on evidence outside the conversation.
Business advice needs an extra check. Even a factually accurate answer can recommend the wrong action for your business. A channel can produce more enquiries while leaving less money. An automation can save time on the easy cases and create a sizeable repair job on the difficult ones. A customer quote can be real without representing the customers you need to understand.
The useful question is what supports the recommendation, how much depends on assumptions and what happens if those assumptions are wrong. You don't have to investigate every sentence. You do have to investigate the sentences that would make you do something different.
Start by telling the AI which decision it is helping you make
"What should I do about marketing?" is an invitation to receive the entire marketing department in bullet points. By the time you've read it, you have seventeen new priorities and the original problem is still there.
A decision gives the answer somewhere to go. For example: "I have one additional campaign slot next month. Should I repeat campaign A, repeat campaign B or leave the slot unused? My aim is to increase contribution after campaign costs, and I can handle no more than six additional clients."
That last sentence changes the job. An answer that maximises leads could be a nuisance when you are already short of delivery capacity. An answer that maximises revenue could also be wrong if the work costs more to fulfil than it brings in. Before opening the chat, write the decision, the available options, the outcome that matters and the limit you cannot exceed.
My guide to writing an AI brief for your business covers the fuller briefing process. For a single decision, keep the opening brief compact enough that you can notice an omission. Include the dates your records cover and the definitions behind the numbers. "Leads" might mean everyone who completed a form, people who meet your qualification criteria or people who asked for a quote. Those are different populations with an inconveniently similar label.
If you need help organising the request, the prompt engineering cheat sheet is a supporting reference. The wording will not rescue a brief that asks the model to increase sales without mentioning that you cannot fulfil another order until November.
Worked example: the campaign with more leads is not the obvious winner
The following figures are invented for a small service business. They are teaching examples, not my results or a client's. Both campaigns have completed the same follow-up period. A customer counts only after buying, and each is counted once.
| Campaign measure | Campaign A | Campaign B |
|---|---|---|
| Campaign cost | $600 | $900 |
| Enquiries | 30 | 60 |
| New customers | 6 | 6 |
| Revenue per new customer | $500 | $500 |
| Variable delivery cost per customer | $200 | $200 |
An AI looking mainly at lead generation might recommend B. It produced twice as many enquiries, and its cost per enquiry was $15 rather than A's $20. The cost-per-lead calculation is correct. The recommendation still needs work.
Each customer leaves $300 after variable delivery costs. Six customers therefore leave $1,800 before campaign costs. Deduct those costs and A leaves $1,200 while B leaves $900. This is contribution under the stated assumptions, before fixed overheads and tax. It is not the business's final profit.
There is another cost hiding in the first table. Someone had to handle B's extra thirty enquiries. If each took ten minutes, B required five additional hours. We have not priced those hours into the calculation because the example has not established whether they create an extra cash cost, displace other useful work or fit into spare capacity. Write down the hours anyway. Unpaid owner time still appears in your week, usually when you had planned to eat.
The marketing metrics cheat sheet is useful when the problem is choosing the right measure. Here, the decision needs customers, contribution and handling time alongside enquiries. A lower cost per lead answers only one part of that question.
Ask the AI to show the calculation in a form you can reproduce. For this example, it should be able to write:
Campaign contribution = customers x (revenue per customer - variable delivery cost per customer) - campaign cost.
For A, that is 6 x ($500 - $200) - $600 = $1,200. For B, it is 6 x ($500 - $200) - $900 = $900. Recalculate independently in a calculator or spreadsheet. A paragraph saying the figures have been checked is not the check.

Find out what the table does not tell you before moving the budget
It would now be tempting to declare A the winner and double its budget. We still haven't earned that conclusion.
The table records what happened at two particular spending levels. It does not show that another $600 on A would buy another six customers. Perhaps A reached a small warm audience that has now been exhausted. Perhaps B brought in customers who renew more often. Perhaps one campaign ran during a period when the owner answered enquiries promptly and the other did not. Those possibilities need evidence before they belong in the recommendation.
First check the measurement. Were both campaigns allowed the same time for enquiries to become customers? Were refunds removed consistently? Did someone count the same buyer in two channels? A tidy report can conceal different definitions in adjacent columns. If the campaign labels themselves are inconsistent, use the UTM and campaign naming template to organise future tracking. It won't reconstruct data you never collected, and it doesn't make an attribution report a controlled experiment.
Then check what changed. In my newsletter design experiment, several things changed together, which limited what I could attribute to any single element. That is a useful distinction to keep in your own reporting: you can describe a result without knowing precisely which change caused it.
For the invented service business, a defensible provisional answer would be: "A produced more contribution at the observed spend and needed fewer enquiries to produce the same number of customers. Before allocating the next slot, check whether A has a comparable audience available and whether delivery and follow-up conditions will remain similar. Do not assume the result scales in proportion to spending."
That is less exciting than "Double down on your winning channel." It is also an answer you could use without first pretending the missing information has arrived.
Separate facts, calculations, assumptions and forecasts
AI can move between these categories so smoothly that you barely notice. Your sales export is a record. The contribution calculated from that export is a calculation. The belief that a new audience will behave like the old one is an assumption. Next month's customer total is a forecast. They should not arrive dressed as four equally certain facts.
Give the model a table like this and require it to fill the evidence column. Check its labels afterwards; asking for labels does not guarantee sensible labelling.
| Part of the recommendation | Status | Evidence or check needed |
|---|---|---|
| A generated six customers | Record supplied | Open the customer export and confirm dates and campaign IDs |
| A left $1,200 after stated costs | Calculation | Recalculate and confirm which costs are included |
| A has another similar audience available | Assumption | Inspect audience availability and previous exposure |
| Another campaign will produce six customers | Forecast | Treat as uncertain; state a range and what supports it |
A source link only helps if the source supports the claim beside it. Open the original page or record, find the relevant passage and check the date, population and conditions. A study of a national retailer might suggest something worth testing. It cannot establish the likely conversion rate of your consultancy's next email.
For software recommendations, check the specific plan and feature in the current product documentation. A help page describing something the product can do does not prove your account includes it. Before entering client or customer information, use the account and data rules appropriate to your business. Often a small anonymised sample is enough to work through the decision.
The same discipline applies when AI summarises customer material. The voice-of-customer mining cheat sheet can help you organise what customers said, but the original message remains the evidence. Keep a reference to it so you can reopen the passage rather than relying on the model's increasingly elegant interpretation.
Worked example: does the proposed automation save enough time?
Now suppose an owner asks whether an AI inbox assistant is worth adopting. This is a separate hypothetical example. The owner handles 200 routine enquiries a month and currently spends an average of six minutes drafting each response. That is twenty hours of drafting.
The proposed system produces drafts for review. Assume, for the calculation, that checking and finishing each draft takes two minutes, and that twenty drafts a month require four additional minutes of repair. Those are pilot assumptions, not claims about any product.
The new workload would be 400 minutes of review plus eighty minutes of repair: eight hours. Against the original twenty hours, the potential saving is twelve hours a month. If setup takes eight hours and the recurring costs total $60 a month, the first month looks different from later months.
At an illustrative value of $30 per recovered hour, twelve hours represent $360 of capacity. Subtract the $60 recurring cost and the ongoing net value estimate is $300 a month. Include the eight setup hours at the same assumed value and the first month's estimate falls to $60. None of this means $300 will arrive in the bank. That requires the released time to reduce a cash expense or produce work somebody pays for.
My AI Savings Calculator helps put an initial value on time. Keep your own measurement sheet beside the estimate so you can replace assumptions with observed review and repair times. Keep capacity, cash savings and additional revenue separate. Adding all three when they describe the same benefit is how a modest improvement becomes a suspiciously magnificent spreadsheet.
The next check is a pilot using representative messages. Include the awkward cases, not just the five enquiries the owner could answer half asleep. Record drafting, review and repair time separately, and record any error that would matter to the customer. Use the same task boundary before and after: if the old timing included reading the enquiry, the new timing must include it too.
Want AI doing the heavy lifting in your marketing?
I build the systems that handle the boring 80 percent, so you get your week back. Done properly, with the human kept in.
The AI tool cost guide is relevant here because the subscription is only one part of the work. My AI inbox management experiment also shows the distinction between organising and preparing a response and allowing a system to act. For this pilot, drafts stay drafts until the owner has checked them.
Test the assumption that could overturn the recommendation
You could spend days researching every uncertainty in either example. Start with the one that changes the decision most.
For the inbox assistant, imagine average review time is four minutes rather than two. The workload becomes 800 minutes of review plus eighty minutes of repair, or fourteen hours and forty minutes. The monthly saving falls to five hours and twenty minutes. At the same $30 valuation, that is $160 of capacity before the $60 recurring cost. The first-month estimate becomes negative once setup time is included.
That calculation tells you exactly what to measure in the pilot. Review time is doing much of the work in the business case. You do not need another article promising that AI will transform your productivity; you need a timer and a representative sample of your own enquiries.
If the pilot reveals repeated factual mistakes, treat those as failures even when the average time improves. The AI Agent Guardrails Checklist supports the permission side of this decision. The AI Agent Failure-Mode Playbook helps you think about what happens when a tool encounters missing inputs or other failures. Neither replaces inspecting the errors your own pilot produces.
For campaign spending, the uncertain assumption might be incremental demand: how much business happened because of the campaign rather than being credited to it. Where volume and logistics allow a fair comparison, the incrementality test design template is a useful planning resource. Decide the audience, comparison, measurement window and stopping conditions before seeing the result. Small counts can leave the answer inconclusive, and extending a test until the result looks pleasing does not resolve that problem.
If you are changing the creative itself, separate that question from the channel decision. The ad creative testing cheat sheet belongs at that stage. Testing a new audience, message, offer and landing page together may tell you whether the package worked; it won't isolate which ingredient made the difference.
Worked example: five customer complaints are not a market survey
Consider a hypothetical business with five recent messages mentioning price. The owner asks AI why sales have slowed, pastes those messages and receives a recommendation to lower prices. The messages are real. The leap from those messages to the whole sales problem is the issue.
Before acting, identify who wrote them and who is missing. Were they five qualified prospects out of twenty, or five unqualified enquiries out of two hundred? Were they discussing the same offer? Did buyers who went ahead describe a different concern? Did the enquiries fall before the price conversations began? A collection selected because it contains price complaints will, with remarkable consistency, contain price complaints.
Use the ideal client profile template to keep the relevant customer group explicit. Then ask the AI to extract each objection with its source reference and distinguish direct statements from your interpretation. "Too expensive for me this month" does not establish that the service is overpriced for its intended buyer.
The next step might be speaking with recent buyers, non-buyers and people who stopped responding, while keeping those groups separate in your notes. The customer interview prompt pack can help prepare the conversation. Ask what they were trying to do, what they considered and what stopped the purchase. "Would a discount make you buy?" invites a hypothetical answer that is very easy to give and costs the respondent nothing.
The AI can organise the responses and identify questions worth investigating. It cannot supply the absent customers or turn stated interest into a purchase record. Your decision may be to clarify the offer, improve qualification, investigate another cause or leave the price alone. A useful review keeps those possibilities open until the evidence narrows them.
What to ask when AI keeps agreeing with you
Excessive agreement is often called sycophancy. The practical problem is that you can mistake support for your preferred answer for evidence that the answer is sound. Asking the model to be brutally honest may change its tone; it does not give it access to the records you omitted.
Try removing your preference from the brief. Replace "Help me justify this automation" with the decision, the alternatives and the measured constraints. Ask what would make the recommendation fail, which evidence supports that concern and what finding would change the answer. This also prevents a theatrical list of objections from passing as a serious review.
For a consequential decision, use one fresh review with the original records and definitions. Compare the reasons, not the confidence of the prose. If the answers differ, identify the assumption or calculation responsible. A third vote won't settle a missing fact. Stop when the next useful step is opening a record, measuring a task or speaking with someone who has the relevant expertise.
Copy this prompt
Help me evaluate a business decision. Do not write an implementation plan yet.
Decision and alternatives, including no change: [insert].
Outcome I want, measurement period and capacity limit: [insert].
Records supplied, their dates and the definitions of the measures: [insert].
Costs, available time and anything I cannot risk: [insert].First identify missing information that could materially change the decision. Separate supplied records, reproducible calculations, assumptions and forecasts. For each important claim, point to the supplied record or mark it unverified. Do not invent customer behaviour, sources or results.
Compare the options against the stated outcome. Show calculations with inputs and units. Check for omitted setup, review, repair and delivery work. Keep cash savings, released capacity and projected revenue separate.
Give a provisional recommendation and the strongest supported reason against it. Identify the one assumption most likely to overturn the answer. Show what happens if that assumption is less favourable.
Propose a limited check with an owner, measurement method, effort limit and stopping condition. Explain what an inconclusive result would look like. End with what I can decide now and what remains unresolved.
Fill the gaps before using it. A prompt cannot force a model to follow every instruction, and it cannot authenticate evidence for you. Read the calculations and open the references it returns. If it reports confidence as a percentage, ask how that figure was obtained; an unsupported number is not a measurement of reliability.
Turn a checked recommendation into a job someone can carry out
The final decision needs an owner and a boundary. For the inbox example, a usable record might read: "Test drafts on a representative sample of anonymised historical enquiries. Record review time and material errors. No messages will be sent. Decide whether to continue only after checking the sample against the original replies and the full time calculation."
If the recommendation becomes a repeatable workflow, the AI Agent Brief Template helps document inputs, outputs and limits. My guide to building an AI workflow for a small business takes the implementation question further. Keep that work downstream of the decision; it is possible to build an impressively organised process for something you should never have started.
Also decide what will count as completion. My WordPress publishing automation experiment concerns a different job, but it illustrates why the finished result needs inspection. A message saying the task succeeded is weaker evidence than opening the thing it was supposed to produce.
Keep a short decision record: the date, options, original evidence, key assumption, chosen action, spending or effort limit and review date. At review, compare the outcome with the reason you acted. The weekly founder review prompts can support that habit, but keep the original record unchanged so hindsight cannot improve your prediction for you.
Common questions about checking AI business advice
Can I use the same AI to check its answer?
Yes, for a first review, but give it something independent to check against: the original records, a calculation you can reproduce or the relevant source. Asking "Are you sure?" supplies no new evidence. For an important decision, a fresh review can help identify a different interpretation, but agreement between models does not settle the facts.
What if I do not have enough data for a test?
Check what you can establish without pretending a small sample gives a precise forecast. You may be able to confirm the costs, inspect the workflow or reject an option whose best plausible outcome is too weak. If the remaining uncertainty is decisive and cannot be reduced affordably, keep the commitment small or leave the decision unchanged. "We don't know yet" is a legitimate result with a practical consequence.
Should I follow advice that contradicts my experience?
Treat the contradiction as a question to investigate. Your experience may contain context the model lacks, or it may reflect an old situation that has changed. Write the competing explanation and identify the record or observation that could distinguish them. Neither the model's confidence nor your familiarity should be the only evidence.
How much checking is enough?
Match the effort to the consequence. Brainstorming headlines for a private draft needs a lighter check than committing to a recurring expense or making a promise to a customer. For regulated or specialist questions, use the relevant qualified adviser; AI can help organise the question and documents without becoming the person responsible for the answer.
For an ordinary reversible business decision, you should be able to name the evidence, reproduce the important calculation, explain the main uncertainty and state the limit on the next step. If you cannot do one of those, narrow the action until you can. That may mean investigating rather than buying, drafting rather than sending, or testing a smaller part of the process.
Open one recommendation you already have and apply those checks. You may discover that the advice is useful. You may discover that the right next step is a ten-minute measurement instead of a three-week project. Either is a better outcome than giving a confident paragraph responsibility for your calendar and bank account.