The data on this is now hard to argue with. RAND Corporation has found that more than 80% of AI projects fail, roughly twice the failure rate of non-AI IT projects. BCG's 2024 research found 74% of companies have yet to see tangible value from AI. MIT's 2025 State of AI in Business report found that 95% of generative AI pilots produce no measurable P&L impact. This isn't a technology problem. It's a sequencing problem.
Where pilots actually die
In our experience, AI pilots stall for three repeatable reasons: the tool was chosen before the problem was diagnosed, so it solves something adjacent to the real constraint. There was no baseline measured before the build, so nobody can prove the pilot improved anything, even when it did. And there was no governance model, so leadership doesn't trust letting the system touch a real customer, which quietly caps the pilot at "interesting demo" forever.
Never automate chaos
Our rule, and the reason we structure engagements the way we do: map how work actually moves through the business first, rank every opportunity in dollars, and only automate what's proven to leak. Before any full build, we build a working prototype on the client's own data, so results are measured against a frozen baseline before anyone scales anything. Most AI projects fail before the first prompt is ever written, because the diagnostic step got skipped.
The Three Levels of AI Authority
The governance question is where most template agencies go quiet, and it's usually the real fear behind a stalled pilot: what happens when the AI is wrong, and who's watching? We use three explicit levels. Level 1: AI acts directly, for narrow, high-confidence, low-stakes situations, general questions, capturing basic details, with no human in the loop needed. Level 2: AI acts, but every action is logged for the owner to review, reminders, follow-ups, tagging, so nothing happens invisibly even though nothing waits on approval either. Level 3: AI drafts a recommendation and a human approves before anything happens, reserved for cancellations, complaints, refunds, anything where a wrong action is expensive to undo. Deciding, explicitly, what belongs at each level before you build anything is what turns "a robot talking to my customers" from a fear into a designed, reviewable system.
Facts, not memory
Day one, we write down the numbers that matter for this engagement, response times, conversion rates, no-shows, whatever's relevant. Day 90, we measure the same numbers again. If they didn't move, that's the answer too, and it's one we'll say out loud. A pilot that never reaches production usually failed at the map, not at the model.