Somewhere around the third "Top 10 AI Automation Consulting Firms" list, they all start to blur. Same stock photos. Same promises about efficiency. Same client logos you can't verify. You close the tab knowing exactly as much as when you opened it.
And if you're like most of the executives I talk to, this isn't your first attempt at getting outside help with AI. Maybe you paid for a strategy engagement that ended with a beautiful deck and a roadmap nobody funded. Or ran a pilot that never left the sandbox. Or you're still paying for licenses your team opens twice a month.
Some of the firms on those lists are good. That's not the problem. The problem is the list can't tell you which ones, because it ranks the wrong things. Rankings measure marketing budgets and backlink profiles. They tell you nothing about how a firm actually works once the contract is signed, and that's the thing that decides whether you get results.
The numbers back up the skepticism. A Gartner survey of 835+ midmarket CIOs, presented at the Midsize Enterprise Summit this spring, found that 59% say their AI investment has delivered little or unclear value, with returns typically landing around 2-3%. Most of those leaders had help. They bought tools, ran pilots, hired advisors. The spend was real. The results weren't. And the grace period is closing: midmarket technology leaders now expect to be judged on measurable outcomes by the end of the year, not on how many pilots they ran.
Look at what most of those disappointments have in common, though. They were structured as events. A big assessment, a big roadmap, a big build, a big bill. Then everyone goes home, the business keeps changing, and the shiny new thing starts aging the day it ships.
We watched software make this exact mistake for decades. There's a reason we stopped.
Operations is having its waterfall moment
Twenty years ago, serious software was built in one motion. Spend a year writing the spec, spend another year building it, release everything at once, and pray. By launch day the market had moved, the requirements were stale, and the team was locked into decisions made back when everyone knew the least.
The industry's answer was agile: stop betting everything on one big release. Ship the most valuable slice first. Measure. Re-rank the backlog. Ship the next slice. The teams that made that shift didn't win because they worked harder. They won because they were never more than a few weeks from correcting course.
Operational improvement is arriving at the same fork right now, and AI is the reason. The cost of fixing a single workflow has collapsed. An automation that would have been a six-figure line item in a systems integrator's proposal five years ago is now a few weeks of focused work. Which means the million-dollar, rip-everything-out transformation program is no longer the price of admission. It's a choice. And usually the wrong one, because it has the same failure mode as the waterfall release: by the time it ships, the business it was designed for doesn't exist anymore.
Here's the part the big-program pitch never mentions: bottlenecks don't stay fixed. Clear the worst one and the constraint moves somewhere else. Win a big customer and your fulfillment process becomes the new chokepoint. Lose a key hire and a workflow that ran fine for years suddenly doesn't. Operations is a moving target by nature.
So the question to ask about any AI automation consulting partner is not "how big is their transformation practice?" It's "when they leave, will we have a system that keeps improving, or a monument that starts decaying?"
That's the test I'd use if I were sitting on your side of the table. After ten years and 100+ projects, I've come to judge partners by what's still working 90 days after they leave. There should be three things.
Leave-behind #1: A living backlog of your bottlenecks
The most expensive mistake in AI automation happens before anyone builds anything. It's picking the wrong thing to improve.
Every operation has dozens of candidates. Order entry that takes three systems and a spreadsheet. Invoices that sit in an inbox for a week. Reports someone rebuilds by hand every Monday. A partner who starts executing before mapping those is guessing with your budget.
A real engagement starts with an automation opportunity assessment. It's unglamorous work: sit with the people doing the jobs, trace where hours actually go, and put a number on each bottleneck. Then rank them across dimensions. Workflow bottleneck prioritization sounds like consulting jargon, but the discipline behind it is simple. For each candidate workflow, a serious partner will answer three questions:
Value: what is this bottleneck costing you? Hours per week, error rates, delayed cash, deals that slip. If they can't attach a number, they can't defend the priority.
Feasibility: can automation reliably improve this today? Some workflows are ready right now. Others need a human in the loop for years. A partner who claims everything is automatable hasn't done this before.
Readiness: is your data and process clean enough to build on? Automating a broken process gets you a faster broken process. Sometimes the honest first recommendation is to fix the workflow before you automate it.
The output is a ranked backlog your leadership team can argue with. But notice the word: backlog, not report. A static list goes stale in a quarter, because the business it describes keeps moving. What you want is the ranking discipline itself, installed and repeatable, so that when a new bottleneck surfaces in month four it gets scored and slotted instead of ignored.
Here's the tell in the first meeting: watch what they ask about. A partner doing real operations consulting will ask where your team loses time. A vendor will show you a demo. If the first meeting is a product walkthrough, you're not buying consulting. You're buying a license with a services markup.
Leave-behind #2: A cadence, not a launch date
The old model of operational improvement was the launch: months of design, a big cutover weekend, a war room, and then years of living with the result. The model that actually works now looks like a good product team's rhythm. Take the top item on the backlog. Ship the smallest version that pays for itself, in weeks. Measure what changed. Re-rank the backlog with what you learned. Take the next item.
Each increment does three jobs at once. It returns value immediately instead of parking it inside an 18-month program. It generates real data about what your operation needs next, which no upfront assessment can fully predict. And it keeps your risk small: the most you can lose on any single decision is a few weeks, not a seven-figure budget.
That's the automation strategy worth buying. Not a grand design, a working loop.
And within the loop, look for triage. Not every item on the backlog deserves the same treatment. Some bottlenecks are quick wins your own people can clear in days with a little guidance. Others are complex, production-grade builds that touch core systems and need senior engineers who will stand behind the work and keep it running as everything around it changes. A good partner sorts each item by complexity, impact, and what your team can carry, then works both lanes at once. One that routes everything to its own bench is billing you for work your team could own. One that hands everything to your team is leaving the hardest, highest-value items on the table.
So when you're evaluating a partner, ask what the first 90 days look like. The answer tells you which model they run. A continuous-improvement firm will describe something concrete and small: the first workflow automation live inside weeks, a measurement against the baseline, a re-ranked backlog. A big-program firm will describe phases. Discovery that lasts a quarter. A platform to build before any workflow improves. Alignment workshops. If nothing improves until everything improves, walk.
There's a budget corollary worth saying plainly. You should never need a million dollars to find out whether a partner is any good. In a continuous model, the first increment is the audition. It's small, it's priced like something small, and it either moves a number or it doesn't. A firm that's confident in its work will happily be judged on that. A firm that needs the full program committed up front is telling you where its confidence actually sits.
Leave-behind #3: A team that can carry its own lane
If every future improvement requires calling the consultant back, you didn't buy a capability. You bought a subscription with extra steps. The partners worth hiring treat internal AI capability building as a deliverable, not a threat to their business model, because a continuous-improvement loop only works if your people can turn it.
What that looks like in practice: a handful of your own people, trained inside the engagement, on your real workflows. Not a lunch-and-learn. Not a prompt cheat sheet. Working sessions where your operations lead builds the improvement for their own bottleneck with an expert beside them, so the skill stays when the expert leaves.
We learned how much this matters the hard way, on our own team. When we first rolled out AI coding tools internally, we gave everyone access and assumed adoption would follow. It didn't. A few people took off; most used the tools like expensive autocomplete. The gap closed only when we built structured training. After we ran our largest team through it, we measured 43% more tickets and 93% more story points completed per sprint. Same people, same codebase, same backlog. The variable was training, not tooling.
The lesson transfers directly. Access to AI is cheap and everywhere. The capability to point it at your own bottlenecks, week after week, is scarce, and you can't hire your way to it in this market. It has to be built into the team you already have.
So ask the third question: "Which of these could our own people handle, and how do you get them there?"
A good partner has a specific answer. Names, roles, which backlog items they'll own by the end. To be clear, capability doesn't mean your team does everything; the complex builds should still go to people who build production software for a living. It means the wins within reach belong to your people, and the prioritization muscle lives in-house. Watch for the structural version of this too: the second improvement should cost less than the first, and the third less than the second, because your team is carrying more each cycle. If the pricing model assumes every project starts from zero, the incentives are pointed at dependency.
Follow the incentives before you sign
None of the failure patterns above happen because consultants are bad people. They happen because the traditional consulting model makes them rational. Time-and-materials billing rewards duration. A big-bang proposal rewards big budgets. A junior-heavy bench rewards selling senior faces and delivering cheaper ones.
So before you sign anything, check the structure:
Who does the work? Ask to meet the actual delivery team, not the pitch team. If the people in the sales meetings disappear after the signature, everything you evaluated walks out the door with them.
How is it priced? Small, fixed-price increments put the overrun risk on the firm, where it belongs, and give you a real decision point after every cycle. Open-ended hourly against a giant scope puts all the risk on you.
Can you stop? In a continuous model, you should be able to stop after any increment and keep everything it produced: the working improvements, the backlog, the trained people. If stopping early means forfeiting value, the model is built for lock-in.
Is a principal in the room? In a firm of any size, find out whether a founder or partner stays involved past kickoff. The prioritization calls ripple through your operation for years. You want someone with skin in the game making them with you.
None of these are gotchas. They're structural facts, and structure predicts behavior better than testimonials do.
The five questions, in one place
If you take nothing else from this, take these into your next first meeting:
- Priorities: "Which of our workflows would you improve first, and how did you decide?"
- Cadence: "What is live and measured inside the first 90 days?"
- The loop: "When a new bottleneck shows up in month four, what happens?"
- Capability: "Which of these could our own people handle, and how do you get them there?"
- Exit: "If we stop after the first increment, what do we keep?"
Any firm worth hiring can answer all five without flinching. Most can't. That's the point.
The lists will keep publishing, and the firms on them will keep looking interchangeable, because the things that separate them don't fit in a ranking. A living backlog. A cadence of small wins. People who get stronger every cycle.
Software stopped betting the company on the big release twenty years ago, and the teams that made the switch never went back. Your operation gets the same choice now. Pick the partner who's built for it.
If you're trying to figure out where improvement would pay off first in your own operation, our free AI assessment is a ten-minute place to start.

Mike is Co-Founder of The Gnar Company, a Boston-based software development agency where he leads project delivery for clients like Whoop, Kolide (acquired by 1Password), LevelUp (acquired by GrubHub), Qeepsake (feaured on Shark Tank), and AARP. With over a decade of experience building impactful software solutions for startups, SMBs, and enterprise clients, Mike brings an unconventional perspective having transitioned from professional lacrosse to software engineering, applying an athlete's mindset of obsessive preparation and relentless iteration to every project. As AI reshapes software development, Mike has become a leading practitioner of agentic development, leveraging the latest AI-assisted practices to deliver high-quality, production-ready code in a fraction of the time traditionally required.



