You can tell whether a software partner actually knows how to use AI by asking process questions instead of tool questions. A team with real AI maturity can show you where AI sits in its delivery workflow, who reviews what it produces, and how it keeps your code and data out of someone else's training set. A team that merely claims AI can name products. The difference surfaces within ten minutes if you ask the right things, and this guide gives you the exact questions.
"AI maturity" has joined price, portfolio, and communication as a standard line on 2026 vendor scorecards, and for good reason. Used well, AI genuinely changes how fast certain work gets done. Used carelessly, it ships unreviewed code into your product and your source into a consumer chatbot. When you evaluate an agency's AI maturity, you're really measuring two things at once: whether they capture the benefit, and whether they manage the risk.
Why this became a scorecard line
Two years ago, asking an agency about AI got you a shrug or a sales pitch. Now the buyer side evaluation guide lists AI maturity next to references and pricing, because the spread between agencies has become enormous and mostly invisible from the outside. Two firms can quote the same project, show similar portfolios, and run completely different engine rooms: one has senior engineers directing AI through a reviewed pipeline, the other has juniors shipping whatever the chatbot suggested. You'll pay a similar price either way. You will get very different software.
There's a second reason buyers started asking. The failure modes are new. A traditional bad agency shipped late and buggy, which you noticed. An AI careless agency can ship on time, with code nobody on their team fully read, trained on your data, and licensing questions nobody can answer. Those problems surface months later, usually during a security review or a fundraise, long after the invoice cleared.
What using AI well actually looks like
Inside a disciplined engineering team, AI shows up in specific, named places. Engineers use it to scaffold boilerplate, draft test suites, generate first pass documentation, and flag suspicious patterns during code review. Each of those is a bounded task with a human on the other end. The output enters the same pipeline as human written code: a senior engineer reads it, tests exercise it, and someone with their name on the project approves the merge.
Mature teams also have rules about what stays out of the tools. Client source code never goes into consumer tier accounts. Contracts with AI vendors include zero retention terms, so prompts and code aren't kept or used for training. Somebody at the agency can tell you, without checking, which AI subprocessors touch client material and under what terms. Boring answers like these are the strongest signal you'll get. Competence in this area sounds like policy and paperwork. It never sounds like magic.
Bolted on versus built in
The superficial version looks different up close. A few developers paste snippets into a chatbot when they get stuck, the proposal template gained an "AI powered delivery" bullet, and nobody can describe a review standard because there isn't one. Ask what the firm refuses to use AI for and you'll get a blank look, because a refusal list only exists where someone has actually thought about failure modes.
That's the core distinction. An agency that bolted AI on adopted a tool. An agency with real maturity redesigned part of its process around the tool, then wrote down where the tool stops. If you've read our guide on choosing a software development partner, this is the same principle applied to a new criterion: you're buying a process you can inspect, and a slide is never a process.
Seven questions to ask before you sign
Ask these in a live conversation, and pay attention to how fast the specifics arrive.
- Where does AI sit in your delivery process today? You want named stages: scaffolding, tests, documentation, review assistance. Vague answers ("our developers use it throughout") mean nobody has mapped it.
- What do you refuse to use AI for? A mature shop has a refusal list: architecture decisions, security-critical code without extra review, anything touching regulated data. No refusal list means no thinking about limits.
- Who reviews AI generated code before it ships? The answer should be a senior human applying the same standard as the hand written code. "The AI checks itself" or "we spot-check" ends the conversation.
- How do you keep our code and data out of model training? Listen for business tier accounts, zero retention agreements, and awareness of which subprocessors sit behind each tool. Then ask for it in writing.
- How does AI change your estimate? An honest answer says some tasks got faster while judgment work costs the same. A promise of dramatic discounts because of AI is a red flag dressed as a benefit.
- Tell me about a task where you turned AI off. Real experience includes failures. Teams that have actually used these tools at production standard can describe where the output wasted more time than it saved.
- Will any of this appear in our contract? Data handling, IP ownership of AI assisted code, and review responsibility should all survive contact with legal. Verbal assurances about AI policy are worth what they cost.
Red flags that should end the evaluation
Some answers tell you everything you need to know.
- The AI story is a list of tool names with no process attached.
- Productivity multiples pitched as guarantees. "We're 10x faster now" is a marketing claim, and nobody seriously prices software that way.
- No human review gate between AI output and your production code.
- Silence, or improvisation, on the training data question.
- The proposal itself reads like unedited AI output: generic, confident, and wrong about your business.
- AI framed as a replacement for senior engineers rather than a tool senior engineers direct.
One caution in the other direction: don't reward theater. An agency that shows you an impressive internal "AI platform" but can't connect it to review standards or data terms has built a demo for buyers, and you're the audience.
What AI doesn't change
AI has made it cheap to produce plausible looking code, and that has made judgment more valuable, since someone still has to decide what to build, how to structure the data, where the security boundaries sit, and what happens when things fail. We see the gap every week in codebases that arrived through AI app builders. The interfaces look finished while authorization, data modeling, and error handling are missing underneath, which is why we wrote a separate guide on what to do with an AI built MVP and why turning an AI prototype into a real product is one of the most common questions we're asked.
The same logic applies to the partner you hire. AI maturity is one criterion stacked on top of the fundamentals: an accountable team, a transparent process, and ownership of outcomes. A shop with brilliant AI workflows and weak engineering judgment will simply ship the wrong thing faster.
Get the answers into the contract
Good verbal answers are the filter. The contract is the proof. Three clauses separate an agency that means it from one that improvised well on the call.
- Data handling. Your code, credentials, and business data don't enter tools that retain or train on them. The agency names its AI subprocessors and the terms it holds with each, and updates you when the list changes.
- IP ownership. You own the delivered work regardless of how it was produced, and the agency stands behind its right to assign it. AI assistance doesn't dilute your ownership, and the contract says so explicitly.
- Review responsibility. The agency warrants that delivered code was reviewed by its engineers to its normal standard. This one sentence removes the "the AI wrote it" excuse from every future conversation.
An agency that already has language like this on file has thought the problem through. An agency that resists it just told you the marketing slide was doing all the work.
How much should this criterion weigh?
Less than the fundamentals, more than zero. AI maturity is a tiebreaker and a risk screen. It shouldn't outrank track record, communication quality, or whether the team has shipped products like yours. A shop that's mediocre at engineering and excellent at AI tooling is still mediocre at engineering. Where the criterion earns its place is at the margin: between two otherwise comparable partners, the one with a real AI process will likely deliver somewhat faster, and will definitely expose you to fewer of the new failure modes above.
How to test all of this in practice
Questions filter the field. A small first engagement proves the answer. Our fixed price Discovery exists partly for this reason: over two to three weeks you watch the team work, question decisions, and see exactly how recommendations get made and reviewed, including where AI helped and where it didn't. You finish with a requirements document, API recommendations, development considerations, and line item pricing, deliverables any team could execute. That's a far cheaper way to verify a partner's process than a signed build contract.
Wondering whether your shortlisted agency would survive these questions? Bring the list to your next vendor call, or start with a Discovery engagement and watch how we answer them ourselves. You can also reach out and ask us anything on this page directly.
Written by the team at Iron Forge Development, a U.S. based software commercialization firm that has helped launch 100+ products from idea to market.