Development
Web

How to Tell If a Software Partner Actually Knows How to Use AI (Not Just Claims To)

You can tell whether a software partner actually knows how to use AI by asking process questions instead of tool questions. A team with real AI maturity can show you where AI sits in its delivery workflow, who reviews what it produces, and how it keeps your code and data out of someone else's training set. A team that merely claims AI can name products. The difference surfaces within ten minutes if you ask the right things, and this guide gives you the exact questions.

"AI maturity" has joined price, portfolio, and communication as a standard line on 2026 vendor scorecards, and for good reason. Used well, AI genuinely changes how fast certain work gets done. Used carelessly, it ships unreviewed code into your product and your source into a consumer chatbot. When you evaluate an agency's AI maturity, you're really measuring two things at once: whether they capture the benefit, and whether they manage the risk.

Why this became a scorecard line

Two years ago, asking an agency about AI got you a shrug or a sales pitch. Now the buyer side evaluation guide lists AI maturity next to references and pricing, because the spread between agencies has become enormous and mostly invisible from the outside. Two firms can quote the same project, show similar portfolios, and run completely different engine rooms: one has senior engineers directing AI through a reviewed pipeline, the other has juniors shipping whatever the chatbot suggested. You'll pay a similar price either way. You will get very different software.

There's a second reason buyers started asking. The failure modes are new. A traditional bad agency shipped late and buggy, which you noticed. An AI careless agency can ship on time, with code nobody on their team fully read, trained on your data, and licensing questions nobody can answer. Those problems surface months later, usually during a security review or a fundraise, long after the invoice cleared.

What using AI well actually looks like

Inside a disciplined engineering team, AI shows up in specific, named places. Engineers use it to scaffold boilerplate, draft test suites, generate first pass documentation, and flag suspicious patterns during code review. Each of those is a bounded task with a human on the other end. The output enters the same pipeline as human written code: a senior engineer reads it, tests exercise it, and someone with their name on the project approves the merge.

Mature teams also have rules about what stays out of the tools. Client source code never goes into consumer tier accounts. Contracts with AI vendors include zero retention terms, so prompts and code aren't kept or used for training. Somebody at the agency can tell you, without checking, which AI subprocessors touch client material and under what terms. Boring answers like these are the strongest signal you'll get. Competence in this area sounds like policy and paperwork. It never sounds like magic.

Bolted on versus built in

The superficial version looks different up close. A few developers paste snippets into a chatbot when they get stuck, the proposal template gained an "AI powered delivery" bullet, and nobody can describe a review standard because there isn't one. Ask what the firm refuses to use AI for and you'll get a blank look, because a refusal list only exists where someone has actually thought about failure modes.

That's the core distinction. An agency that bolted AI on adopted a tool. An agency with real maturity redesigned part of its process around the tool, then wrote down where the tool stops. If you've read our guide on choosing a software development partner, this is the same principle applied to a new criterion: you're buying a process you can inspect, and a slide is never a process.

Seven questions to ask before you sign

Ask these in a live conversation, and pay attention to how fast the specifics arrive.

  • Where does AI sit in your delivery process today? You want named stages: scaffolding, tests, documentation, review assistance. Vague answers ("our developers use it throughout") mean nobody has mapped it.
  • What do you refuse to use AI for? A mature shop has a refusal list: architecture decisions, security-critical code without extra review, anything touching regulated data. No refusal list means no thinking about limits.
  • Who reviews AI generated code before it ships? The answer should be a senior human applying the same standard as the hand written code. "The AI checks itself" or "we spot-check" ends the conversation.
  • How do you keep our code and data out of model training? Listen for business tier accounts, zero retention agreements, and awareness of which subprocessors sit behind each tool. Then ask for it in writing.
  • How does AI change your estimate? An honest answer says some tasks got faster while judgment work costs the same. A promise of dramatic discounts because of AI is a red flag dressed as a benefit.
  • Tell me about a task where you turned AI off. Real experience includes failures. Teams that have actually used these tools at production standard can describe where the output wasted more time than it saved.
  • Will any of this appear in our contract? Data handling, IP ownership of AI assisted code, and review responsibility should all survive contact with legal. Verbal assurances about AI policy are worth what they cost.

Red flags that should end the evaluation

Some answers tell you everything you need to know.

  • The AI story is a list of tool names with no process attached.
  • Productivity multiples pitched as guarantees. "We're 10x faster now" is a marketing claim, and nobody seriously prices software that way.
  • No human review gate between AI output and your production code.
  • Silence, or improvisation, on the training data question.
  • The proposal itself reads like unedited AI output: generic, confident, and wrong about your business.
  • AI framed as a replacement for senior engineers rather than a tool senior engineers direct.

One caution in the other direction: don't reward theater. An agency that shows you an impressive internal "AI platform" but can't connect it to review standards or data terms has built a demo for buyers, and you're the audience.

What AI doesn't change

AI has made it cheap to produce plausible looking code, and that has made judgment more valuable, since someone still has to decide what to build, how to structure the data, where the security boundaries sit, and what happens when things fail. We see the gap every week in codebases that arrived through AI app builders. The interfaces look finished while authorization, data modeling, and error handling are missing underneath, which is why we wrote a separate guide on what to do with an AI built MVP and why turning an AI prototype into a real product is one of the most common questions we're asked.

The same logic applies to the partner you hire. AI maturity is one criterion stacked on top of the fundamentals: an accountable team, a transparent process, and ownership of outcomes. A shop with brilliant AI workflows and weak engineering judgment will simply ship the wrong thing faster.

Get the answers into the contract

Good verbal answers are the filter. The contract is the proof. Three clauses separate an agency that means it from one that improvised well on the call.

  • Data handling. Your code, credentials, and business data don't enter tools that retain or train on them. The agency names its AI subprocessors and the terms it holds with each, and updates you when the list changes.
  • IP ownership. You own the delivered work regardless of how it was produced, and the agency stands behind its right to assign it. AI assistance doesn't dilute your ownership, and the contract says so explicitly.
  • Review responsibility. The agency warrants that delivered code was reviewed by its engineers to its normal standard. This one sentence removes the "the AI wrote it" excuse from every future conversation.

An agency that already has language like this on file has thought the problem through. An agency that resists it just told you the marketing slide was doing all the work.

How much should this criterion weigh?

Less than the fundamentals, more than zero. AI maturity is a tiebreaker and a risk screen. It shouldn't outrank track record, communication quality, or whether the team has shipped products like yours. A shop that's mediocre at engineering and excellent at AI tooling is still mediocre at engineering. Where the criterion earns its place is at the margin: between two otherwise comparable partners, the one with a real AI process will likely deliver somewhat faster, and will definitely expose you to fewer of the new failure modes above.

How to test all of this in practice

Questions filter the field. A small first engagement proves the answer. Our fixed price Discovery exists partly for this reason: over two to three weeks you watch the team work, question decisions, and see exactly how recommendations get made and reviewed, including where AI helped and where it didn't. You finish with a requirements document, API recommendations, development considerations, and line item pricing, deliverables any team could execute. That's a far cheaper way to verify a partner's process than a signed build contract.

Wondering whether your shortlisted agency would survive these questions? Bring the list to your next vendor call, or start with a Discovery engagement and watch how we answer them ourselves. You can also reach out and ask us anything on this page directly.

Written by the team at Iron Forge Development, a U.S. based software commercialization firm that has helped launch 100+ products from idea to market.

FAQs

What does AI maturity mean for a software agency?
AI maturity is how deliberately an agency has integrated AI into its delivery process: named use cases, human review gates on everything AI produces, and written rules about what client code and data can touch which tools. It's a process quality, so measure it with process questions. Ask any agency where AI sits in its workflow and what it refuses to use AI for. Mature teams answer in specifics within seconds, while teams that only adopted the marketing language stall on the second question.
What questions should I ask a software agency about how it uses AI?
Ask where AI sits in their delivery process, what they refuse to use it for, who reviews AI-generated code before it ships, how they keep your code and data out of model training, how AI changes their estimates, and whether any of it appears in the contract. The pattern across the answers matters more than any single one. Specific, boring, policy-shaped answers signal real experience; tool names and productivity claims signal a slide. Get the data-handling and review commitments in writing before you sign.
Is AI-generated code safe to use in production?
Yes, when it goes through the same review pipeline as human-written code: a senior engineer reads it, tests exercise it, and someone accountable approves the merge. Unreviewed AI code is where the risk lives, because the output looks plausible while skipping the things that break under real users, like authorization, error handling, and data modeling. Ask any team who reviews AI output and to what standard. If the answer isn't a named human applying their normal bar, keep looking.
Will AI make my software project cheaper to build?
Somewhat, on some work. AI speeds up scaffolding, tests, and documentation, while the expensive parts of software (deciding what to build, architecture, security, integration) still run on senior judgment. Treat dramatic AI discounts as a red flag rather than a bargain, since they usually mean review steps got skipped. A more reliable way to capture the savings is a fixed price set after discovery, where efficiency shows up as more scope for the same number instead of an estimate that grows later.
How do I keep my code and data out of AI training sets?
Require it contractually. Your agency should use business-tier AI accounts with zero-retention terms, name the AI subprocessors that touch your material, and commit in writing that your code, credentials, and business data never enter tools that retain or train on them. Verbal assurances aren't enforcement. Ask any partner to show you the clause they already use. A team that handles this well has the language on file, and a team that improvises it on the call probably hasn't thought about it before.
Can I turn an AI-generated prototype into a real MVP?
Often, yes. AI app builders are genuinely good at producing a working demo fast, and that demo is useful evidence of what you want. What it usually lacks is what production requires: security, a sound data model, testing, and an architecture that survives real users. The practical path is a code review of what you have, then a plan for what to keep, what to rebuild, and what it will cost. That's exactly the kind of question a discovery engagement answers.
What should I look for in a software development partner?
Look for in-house multidisciplinary teams, a transparent process, relevant portfolio work, clear communication, transparent pricing, and ownership of outcomes rather than just tasks. Ask how they handle scope changes and post-launch support.
What technologies do you use?
Our main stack includes React, Node.js, Python, and TypeScript, with databases like PostgreSQL, MongoDB, and Firestore, deployed on AWS, or Google Cloud. We choose the tools that fit your product's needs and your team's future, not a one-size-fits-all template.

Recommended read

Seven-question AI maturity checklist for evaluating software agencies

How to Tell If a Software Partner Actually Knows How to Use AI (Not Just Claims To)

AI maturity is now a standard line on vendor scorecards. Here are the seven questions, the red flags, and the contract clauses that show whether a software partner actually knows how to use AI.

Development
Web
Four criteria for choosing an AI app rescue and rebuild agency

Best Agencies for AI App Rescue and Rebuilds

Six agencies compared for rescuing or rebuilding an AI-generated app, scored on the four criteria that decide the outcome: code quality, architecture, delivery discipline, and post launch support. Plus the five questions that sort any shortlist.

Development
AI

How Enterprise Teams Should Choose a Custom Web Development Partner

Most enterprise software projects don't fail in the build. They fail in the security review, or when the pilot can't scale. Here's how enterprise IT, product, and innovation teams should evaluate a custom web development partner, and what enterprise grade requires in practice.

Development
Web
Front-end dev