The AI development market right now has a real AI-washing problem: a lot of firms added "AI" to their service list without deep engineering experience in the specific ways AI systems actually fail. Here's what we'd tell a friend evaluating agencies, not a list of things that make us look good.
Ask about specific failure modes, not just successes
Ask what went wrong on a recent project and how it got fixed. Every real engagement hits something unexpected: messier data than assumed, a vendor API limitation discovered mid-build, an architecture decision that needed revisiting. A firm that only has success stories with no rough edges is either not being honest or hasn't shipped enough real projects to have hit the normal bumps. Look for specific, technical detail in the answer, not a vague deflection.
Ask what they'd recommend against building
This is one of the more revealing questions you can ask. A team that says yes to every proposed use case is optimizing for closing the deal, not for what will actually work for your specific situation. A team with real judgment will tell you when an idea isn't worth building yet, or isn't the right first priority, even if it costs them the engagement.
Look for real technical depth in their case studies, not just outcomes
"We built an AI agent and it increased efficiency by 40%" is a headline, not evidence of engineering competence. Look for case studies that explain the actual architecture, the specific tradeoffs made and why, and what didn't work initially. That level of detail is much harder to fake than a percentage improvement, and it tells you whether the team actually understands what they built or is describing it from a distance.
Ask how they handle the AI-specific risks, not just "do you test"
Ask specifically about hallucination prevention (how do they ground responses?), prompt injection and security testing (do they test indirect vectors, not just direct chat input?), and cost predictability (how do they estimate and monitor spend?). Generic answers about "quality assurance" without addressing these AI-specific risks are a signal the team may not have deep experience with where AI systems actually break.
Check whether the scoping team and the build team are the same people
A handoff between the people who scoped your project and the people who actually build it is a real risk point where context gets lost. Ask directly whether that's the case, and if it is, ask how they prevent that loss of context.
Ask about their approach to model and architecture selection
A team that defaults to the same model and architecture for every client, regardless of the use case, is applying a template rather than engineering a solution. Ask how they'd decide between approaches for your specific situation, and listen for whether the answer references your actual constraints or sounds generic.
Red flags worth taking seriously
- A quote or timeline given before any real discovery conversation about your specific systems and data.
- No mention of evaluation, testing, or how they measure whether the system is actually working, beyond "it works well."
- Case studies with only outcome metrics and no technical detail about how they got there.
- An inability to describe a project that didn't go perfectly, or what they'd do differently.
- Pressure to commit quickly, before you've had a chance to evaluate their actual technical judgment.
What we'd want you to check about us
Read our case studies for the technical depth, not just the headline metrics. Ask us what we'd recommend against building for your specific situation. Ask how we handle grounding, security testing, and cost monitoring. We'd rather you evaluate us with real scrutiny than take a pitch at face value. That's the standard we think you should hold every AI development agency to.