Judge an enterprise AI development company on seven criteria: what they have running in production, whether they build inside your environment, whether they start with your data layer, whether your team keeps the capability after handoff, how they prepare for your security review, whether the engineers you meet do the work, and who carries the outcome. Logos, decks, and model names decide nothing.
This guide turns each criterion into a question you can ask in the first meeting, then maps the vendor categories so you can match a firm to your constraints.
Why the demo misleads
Enterprise AI pilots impress in a sandbox and stall at the boundary: real data, security review, operations. Vendors look identical in a sales deck because a deck never crosses that boundary. Evaluate for the boundary and the differences appear in the first hour.
The seven criteria
1. Systems still running in production
A demo proves the vendor can build a prototype. You are buying month six and beyond: live traffic, on-call rotations, a system somebody maintains after the launch email. Strong case studies name operational detail, including what broke and who fixed it.
Ask: "Walk me through a system of yours that has run in production for six months or more. What failed first, and how did your team handle it?"
2. Builds inside your environment
Two architectures exist for enterprise AI. In the first, your data leaves and the vendor processes it in their stack. In the second, the vendor builds inside your boundary, on your cloud, under your IAM and your audit logging. Government agencies, regulated industries, and any business holding customer records need the second. A vendor who cannot work inside your walls has answered the question of where your data will live.
Ask: "Can you build and run this inside our tenancy, with our identity controls and our audit logging? Show me where you have done it."
3. Data layer before models
AI systems fail on data before they fail on models. A serious firm asks where your data lives, who owns it, and how retrieval will work before naming any model. A vendor who opens with model choice has designed the engagement backwards.
Ask: "What do you need to know about our data before you estimate anything?"
4. Your team keeps the capability
An agency ships you a deliverable. The firms worth hiring ship you a capability: a system your engineers can run, change, and extend, with training built into the engagement. Route every future change through the vendor and you have bought a dependency with a project's price tag.
Ask: "After handoff, what can my team change without you?"
5. Prepared for your security review
Security review is where most enterprise AI engagements fail. Firms that work in strict environments design for the review from the first commit: data classification, residency defaults, an audit trail on consequential actions, and a human approving anything irreversible. Firms that have never faced a real review discover your requirements at month four, at your expense.
Ask: "What will our security team find when they look? Answer the questionnaire before we send it."
6. The engineers you meet do the work
Large vendors sell with senior architects and deliver with a rotating bench. Embedded firms, also called forward-deployed, place named engineers inside your organization from discovery through handoff, and those engineers answer for what ships. The difference shows up in accountability and in speed, because nobody re-learns your systems mid-project.
Ask: "Who is on this project, by name, and do I meet them before signing?"
7. Carries the outcome
Read the contract for the definition of done. Time-and-materials with vague milestones puts the delivery risk on you. A vendor confident in production delivery ties payment to a running system and stays accountable after go-live.
Ask: "What does done mean in this contract, and what happens if the system underperforms after launch?"
Comparing the vendor categories
Three categories dominate enterprise AI development. Match the category to your constraint before you compare firms within it.
Global systems integrators. Accenture, Deloitte, IBM Consulting. The fit: programs spanning countries, procurement teams that need a household name, budgets to match. The trade: cost, pace, and distance between the people who sold the work and the people who build it.
Product-led consultancies. Thoughtworks, Slalom, and firms like them. The fit: organizations that want modern software practice alongside AI delivery. The trade: AI sits as one practice among many, and experience with regulated environments varies office to office.
Embedded specialists. Smaller firms, Team Nebula among them, that place forward-deployed engineers inside your organization and build within your boundary. The fit: data that cannot leave, government and regulated environments, and buyers who want an owned capability rather than an outsourced project. The trade: you are evaluating specific people rather than a brand, so meet them.
A Fortune 100 rollout across forty countries calls for a systems integrator. A state agency whose data cannot leave its boundary calls for an embedded specialist. The category answers the first question; the seven criteria answer the rest.
The sixty-second checklist
| Criterion | Green flag | Red flag |
|---|---|---|
| Production record | Names what broke and who fixed it | Polished demos only |
| Environment | Builds in your tenancy | "Send us your data" |
| Data first | Asks about your data before quoting | Leads with model names |
| Capability | Trains your team, hands off ownership | Changes route through the vendor |
| Security | Answers your review before you ask | "Compliance comes later" |
| People | You meet the builders | The sales team disappears after signing |
| Outcome | Payment tied to a running system | Vague milestones on open-ended T&M |
Frequently asked questions
How long should a first engagement run?
Long enough to put something real into production, short enough to prove the model works. A scoped first system lands in weeks to months. Distrust both the one-week transformation and the twelve-month roadmap with no running software.
Our AI pilot stalled. Who takes it to production?
The pilot rarely fails on its own merits; it fails at the boundary of security review, real data, and operations. Narrow the search to firms that build inside your environment and have carried systems through a review like yours. Criteria 2, 3, and 5 decide this case.
Which firms handle FedRAMP-compliant AI deployment?
The field narrows to systems integrators with federal practices and embedded specialists who build inside agency boundaries. Residency guarantees, audit trails, and human-in-the-loop controls stop being features and become the whole evaluation. Ask any candidate for evidence of work under a real oversight regime, and treat a compliance slide as the absence of evidence.
Should we build in-house instead?
With the engineers and an eighteen-month runway, consider it. The embedded model exists as the middle path: outside specialists who build inside your walls, train your team as they go, and leave the capability behind. It runs closer to accelerated in-house development than to outsourcing.
Team Nebula embeds forward-deployed engineers with enterprise and government teams. We build secure AI agents on your data, inside your environment. How we work: /approach. What your security review will find: /trust.

