Aurora AITell us your case

Offering

ServicesProductsCase studies

For whom

Private EquityEnterpriseSMB
ServicesProductsCase studiesAboutBlogContact

Knowledge base

Start hereWikiGlossaryGuides

Guide

Decisions & comparisons

How to choose an AI agent implementation partner: a scorecard

A scorecard for choosing an AI agent implementation partner: ten criteria, what good looks like, what to ask and which red flags should stop you.

Choose an AI agent implementation partner on evidence, not on labels: give every candidate the same description of one workflow and score their answers on ten criteria. These run from whether the partner will tell you an agent isn't needed, through evaluation before build and an audit trail, to the cost of a single task and whose repository the work lives in.

Before you look for a partner

This guide assumes you have decided to work with an outside partner (if not, start with build vs buy) and have one concrete workflow for agentic AI. The scorecard only works if every candidate answers about the same workflow; otherwise you are comparing pitches, not answers. If you're not sure the task needs an agent at all, first check whether an assistant or a workflow is enough and run through the agent or script checklist.

The partner scorecard: ten criteria

The scorecard has ten rows, and each one ends with a request for evidence. Zero points for a claim, one for a description, two for an artefact you have seen. Before you add anything up, mark the must-have criteria: a zero on any of them rules the candidate out.

Why pilots stall is covered in why AI pilots don't scale. Here, each row says what to require from a partner and what evidence to collect before you sign.

Scoring scale: 0 = no answer or only a claim · 1 = described but no artefact · 2 = backed by an artefact you have seen (a sample, a redacted document, a demo on your data).

The guide doesn't prescribe weights. Before scoring, write "M" next to the rows that are must-haves for you.

CriterionWhat good looks likeEvidence to ask forScore 0-2
1. Tests whether the task needs an agentSays in writing when a workflow or a script is enough, and whyA redacted example where they recommended the simpler pattern0 / 1 / 2
2. Real data pathThe pilot uses the real, permissioned access path, not data pasted in onceThe planned access method, the permission list and how access is revoked0 / 1 / 2
3. Evaluation before buildA test set with expected outputs and a quality threshold agreed before build, plus a stop rule (how to tell whether an agent works)A redacted evaluation set or report from previous work0 / 1 / 2
4. Human in the loop and guardrailsNamed approval points where a human in the loop decides with context rather than clicking "OK", plus rules that block misuseThe approval map for your workflow (agentic architecture that passes audit)0 / 1 / 2
5. Audit trail and observabilityEvery agent action logged: who asked, what the model returned, who approved. Observability and monitoring continue after go-liveA redacted log or trace, and who reads it after launch (monitoring AI agents)0 / 1 / 2
6. Integrations and permissionsAccess through a controlled interface, such as MCP or a gateway, with least privilege. Credentials never sit in prompts or spreadsheets (MCP and agent integrations)The integration diagram and the agent's permission list0 / 1 / 2
7. Model and vendor independenceThe model can be swapped without breaking controls. Commercial ties to model, cloud or platform vendors are disclosed (build vs buy)A written disclosure and a description of how a model swap would work0 / 1 / 2
8. Cost per unit of workRunning cost estimated per task before start, including usage-based lines, not only the build costA cost model for your workflow, with its assumptions (what AI in production really costs)0 / 1 / 2
9. Ownership and handoverA named process owner on your side. Code, prompts, evaluations and runbooks in your repository. A handover plan existsThe contract clause on IP and repository location, and the handover checklist0 / 1 / 2
10. Governance and data protectionData classification, permission limits and GDPR and AI Act questions handled from day one, not bolted on later (governing AI agents)The data-classification approach for your workflow and who signs it off (AI and GDPR, EU AI Act for businesses, data privacy)0 / 1 / 2
Operator's rule: a partner who can't tell you when an agent is unnecessary won't tell you when to stop building either.

Partner types, in brief

A software house, a large consultancy (Big 4), an independent architect, a platform vendor's implementation partner and an in-house team can all deliver an agent. The label doesn't predict the result: score each one on the same ten criteria.

Partner typeWhat they usually bringScorecard rows to check first
Software houseAn engineering team that builds and ships code1, 3, 8
Large consultancy (Big 4) or integratorCoordination across many areas and organisation-wide programmes2, 8, 9
Independent architect or boutiqueArchitecture and control design, a small team5, 6, 9
A platform vendor's implementation partnerDepth in one platform and its tooling6, 7, 8
In-house team plus freelancersInside knowledge of the workflow and the data3, 5, 10

The middle column describes a role, not quality.

Questions for the first meeting

Each question maps to one scorecard row, so you can score the answer straight away.

  1. "When did you last advise a client against an agent, and what did you suggest instead?" (row 1)
  2. "What data will the pilot run on, and how do we revoke your access?" (row 2)
  3. "What quality threshold would stop the rollout, and who sets it?" (row 3)
  4. "Where exactly does a person approve what the agent does?" (row 4)
  5. "Who watches the logs after launch, and what do they do when something looks wrong?" (row 5)
  6. "What permissions will the agent get, and where will the credentials live?" (row 6)
  7. "What happens when we change the model?" (row 7)
  8. "What does one completed task cost?" (row 8)
  9. "Whose repository will the code and evaluations live in?" (row 9)
  10. "Which data in this workflow do you treat as sensitive, and who on our side signs that off?" (row 10)

Red flags

The most serious red flag is a partner who proposes an agent for everything and can't show an evaluation before build. The second is running costs that only appear after go-live.

How to run the selection, step by step

Choosing a partner takes five steps: one workflow, a short list, the same material for everyone, scoring with must-haves first, and a narrow first engagement.

  1. Describe one workflow on one page. Inputs, outputs, who does it today, which systems and data are involved, and what an error costs you.
  2. Shortlist a few candidates, ideally of different types, so you compare approaches, not just proposals.
  3. Send all of them the same description and scorecard. Ask for answers row by row, with evidence or a note that there is none.
  4. Score the answers and the evidence. Apply the must-haves first, then add up the remaining rows.
  5. Start with a narrow, time-boxed first engagement with stop conditions. What to buy first is in AI audit vs proof of concept, and how to set stop conditions and a single metric in small pilot, hard proof.

If you'd like to run this scorecard against one of your own workflows, tell us about it.

Terms in this guide

Have a concrete process, deal or bottleneck? Tell us your case.

Tell us your case See how we help

Frequently asked questions

How do you choose an AI agent implementation partner?
Write a description of one workflow, send it to a few candidates together with the scorecard, and compare their answers on ten criteria, asking for evidence on each. What matters most is evaluation before build, real data access, an audit trail, cost per task, and code and evaluations staying in your repository.
A software house, a large consultancy (Big 4) or an independent architect?
Any of them can implement an agent well. The partner type says more about scale and working style than about quality. Decide on the scorecard, not on the label.
What should you require from a partner before they start building?
A written scope for one workflow, an agreed test set with a quality threshold, a data-access plan with permissions, and an estimate of the cost of one task. Without those four, the build is a bet, not a project.
Who should own the code, prompts and evaluations?
The company deploying the agent. Code, prompts, evaluation sets and runbooks should go into its repository, with that written into the contract, so that changing partners doesn't mean starting from scratch.
How do you compare proposals that describe different things?
Bring them to a common baseline: the same workflow description for everyone and the same scorecard. A proposal that doesn't answer a must-have criterion drops out, whatever the price.