Guide
Decisions & comparisons
How to choose an AI agent implementation partner: a scorecard
A scorecard for choosing an AI agent implementation partner: ten criteria, what good looks like, what to ask and which red flags should stop you.
- Give every candidate the same description of one workflow, otherwise the answers can't be compared.
- Score evidence, not claims: evaluation before build, an audit trail, cost per task, ownership of the repository.
- A good partner can tell you that an agent isn't needed.
Choose an AI agent implementation partner on evidence, not on labels: give every candidate the same description of one workflow and score their answers on ten criteria. These run from whether the partner will tell you an agent isn't needed, through evaluation before build and an audit trail, to the cost of a single task and whose repository the work lives in.
Before you look for a partner
This guide assumes you have decided to work with an outside partner (if not, start with build vs buy) and have one concrete workflow for agentic AI. The scorecard only works if every candidate answers about the same workflow; otherwise you are comparing pitches, not answers. If you're not sure the task needs an agent at all, first check whether an assistant or a workflow is enough and run through the agent or script checklist.
The partner scorecard: ten criteria
The scorecard has ten rows, and each one ends with a request for evidence. Zero points for a claim, one for a description, two for an artefact you have seen. Before you add anything up, mark the must-have criteria: a zero on any of them rules the candidate out.
Why pilots stall is covered in why AI pilots don't scale. Here, each row says what to require from a partner and what evidence to collect before you sign.
Scoring scale: 0 = no answer or only a claim · 1 = described but no artefact · 2 = backed by an artefact you have seen (a sample, a redacted document, a demo on your data).
The guide doesn't prescribe weights. Before scoring, write "M" next to the rows that are must-haves for you.
| Criterion | What good looks like | Evidence to ask for | Score 0-2 |
|---|---|---|---|
| 1. Tests whether the task needs an agent | Says in writing when a workflow or a script is enough, and why | A redacted example where they recommended the simpler pattern | 0 / 1 / 2 |
| 2. Real data path | The pilot uses the real, permissioned access path, not data pasted in once | The planned access method, the permission list and how access is revoked | 0 / 1 / 2 |
| 3. Evaluation before build | A test set with expected outputs and a quality threshold agreed before build, plus a stop rule (how to tell whether an agent works) | A redacted evaluation set or report from previous work | 0 / 1 / 2 |
| 4. Human in the loop and guardrails | Named approval points where a human in the loop decides with context rather than clicking "OK", plus rules that block misuse | The approval map for your workflow (agentic architecture that passes audit) | 0 / 1 / 2 |
| 5. Audit trail and observability | Every agent action logged: who asked, what the model returned, who approved. Observability and monitoring continue after go-live | A redacted log or trace, and who reads it after launch (monitoring AI agents) | 0 / 1 / 2 |
| 6. Integrations and permissions | Access through a controlled interface, such as MCP or a gateway, with least privilege. Credentials never sit in prompts or spreadsheets (MCP and agent integrations) | The integration diagram and the agent's permission list | 0 / 1 / 2 |
| 7. Model and vendor independence | The model can be swapped without breaking controls. Commercial ties to model, cloud or platform vendors are disclosed (build vs buy) | A written disclosure and a description of how a model swap would work | 0 / 1 / 2 |
| 8. Cost per unit of work | Running cost estimated per task before start, including usage-based lines, not only the build cost | A cost model for your workflow, with its assumptions (what AI in production really costs) | 0 / 1 / 2 |
| 9. Ownership and handover | A named process owner on your side. Code, prompts, evaluations and runbooks in your repository. A handover plan exists | The contract clause on IP and repository location, and the handover checklist | 0 / 1 / 2 |
| 10. Governance and data protection | Data classification, permission limits and GDPR and AI Act questions handled from day one, not bolted on later (governing AI agents) | The data-classification approach for your workflow and who signs it off (AI and GDPR, EU AI Act for businesses, data privacy) | 0 / 1 / 2 |
Operator's rule: a partner who can't tell you when an agent is unnecessary won't tell you when to stop building either.
Partner types, in brief
A software house, a large consultancy (Big 4), an independent architect, a platform vendor's implementation partner and an in-house team can all deliver an agent. The label doesn't predict the result: score each one on the same ten criteria.
| Partner type | What they usually bring | Scorecard rows to check first |
|---|---|---|
| Software house | An engineering team that builds and ships code | 1, 3, 8 |
| Large consultancy (Big 4) or integrator | Coordination across many areas and organisation-wide programmes | 2, 8, 9 |
| Independent architect or boutique | Architecture and control design, a small team | 5, 6, 9 |
| A platform vendor's implementation partner | Depth in one platform and its tooling | 6, 7, 8 |
| In-house team plus freelancers | Inside knowledge of the workflow and the data | 3, 5, 10 |
The middle column describes a role, not quality.
Questions for the first meeting
Each question maps to one scorecard row, so you can score the answer straight away.
- "When did you last advise a client against an agent, and what did you suggest instead?" (row 1)
- "What data will the pilot run on, and how do we revoke your access?" (row 2)
- "What quality threshold would stop the rollout, and who sets it?" (row 3)
- "Where exactly does a person approve what the agent does?" (row 4)
- "Who watches the logs after launch, and what do they do when something looks wrong?" (row 5)
- "What permissions will the agent get, and where will the credentials live?" (row 6)
- "What happens when we change the model?" (row 7)
- "What does one completed task cost?" (row 8)
- "Whose repository will the code and evaluations live in?" (row 9)
- "Which data in this workflow do you treat as sensitive, and who on our side signs that off?" (row 10)
Red flags
The most serious red flag is a partner who proposes an agent for everything and can't show an evaluation before build. The second is running costs that only appear after go-live.
- Everything is an agent. Every problem ends in an agent proposal, even when a rule or a script would do.
- The demo never runs on your data. Questions about your data get "after signing".
- Evaluation only after build. Quality is to be judged "in use", with no test set and no threshold.
- "Human in the loop" means an OK button. A person approves without context and with no real way to change the outcome.
- A closed platform with no exit. Prompts, evaluations and logs can't be exported.
- Running costs "we'll see". The proposal covers the build; the cost of one task comes after launch.
- The IP stays with the partner. Code and evaluations are licensed to you, not handed over.
- One team in the pitch, another in the build. Nobody can say who will actually build and run the agent.
How to run the selection, step by step
Choosing a partner takes five steps: one workflow, a short list, the same material for everyone, scoring with must-haves first, and a narrow first engagement.
- Describe one workflow on one page. Inputs, outputs, who does it today, which systems and data are involved, and what an error costs you.
- Shortlist a few candidates, ideally of different types, so you compare approaches, not just proposals.
- Send all of them the same description and scorecard. Ask for answers row by row, with evidence or a note that there is none.
- Score the answers and the evidence. Apply the must-haves first, then add up the remaining rows.
- Start with a narrow, time-boxed first engagement with stop conditions. What to buy first is in AI audit vs proof of concept, and how to set stop conditions and a single metric in small pilot, hard proof.
If you'd like to run this scorecard against one of your own workflows, tell us about it.
Terms in this guide
Frequently asked questions
- How do you choose an AI agent implementation partner?
- Write a description of one workflow, send it to a few candidates together with the scorecard, and compare their answers on ten criteria, asking for evidence on each. What matters most is evaluation before build, real data access, an audit trail, cost per task, and code and evaluations staying in your repository.
- A software house, a large consultancy (Big 4) or an independent architect?
- Any of them can implement an agent well. The partner type says more about scale and working style than about quality. Decide on the scorecard, not on the label.
- What should you require from a partner before they start building?
- A written scope for one workflow, an agreed test set with a quality threshold, a data-access plan with permissions, and an estimate of the cost of one task. Without those four, the build is a bet, not a project.
- Who should own the code, prompts and evaluations?
- The company deploying the agent. Code, prompts, evaluation sets and runbooks should go into its repository, with that written into the contract, so that changing partners doesn't mean starting from scratch.
- How do you compare proposals that describe different things?
- Bring them to a common baseline: the same workflow description for everyone and the same scorecard. A proposal that doesn't answer a must-have criterion drops out, whatever the price.